Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

PyDP silently truncates and converts `np.float32` into `int64`, leading to sensitivity underestimation

未关闭 适合新手
#499 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
83/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
cpp, python
领域
backend

调研方向

从第 52 行附近 src/bindings/PyDP/mechanisms/mechanism.cpp 中的绑定声明开始,检查 integer 和 double 重载是如何注册的。使用提供的 np.float32 示例复现这一行为,然后验证 np.float32 值会选择 double 绑定且不会被截断;重新运行与 mechanism 绑定相关的测试。

由索引模型根据 Issue 内容生成。

描述

Type: Bug :bug:

In the code that registers the bindings with the C++ building blocks library, the int64 binding is defined before the double binding. This has a surprising consequence: pybind11, after doing a pass that sees if the input type matches the declared type, it then tries overloads in order with an implicit conversion path. The int caster apparently accepts a np.float32 and silently truncates its input, so values of this type automatically use the int64 binding.

Silent truncation predictably negative consequences on the sensitivity analysis: 0.95 becomes 0 and 1.05 becomes 1, so sensitivity increases unexpectedly. This is observable in PyDP:

import numpy as np
from pydp.algorithms.numerical_mechanisms import LaplaceMechanism

mechanism = LaplaceMechanism(epsilon=1.0, sensitivity=0.1)
print("np.float32, D :", [mechanism.add_noise(np.float32(0.95)) for _ in range(10)])
print("np.float32, D':", [mechanism.add_noise(np.float32(1.05)) for _ in range(10)])

and also affects PipelineDP via add_dp_noise, or VECTOR_SUM with float32 values. This happens even though the value is multiplied by 1.0, maybe to try and cast it to float? Sadly this doesn't work, since 1.0 * np.float32(1.7) still has type np.float32.

Here's a repro for add_dp_noise:

import numpy as np
import pipeline_dp

for name, value in (("D ", np.float32(0.95)), ("D'", np.float32(1.05))):
    releases = []
    for _ in range(10):
        accountant = pipeline_dp.NaiveBudgetAccountant(total_epsilon=1, total_delta=0)
        engine = pipeline_dp.DPEngine(accountant, pipeline_dp.LocalBackend())
        params = pipeline_dp.aggregate_params.AddDPNoiseParams(noise_kind=pipeline_dp.NoiseKind.LAPLACE, l0_sensitivity=1,
                                                               linf_sensitivity=0.1)
        result = engine.add_dp_noise([("partition", value)], params)
        accountant.compute_budgets()
        releases.append(list(result)[0][1])
    print(name, releases)

other PipelineDP aggregations (sum, mean, variance) convert to float64 and so aren't vulnerable to this.

This issue doesn't happen with np.float64, because np.float64 is a subclass of Python's native floats (while np.float32 is not). Therefore, the right binding is selected by pybind11 during the first pass.

Registering the double binding first would fix it.

主要语言
Python
星标
550
派生
142
PR 合并指标
30 天内没有已合并 PR

环境准备

  • 提供 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

OpenMined/PyDP 的其他 Issue

查看 OpenMined/PyDP 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。