PyDP silently truncates and converts `np.float32` into `int64`, leading to sensitivity underestimation
还没有人认领这个 Issue。
评估
调研方向
从第 52 行附近 src/bindings/PyDP/mechanisms/mechanism.cpp 中的绑定声明开始,检查 integer 和 double 重载是如何注册的。使用提供的 np.float32 示例复现这一行为,然后验证 np.float32 值会选择 double 绑定且不会被截断;重新运行与 mechanism 绑定相关的测试。
由索引模型根据 Issue 内容生成。
描述
In the code that registers the bindings with the C++ building blocks library, the int64 binding is defined before the double binding. This has a surprising consequence: pybind11, after doing a pass that sees if the input type matches the declared type, it then tries overloads in order with an implicit conversion path. The int caster apparently accepts a np.float32 and silently truncates its input, so values of this type automatically use the int64 binding.
Silent truncation predictably negative consequences on the sensitivity analysis: 0.95 becomes 0 and 1.05 becomes 1, so sensitivity increases unexpectedly. This is observable in PyDP:
import numpy as np
from pydp.algorithms.numerical_mechanisms import LaplaceMechanism
mechanism = LaplaceMechanism(epsilon=1.0, sensitivity=0.1)
print("np.float32, D :", [mechanism.add_noise(np.float32(0.95)) for _ in range(10)])
print("np.float32, D':", [mechanism.add_noise(np.float32(1.05)) for _ in range(10)])
and also affects PipelineDP via add_dp_noise, or VECTOR_SUM with float32 values. This happens even though the value is multiplied by 1.0, maybe to try and cast it to float? Sadly this doesn't work, since 1.0 * np.float32(1.7) still has type np.float32.
Here's a repro for add_dp_noise:
import numpy as np
import pipeline_dp
for name, value in (("D ", np.float32(0.95)), ("D'", np.float32(1.05))):
releases = []
for _ in range(10):
accountant = pipeline_dp.NaiveBudgetAccountant(total_epsilon=1, total_delta=0)
engine = pipeline_dp.DPEngine(accountant, pipeline_dp.LocalBackend())
params = pipeline_dp.aggregate_params.AddDPNoiseParams(noise_kind=pipeline_dp.NoiseKind.LAPLACE, l0_sensitivity=1,
linf_sensitivity=0.1)
result = engine.add_dp_noise([("partition", value)], params)
accountant.compute_budgets()
releases.append(list(result)[0][1])
print(name, releases)
other PipelineDP aggregations (sum, mean, variance) convert to float64 and so aren't vulnerable to this.
This issue doesn't happen with np.float64, because np.float64 is a subclass of Python's native floats (while np.float32 is not). Therefore, the right binding is selected by pybind11 during the first pass.
Registering the double binding first would fix it.
- 主要语言
- Python
- 星标
- 550
- 派生
- 142
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
OpenMined/PyDP 的其他 Issue
-
Type: New Feature :heavy_plus_sign:
难度 4/5 3-5 天 新手友好度 45/100
-
Type: Question :grey_question:
难度 3/5 1-2 天 新手友好度 25/100
-
Type: Question :grey_question:
难度 4/5 3-5 天 新手友好度 32/100
-
Raise different errors based on the error code returned from the DP Library可能已有人在做 @kalra-mohit 于 77 天前认领。 未关闭Type: Improvement :chart_with_upwards_trend:
难度 3/5 1-2 天 新手友好度 35/100
-
Type: New Feature :heavy_plus_sign:
难度 3/5 1-2 天 新手友好度 38/100
相似的 Issue
-
area/install reliability
难度 2/5 1-3 小时 新手友好度 75/100
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 83/100
FluidNumerics/fluid-walk-blocker#191 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
TransformerLensOrg/TransformerLens#1868 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复
-
难度 1/5 1 小时以内 新手友好度 85/100
climate-analytics-lab/jax-gcm#1057 ·
维护者通常 1 天内回复