Pipeline SWIFT query selection appears to use exact marginal errors
Maintainers usually reply within 2 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
Research direction
Start in dpsynth/pipeline_transformations/swift.py by tracing compute_exact_marginals, compute_errors, the “Swift Select Queries” budget, and swift.select_queries; compare this path with the local discrete_mechanisms.swift selection logic. Review draft PR #31, and consider the issue complete when selection scores are protected by a separate budget, measurement uses the remaining budget, and diagnostics do not publish exact errors.
Written by the indexing model from the issue text.
Description
Problem
The scalable pipeline SWIFT path appears to use exact-error-driven query selection.
In dpsynth/pipeline_transformations/swift.py, the pipeline computes exact candidate marginals, converts them into errors with marginals_computations.compute_errors(...), requests a budget named Swift Select Queries, and then passes the errors into swift.select_queries(...).
However, the selection scores do not appear to be noised before swift.select_queries(...); noise is added only later to the selected marginal measurements.
Why this matters
The selected clique tree / selected workload is itself data-dependent output. Concretely, the junction-tree topology, the selected clique set, and (when diagnostics are enabled) the exact error scores are all released and all depend on exact high-order marginals. If selection is driven by exact marginal errors, the later noisy measurement step does not protect the information leaked by which queries were selected.
This is separate from the local discrete_mechanisms.swift path, which has its own score-noising logic (_compute_initial_errors adds noise funded by a dedicated selection budget). The issue here is the scalable pipeline transformation path, which has no equivalent noising step.
Local evidence
Reviewed at commit 18c2c951bd2923f889f6e3b2b757e01aaae398ee; re-verified still present at current main (91e9181) — the pipeline path still feeds unnoised errors from compute_errors into swift.select_queries.
Relevant lines in the current tree:
dpsynth/pipeline_transformations/swift.py:exact_marginals = marginals_computations.compute_exact_marginals(...)dpsynth/pipeline_transformations/swift.py:errors = marginals_computations.compute_errors(...)dpsynth/pipeline_transformations/swift.py: budget request namedSwift Select Queriesdpsynth/pipeline_transformations/swift.py:return swift.select_queries(errors_dict, ...)dpsynth/pipeline_transformations/swift.py: noise is added at the laterAdd noise to selected marginalsstagedpsynth/pipeline_transformations/marginals_computations.py:compute_errors(...)usesexact_valsfrom exact marginals
Possible fix
Account separately for selection and measurement. Add DP noise to the vector of SWIFT candidate error scores before clique-tree/query selection, and use the remaining measurement budget only for selected marginal measurement. Diagnostic output should avoid publishing exact errors.
Draft PR
I opened a draft fix here: https://github.com/google/dpsynth/pull/31
- Dominant language
- Python
- Stars
- 32
- Forks
- 13
- Avg merge
- 1d 19h
- Merged PRs (30d)
- 20
Getting set up
- No Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from google/dpsynth
-
import dpsynth fails because mbi.Dataset is registered as a JAX dataclass twicePossibly taken @hanzalaareeb claimed this 9 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 2 days
-
Clarify installation requirements in quickstart.ipynbPossibly taken @hanzalaareeb claimed this 8 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
google/dpsynth#194 · 1 comment ·
Maintainers usually reply within 2 days
-
`IndependentConfig` synthesis raises "Cliques must be unique."May be free again A pull request for this issue was closed without being merged. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 2 days
-
Windows install of pylock.toml fails on the pipeline extra due to missing Windows wheel for python-dpPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 3/5 1-2 days Newbie friendliness 72/100
Maintainers usually reply within 2 days
-
Add an option to control the maximum marginal degree in AIM workload constructionPossibly taken A pull request linked to this issue is open or already merged. Open
Difficulty 3/5 1-2 days Newbie friendliness 65/100
google/dpsynth#199 · 3 comments ·
Maintainers usually reply within 2 days
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Maintainers usually reply within 3 days
-
Negation with "not" and "no" is ignored during sentiment analysisPossibly taken @vivek-3728 claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
techcsispit/mess-mood#11 · 1 comment ·
-
changelog investigate
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
ramnes/notion-sdk-py#408 ·
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 83/100
btclib-org/btclib-wallet#267 ·
Maintainers usually reply within 1 day
-
good first issue tech-debt
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
knnmelprop/YAADO#111 ·