SOTA privacy experiment: decouple stored embeddings with shadow queries
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- wasm
- Lĩnh vực
- databases, machine-learning, performance, security
Hướng nghiên cứu
Không có tệp triển khai hoặc bài kiểm thử nào được nêu tên. Hãy bắt đầu bằng cách lập bản đồ năm điều kiện cố định và cuộc tấn công thích ứng ở cấp tập hợp đối với hoạt động lập chỉ mục RuVector hiện có, sau đó xác định các đầu vào benchmark và phép đo được liệt kê trong báo cáo. Được coi là hoàn tất khi đáp ứng mọi promotion gate, duy trì việc xóa và cô lập tenant, hoặc ghi lại sự phản chứng và từ chối thiết kế mà không thay đổi các giá trị mặc định của production.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Finding
Shadow Queries for Private Retrieval in Vector Databases, submitted 2026-09-04, targets embedding inversion rather than query privacy. Instead of storing a document embedding directly, SHAQ generates diverse semantically relevant shadow queries per document and stores embeddings of those queries. The originating team reports text recovery as low as 0.2104, up to 19.50% more defended tokens than baseline defenses, MAP@10 up to 0.7967, and up to 5.53% utility improvement over the compared defense.
Evidence class: originating-team measured, not independently reproduced by RuV. The arXiv manuscript is under the arXiv perpetual non-exclusive license. No implementation license suitable for code reuse was verified in this cycle, so this issue imports no source code.
RuV implication
This is orthogonal to issue #967. #967 addresses outsourced query privacy under a two-server non-collusion model. This issue addresses stored-embedding inversion if an attacker obtains or queries the vector representation itself.
Potential reuse: RuVector hosted indexes, Core Memory enterprise memory, Cognitum RAG, MCP retrieval, RVF provenance, and bounded RuVector WASM stores.
Reversible experiment
Compare five frozen conditions:
A. ordinary document embeddings
B. additive-noise defense at matched retrieval utility
C. vector scaling or normalization defense at matched retrieval utility
D. one shadow-query embedding per document
E. diverse multi-shadow-query indexing with a fixed generation budget
Use at least two embedding models and three corpora with materially different document length and semantic density.
Attack model
Reproduce a modern embedding inversion baseline such as vec2text, then add an adaptive attacker that knows the defense architecture and generation prompt family but not secret tenant data.
Required benchmark report
Record corpus digest, embedding model and version, shadow generator and version, prompts, seeds, document count, query count, index size, construction latency, generation tokens and cost, MAP@10, recall@10, p50/p95/p99 query latency, storage multiplier, inversion recovery, defended-token rate, CPU, memory, and energy where measurable. Include malformed documents, low-information documents, duplicate content, updates, deletes, distribution shift, and adversarial query patterns.
Promotion gate
A shadow-query design advances only if all are true:
- inversion recovery falls by at least 50% relative to ordinary document embeddings
- retrieval quality loses no more than 1 absolute point of MAP@10 or recall@10 against the stronger baseline
- p95 query latency regresses by less than 10%
- index storage stays below 3 times the document-embedding baseline
- generation cost is amortized within the declared customer workload horizon
- deletion and tenant isolation semantics remain exact
Falsification
The defense may simply move sensitive information from a document embedding into several semantically revealing query embeddings. Test an adaptive attacker over the entire per-document shadow set, not one vector at a time. If privacy gain disappears under set-level attacks, reject the design.
A cheaper dimensionality reduction or quantization baseline must also be included. If it performs within variance at lower cost, prefer the simpler defense.
Security and governance
Shadow queries are derived sensitive artifacts and inherit the source document's tenant, retention, deletion, and access policy. They cannot be logged or reused across tenants. Retrieval quality is not evidence of privacy. Privacy measurements cannot authorize release or declassification.
Existing RuVector indexing remains the rollback path. No production format migration or default change is authorized.
- Ngôn ngữ chính
- Rust
- Star
- 4.5k
- Fork
- 603
- Merge trung bình
- 2 ngày 9 giờ
- Pull request đã merge (30 ngày)
- 34
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của ruvnet/RuVector
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 83/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
Maintainer thường phản hồi trong vòng 1 ngày
-
adr phase-w4-3 pir stretch wave-4
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của ruvnet/RuVector
Issue tương tự
-
`sysknife history --help` says --since takes ISO-8601, and the parser refuses offsets and bare datesĐang mởbug easy good first issue help wanted
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 94/100
lacs-project/sysknife#519 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
enhancement
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 72/100
-
area:breg bug criticality:p3 triage:needs-implementation
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
registrystack/registry-stack#1699 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
documentation
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 84/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
lbjlaq/Antigravity-Manager#3539 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày