Spatial memory: benchmark and implement receipt preserving trajectory compression
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
- Tipo de issue
- Nueva funcionalidad
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- rust
- Área
- ai, backend, data, performance
Línea de trabajo
Comienza en crates/ruvector-agent-memory leyendo las APIs existentes de compactación de memoria temporal o de agente e identificando el límite de trayectoria espacial adecuado. Revisa las baselines propuestas, los requisitos de receipt y el promotion gate antes de redactar el ADR. El trabajo requiere una implementación controlada por feature flag, pruebas unitarias, de propiedades y adversariales, benchmarks repetidos deterministas, validación de plataforma y el informe de benchmarks especificado.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Problem
Long lived spatial systems currently retain rich observations and vector memories, but dynamic object trajectories can grow with observation duration. This creates avoidable storage, retrieval, and context costs, while aggressive forgetting can destroy the exact state transitions an agent later needs to answer spatial questions.
Fresh research: Linguistic Trajectory Encoding, arXiv:2609.04802, submitted 2026-09-04, reports a hybrid per object representation using linguistic motion phases plus sparse spatial and visual anchors. On its Spatial Memory Benchmark it reports 45.3% semantic trajectory retrieval and 48.7% long horizon object retrieval versus 31.9% and 34.4% for the strongest reported baseline, with 8.7x to 26.1x trajectory compression and sub second queries over 24 hour video.
Source: https://arxiv.org/abs/2609.04802
This should not be copied literally into RuVector. Authoritative geometry, provenance, and witnessed source state must remain symbolic and verifiable. Language or latent summaries should be residual indexes over exact evidence, never replacements for it.
Proposed architecture
Target: crates/ruvector-agent-memory plus a small spatial trajectory module or namespace if existing boundaries require it.
Add a receipt preserving SpatialTrajectory representation with:
- object or track identifier
- ordered exact anchor points with timestamp and covariance
- source receipt or state root references per retained interval
- optional semantic phase description and embedding
- compression error bound in metres and time
- observation gap semantics that distinguish last seen from interpolated motion
- deterministic simplification policy, initially Douglas Peucker or equivalent geometric simplification
- exact rollback path to the source evidence references
Keep exact source events and geometry outside the compact representation according to retention policy. Never allow semantic summaries to raise confidence or overwrite authoritative coordinates.
Baselines
- Dense per sample trajectory storage
- Existing RuVector temporal or agent memory compaction
- Deterministic geometric trajectory simplification only
- Hybrid semantic plus spatial trajectory compression
Benchmark plan
Use fixed seeds and identical source trajectories. Prefer RuView real captured tracks when available; add an external SMB compatible evaluation only if licensing and dataset access permit.
Measure:
- bytes per trajectory and compression ratio
- semantic trajectory retrieval success
- last occurrence retrieval success
- spatial query error in metres
- temporal retrieval error
- p50 and p95 query latency
- index build latency
- peak memory
- receipt verification coverage
- deterministic reproduction across repeated runs
Promotion gate:
- at least 8x storage reduction on long horizon tracks
- no more than 5 percentage points absolute loss versus dense authoritative retrieval on geometric and temporal queries
- p95 query latency below 100 ms on a 24 hour equivalent workload
- 100% retained interval coverage by source receipt or state root references
- zero confidence amplification from summaries
- deterministic byte identical compact output for identical inputs and configuration
Security and privacy
Threats to test:
- prompt injection in generated trajectory descriptions
- malicious or oversized semantic text
- forged source receipt references
- object identity leakage
- location privacy leakage
- adversarial trajectories designed to trigger pathological anchor counts
- resource exhaustion from oscillatory motion
Semantic text must be treated as untrusted data. Cap lengths and counts. Do not execute or interpolate instructions from memory content. Validate all source receipt bindings before returning an authoritative answer.
Compatibility and migration
Additive feature behind a feature flag first. Existing dense trajectory and agent memory APIs remain the fallback. No RVF wire change until an explicit profile and golden fixtures exist.
MetaHarness plan
Use the latest stable ruvnet/metaharness with repository specific objective functions for storage ratio, retrieval quality, latency, and provenance coverage. Darwin mutations are permitted only for simplification thresholds or bounded policy parameters and must fail closed on any provenance, determinism, or retrieval regression.
Definition of done
- ADR with context, alternatives, privacy, security, benchmark and rollback
- production Rust implementation behind a feature flag
- unit, property and adversarial tests
- deterministic benchmark with at least three repeated runs
- ARM64 and x86 validation
- WASM evaluation if the representation is exposed to browser spatial queries
- before and after benchmark report
- focused PR with reproduction instructions
Production classification
Production candidate only after the benchmark gate passes. Until then: Experimental.
- Lenguaje dominante
- Rust
- Estrellas
- 4.5k
- Forks
- 603
- Merge medio
- 2 d 11 h
- PR fusionados (30 d)
- 33
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de ruvnet/RuVector
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
Los mantenedores suelen responder en 1 día
-
adr phase-w4-3 pir stretch wave-4
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 4/5 3-5 días Aptitud para principiantes 45/100
Los mantenedores suelen responder en 1 día
-
Dificultad 5/5 Más de una semana Aptitud para principiantes 32/100
Los mantenedores suelen responder en 1 día
Todos los issues de ruvnet/RuVector
Issues similares
-
`categorize_command` has no `uv` arm, so every `rtk uv …` row counts as `other` in the ecosystem mixAbiertoarea:api bug good first issue priority:low
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
rtk-ai/rtk#4316 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 92/100
Los mantenedores suelen responder en 1 día
-
area/cli kind/bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 90/100
Los mantenedores suelen responder en 1 día
-
enhancement
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
-
good first issue open-endedness: low type: new feature
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100