Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

bench(spatial): add trajectory, measured pose, and calibration robustness gates for world models

Abierto
#951 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
rust

Línea de trabajo

Revisa crates/ruvector-robotics y sus espacios de nombres existentes de Pose, Trajectory, SpatialIndex y benchmarks antes de decidir si encaja mejor src/bridge/pose_coverage.rs o src/eval/spatial_robustness.rs. Empieza con las pruebas unitarias, de propiedades, de regresión y de rendimiento descritas en el issue, usando fixtures sintéticos en lugar de datos PIVOT incluidos en el repositorio. Se considera terminado cuando los informes reproducibles distinguen entre trayectorias representadas y no vistas, además de poses medidas y optimizadas, con digests deterministas y fallos explícitos para datos no válidos.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Problem

Spatial and Gaussian world models can look strong when evaluated on held out views that come from the same capture trajectory family, after offline pose optimization, or with scene optimized camera intrinsics. Those conditions are materially easier than deployed robotics, mobile sensing, and multimodal RuView plus LiDAR pipelines.

PIVOT, submitted 2026 08 26, makes this gap explicit. It evaluates complete unseen trajectories separately from held out views on represented trajectories, preserves sensor measured poses alongside COLMAP optimized poses, preserves reusable calibrated intrinsics alongside scene optimized intrinsics, and introduces a directed pose space Chamfer distance that describes evaluation trajectory coverage relative to training poses.

Research:

https://arxiv.org/abs/2608.25401

Source code is MIT licensed and the dataset is CC BY NC 4.0. The methodology can therefore be implemented in RuVector without taking a runtime dependency on Nerfstudio or the dataset.

Why it matters

ruvector-robotics already exposes Pose, Trajectory, SpatialIndex, SceneGraph, perception pipelines, memory, and a WorldModel. RuVector is also becoming the persistence layer for Gaussian scene state and multimodal spatial evidence.

A spatial model should not promote because it performs well only under reconstruction friendly camera paths or optimized pose metadata. The benchmark needs to quantify the deployment gap directly.

Target repository and package

Primary:

crates/ruvector-robotics

Likely new modules:

src/bridge/pose_coverage.rs

src/eval/spatial_robustness.rs

or an equivalent existing benchmark namespace after repository review.

Avoid creating a new crate unless the existing package boundary proves unsuitable.

Potential later integration:

RuView LiDAR and Gaussian world model fixtures

RVF spatial artifacts

MetaHarness spatial evaluator

RuVector WASM for browser side metric computation if the pure Rust metric compiles cleanly without platform dependencies

Current architecture

ruvector-robotics currently contains self contained types for point clouds, robot state, pose, scene graphs, trajectories, a spatial index, a perception stack, world model, memory, and MCP tools.

The current published architecture describes point and vector performance targets but does not expose an evaluation contract that distinguishes spatial interpolation from trajectory level generalization or measured pose from offline optimized pose.

Observed limitation

A benchmark can accidentally reward favorable preprocessing rather than world model robustness when:

  1. test frames come from a path already represented in training
  2. test poses use offline optimized values unavailable online
  3. camera intrinsics are optimized per scene
  4. trajectory coverage is not reported
  5. failures caused by unregistered or poorly calibrated poses are silently dropped

These are exactly the conditions PIVOT was designed to isolate.

Proposed architecture

1. Directed pose coverage metric

Implement a deterministic pose coverage descriptor over Pose and Trajectory.

Conceptually:

D(eval to train) = mean over evaluation poses of nearest pose distance to training set

The pose distance must expose translation and rotation components separately as well as a combined normalized score. Do not claim it is a complete visibility or scene difficulty metric.

Requirements:

  1. directed, not symmetric
  2. deterministic ordering and floating point handling
  3. configurable but explicit translation and rotation normalization
  4. bounded rejection of NaN, infinity, invalid quaternions, or empty sets
  5. no hidden pose optimization
2. Spatial evaluation manifest

Add a serializable manifest that records at minimum:

  1. scene id
  2. trajectory ids and trajectory family
  3. training versus evaluation membership
  4. pose source: measured or optimized
  5. intrinsic source: calibrated or optimized
  6. sensor or capture device id
  7. model or world artifact id
  8. random seed
  9. exact metric configuration
  10. evidence and calibration receipts where available
3. Benchmark families

Family A: represented versus unseen trajectory

Train or construct the world state with one or more named trajectories. Evaluate both held out samples on represented trajectories and complete trajectories absent from training.

Family B: measured versus optimized pose

Keep scene observations fixed and vary translation and rotation source independently when both representations exist.

Family C: calibrated versus optimized intrinsics

Keep pose source fixed and compare reusable physical calibration with scene optimized intrinsics.

Family D: trajectory coverage curve

Report error against directed pose coverage distance. This is descriptive evidence, not proof of causality.

4. RuView and Gaussian extension

The same manifest should support future RuView fixtures where pose source comes from ARKit, LiDAR, radar SLAM, IMU, or RF localization rather than a camera only pipeline.

For Gaussian scene state report at minimum:

  1. geometry error where ground truth exists
  2. localization or pose error
  3. semantic query consistency if applicable
  4. primitive count and resident bytes
  5. update latency
  6. query latency
  7. provenance coverage
  8. uncertainty calibration

Image rendering metrics such as PSNR, SSIM, or LPIPS are optional adapters, not core RuVector dependencies.

External evidence

PIVOT v1 contains five real scenes and reports a consistent degradation on unseen trajectories for both Nerfacto and Splatfacto. Its paper also reports substantial degradation when measured poses replace COLMAP optimized poses and sensitivity to fixed calibrated versus scene optimized intrinsics.

This is valuable as a benchmark design result. It is not evidence that RuVector has the same error magnitude.

Expected measurable improvement

This issue improves evaluation quality, not model accuracy by itself.

Acceptance targets:

  1. the benchmark can detect a deliberately overfit world model that performs well on represented paths and poorly on unseen paths
  2. all benchmark reports state trajectory, pose source, and intrinsic source explicitly
  3. directed pose coverage is deterministic across repeated process runs
  4. benchmark fixtures with invalid pose data fail closed rather than becoming silently optimized or dropped
  5. metric computation on 100,000 poses stays below 100 ms p95 on a reference x86 machine or an alternative target justified by measured scaling
  6. pure metric memory stays O(number of poses), with no dense all pairs matrix required
  7. a WASM build is evaluated if the implementation remains dependency light; no WASM claim if not measured

The first model promotion using this gate must report a before and after represented versus unseen trajectory gap. No improvement is claimed until that experiment exists.

Dependencies

Prefer existing RuVector pose, trajectory, spatial index, graph, RVF, and receipt primitives.

PIVOT source or Nerfstudio must not become a production runtime dependency.

Dataset licensing is noncommercial, so any checked in PIVOT data must be limited to redistribution permitted by its license or replaced with small synthetic fixtures plus user downloaded benchmark data. Legal provenance must be explicit.

Security review

  1. Treat imported trajectory manifests as untrusted data.
  2. Bound frame, trajectory, and pose counts before allocation.
  3. Reject nonfinite coordinates and invalid quaternion norms.
  4. Normalize file paths and reject traversal if filesystem import is added.
  5. Avoid parser recursion or unbounded JSON objects.
  6. Do not let benchmark metadata change authoritative world state.
  7. Do not allow evaluation labels or optimized poses into the online model path being scored.
  8. Preserve train and evaluation split digests so MetaHarness cannot reward a candidate using gold evaluation state.
  9. Hash exact benchmark manifest and output receipts for replay.

Privacy

Real spatial trajectories can reveal precise location and movement. Benchmark receipts should contain hashes and bounded metadata by default, not raw coordinate histories, unless the evaluation artifact is explicitly authorized for storage.

Backward compatibility

Additive benchmark and metric APIs only. Existing world model and robotics interfaces remain unchanged.

If later wired into a promotion gate, start as advisory and move to hard gate only after one stable baseline corpus is frozen.

Testing plan

Unit tests:

  1. identical pose sets produce zero directed distance
  2. directionality differs for asymmetric pose sets
  3. translation only and rotation only controls
  4. invalid quaternions reject
  5. NaN and infinity reject
  6. empty training or evaluation set has explicit error semantics
  7. deterministic serialization and report hash

Property tests:

  1. nonnegative distance
  2. duplicating an evaluation pose leaves the mean within numerical tolerance expected from weighting
  3. adding an identical training pose cannot increase nearest pose distance
  4. translation only controls are invariant to common global translation

Regression tests:

  1. seen path synthetic fixture
  2. unseen path synthetic fixture
  3. measured versus optimized perturbation fixture
  4. calibrated versus optimized intrinsics metadata fixture

Performance:

At least five repeated runs across 1k, 10k, and 100k pose sets with hardware, OS, rustc, configuration, mean, p50, p95, memory, and variance recorded.

MetaHarness plan

Use stable metaharness v0.4.4 with a RuVector specific spatial evaluator. Freeze train and evaluation manifests before candidate mutation. Darwin may optimize a model or compaction policy only when:

  1. unseen trajectory performance is part of fitness
  2. measured pose performance is a hard nonregression gate
  3. provenance and privacy tests pass
  4. benchmark receipts replay successfully
  5. no candidate can read optimized test poses or gold evaluation state during inference

A candidate that improves represented trajectory quality while materially worsening unseen trajectory quality must not promote.

Rollback

The benchmark is additive. Remove the evaluator or disable its promotion gate. No stored vector or RVF migration is required.

Definition of done

Another engineer can take one frozen spatial dataset, run the same model under represented versus unseen trajectories and measured versus optimized pose conditions, reproduce the exact report digest, and see an explicit generalization gap rather than a single aggregate score.

Lenguaje dominante
Rust
Estrellas
4.5k
Forks
603
Merge medio
1 d 11 h
PR fusionados (30 d)
56

Preparar el entorno

Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de ruvnet/RuVector

Todos los issues de ruvnet/RuVector

Issues similares

Más issues de Rust

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.