Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Add entity resolution to deduplicate and merge cross-source observations

Abierto
#250 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
go, kafka

Línea de trabajo

Read docs/_drafts/06-roadmap.md and inspect the existing Upsert behavior first, focusing on the Tier 1 exact-URN scope. Done means observations merge idempotently into an existing entity, provenance is tracked, and the merge behavior is defined; Tier 2 and Tier 3 are follow-ups.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Context

Compass stores entities from multiple sources as separate records. A Kafka topic ingested from two different systems creates two unrelated entities with different URNs. The graph is fragmented — context assembly, impact analysis, and search all operate on disconnected duplicates.

Entity resolution is the mechanism that matches incoming observations against existing entities, merges properties, and maintains unified identity. This is the prerequisite for a coherent knowledge graph.

Scope

Tier 1: Exact URN Match
  • When an observation arrives with a URN that already exists, merge properties into the existing entity
  • Track provenance: which source contributed which properties
  • Idempotent — re-sending the same observation must not create duplicates or mutate state unexpectedly
  • This is what Upsert partially does today, but without provenance tracking or merge strategy
Tier 2: Heuristic Matching
  • Match observations where URN differs but type + name + source pattern suggests the same logical entity
  • Configurable matching rules (e.g., "bigquery table names map to dbt model names via this pattern")
  • Candidate scoring with a confidence threshold
Tier 3: Semantic Similarity (follow-up)
  • Use embedding similarity to catch non-obvious matches
  • Only viable after the embedding pipeline has indexed sufficient entities
  • Should be a signal fed into Tier 2 scoring, not a standalone matcher
Merge Strategy
  • When a match is found, merge properties from the new observation into the existing entity
  • Default: last-write-wins per field
  • Track which source contributed which properties (provenance)
  • Resolution audit log: record what was matched, merged, and why

Design Considerations

  • Resolution must be idempotent
  • Meteor sends raw observations, Compass resolves — keep the interface simple
  • Start with Tier 1 (exact URN match with provenance). Ship it. Tier 2 and 3 are follow-ups.
  • Graph-aware ranking (#237) depends on a coherent, deduplicated graph — this should ship first

References

Lenguaje dominante
Go
Estrellas
72
Forks
9
Métricas de merge de PR
Sin PR fusionados en 30 d

Guía de contribución

No hay ninguna guía de contribución indexada para este repositorio

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de raystack/compass

Todos los issues de raystack/compass

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.