Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Python: relevance-score gates (min_relevance_score / DEFAULT_RELEVANCE) have no polarity cases in tests: a memory that says the opposite of the query is recalled as highly relevant.

Abierto
#14,295 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 2 días

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
55/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
python
Área
ai, testing

Línea de trabajo

Revisa python/tests/integration/embeddings/test_embedding_service_with_memory.py y las rutas de recall en text_memory_plugin.py y semantic_text_memory.py. Añade cobertura de comportamiento para los casos de polaridad y de baja superposición en umbrales distintos de cero, utilizando los requisitos de procedencia de vectores propuestos en el issue; se considera terminado cuando los gates incluidos en la versión se ejercitan más allá de las llamadas de integración actuales con min_relevance_score=0.0.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Language: Python. The .NET defaults differ and are recorded at the end.
Verified against main at b39d95a on 2026-08-25.

Hi SK team. I just finished a construct-validity audit of embedding-cosine gates across agent
frameworks, and SK's memory-recall path carries the same test blind spot I found in four other
suites. Preprint: https://arxiv.org/abs/2608.10216.

The pattern. TextMemoryPlugin.recall gates on DEFAULT_RELEVANCE = 0.75
(text_memory_plugin.py#L22),
and SemanticTextMemory.search gates on min_relevance_score, defaulting to 0.0
(semantic_text_memory.py#L129),
which it passes through to the store at L153. Both implement one decision, "is this stored text
relevant to, meaning does it answer or agree with, this query?", and both answer it by thresholding
embedding similarity. That score measures shared wording rather than shared meaning, and in the pairs
a recall gate most needs to rank correctly the two run opposite ways: reversing an instruction is a
one-token edit that keeps similarity near maximal, while a faithful restatement in fresh words drops
it substantially. To be precise about framing, nothing here is a bug in SK's code, and the gate does
exactly what it says. The blind spot sits in what the test suite asserts the gate means.

Specimen from my audit corpus (cosine under nomic-embed-text-v1.5, MRL-256, frozen in the paper's
artifact; the mutation class scored 0.83 to 0.9997 across all nine encoders tested, so this is not
one model's quirk):

  • Stored memory: "Withhold the study drug from any participant who reports chest tightness."
  • Query: "Administer the study drug to any participant who reports chest tightness."
  • Cosine 0.9608. It clears 0.75 comfortably, and clears any plausible relevance floor. At the 0.0
    default it is not even a speed bump.
  • Mirror case: faithful low-overlap restatements in the same corpus average about 0.76 under the same
    encoder, for example "Retry the request at most three times." against "Give the call up to three
    attempts, then stop.". That average lands right at the plugin cut, so faithful restatements sit on
    the knife edge while reversals clear it comfortably.

What the suite covers today. recall has no behavioral unit test. Across python/tests,
TextMemoryPlugin appears only in test_serialization.py, and the one test that exercises the gate
is
test_embedding_service_with_memory.py#L155-L169,
an integration test needing live credentials, which calls memory.search(...) four times and passes
min_relevance_score=0.0 at every call site. The threshold is never exercised at a non-zero value
anywhere in the suite, and the fixtures are taxonomic facts, so no pair shares wording while
disagreeing about what it decides.

Proposal, sized to be PR-able, and I am happy to write it:

  1. Add polarity cases: pairs sharing wording with opposite decisions (negation, must to may,
    quantity changes, scope inversion) asserted at the shipped defaults, and pairs sharing the
    decision with no shared wording. Section 11.3 of the paper gives the two families. The corpus
    pairs are released and were authored blind to any encoder.
  2. Record the vectors from a named encoder and check them in with provenance (model id, dimension,
    date, generating script) rather than hand-authoring them. A hand-authored vector asserts only the
    fixture author's arithmetic, and passes whatever the sentences happen to say.
  3. Optionally, one line in the relevance and min_relevance_score docstrings noting that the score
    is surface-form similarity, so "relevant" should not be read as "in agreement".

.NET, and a correction to my original filing. @journaltraces checked the C# side and found two things
I had wrong.
TextMemoryPlugin.cs#L47
has private const double DefaultRelevance = 0.0;, so the .NET plugin ships an open gate where the
Python one ships 0.75. And VectorStoreTextSearch carries no minRelevanceScore at all:
ExecuteVectorSearchCoreAsync builds a VectorSearchOptions<TRecord> with Filter only and passes
top. I originally wrote that the deprecation migration ports the blind spot forward. It does not
port it, it drops the gate: a caller who migrates moves from a threshold that can at least be raised
to a top-k that always returns its k, with no floor left to argue about. Thanks to them for both.

The audit found the identical gap in LlamaIndex, LangChain, and GPTCache-style caches. SK is in good
company, which is rather the point: the failure mode ships because no framework's tests can see it.

Lenguaje dominante
C#
Estrellas
28.6k
Forks
4.8k
Merge medio
13 h 24 min
PR fusionados (30 d)
11

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de microsoft/semantic-kernel

Todos los issues de microsoft/semantic-kernel

Issues similares

Más issues de C#

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.