Python: relevance-score gates (min_relevance_score / DEFAULT_RELEVANCE) have no polarity cases in tests: a memory that says the opposite of the query is recalled as highly relevant.
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 55/100
Direzione di ricerca
Esamina python/tests/integration/embeddings/test_embedding_service_with_memory.py e i percorsi di recall in text_memory_plugin.py e semantic_text_memory.py. Aggiungi una copertura comportamentale per i casi di polarità e di bassa sovrapposizione a soglie diverse da zero, utilizzando i requisiti di provenienza dei vettori proposti nell’issue; il lavoro è completato quando i gate distribuiti vengono esercitati oltre le attuali chiamate di integrazione con min_relevance_score=0.0.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Language: Python. The .NET defaults differ and are recorded at the end.
Verified against main at b39d95a on 2026-08-25.
Hi SK team. I just finished a construct-validity audit of embedding-cosine gates across agent
frameworks, and SK's memory-recall path carries the same test blind spot I found in four other
suites. Preprint: https://arxiv.org/abs/2608.10216.
The pattern. TextMemoryPlugin.recall gates on DEFAULT_RELEVANCE = 0.75
(text_memory_plugin.py#L22),
and SemanticTextMemory.search gates on min_relevance_score, defaulting to 0.0
(semantic_text_memory.py#L129),
which it passes through to the store at L153. Both implement one decision, "is this stored text
relevant to, meaning does it answer or agree with, this query?", and both answer it by thresholding
embedding similarity. That score measures shared wording rather than shared meaning, and in the pairs
a recall gate most needs to rank correctly the two run opposite ways: reversing an instruction is a
one-token edit that keeps similarity near maximal, while a faithful restatement in fresh words drops
it substantially. To be precise about framing, nothing here is a bug in SK's code, and the gate does
exactly what it says. The blind spot sits in what the test suite asserts the gate means.
Specimen from my audit corpus (cosine under nomic-embed-text-v1.5, MRL-256, frozen in the paper's
artifact; the mutation class scored 0.83 to 0.9997 across all nine encoders tested, so this is not
one model's quirk):
- Stored memory: "Withhold the study drug from any participant who reports chest tightness."
- Query: "Administer the study drug to any participant who reports chest tightness."
- Cosine 0.9608. It clears 0.75 comfortably, and clears any plausible relevance floor. At the 0.0
default it is not even a speed bump. - Mirror case: faithful low-overlap restatements in the same corpus average about 0.76 under the same
encoder, for example "Retry the request at most three times." against "Give the call up to three
attempts, then stop.". That average lands right at the plugin cut, so faithful restatements sit on
the knife edge while reversals clear it comfortably.
What the suite covers today. recall has no behavioral unit test. Across python/tests,
TextMemoryPlugin appears only in test_serialization.py, and the one test that exercises the gate
is
test_embedding_service_with_memory.py#L155-L169,
an integration test needing live credentials, which calls memory.search(...) four times and passes
min_relevance_score=0.0 at every call site. The threshold is never exercised at a non-zero value
anywhere in the suite, and the fixtures are taxonomic facts, so no pair shares wording while
disagreeing about what it decides.
Proposal, sized to be PR-able, and I am happy to write it:
- Add polarity cases: pairs sharing wording with opposite decisions (negation,
musttomay,
quantity changes, scope inversion) asserted at the shipped defaults, and pairs sharing the
decision with no shared wording. Section 11.3 of the paper gives the two families. The corpus
pairs are released and were authored blind to any encoder. - Record the vectors from a named encoder and check them in with provenance (model id, dimension,
date, generating script) rather than hand-authoring them. A hand-authored vector asserts only the
fixture author's arithmetic, and passes whatever the sentences happen to say. - Optionally, one line in the
relevanceandmin_relevance_scoredocstrings noting that the score
is surface-form similarity, so "relevant" should not be read as "in agreement".
.NET, and a correction to my original filing. @journaltraces checked the C# side and found two things
I had wrong.
TextMemoryPlugin.cs#L47
has private const double DefaultRelevance = 0.0;, so the .NET plugin ships an open gate where the
Python one ships 0.75. And VectorStoreTextSearch carries no minRelevanceScore at all:
ExecuteVectorSearchCoreAsync builds a VectorSearchOptions<TRecord> with Filter only and passes
top. I originally wrote that the deprecation migration ports the blind spot forward. It does not
port it, it drops the gate: a caller who migrates moves from a threshold that can at least be raised
to a top-k that always returns its k, with no floor left to argue about. Thanks to them for both.
The audit found the identical gap in LlamaIndex, LangChain, and GPTCache-style caches. SK is in good
company, which is rather the point: the failure mode ships because no framework's tests can see it.
- Lingua principale
- C#
- Stelle
- 28.6k
- Fork
- 4.8k
- Merge medio
- 13h 24m
- PR unite (30g)
- 11
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/semantic-kernel
-
python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
microsoft/semantic-kernel#14491 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
microsoft/semantic-kernel#14490 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Python: [Python] structured_outputs_transform reuses ChatHistory across calls (prompt pollution)Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
microsoft/semantic-kernel#14483 · 2 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
.NET python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
microsoft/semantic-kernel#14482 · 3 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
Python: [Python] as_agent_framework_tool drops parameter defaults (optionals become required)Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
microsoft/semantic-kernel#14481 ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di microsoft/semantic-kernel
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
NethermindEth/nethermind#14012 ·
I maintainer di solito rispondono entro 1 giorno
-
dependencies Status: Triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
json-schema-org/website#2518 ·
I maintainer di solito rispondono entro 1 giorno
-
agentic-workflows
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
builtbybel/Flyoobe#498 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
builtbybel/CrapFixer#112 ·