Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

ExactTextMatching reports a match for an empty or whitespace-only target

Aperta Adatta ai principianti
#2,881 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
91/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
security, testing

Direzione di ricerca

Inizia in pyrit/analytics/text_matching.py, in ExactTextMatching.is_match(), quindi confronta la copertura esistente di test_target_too_short in tests/unit/analytics/test_text_matching.py. Aggiungi la copertura per target vuoti e composti solo da spazi bianchi, preservando il matching per i target non vuoti, ed esegui i test mirati di text-matching. Il lavoro è completo quando i target vuoti non producono più una corrispondenza e il comportamento esistente rimane invariato.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Describe the bug

ExactTextMatching.is_match() guards an empty text (if not text: return False) but not an empty target, and "" in anything is True in Python. With the default ignore_whitespace=True the check is reachable with a whitespace-only target too, because target.strip() turns it into "".

Its sibling in the same module, ApproximateTextMatching, has the opposite behaviour and says why:

if len(target) < self._n:
    return 0.0  # Confidence is too low for short targets

That rule is pinned by tests/unit/analytics/test_text_matching.py::test_target_too_short. So the two implementations of the same TextMatching interface disagree on a degenerate target.

This matters beyond the helper: DecodingScorer defaults to ExactTextMatching(case_sensitive=False) and calls it with no guard on user_piece.original_value and user_piece.converted_value — while guarding the adjacent decoded_text on the very next lines with if decoded_text and .... MessagePiece.converted_value defaults to "", so when no converter ran, the converted_value check matches every response and the scorer reports a successful decoding.

Steps/Code to Reproduce
from pyrit.analytics import ApproximateTextMatching, ExactTextMatching

print(ExactTextMatching().is_match(target="", text="I refuse to help with that."))
print(ExactTextMatching().is_match(target="   \n ", text="I refuse to help with that."))
print(ExactTextMatching(case_sensitive=True).is_match(target="", text="anything"))
print(ExactTextMatching(ignore_whitespace=False).is_match(target="", text="hello"))
print(ApproximateTextMatching().is_match(target="", text="hello world"))  # sibling
Expected Results

The first four should be False: a target that carries no content cannot be found in the text. The sibling already returns False. is_match(target="refuse", text="I refuse to help") should stay True.

Actual Results
True
True
True
True
False
Screenshots

Not applicable.

Versions
  • OS: macOS
  • Python version: 3.11
  • PyRIT version: installed from main in editable mode (pyrit/analytics/text_matching.py)
  • Existing coverage: tests/unit/analytics/test_text_matching.py::test_empty_text covers an empty text only; no test covers an empty target.
Proposed direction

I would add the symmetric guard to ExactTextMatching.is_match — return False when the (whitespace-normalised) target is empty — and a test mirroring test_target_too_short. That keeps the fix in the shared abstraction, so DecodingScorer and any other caller are covered without touching them. Happy to send a PR if that matches your intent; I am also happy to guard the two call sites in DecodingScorer instead if you prefer the narrower change.

Lingua principale
Python
Stelle
4.5k
Fork
896
Merge medio
2g 19h
PR unite (30g)
206

Preparare l'ambiente

Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di microsoft/PyRIT

Tutte le issue di microsoft/PyRIT

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.