BUG PlagiarismScorer tokenizer strips combining marks and skips normalization, so a verbatim copy can score 0.0 and different words can score 1.0
I maintainer di solito rispondono entro 2 giorni
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 25/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- ai, internationalization
Direzione di ricerca
The bug is in PlagiarismScorer._tokenize in pyrit/score/float_scale/plagiarism_scorer.py, around line 73, where the regex strips non-word characters before any Unicode normalisation. Before reading the fix, decide between NFKC and NFC with the maintainers, since the issue leaves that open. Pull request #3053 is already open against this issue, so check it first. Done means the NFC/NFD and fullwidth examples score 1.0 against the reference, the Devanagari pair no longer collides, and the existing tests in tests/unit/score still pass.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Describe the bug
PlagiarismScorer._tokenize (pyrit/score/float_scale/plagiarism_scorer.py:73) lowercases, then strips everything that is neither \w nor \s:
text = re.sub(r"[^\w\s]", "", text)
Python's \w does not include combining marks (category M), and no normalisation happens first. That makes the scorer wrong in both directions.
False negatives. Text that reads identically to the reference scores as unrelated, because the bytes differ in a way the tokenizer does not fold. Verbatim reproductions of the reference:
| response | lcs | lev | jaccard |
|---|---|---|---|
| verbatim | 1.000 | 1.000 | 1.000 |
| same text, NFD form | 0.545 | 0.545 | 0.000 |
| same text, fullwidth | 0.000 | 0.000 | 0.000 |
| same text, math-bold | 0.000 | 0.000 | 0.000 |
NFD affects accented text: café decomposes to cafe + U+0301, the mark is stripped as non-\w, and the token becomes cafe while the NFC reference keeps café. Fullwidth and math-alphanumeric are not encoding accidents but the ordinary homoglyph substitutions, and they take every metric to zero.
False positives. Because marks are stripped rather than normalised, scripts that carry meaning in combining marks collapse into each other. दिन ("day") and दीन ("poor") are different Hindi words; both tokenize to दन:
PlagiarismScorer(reference_text="दिन", metric=PlagiarismMetric.LCS)
._plagiarism_score("दीन", "दिन", metric=PlagiarismMetric.LCS)
# 1.0
Devanagari, Thai and Tamil are mangled wholesale: सिस्टम → ससटम, สวัสดี → สวสด, வணக்கம் → வணககம.
This is distinct from #2971 / #2972, which cover n-gram size and blank reference validation.
Steps/Code to Reproduce
import unicodedata
from pyrit.score.float_scale.plagiarism_scorer import PlagiarismScorer, PlagiarismMetric
REF = ("Il était une fois une très jeune fille naïve qui habitait près "
"d'un café à Genève où l'on servait des crèmes brûlées")
for metric in PlagiarismMetric:
s = PlagiarismScorer(reference_text=REF, metric=metric)
print(metric.value,
s._plagiarism_score(REF, REF, metric=metric, n=5),
s._plagiarism_score(unicodedata.normalize("NFD", REF), REF, metric=metric, n=5))
ASCII = "the quick brown fox jumps over the lazy dog"
fullwidth = "".join(chr(ord(c) - 0x20 + 0xFF00) if 0x21 <= ord(c) <= 0x7e
else (" " if c == " " else c) for c in ASCII)
s = PlagiarismScorer(reference_text=ASCII, metric=PlagiarismMetric.JACCARD)
print(s._plagiarism_score(fullwidth, ASCII, metric=PlagiarismMetric.JACCARD, n=5))
Expected Results
A response that a reader would call a verbatim copy scores as one, and two different words do not score as identical.
Actual Results
lcs 1.0 0.5454545454545454
levenshtein 1.0 0.5454545454545454
jaccard 1.0 0.0
0.0 <- fullwidth
Suggested fix
Normalise before tokenizing, and keep combining marks rather than discarding them:
text = unicodedata.normalize("NFKC", text).lower()
text = "".join(
c for c in text
if c.isspace() or c.isalnum() or c == "_" or unicodedata.category(c).startswith("M")
)
return text.split()
I validated this on four axes: NFC and NFD forms now tokenize identically (0 disagreements over accented Latin, Greek and Vietnamese); the Hindi, Thai and Tamil pairs above no longer collide; all 23 pure-ASCII inputs in my corpus tokenize exactly as before; and no input in a 334-case corpus (20 scripts × NFC/NFD × 8 whitespace shapes, plus edge cases) raises.
With it applied, all four rows of the first table read 1.000 and the Hindi false positive goes to 0.0. tests/unit/score and tests/unit/converter are 5316 passed, 108 skipped.
One choice is yours rather than mine: NFKC or NFC. NFC fixes the NFD row only. NFKC additionally folds the homoglyph rows, which is why I used it, but it is lossier — ½ → 12, Ⅻ → xii, fi → fi, x² → x2. For a scorer whose job is detecting reproduction I think that trade is right, since those are evasion vectors too, but ½ → 12 is a real wart and you may prefer NFC plus an explicit confusable-folding step.
Happy to open the PR with whichever you pick, with regression tests for both directions.
Versions
- OS: macOS 26.5
- Python version: 3.12.15
- PyRIT version: 1.2.0.dev0, installed from
mainin editable mode (08ed8f45)
- Lingua principale
- Python
- Stelle
- 4.6k
- Fork
- 944
- Merge medio
- 3g 1h
- PR unite (30g)
- 278
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Nessuna guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/PyRIT
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 2 giorni
-
BUG PuzzledConverter cannot select words carrying non-ASCII letters, so the mask falls on articles insteadForse già presa @adimalkar l’ha presa 2 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
microsoft/PyRIT#3022 · 2 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
BUG Configuration keeps runtime-status errors after polling recoversForse già presa @rupayon123 l’ha presa 15 giorni fa. ApertaBug: triage GUI help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
microsoft/PyRIT#2868 · 3 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 45/100
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di microsoft/PyRIT
Issue simili
-
Claiming namespace `jft63`Apertanamespace operations
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
EclipseFdn/open-vsx.org#14043 ·
I maintainer di solito rispondono entro 1 giorno
-
netbox status: needs triage type: bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
netbox-community/netbox#23376 ·
I maintainer di solito rispondono entro 1 giorno
-
feedback simulation workshop
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
githubnext/gh-aw-workshop#4455 ·
I maintainer di solito rispondono entro 1 giorno
-
Triage 🩺
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitApertaneeds-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 77/100
krkn-chaos/krkn#1627 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno