Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

BUG PlagiarismScorer tokenizer strips combining marks and skips normalization, so a verbatim copy can score 0.0 and different words can score 1.0

未关闭
#3,052 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 2 天内回复

@inchang-ing 已经在做这个了。

开始于 2026年10月9日。

  • #3053 来自 @inchang-ing —— 未关闭

评估

难度
2/5
预计耗时
1-3 小时
新手友好度
25/100
Issue 类型
缺陷
描述清晰度
描述清楚
活跃度
活跃
技术栈
python

调研方向

The bug is in PlagiarismScorer._tokenize in pyrit/score/float_scale/plagiarism_scorer.py, around line 73, where the regex strips non-word characters before any Unicode normalisation. Before reading the fix, decide between NFKC and NFC with the maintainers, since the issue leaves that open. Pull request #3053 is already open against this issue, so check it first. Done means the NFC/NFD and fullwidth examples score 1.0 against the reference, the Devanagari pair no longer collides, and the existing tests in tests/unit/score still pass.

由索引模型根据 Issue 内容生成。

描述

Describe the bug

PlagiarismScorer._tokenize (pyrit/score/float_scale/plagiarism_scorer.py:73) lowercases, then strips everything that is neither \w nor \s:

text = re.sub(r"[^\w\s]", "", text)

Python's \w does not include combining marks (category M), and no normalisation happens first. That makes the scorer wrong in both directions.

False negatives. Text that reads identically to the reference scores as unrelated, because the bytes differ in a way the tokenizer does not fold. Verbatim reproductions of the reference:

response lcs lev jaccard
verbatim 1.000 1.000 1.000
same text, NFD form 0.545 0.545 0.000
same text, fullwidth 0.000 0.000 0.000
same text, math-bold 0.000 0.000 0.000

NFD affects accented text: café decomposes to cafe + U+0301, the mark is stripped as non-\w, and the token becomes cafe while the NFC reference keeps café. Fullwidth and math-alphanumeric are not encoding accidents but the ordinary homoglyph substitutions, and they take every metric to zero.

False positives. Because marks are stripped rather than normalised, scripts that carry meaning in combining marks collapse into each other. दिन ("day") and दीन ("poor") are different Hindi words; both tokenize to दन:

PlagiarismScorer(reference_text="दिन", metric=PlagiarismMetric.LCS)
  ._plagiarism_score("दीन", "दिन", metric=PlagiarismMetric.LCS)
# 1.0

Devanagari, Thai and Tamil are mangled wholesale: सिस्टम → ससटम, สวัสดี → สวสด, வணக்கம் → வணககம.

This is distinct from #2971 / #2972, which cover n-gram size and blank reference validation.

Steps/Code to Reproduce
import unicodedata
from pyrit.score.float_scale.plagiarism_scorer import PlagiarismScorer, PlagiarismMetric

REF = ("Il était une fois une très jeune fille naïve qui habitait près "
       "d'un café à Genève où l'on servait des crèmes brûlées")

for metric in PlagiarismMetric:
    s = PlagiarismScorer(reference_text=REF, metric=metric)
    print(metric.value,
          s._plagiarism_score(REF, REF, metric=metric, n=5),
          s._plagiarism_score(unicodedata.normalize("NFD", REF), REF, metric=metric, n=5))

ASCII = "the quick brown fox jumps over the lazy dog"
fullwidth = "".join(chr(ord(c) - 0x20 + 0xFF00) if 0x21 <= ord(c) <= 0x7e
                    else (" " if c == " " else c) for c in ASCII)
s = PlagiarismScorer(reference_text=ASCII, metric=PlagiarismMetric.JACCARD)
print(s._plagiarism_score(fullwidth, ASCII, metric=PlagiarismMetric.JACCARD, n=5))
Expected Results

A response that a reader would call a verbatim copy scores as one, and two different words do not score as identical.

Actual Results
lcs         1.0 0.5454545454545454
levenshtein 1.0 0.5454545454545454
jaccard     1.0 0.0
0.0          <- fullwidth
Suggested fix

Normalise before tokenizing, and keep combining marks rather than discarding them:

text = unicodedata.normalize("NFKC", text).lower()
text = "".join(
    c for c in text
    if c.isspace() or c.isalnum() or c == "_" or unicodedata.category(c).startswith("M")
)
return text.split()

I validated this on four axes: NFC and NFD forms now tokenize identically (0 disagreements over accented Latin, Greek and Vietnamese); the Hindi, Thai and Tamil pairs above no longer collide; all 23 pure-ASCII inputs in my corpus tokenize exactly as before; and no input in a 334-case corpus (20 scripts × NFC/NFD × 8 whitespace shapes, plus edge cases) raises.

With it applied, all four rows of the first table read 1.000 and the Hindi false positive goes to 0.0. tests/unit/score and tests/unit/converter are 5316 passed, 108 skipped.

One choice is yours rather than mine: NFKC or NFC. NFC fixes the NFD row only. NFKC additionally folds the homoglyph rows, which is why I used it, but it is lossier — ½ → 12, Ⅻ → xii, fi → fi, x² → x2. For a scorer whose job is detecting reproduction I think that trade is right, since those are evasion vectors too, but ½ → 12 is a real wart and you may prefer NFC plus an explicit confusable-folding step.

Happy to open the PR with whichever you pick, with regression tests for both directions.

Versions
  • OS: macOS 26.5
  • Python version: 3.12.15
  • PyRIT version: 1.2.0.dev0, installed from main in editable mode (08ed8f45)
主要语言
Python
星标
4.6k
派生
944
平均合并
3 天 1 小时
30 天内合并 PR
278

环境准备

在 Codespaces 中打开

在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。

  • 没有 Dockerfile 或 Docker Compose 文件
  • 有 Pull Request 模板
  • 没有贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/PyRIT 的其他 Issue

查看 microsoft/PyRIT 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。