BUG: PlagiarismScorer accepts invalid n-gram size and blank reference text
维护者通常 2 天内回复
评估
调研方向
从 pyrit/score/float_scale/plagiarism_scorer.py 中的 PlagiarismScorer.init 开始,然后将其参数处理与 issue 中描述的 ApproximateTextMatching 验证进行比较。运行现有的 PlagiarismScorer 测试,并为无效的 n、空白参考文本和无效指标添加覆盖;完成标准是构造过程拒绝所列出的每种情况,同时不改变有效评分。
由索引模型根据 Issue 内容生成。
描述
Describe the bug
PlagiarismScorer does not validate its n (n-gram size) or reference_text parameters during initialization.
-
n <= 0creates false-positive 100% plagiarism matches:
Whenn = 0,_ngram_setreturns{()}(an empty tuple). InPlagiarismMetric.JACCARD,len(ref_ngrams & res_ngrams) / len(ref_ngrams)evaluates to1 / 1 = 1.0for any response. A completely unrelated prompt response is therefore reported as 100% plagiarized.
Whenn < 0(e.g.n = -1), negative slicing produces spurious n-grams and invalid partial overlap scores (e.g.0.5).
Whennis a non-integer float (e.g.2.5), construction succeeds, but scoring later crashes withTypeError: 'float' object cannot be interpreted as an integer.
Whennis a boolean (e.g.False), it evaluates as0and yields the same false-positive1.0. -
Empty or whitespace-only
reference_textsilently reports0.0for all responses:
Whenreference_textis""or whitespace-only" ",_tokenize(reference)produces[](reference_len = 0). The conditionresponse_len == 0 or reference_len == 0causes_plagiarism_scoreto silently return0.0for every response, even when the response is identical to the reference text. A static misconfiguration is thus silently swallowed and reported as non-plagiarism.
The sibling class in analytics, ApproximateTextMatching, already validates that its n-gram parameter n must be an integer >= 1.
Steps/Code to Reproduce
from pyrit.score import PlagiarismScorer, PlagiarismMetric
# 1. n=0 scores 1.0 (100% match) for completely unrelated text
scorer = PlagiarismScorer(
reference_text="The quick brown fox jumps over the lazy dog",
metric=PlagiarismMetric.JACCARD,
n=0,
)
score = scorer._plagiarism_score(
response="Completely unrelated content about astrophysics and quantum mechanics",
reference="The quick brown fox jumps over the lazy dog",
metric=PlagiarismMetric.JACCARD,
n=0,
)
print("Score with n=0:", score)
# 2. n=-1 yields spurious partial match
print(
"Score with n=-1:",
scorer._plagiarism_score(
response="Completely unrelated content",
reference="The quick brown fox jumps over the lazy dog",
metric=PlagiarismMetric.JACCARD,
n=-1,
),
)
# 3. Empty reference_text silently evaluates every response as 0.0
empty_scorer = PlagiarismScorer(reference_text="", metric=PlagiarismMetric.LCS)
print("Score with empty reference:", empty_scorer._plagiarism_score(response="", reference=""))
Expected Results
PlagiarismScorer.__init__ should validate parameters at construction time and raise ValueError:
reference_textmust be a non-empty string containing at least one word token.nmust be an integer>= 1(rejecting0, negative values, booleans, and non-integers).metricmust be an instance ofPlagiarismMetric.
Actual Results
Score with n=0: 1.0
Score with n=-1: 0.5
Score with empty reference: 0.0
PlagiarismScorer.__init__ assigns self.reference_text = reference_text and self.n = n directly at pyrit/score/float_scale/plagiarism_scorer.py:54-56 without validation.
Versions
- OS: Windows 11
- Python: 3.14
- PyRIT:
mainatc2ae5eeb
- 主要语言
- Python
- 星标
- 4.6k
- 派生
- 924
- 平均合并
- 2 天 14 小时
- 30 天内合并 PR
- 227
环境准备
在浏览器里用你自己的 GitHub 账号启动这个项目的开发容器。
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 没有贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
microsoft/PyRIT 的其他 Issue
-
PackageHallucinationScorer (Python) misses `from pkg.sub import x` and indented imports可能已有人在做 @barry166 于 4 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
microsoft/PyRIT#2948 · 1 条评论 ·
维护者通常 2 天内回复
-
LiteLLMChatTarget does not flag or survive output-token truncation可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 2 天内回复
-
BUG Configuration keeps runtime-status errors after polling recovers可能已有人在做 @rupayon123 于 10 天前认领。 未关闭Bug: triage GUI help wanted
难度 2/5 1-3 小时 新手友好度 86/100
microsoft/PyRIT#2868 · 1 条评论 ·
维护者通常 2 天内回复
-
ObjectiveScorerEvaluator scores every conversation message as an assistant response可能已有人在做 @feiiiiii5 于 11 天前认领。 未关闭
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 2 天内回复
-
难度 4/5 3-5 天 新手友好度 45/100
维护者通常 2 天内回复
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 85/100
维护者通常 3 天内回复
-
Negation with "not" and "no" is ignored during sentiment analysis可能已有人在做 @vivek-3728 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 68/100
techcsispit/mess-mood#11 · 1 条评论 ·
-
changelog investigate
难度 2/5 1-3 小时 新手友好度 68/100
ramnes/notion-sdk-py#408 ·
-
good first issue
难度 2/5 1-3 小时 新手友好度 83/100
btclib-org/btclib-wallet#267 ·
维护者通常 1 天内回复
-
good first issue tech-debt
难度 2/5 1-3 小时 新手友好度 85/100
knnmelprop/YAADO#111 ·