BUG: ScorerMetrics.to_json() raises TypeError on the trial_scores array ScorerEvaluator attaches
Maintainers usually reply within 2 days
@feiiiiii5 is already working on this.
Since Sep 21, 2026.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 78/100
Research direction
Start in pyrit/score/scorer_evaluation/scorer_metrics.py at to_json() and from_json_file(), then review how scorer_evaluator.py attaches trial_scores. Use the reproduction and add regression coverage alongside tests/unit/score/test_scorer_metrics.py for both metrics subclasses. Done means evaluator-produced metrics serialize to JSON and restore trial_scores with its expected array shape, while unsupported values still fail.
Written by the indexing model from the issue text.
Description
Describe the bug
ScorerMetrics.to_json() raises TypeError on exactly the metrics objects that
ScorerEvaluator returns, because trial_scores is a numpy array and json.dumps cannot
encode one.
to_json() documents itself as "the canonical serialization entry point for ScorerMetrics
and its subclasses", paired with from_json_file() "for round-trip (de)serialization"
(pyrit/score/scorer_evaluation/scorer_metrics.py:51-63). The evaluator deliberately attaches
that array to the object it hands back
(pyrit/score/scorer_evaluation/scorer_evaluator.py:448-450):
# Include trial scores for debugging and future mismatch analysis
# (not persisted to registry - use returned metrics object for detailed analysis)
metrics.trial_scores = all_model_scores
so the object the in-source comment points callers at for "detailed analysis" is the one object
the documented serializer refuses to serialize. from_json_file() filters out only
underscore-prefixed keys, i.e. it is written as though trial_scores were part of what
to_json() emits.
Two things this is not, so the scope is clear:
- The JSONL registry path is unaffected.
_metrics_to_registry_dict()
(pyrit/score/scorer_evaluation/scorer_metrics_io.py:42-58) excludestrial_scores, and
tests/unit/score/test_scorer_metrics_io.py:136pins that exclusion.evaluate_async()and
registry reads work today. - No in-tree code and no documentation example hits it. The only callers of
metrics.to_json()
in the repo are the two round-trip tests intests/unit/score/test_scorer_metrics.py
(:36and:57), and both construct metrics withouttrial_scores— which is why the suite
is green while the pair's documented contract is broken for real evaluator output.
Steps/Code to Reproduce
import numpy as np
from pyrit.score import ObjectiveScorerMetrics
metrics = ObjectiveScorerMetrics(
num_responses=2,
num_human_raters=1,
accuracy=0.9,
accuracy_standard_error=0.05,
f1_score=0.8,
precision=0.85,
recall=0.75,
trial_scores=np.array([[0.2, 0.4], [0.2, 0.4]]),
)
print(metrics.to_json())
Expected Results
A JSON string, "trial_scores": [[0.2, 0.4], [0.2, 0.4]], that from_json_file() reads back
into the same shape — as it already does for every other field.
Actual Results
File ".../python3.14/json/encoder.py", line 182, in default
raise TypeError(f'Object of type {o.__class__.__name__} '
f'is not JSON serializable')
TypeError: Object of type ndarray is not JSON serializable
when serializing dict item 'trial_scores'
Same for HarmScorerMetrics, and reached without hand-building anything: running the real
HarmScorerEvaluator.evaluate_dataset_async(...) with two trials (the mocked-scorer recipe from
tests/unit/score/test_scorer_evaluator.py:75) returns metrics whose fields are
trial_scores: type=ndarray dtype=float64 value=array([[0.2, 0.4],
[0.2, 0.4]])
mean_absolute_error: type=float64 dtype=float64 value=np.float64(0.0)
and then metrics.to_json() raises the identical TypeError. Note that the np.float64 fields
are not the problem: np.float64 subclasses float, so json.dumps handles them; ndarray,
np.int64 and np.bool_ are the numpy types it rejects (measured on numpy 2.4.4).
Versions
- OS: macOS 27.0.0 (arm64)
- Python version: 3.14.5
- PyRIT version: 1.2.0.dev0, from
mainat543c20c - numpy 2.4.4
Proposed fix
If maintainers agree this is worth fixing, I'd send a small PR: give to_json() a default=
hook that encodes numpy arrays/scalars with .tolist(), and have from_json_file() restore
trial_scores as an np.ndarray so the round trip returns the declared type instead of a nested
list. Tests: one regression test per metrics subclass with trial_scores populated, plus one
asserting the hook still raises for values that are neither JSON nor numpy (so the fix does not
turn the crash into silent stringification).
One adjacent thing I am deliberately not folding in, but flagging: == between two
array-carrying metrics raises
ValueError: The truth value of an array with more than one element is ambiguous, because the
dataclass compares trial_scores elementwise. That means assert loaded == metrics — the oracle
the existing two round-trip tests use — cannot be applied to an object with trial scores at all,
and a fix would be a semantics decision (field(compare=False) vs. something else) rather than a
bug fix. Happy to open a separate issue for it if useful.
- Dominant language
- Python
- Stars
- 4.5k
- Forks
- 896
- Avg merge
- 2d 22h
- Merged PRs (30d)
- 230
Getting set up
We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/PyRIT
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 2 days
-
Bug: triage GUI help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
microsoft/PyRIT#2868 · 1 comment ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 2 days
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
mozilla/bedrock#17413 · 1 reaction ·
Maintainers usually reply within 2 days
-
instance instance add
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
searxng/searx-instances#943 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Maintainers usually reply within 1 day
-
bug tools
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
lance-format/lance#9655 ·
Maintainers usually reply within 2 days