Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

krippendorff_alpha reports 1.0 instead of NaN for single-rater identical ratings

Closed Beginner friendly
#2,843 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
86/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
analytics

Research direction

Start at pyrit/score/scorer_evaluation/krippendorff.py and inspect krippendorff_alpha, especially the category-count early return and the existing no-pairable-ratings handling. Run the two NumPy examples from the issue first. Done means every dataset with no pairable ratings returns NaN, including identical single-rater ratings, while existing differing single-rater behavior remains correct.

Written by the indexing model from the issue text.

Description

what happens

krippendorff_alpha returns 1.0 (perfect agreement) when every item is rated by a single rater and all those ratings happen to be identical. with no pairable ratings the statistic is undefined, so this should be NaN like it already is for single-rater data whose ratings differ.

the cause is ordering: the num_categories == 1 early return fires before the check that no item has two usable ratings, so identical-but-unpairable data takes the perfect-agreement exit.

repro

import numpy as np
from pyrit.score.scorer_evaluation.krippendorff import krippendorff_alpha

krippendorff_alpha(np.array([[1.0, 1.0]]))  # 1.0 — undefined statistic reported as perfect
krippendorff_alpha(np.array([[1.0, 2.0]]))  # nan — correct

same unpairable structure, opposite answers depending on whether the values coincide.

why it matters

this feeds the reliability metrics in scorer evaluation, and single-vote gold labels are a common input (cf #2628). reporting perfect agreement from one rater's identical votes overstates reliability.

expected

NaN for any dataset with no pairable ratings, per the docstring's own undefined-statistic rule.

env: main 5503ecb

Dominant language
Python
Stars
4.5k
Forks
896
Avg merge
2d 19h
Merged PRs (30d)
206

Getting set up

We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/PyRIT

All issues in microsoft/PyRIT

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.