Bug: Precision and Recall metrics return the exact same value
Assessment
- Difficulty
- 1/5
- Estimated time
- Under an hour
- Newbie friendliness
- 88/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- machine-learning
Research direction
Read mess_mood/metrics.py and the test_metrics function in tests/test_mess_mood.py; the issue identifies the incorrect precision formula and gives a failing example. Run the metrics test first, then verify precision uses true positives and false positives while recall still passes. Done when the test expects precision 1.0 and the metrics tests pass.
Written by the indexing model from the issue text.
Description
What I did:
I checked the implementation of the precision function inside mess_mood/metrics.py to see how model evaluation metrics were being calculated. To prove my findings, I updated the test_metrics function in tests/test_mess_mood.py to actually test the precision mathematically:
Click to view the test code
def test_metrics():
y_true = ["pos", "pos", "neg", "neg"]
y_pred = ["pos", "neg", "neg", "neg"]
assert accuracy(y_true, y_pred) == 0.75
assert recall(y_true, y_pred, "pos") == 0.5
# I added this line. Mathematically, True Positives = 1, False Positives = 0.
# Therefore Precision = 1 / (1 + 0) = 1.0
assert precision(y_true, y_pred, "pos") == 1.0
assert f1(1.0, 0.5) == 2 / 3
assert confusion_matrix(y_true, y_pred, ["neg", "pos"]) == {"neg": {"neg": 2, "pos": 0}, "pos": {"neg": 1, "pos": 1}}
What I expected:
The README states that precision is "of the reviews the model gave that label, the share that really had it." Mathematically, this means the formula should be True Positives divided by True Positives plus False Positives (tp / (tp + fp)). The new test case should pass because the precision is 1.0.
What happened instead:
The precision function returns tp / (tp + fn). This is the exact formula for recall (True Positives plus False Negatives), meaning the precision metric currently duplicates the recall metric entirely.
Running the test confirms this; it fails because the precision function returns 0.5 instead of 1.0:
Proof (Terminal Output / Screenshot):
- Dominant language
- Python
- Stars
- 0
- Forks
- 3
- Avg merge
- 10h 15m
- Merged PRs (30d)
- 7
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from techcsispit/mess-mood
-
Add more labelled reviewsPossibly taken @p1xl07 claimed this 3 days ago. Opengood first issue
techcsispit/mess-mood#1 · 1 comment · 1 assignee ·
All issues in techcsispit/mess-mood
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NVIDIA/earth2studio#1241 ·
Maintainers usually reply within 3 days
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Opendocumentation priority: low size: S
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Maintainers usually reply within 1 day
-
bug help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 65/100
ansys/pydpf-core#3547 ·
Maintainers usually reply within 1 day
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
OktoLabsAI/okto-pulse#114 ·
Maintainers usually reply within 1 day