Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Bug: "train_test_split" trains and tests on the exact same data

Open Beginner friendly
#9 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
1/5
Estimated time
Under an hour
Newbie friendliness
85/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

The bug is in mess_mood/evaluate.py's train_test_split function, which incorrectly returns rows, rows instead of slicing the input using the precomputed cut index. Start by running the provided test in tests/test_mess_mood.py to confirm the failure, fix the return statement to return the correctly sized non-overlapping train and test splits, then re-run the test to verify the fix works.

Written by the indexing model from the issue text.

Description

What I did:
I reviewed the implementation of the train_test_split function in mess_mood/evaluate.py. To verify my findings, I added a test case in tests/test_mess_mood.py that asserts that the train and test sets should have a length of 80 and 20 respectively (for an input of 100 items), and that they should have zero overlap:

Click to view the test code
def test_train_test_split_does_not_overlap():
    from mess_mood.evaluate import train_test_split
    
    rows = [(str(i), "x") for i in range(100)]
    train, test = train_test_split(rows, test_size=0.2)
    
    # We expect 80% train and 20% test
    assert len(train) == 80
    assert len(test) == 20
    
    # We expect 0 overlapping items
    train_ids = {r[0] for r in train}
    test_ids = {r[0] for r in test}
    assert not train_ids.intersection(test_ids)

What I expected:
According to the README, evaluate should train on 80% of the reviews and test on the other 20%, and the two sets should never overlap. I expected the function to slice the array using the cut index and return correctly sized train and test datasets, and for the test to pass.

What happened instead:
The function calculates the cut index but then completely ignores it, instead returning return rows, rows. This causes the model to train and test on the exact same 100% of the data.

Running the test confirms this, failing immediately because len(train) evaluates to 100 instead of 80:

Proof (Terminal Output / Screenshot):

Image
Dominant language
Python
Stars
0
Forks
3
Avg merge
4h 39m
Merged PRs (30d)
5

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from techcsispit/mess-mood

All issues in techcsispit/mess-mood

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.