Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

MM validation wrappers truncate ground truth when batch_size exceeds one

Open Beginner friendly
#3,199 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@dajiaohuang is already working on this.

Since Oct 2, 2026.

  • #3200 by @dajiaohuang — open

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
72/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active

Research direction

Start with the MMDetDataset.len and MMSegDataset.len validation implementations mentioned in the issue. Reproduce the 8-sample, batch_size=2, one-GPU case and verify that the wrapper length matches the full underlying dataset length so all validation predictions and annotations remain aligned.

Written by the indexing model from the issue text.

Description

Summary

Both MMDetDataset.len and MMSegDataset.len return floor(dataset_length / (batch_size * num_gpus)) * num_gpus in validation mode. These wrappers count samples, but the formula uses the number of batches per GPU and omits multiplication by batch_size. For an 8-sample dataset with batch_size=2 and one GPU, len() returns 4. The validation DataLoader wraps the underlying dataset and still returns all samples when drop_last=False, while the evaluator uses the MMDet/MMSeg wrapper length to load annotations or ground-truth masks. The remaining predictions are therefore omitted or misaligned during evaluation.

Expected behavior

Validation wrapper length should match the number of examples whose predictions are evaluated. Since drop_last is configurable separately and not passed into these wrappers, their sample count should remain the full underlying dataset length.

Dominant language
C++
Stars
9.2k
Forks
725
PR merge metrics
No merged PRs in 30d

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from activeloopai/deeplake

All issues in activeloopai/deeplake

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.