Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Expose per-image hashes and accept precomputed hashes for duplicate detection

Open Beginner friendly
#278 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
80/100
Issue type
Feature
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start in the duplicate_issue_manager module (referenced in the issue's workaround) to find where per-image hashes are computed during find_issues(). First, expose these hashes in the issues output and info["exact_duplicates"]/info["near_duplicates"] entries. Then add a reference_hashes parameter to Imagelab.find_issues() that accepts a precomputed hash table, validates its columns and hash parameters against the current run, and merges it with newly computed hashes before grouping duplicates. Run the existing duplicate detection test suite to confirm no regressions, and add tests for the new functionality.

Written by the indexing model from the issue text.

Description

Expose per-image hashes and accept precomputed hashes in duplicate detection, so a new batch of images can be checked against earlier batches without re-reading them.

Details

Problem: When a dataset grows in batches, finding duplicates between a new batch and the earlier ones currently means running Imagelab over all images again, since it only works on image files and keeps no per-image hashes after find_issues() (only the resulting duplicate sets in info).

Proposal:

  1. Expose the per-image hashes computed in find_issues(), e.g. as columns in imagelab.issues or in info["exact_duplicates"] / info["near_duplicates"].
  2. Accept a table of precomputed hashes (id, hash type and parameters, hash) as extra, read-only members of the exact/near duplicate sets, e.g.
imagelab.find_issues(
    issue_types={"exact_duplicates": {}, "near_duplicates": {}},
    reference_hashes=reference_df,  # columns: id, md5, phash
)

Since duplicates are grouped by hash equality, this gives the same sets as one run over all images, and the hash parameters (hash_type, hash_size) can be checked against the current run.

Workaround: Computing the hashes outside CleanVision with cleanvision.issue_managers.duplicate_issue_manager.get_hash() (same parameters as the defaults) and grouping by equal hash strings with pandas works, but depends on an internal function and duplicates CleanVision's grouping logic.

Alternatives considered: Imagelab.save()/load() keeps a whole instance tied to one data_path, so it doesn't help for comparing against images that are no longer on disk or live elsewhere.

Dominant language
Python
Stars
1.2k
Forks
84
PR merge metrics
No merged PRs in 30d

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from cleanlab/cleanvision

All issues in cleanlab/cleanvision

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.