Expose per-image hashes and accept precomputed hashes for duplicate detection
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 80/100
- Issue type
- Feature
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start in the duplicate_issue_manager module (referenced in the issue's workaround) to find where per-image hashes are computed during find_issues(). First, expose these hashes in the issues output and info["exact_duplicates"]/info["near_duplicates"] entries. Then add a reference_hashes parameter to Imagelab.find_issues() that accepts a precomputed hash table, validates its columns and hash parameters against the current run, and merges it with newly computed hashes before grouping duplicates. Run the existing duplicate detection test suite to confirm no regressions, and add tests for the new functionality.
Written by the indexing model from the issue text.
Description
Expose per-image hashes and accept precomputed hashes in duplicate detection, so a new batch of images can be checked against earlier batches without re-reading them.
Details
Problem: When a dataset grows in batches, finding duplicates between a new batch and the earlier ones currently means running Imagelab over all images again, since it only works on image files and keeps no per-image hashes after find_issues() (only the resulting duplicate sets in info).
Proposal:
- Expose the per-image hashes computed in
find_issues(), e.g. as columns inimagelab.issuesor ininfo["exact_duplicates"]/info["near_duplicates"]. - Accept a table of precomputed hashes (id, hash type and parameters, hash) as extra, read-only members of the exact/near duplicate sets, e.g.
imagelab.find_issues(
issue_types={"exact_duplicates": {}, "near_duplicates": {}},
reference_hashes=reference_df, # columns: id, md5, phash
)
Since duplicates are grouped by hash equality, this gives the same sets as one run over all images, and the hash parameters (hash_type, hash_size) can be checked against the current run.
Workaround: Computing the hashes outside CleanVision with cleanvision.issue_managers.duplicate_issue_manager.get_hash() (same parameters as the defaults) and grouping by equal hash strings with pandas works, but depends on an internal function and duplicates CleanVision's grouping logic.
Alternatives considered: Imagelab.save()/load() keeps a whole instance tied to one data_path, so it doesn't help for comparing against images that are no longer on disk or live elsewhere.
- Dominant language
- Python
- Stars
- 1.2k
- Forks
- 84
- PR merge metrics
- No merged PRs in 30d
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from cleanlab/cleanvision
-
Broken linkPossibly taken @tdashelby-cmyk claimed this 151 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
cleanlab/cleanvision#274 ·
-
Documentation / UI improvements: sidebar behavior, “edit page” link & accessibility enhancementsOpen
Difficulty 5/5 Over a week Newbie friendliness 25/100
cleanlab/cleanvision#260 ·
-
question
Difficulty 4/5 3-5 days Newbie friendliness 25/100
cleanlab/cleanvision#259 · 1 comment ·
-
Difficulty 3/5 1-2 days Newbie friendliness 30/100
cleanlab/cleanvision#246 · 1 reaction ·
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
cleanlab/cleanvision#241 · 1 comment ·
All issues in cleanlab/cleanvision
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
NVIDIA/earth2studio#1241 ·
Maintainers usually reply within 3 days
-
docs(types): update the collection binding note now that typed collections shipped in pycubrid 1.9.0Opendocumentation priority: low size: S
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
cubrid-lab/sqlalchemy-cubrid#768 ·
Maintainers usually reply within 1 day
-
bug help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 65/100
ansys/pydpf-core#3547 ·
Maintainers usually reply within 1 day
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
OktoLabsAI/okto-pulse#114 ·
Maintainers usually reply within 1 day