Normalize and filter dataset
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
Research direction
No files or tests are named. Start by locating the dataset pipeline, the normalization rules from foundations, and the canonical schema; then identify where filtering and quality metrics belong. Done means a cleaned canonical dataset, documented filtering statistics, and before/after quality reports covering every listed criterion.
Written by the indexing model from the issue text.
Description
Summary
Apply normalization rules and filter out low-quality commit messages.
Success Criteria
- Apply all normalization rules from foundations
- Filter criteria defined and applied:
- Remove non-conventional commits
- Remove commits with parsing errors
- Remove duplicates
- Remove auto-generated commits (dependabot, etc.)
- Quality metrics computed and documented
- Before/after statistics reported
Quality Filters
- Minimum subject length
- Valid type (feat, fix, docs, etc.)
- No merge commits
- English language only (v1)
Output
- Cleaned dataset in canonical schema format
- Quality report with filtering statistics
- Dominant language
- Python
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from aRustyDev/ccgram
-
docs experiment
Difficulty 5/5 Over a week Newbie friendliness 35/100
-
docs
Difficulty 4/5 3-5 days Newbie friendliness 55/100
-
Freeze architecture Openinfra model
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Document findings Opendocs experiment
Difficulty 3/5 3-5 days Newbie friendliness 45/100
-
ablation evaluation
Difficulty 4/5 3-5 days Newbie friendliness 30/100
All issues in aRustyDev/ccgram
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
anthropics/skills#1811 · 1 comment ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
speaches-ai/speaches#678 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
datalayer/mcp-compose#42 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
conda-forge/spacy-feedstock#177 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
UKGovernmentBEIS/inspect_evals#2523 ·