save right and wrong file in test_regex.py and improve accuracy stats
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 42/100
Research direction
Start by reading and running test_regex.py to trace how the right and wrong variables, rule_id, rule_group, and original regex are produced. Done means the results are saved in JSON, rule identifiers remain tied to their source regexes, and statistics are printed per rule_id, per rule_group, and overall.
Written by the indexing model from the issue text.
Description
Save away the "right" and "wrong" variables returned in test_regex.py in a json file. Also, tie back the rule_id/rule_group numbers back to the original regex. Print some stats, by rule_id, rule_group and overall regex.
- Dominant language
- Python
- Stars
- 9
- Forks
- 6
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from bigscience-workshop/pii_processing
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
-
add custom lexicon Open
Difficulty 3/5 1-2 days Newbie friendliness 35/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 45/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 40/100
All issues in bigscience-workshop/pii_processing
Similar issues
-
documentation help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
simonw/sqlite-utils#872 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100