Create test criteria.
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- python
- Domain
- machine-learning, search, testing-qa
Research direction
No files, tests, or entry points are named. Start by inspecting the repository for existing evaluation and retrieval code, then review the SelfCheckGPT reference and candidate public datasets; done means agreed metrics and test suites that measure hallucinations, response quality, and classic-versus-graph RAG recall.
Written by the indexing model from the issue text.
Description
Metrics
We need to define the metrics to create test suites and measure results.
I would like to test for hallucinations and response quality according to a grounded set.
Ideally, we would use something like SelfCheckGPT to check for hallucinations.
I'd also like to test recall for document retrieval on a known public dataset with classic vs graph RAG.
- Dominant language
- Python
- Stars
- 103
- Forks
- 16
- Avg merge
- 15h 29m
- Merged PRs (30d)
- 1
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/deepRAG
-
Ontology Open
Difficulty 5/5 Over a week Newbie friendliness 25/100
All issues in microsoft/deepRAG
Similar issues
-
documentation help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
simonw/sqlite-utils#872 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100