how to get the evaluation result of gemma4?
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 55/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- jupyter-notebook, python
- Domain
- ai, documentation, machine-learning
Research direction
Look for evaluation scripts or notebooks in the repository, likely under an 'eval' or 'benchmarks' directory. Check the README for instructions on running evaluations. Compare the output format with the leaderboard image to understand the discrepancy. Running the evaluation on Gemma4 and verifying the results against the paper will confirm the process.
Written by the indexing model from the issue text.
Description
I find that the evaluation result of gemme4 is different from the paper, Could you please share how to evaluate gemma4?
paper:
here:
from leader board
- Dominant language
- Jupyter Notebook
- Stars
- 11.3k
- Forks
- 778
- Avg merge
- 5h 3m
- Merged PRs (30d)
- 6
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from OpenBMB/MiniCPM
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
-
Difficulty 5/5 Over a week Newbie friendliness 15/100
-
feature
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 3/5 1-2 days Newbie friendliness 62/100
Similar issues
-
GeminiUtil placeholder user turn ("Continue output. DO NOT look at this line ...") is flagged by prompt injection filtersPossibly taken @innoprej claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Nano-Collective/nanocoder#1639 ·
Maintainers usually reply within 4 days
-
sidecar: encoder disaggregation skips Anthropic image blocks on `/v1/messages`.Possibly taken @Mohammad-nassar10 claimed this today. Openkind/bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
llm-d/llm-d-router#3222 · 1 comment ·
Maintainers usually reply within 1 day
-
beta
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
cline/cline#14917 · 1 comment ·
Maintainers usually reply within 1 day
-
getRelevantContext can return a section heading with no entries under it when the budget is smallOpenbug
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day