Clarification on how OCR annotations are used during training
@anxiangsir is already working on this.
Since Apr 9, 2026.
Assessment
This issue has not been assessed yet.
Description
Hi, thank you for releasing this excellent work.
While reading the paper, there seems to be one point that is still unclear: how the OCR annotations are actually incorporated into training.
From the paper, the following part is understood:
PaddleOCR is applied to images from OBELICS and Zero250M
the recognized text is tokenized
100 fine-grained tags are constructed for each image
OCR data is introduced in Stage 2 together with video supervision
However, the paper does not seem to explicitly describe how these OCR-derived tags are optimized in the training objective.
- Dominant language
- Python
- Stars
- 403
- Forks
- 20
- PR merge metrics
- No merged PRs in 30d
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EvolvingLMMs-Lab/OneVision-Encoder
-
Difficulty 1/5 Under an hour Newbie friendliness 20/100
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 45/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
EvolvingLMMs-Lab/OneVision-Encoder#116 · 1 comment ·
All issues in EvolvingLMMs-Lab/OneVision-Encoder
Similar issues
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
-
Difficulty 1/5 Under an hour Newbie friendliness 75/100
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
data-umbrella/du-event-board#225 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100