Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Clarification on how OCR annotations are used during training

Open
#105 2 comments 1 reaction 1 assignee View on GitHub

@anxiangsir is already working on this.

Since Apr 9, 2026.

Assessment

This issue has not been assessed yet.

Description

Hi, thank you for releasing this excellent work.

While reading the paper, there seems to be one point that is still unclear: how the OCR annotations are actually incorporated into training.

From the paper, the following part is understood:

PaddleOCR is applied to images from OBELICS and Zero250M
the recognized text is tokenized
100 fine-grained tags are constructed for each image
OCR data is introduced in Stage 2 together with video supervision

However, the paper does not seem to explicitly describe how these OCR-derived tags are optimized in the training objective.

Dominant language
Python
Stars
403
Forks
20
PR merge metrics
No merged PRs in 30d

Getting set up

  • Ships a Dockerfile or Docker Compose file
  • No pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from EvolvingLMMs-Lab/OneVision-Encoder

All issues in EvolvingLMMs-Lab/OneVision-Encoder

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.