Are all I-frame tokens intended to be preserved in the current implementation?
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start by reading ap_dataloader_dali_codec.py, focusing on get_frame_id_list and compute_visible_indices_cpu. Trace how zeroed I-frame residuals affect patch scores, then compare that behavior with the paper's Equation (2); done means confirming whether I-frame tokens are preserved and documenting or correcting the mismatch.
Written by the indexing model from the issue text.
Description
Hi, thanks for the great work on this HEVC-based token selection pipeline. I have a question about how I-frames are handled in ap_dataloader_dali_codec.py.
My understanding from the paper is that all tokens from I-frames are preserved, while Top-K selection is only applied to P-frame patches based on codec-derived saliency. In particular, Equation (2) seems to describe the HEVC input as keeping the full patchified I-frame and applying the visibility mask only to decoded P-frames.
However, in get_frame_id_list, I noticed that residuals at I-frame positions are explicitly zeroed out:
if pos in I_pos_set:
residuals_y[pos] = np.zeros((H0, W0), dtype=dtype0 or np.uint8)
Since patch scores in compute_visible_indices_cpu are computed from residual energy, this seems to imply that all I-frame patches receive a score of 0 and therefore would not be selected by Top-K, except possibly through tie-breaking or the static_fallback path.
So I wanted to check whether I am misunderstanding the implementation, or whether the current code is using a different behavior from what I inferred from the paper. If I-frame tokens are indeed intended to be fully preserved, could you clarify where that happens in the pipeline?
Thanks!
- Dominant language
- Python
- Stars
- 402
- Forks
- 20
- PR merge metrics
- No merged PRs in 30d
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EvolvingLMMs-Lab/OneVision-Encoder
-
Difficulty 1/5 Under an hour Newbie friendliness 20/100
-
Difficulty 3/5 1-2 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 45/100
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
EvolvingLMMs-Lab/OneVision-Encoder#116 · 1 comment ·
All issues in EvolvingLMMs-Lab/OneVision-Encoder
Similar issues
-
adr
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
kristofdegrave/homeassistant-smart-charging#1607 ·
Maintainers usually reply within 1 day
-
namespace operations
Difficulty 2/5 1-3 hours Newbie friendliness 64/100
EclipseFdn/open-vsx.org#13665 ·
Maintainers usually reply within 1 day
-
doc good first issue help wanted
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
collective/icalendar#1865 · 2 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
canonical/opentelemetry-collector-operator#409 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 85/100
mozilla/addons-release-tests#1243 ·
Maintainers usually reply within 1 day