InternVideo-Next Multi-modality probes
还没有人认领这个 Issue。
评估
- 难度
- 5/5
- 预计耗时
- 一周以上
- 新手友好度
- 25/100
- Issue 类型
- 文档
- 描述清晰度
- 需要澄清
- 活跃度
- 冷清
- 技术栈
- python
调研方向
该 issue 没有指出源文件、测试或入口点。首先比较 issue #312 中引用的 InternVideo2 多模态训练设置与 InternVideo-Next 的 paper 和 probe 描述;只有在获得 maintainer 确认的、针对训练、维度对齐和已报告结果相关问题的回答后,才算完成。
由索引模型根据 Issue 内容生成。
描述
Hello @Revliter ,
Thank you so much for your previous response on issue https://github.com/OpenGVLab/InternVideo/issues/312 — I really appreciate the time you've taken to help. I've been digging deeper into the text encoder training for InternVideo-Next and have run into a few questions I'd love your input on.
Question 1 — Correct Training Setup: Paper vs. InternVideo2
In the paper, you mention freezing the ViT backbone and training only the text encoder. However, in issue #312 you pointed me toward the InternVideo2 multi-modality training, which uses a slightly different setup:
- Vision backbone → fully frozen
- Text backbone → fully frozen
- clip-projector (vision side) → unfrozen
- Alignment layer added on the vision side
Could you clarify which approach is correct for reproducing InternVideo-Next zero shot t2v results? Specifically: Should I follow the InternVideo2 setup exactly, or Adapt it to better match the paper — e.g., unfreeze the text backbone, and optionally freeze/unfreeze the clip-projector and add alignment on the text and/or vision side?
Question 2 — Dimension Alignment with SigLIP2 1B Teacher
You mentioned that SigLIP2 1B (giant opt) was used as a teacher in Stage 1 pretraining. However, its embedding dimensionality is quite different from the resulting InternVideo-Next vision encoder. How was dimension alignment handled between the two models?
Additionally — and I'm not sure if you tried this — i would expect the InternVideo-Next vision encoder shift away from SigLIP2's embedding space after Stage 2, making the two spaces incomparable at that point right?
Question 3 — Text-Side Training Settings, Epochs, and Room for Improvement
A few related sub-questions here:
- Training config: Do the text-side training settings (temperature, epochs, weight decay, learning rate) fully follow the InternVideo2 configs?
- Epoch discrepancy: In the paper, zero-shot T2V results are compared against InternVideo2 CLIP-L/14, which was trained for 3 epochs, whereas the InternVideo-Next multi-modality probe (Section: Multi-modality Tasks) mentions 5 epochs. Could you clarify this difference?
- Are these results final? You refer to these experiments as probes — do you believe there is room for improvement with further tuning (e.g., dataset size, text encoder size, hyperparameters), or are the reported numbers the expected ceiling for this configuration?
I find this work incredibly insightful and plan to use the vision encoder in my diploma thesis given its strong potential. These clarifications would really help me move forward.
Thank you so much in advance for your time and help!
- 主要语言
- Python
- 星标
- 2.4k
- 派生
- 160
- PR 合并指标
- 30 天内没有已合并 PR
环境准备
我们还没有检查这个项目的环境配置文件。先看它的 README,通用步骤见我们的新手贡献指南。
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
OpenGVLab/InternVideo 的其他 Issue
-
难度 5/5 一周以上 新手友好度 25/100
OpenGVLab/InternVideo#324 · 1 条评论 ·
-
难度 5/5 一周以上 新手友好度 20/100
OpenGVLab/InternVideo#323 ·
-
难度 5/5 一周以上 新手友好度 25/100
OpenGVLab/InternVideo#322 ·
-
难度 4/5 3-5 天 新手友好度 42/100
OpenGVLab/InternVideo#321 ·
-
难度 4/5 3-5 天 新手友好度 35/100
OpenGVLab/InternVideo#319 · 1 条评论 ·
查看 OpenGVLab/InternVideo 的全部 Issue
相似的 Issue
-
good first issue
难度 2/5 1-3 小时 新手友好度 88/100
-
难度 2/5 1-3 小时 新手友好度 88/100
vllm-project/vllm-metal#822 ·
维护者通常 1 天内回复
-
vector-store
难度 1/5 1-3 小时 新手友好度 90/100
维护者通常 1 天内回复
-
[Bug]: chunk_span_bounds and _validated_chunk_spans reject Pydantic models ChunkSpan and AudioFile未关闭
难度 2/5 1-3 小时 新手友好度 78/100
BasedHardware/omi#19047 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复