Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

InternVideo-Next Multi-modality probes

Aperta
#318 4 commenti 3 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
5/5
Tempo stimato
Più di una settimana
Idoneità per principianti
25/100
Tipo di issue
Documentazione
Chiarezza
Da chiarire
Stato di attività
Tranquilla
Stack tecnologico
python

Direzione di ricerca

L’issue non indica file sorgente, test o punti di ingresso. Inizia confrontando la configurazione di training multimodale di InternVideo2 citata nell’issue #312 con il paper e la descrizione della probe di InternVideo-Next; il lavoro sarebbe completo quando fossero disponibili risposte confermate dai maintainer alle domande sul training, sull’allineamento delle dimensioni e sui risultati riportati.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

Hello @Revliter ,
Thank you so much for your previous response on issue https://github.com/OpenGVLab/InternVideo/issues/312 — I really appreciate the time you've taken to help. I've been digging deeper into the text encoder training for InternVideo-Next and have run into a few questions I'd love your input on.

Question 1 — Correct Training Setup: Paper vs. InternVideo2
In the paper, you mention freezing the ViT backbone and training only the text encoder. However, in issue #312 you pointed me toward the InternVideo2 multi-modality training, which uses a slightly different setup:

  • Vision backbone → fully frozen
  • Text backbone → fully frozen
  • clip-projector (vision side) → unfrozen
  • Alignment layer added on the vision side

Could you clarify which approach is correct for reproducing InternVideo-Next zero shot t2v results? Specifically: Should I follow the InternVideo2 setup exactly, or Adapt it to better match the paper — e.g., unfreeze the text backbone, and optionally freeze/unfreeze the clip-projector and add alignment on the text and/or vision side?

Question 2 — Dimension Alignment with SigLIP2 1B Teacher
You mentioned that SigLIP2 1B (giant opt) was used as a teacher in Stage 1 pretraining. However, its embedding dimensionality is quite different from the resulting InternVideo-Next vision encoder. How was dimension alignment handled between the two models?
Additionally — and I'm not sure if you tried this — i would expect the InternVideo-Next vision encoder shift away from SigLIP2's embedding space after Stage 2, making the two spaces incomparable at that point right?

Question 3 — Text-Side Training Settings, Epochs, and Room for Improvement
A few related sub-questions here:

  • Training config: Do the text-side training settings (temperature, epochs, weight decay, learning rate) fully follow the InternVideo2 configs?
  • Epoch discrepancy: In the paper, zero-shot T2V results are compared against InternVideo2 CLIP-L/14, which was trained for 3 epochs, whereas the InternVideo-Next multi-modality probe (Section: Multi-modality Tasks) mentions 5 epochs. Could you clarify this difference?
  • Are these results final? You refer to these experiments as probes — do you believe there is room for improvement with further tuning (e.g., dataset size, text encoder size, hyperparameters), or are the reported numbers the expected ceiling for this configuration?

I find this work incredibly insightful and plan to use the vision encoder in my diploma thesis given its strong potential. These clarifications would really help me move forward.
Thank you so much in advance for your time and help!

Lingua principale
Python
Stelle
2.4k
Fork
160
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Guida per i contributori

Nessuna guida per i contributori indicizzata per questo repository

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di OpenGVLab/InternVideo

Tutte le issue di OpenGVLab/InternVideo

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.