InternVideo-NeXt clip_projector training procedure: eval only? from scratch?
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 25/100
- Tipo di issue
- Documentazione
- Chiarezza
- Da chiarire
- Stato di attività
- Ferma
- Stack tecnologico
- huggingface, python
- Ambito
- computer-vision, documentation, machine-learning
Direzione di ricerca
Inizia dal file modeling_internvideo_next.py di Hugging Face collegato e da InternVideo2/single_modality/run_linear_probing.py, quindi confronta il comportamento dei relativi proiettori con le sezioni B.2 e 4.2 dell’articolo. Il lavoro è completato quando sono documentati i dati di addestramento di clip_projector, se la valutazione usa pesi addestrati da zero o preaddestrati e quale configurazione ha prodotto i risultati riportati.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Dear authors,
Thanks for your interesting paper!
I see you mention in your paper in section B.2 of the appendix that you train an attention pooling head for evaluation.
At the same time, I see that the model you released on HuggingFace has a clip_projector module (AttentionPoolingBlock) attached to it, and in the forward of the model you set projected=True (HF code here)
What I would like to know is:
- on what data was this release clip_projector trained?
- if I want to evaluate your model on action recognition, should I:
a. train the attentive probe from scratch on the target dataset
b. finetune the attentive probe on the target dataset (if so, from which weights?) - which setup did you use in your paper, 2.a. or 2.b., or none of them?
You released some action recognition code for InternVideo2 and for linear probing / attention probing there is this --open_clip_projector parameter which controls whether you're finetuning the head, or not. But this doesn't say on which data the head was trained before the finetuning.
You mention in the paper in section 4.2 (Video Classification)
We test the model in an ‘Attentive Probing’ setting
where the encoders are frozen and a single-layer attention
pooling head is trained. Such Frozen Encoder settings can
test representation’s quality in an unbiased way. Our methods achieve the best results with only public data and less
computation cost on these foundation tasks.
What would make sense to me is that you're not using the Internvideo2 evaluation script anymore and you're training the attentive probe from scratch (2.a. mentioned above). But if so, I'd like to know on which evaluation experiment the weights of the release clip_projector were obtained.
Thanks in advance for your clarifications!
- Lingua principale
- Python
- Stelle
- 2.4k
- Fork
- 160
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Guida per i contributori
Nessuna guida per i contributori indicizzata per questo repository
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di OpenGVLab/InternVideo
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
OpenGVLab/InternVideo#324 · 1 commento ·
-
InternVideo2 stage-1 weights Aperta
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
OpenGVLab/InternVideo#323 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
OpenGVLab/InternVideo#322 ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 42/100
OpenGVLab/InternVideo#321 ·
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
OpenGVLab/InternVideo#319 · 1 commento ·
Tutte le issue di OpenGVLab/InternVideo
Issue simili
-
essnmx good first issue
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 95/100
-
[Feature] 奇物选择添加优先级 Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
syfoud/Simulated_Scepter#174 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Giskard-AI/giskard-oss#2840 · 1 commento ·
-
A claim comment carrying the issue number is silently declined while the workflow reports success Apertaarea: repo bug perceived difficulty: 2
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
yeti-platform/yeti#1380 ·