Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Questions About InternVideo2clip Training Data and Fine-Tuning Requirements

未关闭
#293 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
15/100
Issue 类型
文档
描述清晰度
需要澄清
活跃度
停滞
技术栈
python

调研方向

从 InternVideo2 论文和链接的 InternVideo2-Stage2_1B-224p-f4 权重页面开始。确定项目是否提供 InternVideo2clip 训练数据集、中文数据相关信息和微调指南;仅记录经维护者或项目材料确认的细节。

由索引模型根据 Issue 内容生成。

描述

Thank you for your work! I have a question:

In the paper, it is stated: "We also learn a CLIP-style InternVideo2 indicated by InternVideo2clip. It is post-pretrained from InternVideo2s2 by only preserving video and text encoders and contrastive loss."

May I ask what training dataset was used for InternVideo2clip? Does it include any Chinese data? Approximately how much data would be required to fine-tune it effectively?

I noticed that only the attnpool of the vision encoder has been released in the official weights. https://huggingface.co/OpenGVLab/InternVideo2-Stage2_1B-224p-f4

主要语言
Python
星标
2.4k
派生
160
PR 合并指标
30 天内没有已合并 PR

贡献指南

这个仓库没有索引到贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

OpenGVLab/InternVideo 的其他 Issue

查看 OpenGVLab/InternVideo 的全部 Issue

相似的 Issue

更多 Python Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。