[Reproducibility] Training script / config for MiniMax-H3-TrainingAdapter (DeCFG adapter) itself
メンテナーはふだん 1 日以内に返信
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 35/100
- issue の種類
- ドキュメント
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- python
調査の方向性
Start with examples/minimax_h3/model_training/train.py and the scripts under examples/minimax_h3/model_training/lora/, then trace how preset_lora_path and preset_lora_model are consumed. Done means documenting or adding the adapter-training command, data selection, objective, hyperparameters, checkpoint choice, and weight conversion well enough to reproduce the named adapter files.
索引モデルが issue の本文から書いたものです。
説明
Hi, and thanks for open-sourcing the MiniMax-H3 work.
I'd like to reproduce the training adapter itself — https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter — not the downstream LoRAs.
What I've already found in the repo
examples/minimax_h3/model_training/train.pyand the two-stagesft:data_process→sft:trainworkflow.- The
.shscripts underexamples/minimax_h3/model_training/lora/(e.g.MiniMax-H3-Pruned-FL2VA.sh, the Ref2VA variants). These train downstream LoRAs and reference the adapter only as an optional--preset_lora_path/--preset_lora_model "dit"to fuse in during training. - The base weights (
Comfy-Org/MiniMax-H3,MiniMax/MiniMax-H3) and the MiniMax-H3-Self-Generated-Dataset.
What I can't find
A script or config that produces the adapter itself (the rank-64 DeCFG LoRA, FL2VA and Ref2VA variants) from the CFG-distilled base + the self-generated dataset. The example scripts consume the adapter rather than create it.
Could you share the following for the adapter's own training?
- The exact training script / command (equivalent
.sh) used to producemodel_for_comfy_dit.safetensorsandmodel_ref2va_for_comfy_dit.safetensors. - The DeCFG / "differential training" details — what the objective is and how it differs from standard SFT LoRA training (loss formulation, any CFG/DeCFG-specific handling, timestep sampling/shift).
- Hyperparameters: LoRA rank (64?) and target modules, learning rate, batch size, number of steps/epochs,
dataset_repeat, resolution /num_frames,audio_loss_weight, optimizer/schedule, and seed if fixed. - Which subset/split of MiniMax-H3-Self-Generated-Dataset was used, and whether the released dataset is the complete training set or a sample.
- Base checkpoint used as the starting point (the pruned bf16 DiT, or another), plus any pre/post-processing of the adapter weights (e.g. the ComfyUI qkv layout conversion).
Happy to open a PR to add a reproduction script/README under examples/minimax_h3/ if that's easier on your side. Thanks!
- 主要言語
- Python
- スター
- 13.2k
- フォーク
- 1.3k
- 平均マージ
- 12時間 52分
- マージ済み PR(30日)
- 20
環境構築
このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
modelscope/DiffSynth-Studio のほかの issue
-
WanToDance: hardcoded `device='cuda'` in music encoder construction crashes model loading on Ascend NPU対応中かも @li-lizhe が 17 日前に担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
modelscope/DiffSynth-Studio#1707 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Partial download surfaces as "Cannot detect the model type" rather than a download error対応中かも @MohammadHijjawi97 が 15 日前に担当しました。 オープン
難易度 1/5 1〜3時間 初心者へのやさしさ 78/100
modelscope/DiffSynth-Studio#1668 ·
メンテナーはふだん 1 日以内に返信
-
bfloat16训练时段错误: 数组越界的一种方案再び着手できるかも このイシューのプルリクエストはマージされずにクローズされました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
modelscope/DiffSynth-Studio#1499 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 78/100
modelscope/DiffSynth-Studio#1373 · コメント 5 件 · リアクション 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 5/5 1週間以上 初心者へのやさしさ 8/100
modelscope/DiffSynth-Studio#1735 · コメント 9 件 ·
メンテナーはふだん 1 日以内に返信
modelscope/DiffSynth-Studio の issue をすべて見る
似ている issue
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitオープンneeds-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 77/100
krkn-chaos/krkn#1627 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
NousResearch/hermes-agent#136483 ·
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last element対応中かも @peterdsharpe が今日担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
pytorch/tensordict#2307 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信