[Reproducibility] Training script / config for MiniMax-H3-TrainingAdapter (DeCFG adapter) itself
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Documentazione
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- documentation, machine-learning
Direzione di ricerca
Start with examples/minimax_h3/model_training/train.py and the scripts under examples/minimax_h3/model_training/lora/, then trace how preset_lora_path and preset_lora_model are consumed. Done means documenting or adding the adapter-training command, data selection, objective, hyperparameters, checkpoint choice, and weight conversion well enough to reproduce the named adapter files.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Hi, and thanks for open-sourcing the MiniMax-H3 work.
I'd like to reproduce the training adapter itself — https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter — not the downstream LoRAs.
What I've already found in the repo
examples/minimax_h3/model_training/train.pyand the two-stagesft:data_process→sft:trainworkflow.- The
.shscripts underexamples/minimax_h3/model_training/lora/(e.g.MiniMax-H3-Pruned-FL2VA.sh, the Ref2VA variants). These train downstream LoRAs and reference the adapter only as an optional--preset_lora_path/--preset_lora_model "dit"to fuse in during training. - The base weights (
Comfy-Org/MiniMax-H3,MiniMax/MiniMax-H3) and the MiniMax-H3-Self-Generated-Dataset.
What I can't find
A script or config that produces the adapter itself (the rank-64 DeCFG LoRA, FL2VA and Ref2VA variants) from the CFG-distilled base + the self-generated dataset. The example scripts consume the adapter rather than create it.
Could you share the following for the adapter's own training?
- The exact training script / command (equivalent
.sh) used to producemodel_for_comfy_dit.safetensorsandmodel_ref2va_for_comfy_dit.safetensors. - The DeCFG / "differential training" details — what the objective is and how it differs from standard SFT LoRA training (loss formulation, any CFG/DeCFG-specific handling, timestep sampling/shift).
- Hyperparameters: LoRA rank (64?) and target modules, learning rate, batch size, number of steps/epochs,
dataset_repeat, resolution /num_frames,audio_loss_weight, optimizer/schedule, and seed if fixed. - Which subset/split of MiniMax-H3-Self-Generated-Dataset was used, and whether the released dataset is the complete training set or a sample.
- Base checkpoint used as the starting point (the pruned bf16 DiT, or another), plus any pre/post-processing of the adapter weights (e.g. the ComfyUI qkv layout conversion).
Happy to open a PR to add a reproduction script/README under examples/minimax_h3/ if that's easier on your side. Thanks!
- Lingua principale
- Python
- Stelle
- 13.2k
- Fork
- 1.3k
- Merge medio
- 22h 15m
- PR unite (30g)
- 31
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di modelscope/DiffSynth-Studio
-
WanToDance: hardcoded `device='cuda'` in music encoder construction crashes model loading on Ascend NPUForse già presa @li-lizhe l’ha presa 12 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
modelscope/DiffSynth-Studio#1707 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Partial download surfaces as "Cannot detect the model type" rather than a download errorForse già presa @MohammadHijjawi97 l’ha presa 9 giorni fa. Aperta
Difficoltà 1/5 1-3 ore Idoneità per principianti 78/100
modelscope/DiffSynth-Studio#1668 ·
I maintainer di solito rispondono entro 1 giorno
-
bfloat16训练时段错误: 数组越界的一种方案Forse di nuovo libera Una pull request per questa issue è stata chiusa senza essere unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
modelscope/DiffSynth-Studio#1499 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 78/100
modelscope/DiffSynth-Studio#1373 · 5 commenti · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 30/100
modelscope/DiffSynth-Studio#1709 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di modelscope/DiffSynth-Studio
Issue simili
-
needs triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Forse già presa @mayureshsonawane17 l’ha presa oggi. ApertaWaiting for: Product Owner
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 5 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
I maintainer di solito rispondono entro 1 giorno
-
Add .devin pluginAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
ayghri/i-have-adhd#249 ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
modelscope/FunASR#3762 ·
I maintainer di solito rispondono entro 1 giorno