[Reproducibility] Training script / config for MiniMax-H3-TrainingAdapter (DeCFG adapter) itself
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 5/5
- Tempo estimado
- Mais de uma semana
- Facilidade para iniciantes
- 35/100
- Tipo de issue
- Documentação
- Clareza
- Razoavelmente clara
- Status de atividade
- Ativa
- Stack de tecnologia
- python
- Domínio
- documentation, machine-learning
Direção de pesquisa
Start with examples/minimax_h3/model_training/train.py and the scripts under examples/minimax_h3/model_training/lora/, then trace how preset_lora_path and preset_lora_model are consumed. Done means documenting or adding the adapter-training command, data selection, objective, hyperparameters, checkpoint choice, and weight conversion well enough to reproduce the named adapter files.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Hi, and thanks for open-sourcing the MiniMax-H3 work.
I'd like to reproduce the training adapter itself — https://modelscope.ai/models/DiffSynth-Studio/MiniMax-H3-TrainingAdapter — not the downstream LoRAs.
What I've already found in the repo
examples/minimax_h3/model_training/train.pyand the two-stagesft:data_process→sft:trainworkflow.- The
.shscripts underexamples/minimax_h3/model_training/lora/(e.g.MiniMax-H3-Pruned-FL2VA.sh, the Ref2VA variants). These train downstream LoRAs and reference the adapter only as an optional--preset_lora_path/--preset_lora_model "dit"to fuse in during training. - The base weights (
Comfy-Org/MiniMax-H3,MiniMax/MiniMax-H3) and the MiniMax-H3-Self-Generated-Dataset.
What I can't find
A script or config that produces the adapter itself (the rank-64 DeCFG LoRA, FL2VA and Ref2VA variants) from the CFG-distilled base + the self-generated dataset. The example scripts consume the adapter rather than create it.
Could you share the following for the adapter's own training?
- The exact training script / command (equivalent
.sh) used to producemodel_for_comfy_dit.safetensorsandmodel_ref2va_for_comfy_dit.safetensors. - The DeCFG / "differential training" details — what the objective is and how it differs from standard SFT LoRA training (loss formulation, any CFG/DeCFG-specific handling, timestep sampling/shift).
- Hyperparameters: LoRA rank (64?) and target modules, learning rate, batch size, number of steps/epochs,
dataset_repeat, resolution /num_frames,audio_loss_weight, optimizer/schedule, and seed if fixed. - Which subset/split of MiniMax-H3-Self-Generated-Dataset was used, and whether the released dataset is the complete training set or a sample.
- Base checkpoint used as the starting point (the pruned bf16 DiT, or another), plus any pre/post-processing of the adapter weights (e.g. the ComfyUI qkv layout conversion).
Happy to open a PR to add a reproduction script/README under examples/minimax_h3/ if that's easier on your side. Thanks!
- Linguagem predominante
- Python
- Estrelas
- 13.2k
- Forks
- 1.3k
- Merge médio
- 22h 15min
- PRs com merge (30d)
- 31
Preparar o ambiente
Este projeto não oferece contêiner de desenvolvimento, Dockerfile nem guia de contribuição, então a configuração fica por sua conta: comece pelo README e veja nosso guia da primeira contribuição para os passos gerais.
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de modelscope/DiffSynth-Studio
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
modelscope/DiffSynth-Studio#1707 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 1-3 horas Facilidade para iniciantes 78/100
modelscope/DiffSynth-Studio#1668 ·
Mantenedores costumam responder em até 1 dia
-
bfloat16训练时段错误: 数组越界的一种方案Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 72/100
modelscope/DiffSynth-Studio#1499 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 78/100
modelscope/DiffSynth-Studio#1373 · 5 comentários · 1 reação ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 30/100
modelscope/DiffSynth-Studio#1709 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
Todas as issues de modelscope/DiffSynth-Studio
Issues semelhantes
-
repo-audit
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
scverse/repo-health#20 ·
Mantenedores costumam responder em até 1 dia
-
/context/prime scope override double-prefixes an entity-ref project and drops its scoped memoriesAberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
phasespace-labs/palinode#232 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 82/100
collective/icalendar#1858 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
lfx-mcp cannot supply global variables: LangflowClient drops X-LANGFLOW-GLOBAL-VAR-* from envAbertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 78/100
langflow-ai/langflow#15496 ·
Mantenedores costumam responder em até 1 dia