How were the VAE latent mean/std values obtained, and why are they used for latent scaling?
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Documentation
- Clarity
- Needs clarification
- Activity status
- Quiet
- Tech stack
- python
- Domain
- machine-learning
Research direction
Start by tracing the Qwen-Image and Wan pipeline implementations where the channel-wise latent normalization is applied. Check the available VAE and training documentation or statistics-collection code to confirm how mean and standard deviation were obtained and why scaling is used; done means providing a clear, referenced explanation.
Written by the indexing model from the issue text.
Description
Hi DiffSynth-Studio team,
Thank you for open-sourcing this project.
I noticed that the Qwen-Image/Wan pipeline applies fixed channel-wise normalization to the VAE latent:
latents = (latents - mean) / std
Could you clarify:
How were these mean and std values calculated?
Were they computed by encoding the training data, taking the posterior mean for each sample, and then calculating channel-wise statistics over all samples and spatial-temporal positions?
Why is this scaling necessary?
My guess is that the raw latent channels have different means and variances, so this normalization makes their scales more balanced, better matches the scale of Gaussian noise, and prevents high-variance channels from dominating the flow-matching loss.
It would be very helpful if you could confirm this or share the statistics-collection method. Thanks!
- Dominant language
- Python
- Stars
- 13.2k
- Forks
- 1.3k
- Avg merge
- 22h 15m
- Merged PRs (30d)
- 31
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from modelscope/DiffSynth-Studio
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
modelscope/DiffSynth-Studio#1707 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 1-3 hours Newbie friendliness 78/100
modelscope/DiffSynth-Studio#1668 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
modelscope/DiffSynth-Studio#1499 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 78/100
modelscope/DiffSynth-Studio#1373 · 5 comments · 1 reaction ·
Maintainers usually reply within 1 day
-
[Reproducibility] Training script / config for MiniMax-H3-TrainingAdapter (DeCFG adapter) itselfOpen
Difficulty 5/5 Over a week Newbie friendliness 35/100
modelscope/DiffSynth-Studio#1713 · 1 comment ·
Maintainers usually reply within 1 day
All issues in modelscope/DiffSynth-Studio
Similar issues
-
tool-calling
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
vllm-project/vllm#59838 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
raullenchai/Rapid-MLX#4042 ·
Maintainers usually reply within 1 day
-
documentation
Difficulty 1/5 Under an hour Newbie friendliness 92/100
transitmatters/mbta-slow-zone-bot#70 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
litestar-org/advanced-alchemy#811 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day