performance regression on Ring SP
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python, pytorch
Research direction
Start with diffsynth_engine/layers/attention/backends/sdpa.py around lines 99–113 and trace how Ring SP selects the SDPA backend. Reproduce the comparison with the provided Qwen Image 2512, GPU, and kernel benchmarks, then verify that the affected Ring SP/CP paths no longer incur the slower kernel without changing other backends.
Written by the indexing model from the issue text.
Description
Description
In v1 branch, deploying DiffSynth engine with 2 GPUs without NVLink and inferencing model Qwen Image 2512, ring SP2 is much slower than Ulysses SP2.
However, the total communication volume is identical for both strategies when using 2 GPUs while ring CP can overlap same comunication time spend with computation. Which means CP would be faster theoretically.
Reason Explanations
In v1 branch, when use ring sp (or cp) with sdpa, the attention backend would be chose as _scaled_dot_product_efficient_attention and this kernel is far slower than _scaled_dot_product_flash_attention. Other backends would not be influenced.
Detailed comparison
On 4 RTX Pro 5000 Blackwell, for one 1024x1024 picture with 5 steps, the benchmarks table can be concluded as below.
| kernels | Efficient/kernel | Torch Flash/kernel | FA4/kernel | FA4 vs Flash | Torch Flash/step | FA4/step |
|---|---|---|---|---|---|---|
| cp2cfg | 733.068 us | 287.672 us | 279.057 us | -2.99% | 276.576 ms | 275.910 ms |
| cp4 | 197.949 us | 78.009 us | 75.263 us | -3.52% | 442.760 ms | 425.265 ms |
As we can see, the Torch Flash kernel is much faster than Efficient kernel.
- Dominant language
- Python
- Stars
- 432
- Forks
- 51
- Avg merge
- 3d 5h
- Merged PRs (30d)
- 1
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from modelscope/DiffSynth-Engine
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
modelscope/DiffSynth-Engine#238 · 1 comment ·
-
Difficulty 5/5 Over a week Newbie friendliness 20/100
modelscope/DiffSynth-Engine#235 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
modelscope/DiffSynth-Engine#224 · 2 comments ·
-
Difficulty 3/5 1-2 days Newbie friendliness 32/100
modelscope/DiffSynth-Engine#223 · 2 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 55/100
modelscope/DiffSynth-Engine#221 ·
All issues in modelscope/DiffSynth-Engine
Similar issues
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 90/100
learningequality/ricecooker#747 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
BSData/horus-heresy-3rd-edition#3171 ·
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
run-llama/llama_index#23199 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
KhronosGroup/glTF-Blender-IO#2769 ·