[BUG]: FFT sample launches `Batch` blocks that each process the full batch
Nobody has claimed this yet.
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- python
- Domain
- performance
Research direction
Start in samples/FFT.py at cutile_fft() around the BS assignment on line 274 and grid launch on line 315, then read fft_kernel to trace how bid selects its tile. Run the FFT sample with the reported N=512 and batch=64 configuration, and verify that the output remains numerically correct while kernel work and timing scale linearly with the batch.
Written by the indexing model from the issue text.
Description
Version
1.5.0
Version
13.3
Describe the bug.
In samples/FFT.py, cutile_fft() sets the kernel constant BS to the full batch size (BS = x.shape[0], line 274) and also launches grid = (BS, 1, 1) (line 315). Inside fft_kernel every block loads a (BS, N*2//D, D) tile at index (bid, 0, 0) — i.e. every block loads and transforms the entire batch, then writes it out. The result is numerically correct, but the work is O(Batch²) instead of O(Batch), and the kernel spills registers / shared memory at modest batch sizes.
Expected: one block per batch item (or per fixed-size minibatch), with the grid sized Batch // BS, so cost scales linearly with batch.
Measured on a DGX Spark, N=512, batch=64, factors=(8,8,8), twiddles precomputed: kernel time 2376 µs -> 12 µs (~200x) after fixing the grid/BS relationship.
Contributing Guidelines
- I agree to follow cuTile Python's contributing guidelines
- I have searched the open bugs and have found no duplicates for this bug report
- Dominant language
- Python
- Stars
- 2.2k
- Forks
- 155
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/cutile-python
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA/cutile-python#105 · 2 comments ·
-
Difficulty 4/5 3-5 days Newbie friendliness 68/100
NVIDIA/cutile-python#101 ·
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
NVIDIA/cutile-python#97 · 1 comment ·
-
bug
Difficulty 4/5 3-5 days Newbie friendliness 52/100
NVIDIA/cutile-python#96 · 1 comment ·
-
bug
NVIDIA/cutile-python#95 · 1 comment · 1 assignee ·
All issues in NVIDIA/cutile-python
Similar issues
-
essnmx good first issue
Difficulty 1/5 Under an hour Newbie friendliness 95/100
-
[Feature] 奇物选择添加优先级 Open
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
syfoud/Simulated_Scepter#174 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Giskard-AI/giskard-oss#2840 · 1 comment ·
-
A claim comment carrying the issue number is silently declined while the workflow reports success Openarea: repo bug perceived difficulty: 2
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
yeti-platform/yeti#1380 ·