Use of cuda::std::atomic is surprising
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 35/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- cpp
- Domain
- performance
Research direction
Start by locating the atomics-selection logic that switches to <cuda/std/atomic> when the header is available, then inspect cuda::std::atomic::wait behavior on Windows versus WaitOnAddress. Define the chosen policy for avoiding or explicitly controlling CUDA atomics, and verify that globally installed CUDA no longer causes surprising selection or high-latency waits.
Written by the indexing model from the issue text.
Description
The code switches its whole atomics implementation to the one in CUDA if the <cuda/std/atomic> header is available. This can already be the case if CUDA is just installed globally or if the application uses CUDA in different parts and not for stdexec.
The problem is that on Windows, the cuda implementation is worse, for instance cuda::std::atomic<T>::wait seems to fall back to polling, instead of WaitOnAddress which leads to very high latencies (on the order of the scheduler tick of ~15 ms).
This is quite surprising and it would be good if it did not happen at all. Failing that it would be good if the switch were more explicit (maybe opt-in and fail if cuda atomics are needed for something?) and failing that it would be good if there was a way to easily disable the switch to cuda atomics.
- Dominant language
- C++
- Stars
- 2.4k
- Forks
- 270
- Avg merge
- 2d 16h
- Merged PRs (30d)
- 43
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA/stdexec
-
inline_scheduler's namespace-scope static_assert fails under nvcc (private nested __sender access) Open
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
-
Difficulty 4/5 3-5 days Newbie friendliness 48/100
-
Difficulty 4/5 3-5 days Newbie friendliness 65/100
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
google/libultrahdr#485 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
godotengine/godot#123776 ·
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 60/100
-
good first issue
Difficulty 1/5 Under an hour Newbie friendliness 90/100
-
good first issue
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
ros2/common_interfaces#344 ·