关于 07/03_prefetch/06 运行结果的疑问
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- cpp
- Domain
- hpc, performance
Research direction
Start with the 07/03_prefetch/06 example and reproduce the two benchmark runs on the reported Intel i5-13500, Ubuntu 24.04, and GCC 13.2. Compare the results with and without #pragma omp parallel for; the issue is resolved when the effect of the pragma and the differing benchmark timings are explained or documented.
Written by the indexing model from the issue text.
Description
hi, 小彭老师好。关于 07/03_prefetch/06 例子运行结果我有一些疑问,望指正。
我的平台是 Intel i5-13500, Ubuntu 24.04, gcc version 13.2.0
在运行 07/03_prefetch/06 这个例子时,
去掉例子中的 #pragma omp parallel for 才能得到与课程中类似的结果。我不清楚 #pragma omp parallel for 是否除了并行之外还有其他的优化?
原始版本运行结果
从运行结果可以看到,BM_write_stream_then_read 跟 BM_write_streamed 运行耗时相近,似乎读对 stream 指令并没有影响
-----------------------------------------------------------------------
Benchmark Time CPU Iterations
-----------------------------------------------------------------------
BM_read 25228152 ns 18180668 ns 38
BM_write 32696238 ns 25309548 ns 33
BM_write_streamed 19530899 ns 17132181 ns 36
BM_write_stream_then_read 19586335 ns 17525509 ns 43
BM_write_streamed_ps 19550735 ns 14485110 ns 39
BM_write_streamed_ps_skipped 37094026 ns 26238143 ns 26
BM_read_and_write 36829027 ns 33520956 ns 22
去除 #pragma omp parallel for 版本运行结果
从运行结果可以看到,BM_write_stream_then_read 运行耗时显著比 BM_write_streamed 长
-----------------------------------------------------------------------
Benchmark Time CPU Iterations
-----------------------------------------------------------------------
BM_read 38213301 ns 38207623 ns 19
BM_write 52209723 ns 52203705 ns 13
BM_write_streamed 34738316 ns 34735390 ns 20
BM_write_stream_then_read 40930259 ns 40927256 ns 17
BM_write_streamed_ps 17725541 ns 17724305 ns 36
BM_write_streamed_ps_skipped 36891533 ns 36889477 ns 19
BM_read_and_write 44972351 ns 44969916 ns 12
- Dominant language
- C++
- Stars
- 4.2k
- Forks
- 555
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from parallel101/course
-
访问者设计模式缺少继承Open
Difficulty 1/5 1-3 hours Newbie friendliness 68/100
parallel101/course#35 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 62/100
parallel101/course#29 ·
-
Difficulty 1/5 Under an hour Newbie friendliness 45/100
parallel101/course#39 ·
-
C++17之后才保证new一定对齐。Open
Difficulty 2/5 1-3 hours Newbie friendliness 42/100
parallel101/course#38 ·
-
函数返回匿名结构体无法编译Open
Difficulty 2/5 1-3 hours Newbie friendliness 45/100
parallel101/course#36 · 2 comments ·
All issues in parallel101/course
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Maintainers usually reply within 1 day
-
hipRTC lit tests compile against /opt/rocm's LLVM instead of the ROCm under test (ci/ hardcodes LLVM_PATH)Possibly taken @bernardogv claimed this today. Open
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
請增加教學:數字後的句號Open
Difficulty 1/5 Under an hour Newbie friendliness 70/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 79/100
Maintainers usually reply within 1 day