Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

DAXPY outperforms DSCAL in multi-threaded environments

未關閉
#5,328 2 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

維護者通常 1 天內回覆

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
35/100
Issue 類型
缺陷
描述清晰度
基本清楚
活躍度
停滯
技術堆疊
c
領域
performance

研究方向

首先,在長度為 80,000 的向量上將 OPENBLAS_NUM_THREADS 分別設為 1 和 16,重現回報的 dscal 和 daxpy 執行時間。比較 dscal 和 daxpy 的執行路徑,並記錄它們的多執行緒行為為何不同;當差異得到解釋,且任何變更都已針對這兩種執行緒設定完成驗證時,即視為完成。

由索引模型根據 Issue 內容生成。

描述

Issue Description:

I'm observing a significant performance disparity between dscal and daxpy when performing vector-scalar multiplication on an Intel(R) Xeon(R) Platinum 8378C CPU @ 2.80GHz. My code involves the operation y=ax, where x is a vector of length 80,000.

Observed Behavior:

Despite setting the OPENBLAS_NUM_THREADS environment variable to either 1 or 16, the execution time for dscal remains unchanged, indicating no utilization of multiple cores.

However, when I replace dscal with an equivalent operation using daxpy, specifically y=(a−1)x+x (having a loss in precision), I observe a multi-fold performance improvement in the multi-core scenario.

Problem:

Given that dscal and daxpy have very similar computational patterns, I'm seeking to understand why there's such a substantial difference in their multi-core performance. This behavior suggests that dscal is not effectively leveraging the available CPU cores, unlike daxpy.

主要語言
C
星號
7.6k
分支
1.7k
平均合併
1 天 8 小時
30 天內合併 PR
48

環境準備

這個專案沒有提供開發容器、Dockerfile 或貢獻指南,環境需要你自己搭建:先看它的 README,通用步驟見我們的新手貢獻指南。

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

OpenMathLib/OpenBLAS 的其他 Issue

查看 OpenMathLib/OpenBLAS 的全部 Issue

相似的 Issue

更多 C Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。