Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Parallelization along K-dimension (Parallel Reduction) for GEMM with small M/N and large K

未关闭
#5,629 2 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
30/100
Issue 类型
功能
描述清晰度
需要澄清
活跃度
冷清
技术栈
c
领域
performance

调研方向

从 zgemm 当前的 threading 路径和 GEMM K-loop 入手,使用报告中的 small-M/N, large-K 形状来复现单线程行为。检查是否已经支持 K-dimension 并行归约;要算 done,需要针对这一情况做出明确的实现决策并提供性能证据。

由索引模型根据 Issue 内容生成。

描述

Hi OpenBLAS team,

I noticed that zgemm (and other GEMM functions) falls back to single-threaded execution when M and N are small (e.g., 32) but K is extremely large (e.g., 1,000,000).

On my many-core system, this leaves most cores idle. Given the large K size, parallelizing the K-loop (via parallel reduction) should theoretically offer significant speedup. I perform the matrix partitioning (of k) externally, and then use multithreading to call zgemm, but the performance is only average.

Questions:

Does OpenBLAS currently support threading along the K-dimension for this shape?

If not, are there any plans to implement parallel reduction for large K?

My Machine Info:

Image Image

Thanks!

主要语言
C
星标
7.6k
派生
1.7k
平均合并
1 天 6 小时
30 天内合并 PR
46

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

OpenMathLib/OpenBLAS 的其他 Issue

查看 OpenMathLib/OpenBLAS 的全部 Issue

相似的 Issue

更多 C Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。