Request: support add perf testing of matrix multiply "accelerator units" in benchmark..
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
- Issue type
- Feature
- Clarity
- Mostly clear
- Activity status
- Stale
- Tech stack
- cpp
- Domain
- performance
Research direction
Start by locating the existing int8 dot-product benchmark and the project’s OpenCL extension handling. Review the referenced ARM and Intel matrix-multiply extensions and compare their availability across vendors. Done means the benchmark has a documented cross-vendor matrix-multiply accelerator performance test, with results that can be run on supported hardware.
Written by the indexing model from the issue text.
Description
Hi,
seems after dot product (int8) support next "hot" feature could add is "matrix multiply accelerator" perf testing support..
aka tensor cores on nvidia etc...
searching on opencl.gpuinfo.org I see ARM and Intel extensions (not Nvidia nor AMD nor Qualcomm which is sad to see):
*cl_arm_matrix_multiply 0.4%
*cl_intel_subgroup_matrix_multiply_accumulate 2.2%
*cl_intel_subgroup_matrix_multiply_accumulate_tf32 0.5%
*cl_intel_subgroup_split_matrix_multiply_accumulate 1.6%
of intel the one to test seems "cl_intel_subgroup_matrix_multiply_accumulate" (as tf32 is the tensorfloat32 variant) supported on Intel ARC on Windows and Linux and also..
the ARM one is only supported on Pixel products which is sad:
https://opencl.gpuinfo.org/listreports.php?extension=cl_arm_matrix_multiply
more sad is that no ext seems published altough seems to be supported in ARM Compute library which is open source I think so maybe there is code to learn from..:
https://android.googlesource.com/platform/external/ComputeLibrary/+/refs/heads/android14-qpr1-s2-release%5E1..refs/heads/android14-qpr1-s2-release/
so in brief.. do you plan on investigating adding a cross vendor "matrix multiply accelerator units" test at least for Intel and ARM GPUs..
NOTE: vkpeak does the same with equivalent cooperative matrix ext, but there HW support is broader (AMD,NV,Intel, etc..)
thanks..
- Dominant language
- C++
- Stars
- 333
- Forks
- 37
- PR merge metrics
- No merged PRs in 30d
Getting set up
This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from ProjectPhysX/OpenCL-Benchmark
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
ProjectPhysX/OpenCL-Benchmark#38 · 1 comment ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
ProjectPhysX/OpenCL-Benchmark#35 · 1 comment ·
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
ProjectPhysX/OpenCL-Benchmark#32 · 1 reaction ·
-
Difficulty 5/5 Over a week Newbie friendliness 18/100
All issues in ProjectPhysX/OpenCL-Benchmark
Similar issues
-
Feature
Difficulty 1/5 Under an hour Newbie friendliness 65/100
Narezzurri/OpenVPN-Config-Manager#95 ·
Maintainers usually reply within 1 day
-
bot-found bug priority: P1
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
madenvel/KalinkaPlayer#251 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
sqlitebrowser/sqlitebrowser#4208 ·
-
ROSES ROSES - Student Review
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Maintainers usually reply within 1 day
-
area/ysql kind/bug priority/medium
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
yugabyte/yugabyte-db#34552 ·
Maintainers usually reply within 1 day