Wrong results for overflowing Int16 arithmetic in a reduction kernel
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 45/100
Hướng nghiên cứu
The report points to AcceleratedKernels’ generic mapreducedim! kernel for sources without strides and the reductions/mapreducedim! test in GPUArrays 12. Start by reproducing the failing Complex{Int16} abs2 reduction and comparing the generic kernel’s generated code with the handwritten oneAPI kernels described here. Done means identifying and fixing the cause of the incorrect results, with a regression test covering overflowing values.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
On an Iris Xe (i7-1165G7), a reduction of abs2 over Complex{Int16} values gives wrong results when abs2 overflows Int16, in AcceleratedKernels' generic mapreducedim! kernel (the one used for sources without strides). Julia's IR is correct (wrapping mul i16/add i16 feeding llvm.smax.i16), so this looks like a miscompile further down, in the SPIR-V translation or IGC.
GPUArrays 12's testsuite hit it (reductions/mapreducedim!, Base.mapreducedim!(abs2, max, R, adjoint(A)) with Complex{Int16}). GPUArrays 12.0.1 keeps those values small.
using oneAPI
function check(label, gen, f)
bad = 0
for _ in 1:50
A = gen()
R0 = zeros(Int16, 1, 3)
mk = A -> view(permutedims(A), [1, 2, 3, 4], :) # not strided: AK's generic kernel
c = Base.mapreducedim!(f, max, copy(R0), mk(A))
g = Array(Base.mapreducedim!(f, max, oneArray(R0), mk(oneArray(A))))
c == g || (bad += 1)
end
println(rpad(label, 30), bad, "/50")
end
function main()
small = () -> complex.(rand(Int16(-100):Int16(100), 3, 4), rand(Int16(-100):Int16(100), 3, 4))
full = () -> rand(Complex{Int16}, 3, 4)
check("small abs2", small, abs2)
check("full abs2", full, abs2)
check("full re*re", full, x -> real(x) * real(x))
check("full re*re+im*im", full, x -> real(x) * real(x) + imag(x) * imag(x))
check("full re*im", full, x -> real(x) * imag(x))
end
main()
small abs2 0/50
full abs2 50/50
full re*re 0/50
full re*re+im*im 49/50
full re*im 0/50
The kernel's loop body, from @device_code_llvm:
%.unpack = load i16, ptr addrspace(1) %16, align 2
%.unpack73 = load i16, ptr addrspace(1) %.elt72, align 2
%17 = mul i16 %.unpack, %.unpack
%18 = mul i16 %.unpack73, %.unpack73
%19 = add i16 %18, %17
%20 = call i16 @llvm.smax.i16(i16 %19, i16 %value_phi24)
Hand-written @oneapi kernels with the same arithmetic (max over abs2 of Complex{Int16} loaded from a 2D array, with the seed passed as an argument or as a constant) compute the right result, so I couldn't reduce it below AK's kernel yet.
oneAPI 2.9.2 and the GPUArrays 12 port (#665), AcceleratedKernels 0.5.1, Julia 1.12.7 and 1.13.
- Ngôn ngữ chính
- Julia
- Star
- 217
- Fork
- 37
- Merge trung bình
- 18 giờ 12 phút
- Pull request đã merge (30 ngày)
- 26
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của JuliaGPU/oneAPI.jl
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 35/100
JuliaGPU/oneAPI.jl#663 · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 45/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
JuliaGPU/oneAPI.jl#607 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của JuliaGPU/oneAPI.jl
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
oxfordcontrol/COSMO.jl#211 ·
-
documentation
Độ khó 2/5 Nửa ngày Mức phù hợp với người mới 65/100
Maintainer thường phản hồi trong vòng 6 ngày
-
Out-of-place JLArray/GPU problem with VectorContinuousCallback scalar-indexes (callback cache built with CPU zeros)Có thể đã có người làm @ChrisRackauckas-Claude đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
SciML/OrdinaryDiffEq.jl#4813 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
ARKODE: callbacks that modify `u` throw MethodError on reinitCó thể đã có người làm @devmotion đã nhận 1 ngày trước. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 79/100
SciML/Sundials.jl#575 ·
-
broken links in docsĐang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày