Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

FP16 simulation execution context + FP32-accumulating LayerNorm lowering rule (variance overflow repro)

Open
#1,281 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Feature
Clarity
Mostly clear
Activity status
Active
Tech stack
kotlin

Research direction

Start by locating Fp16SimulationExecutionContext or the DirectCpuExecutionContext option and the Fp16Codec, then identify the reduction tests and the layerNorm lowering rule in skainet-compile-hlo. Reproduce the overflow on the JVM with a synthetic [1500,1280] input, and check that the FP16 lowering emits an F32 variance reduction with a TraceEvent naming the rule. Done also includes the documentation page and, when an FP16 target exists, passing the IREE parity harness (#1148).

Written by the indexing model from the issue text.

Description

compiler enhancement skill:numerics

Context

An FP16 LayerNorm computes variance as a mean over 1280 squared, centred activations. For real Whisper encoder activations the sum exceeds the FP16 maximum before the division, variance becomes +Inf, rsqrt becomes 0, and the block silently emits zeros for those frames — 6 of 1500 frames in the measured case, enough to drop encoder cosine from 0.9999 to 0.968. The fix on the LiteRT side was a graph rewrite ((0.25·δ)² · 16) because the runtime offered no lever.

SKaiNET will meet the same bug the moment an FP16 GPU or NPU lowering exists. It should be reproducible on a desktop CPU before that, so the lowering rule (accumulate variance in FP32, or apply the rescale) can be written and regression-tested against a known-failing input.

Scope

  • Fp16SimulationExecutionContext (or a DirectCpuExecutionContext option): after every op, round the output through Fp16Codec (round-to-nearest-even, overflow to ±Inf, gradual underflow) — storage-only simulation. A second mode also rounds accumulators inside reductions (sum/mean/variance/matmul) to model true half-precision arithmetic.
  • A regression test with a synthetic [1500,1280] input scaled so that naive FP16 variance overflows, asserting +Inf under accumulate-in-half and a finite, correct value under the FP32-accumulate lowering.
  • layerNorm lowering rule for FP16 targets in skainet-compile-hlo: variance reduction in F32 (or the documented rescale), with a TraceEvent naming which rule fired.
  • Docs: an explanation page on narrow-float numerics pitfalls with this as the worked example.

Acceptance

  • The overflow reproduces in a unit test on a JVM with no GPU.
  • The StableHLO emitted for an FP16 LayerNorm shows the F32 reduction, and the test passes through the IREE parity harness (#1148) when an FP16 target exists.

Related

  • #884, #885 — narrow-float layer and FP16 kernels
  • #1148 — parity acceptance
Dominant language
Kotlin
Stars
52
Forks
15
Avg merge
1d 15h
Merged PRs (30d)
36

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from SKaiNET-developers/SKaiNET

All issues in SKaiNET-developers/SKaiNET

Similar issues

More Kotlin issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.