Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

test(ad): GPU/packed-field AD parity for the exchange rules

Abierto
#43 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
48/100
Tipo de issue
Nueva funcionalidad
Claridad
Bastante claro
Estado de actividad
Tranquilo
Stack tecnológico
julia
Área
backend, testing

Línea de trabajo

Start with test/enzyme_rules.jl, test/device.jl, and test/device_gpu.jl, then inspect the packed-field exchange entry points in src/packedfield.jl and src/transfer_kernels.jl. Add CPU packed-field AD coverage and GPU coverage under MFO_TEST_GPU, checking the declared gradient and ENZ_EXT.rule_hits(). Document whether GPU reverse mode is supported or refused.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

enhancement

ext/MatrixFreeOperatorsEnzymeCoreExt.jl registers rules on _exchange_storage!/_bc_storage!, which are layout-agnostic: they take raw storage plus a BlockLayout tag, so PackedBlockField goes through the same rules as BlockField. That is the mechanism DESIGN.md §6 calls mandatory for @kernel-authored leaves — the packed GPU exchange must never be differentiated through, and now it structurally cannot be.

None of that is tested on a GPU. Specifically:

  • No AD test uses a packed field at all, on any backend. src/packedfield.jl:17-18 still designates BlockField as "the reference layout for AD".
  • test/device.jl / test/device_gpu.jl never differentiate, so MFO_TEST_GPU=true adds no AD coverage.
  • The device exchange path (_run_exchange!(::PackedBlockField, ...) in src/transfer_kernels.jl:194-198) dispatches to batched kernels below the rule seam only when the backend is a real GPU. On CPU it falls back to _run_exchange_host!, which is what the current tests exercise — so the rule is verified against the fallback, not against the kernels.

What to add

  1. CPU-side first — the quick, self-contained first deliverable. It needs no hardware and can land as its own PR ahead of steps 2-3: a packed-field AD test asserting the gradient matches the declared adjoint and that the rules fire (ENZ_EXT.rule_hits() counters increase across the call), mirroring test/enzyme_rules.jl.
  2. Under MFO_TEST_GPU, the same against CUDA arrays — the real question being whether the rule's reverse body (_exchange_storage_adjoint!, which stays on host descriptor loops on all backends by design, see transfer_kernels.jl:16-23) composes correctly with a forward pass that ran as batched kernels.
  3. Decide and document whether GPU reverse mode is supported or explicitly refused. Right now it is neither — it is untested.

—
🤖 Claude fixed a line number that had wandered off.
claude-opus-5-5[1m] (high) · claude-code 2.1.285 · unreviewed

Lenguaje dominante
Julia
Estrellas
3
Forks
1
Merge medio
23 h 5 min
PR fusionados (30 d)
16

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de RallypointOne/MatrixFreeOperators.jl

Todos los issues de RallypointOne/MatrixFreeOperators.jl

Issues similares

Más issues de Julia

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.