Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

perf(amr): batch the coarse-fine flux rewrite into one device launch for packed Diffusion

Abierto
#78 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
52/100
Tipo de issue
Refactorización
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
julia

Línea de trabajo

Start with _cf_flux_rewrite! and _cf_flux_rewrite_adjoint!, tracing their host loops over cfflux descriptors and the adjoint zero_ghosts! path. Use test/forest_diffusion.jl's KA-CPU direct-launch tests first, then run the CUDA leg with MFO_TEST_GPU=true. Done means bit-parity with the host loop and no change to the exchange count.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

performance

Follow-up to #59 / #76.

The packed Diffusion sweep on a GPU backend is one KernelAbstractions launch over (blocksize..., nleaves) for the stencil, but the conservative coarse–fine seam (_cf_flux_rewrite! / _cf_flux_rewrite_adjoint!) still runs as a host loop over the cfflux descriptors — a few small broadcast launches on device views per coarse–fine face, ahead of (forward) or after (adjoint) the stencil launch. The launch count therefore still scales with the number of coarse–fine faces, which only partly meets #59's "leaf-count-independent launch structure" acceptance item. The adjoint's per-leaf zero_ghosts! before the gather kernel is the same pattern (pre-existing on the other adjoint kernels).

Proposed: a device-resident twin of the cfflux table (SoA of leaf/face/offset descriptors) and one @kernel per direction that applies every face rewrite in a single launch, keeping the per-face body identical to the host loop so the numerical definition does not fork. Same for a batched ghost-zeroing kernel on the adjoint path. Acceptance: bit-parity with the host loop on the test/forest_diffusion.jl KA-CPU direct-launch tests and on the CUDA leg (MFO_TEST_GPU=true), and no change to the exchange count.

🤖 Beep boop — filed by Claude, not Kyle leaving himself homework.

Lenguaje dominante
Julia
Estrellas
3
Forks
1
Merge medio
23 h 5 min
PR fusionados (30 d)
16

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de RallypointOne/MatrixFreeOperators.jl

Todos los issues de RallypointOne/MatrixFreeOperators.jl

Issues similares

Más issues de Julia

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.