Fully lazy imperative function evaluation
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 25/100
Línea de trabajo
Lee la implementación enlazada de cumsum/cumprod con axis=None y sigue la maquinaria de cálculo de lazyexpr, LazyUDFs y el manejo de chunks de matmul. Consulta issue #441 para conocer el contexto de la indexación avanzada. La tarea está terminada cuando las funciones evalúan perezosamente los chunks de salida solicitados, incluidos los expresiones anidadas de matmul/sum, sin temporales del resultado completo, manteniendo la preferencia por numexpr y evitando repetir el trabajo de reducción.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
In order to make function evaluation properly lazy, all functions (e.g. matmul) should be implemented as LazyUDFs.
Secondly, the lazyexpr machinery of compute should loop over chunks of the result. Each function must then decide what slices of the operands are necessary to form the corresponding result chunk (as matmul currently does internally). Then, upon evaluation, although the expression evaluates term-by-term, it does not compute the full result for e..g matmul before proceeding to the next term, but only for the necessary chunk(s) of output. Thus there is a higher chance of cache hits (not the case currently for eager execution of linalg and reductions).
Something in this spirit has been implemented for cumsum/cumprod when axis=None (see this code ):
# Special case for cumulative operations with axis = None
if reduce_args["axis"] is None and reduce_op in {ReduceOp.CUMULATIVE_PROD, ReduceOp.CUMULATIVE_SUM}:
# res_out_ is just None, out set to all 0s (sum) or 1s (prod)
out, res_out_ = convert_none_out(dtype, reduce_op, reduced_shape)
# reduced_shape is just one-element tuple
chunklen = out.chunks[0] if hasattr(out, "chunks") else chunks[-1]
carry = 0
for cidx in range(0, reduced_shape[0] // chunklen):
slice_starts = np.unravel_index(cidx * chunklen, shape)
slice_stops = np.unravel_index((cidx + 1) * chunklen, shape)
cslice = tuple(
slice(start, stop) for start, stop in zip(slice_starts, slice_stops, strict=True)
)
_get_chunk_operands(operands, cslice, chunk_operands, shape)
result, _ = _get_result(expression, chunk_operands, ne_args, where)
result = np.require(result, requirements="C")
if reduce_op == ReduceOp.CUMULATIVE_SUM:
res = np.cumulative_sum(result, axis=None) + carry
else:
res = np.cumulative_prod(result, axis=None) * carry
carry = res[-1]
out[cidx * chunklen + include_initial : (cidx + 1) * chunklen + include_initial] = res
It should also be implemented for fancy-indexing with slice (see https://github.com/Blosc/python-blosc2/issues/441).
Should be possible to handle even something like "matmul(sum(a,axis=1), b)" by passing the desired slice from matmul->sum, which treats the asked-for-slice as a desired output slice and handles accordingly.
This would avoid large in-memory temporaries for example when calculating from eagerly executed linear algebra functions e.g. in "matmul(a, b) + b".
Problems:
1 - reductions could be a problem since for example "sum(a) + a" for the chunks of output would recalculate the scalar "sum(a)" for each chunk of output. This could be avoided by using some kind of result cache for reductions.
2 - naturally, numexpr would always be faster since it compiles the expression into bytecode. Thus we should make sure that, when possible, we still use numexpr preferentially (for most elementwise funcs).
- Lenguaje dominante
- Python
- Estrellas
- 212
- Forks
- 66
- Merge medio
- 1 d 17 h
- PR fusionados (30 d)
- 11
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Blosc/python-blosc2
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
Blosc/python-blosc2#731 ·
Los mantenedores suelen responder en 1 día
-
documentation sustain-2026
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Blosc/python-blosc2#720 · 3 comentarios ·
Los mantenedores suelen responder en 1 día
-
sustain-2026
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Blosc/python-blosc2#717 ·
Los mantenedores suelen responder en 1 día
-
documentation sustain-2026
Dificultad 2/5 1-3 horas Aptitud para principiantes 80/100
Blosc/python-blosc2#714 ·
Los mantenedores suelen responder en 1 día
-
documentation sustain-2026
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
Blosc/python-blosc2#710 ·
Los mantenedores suelen responder en 1 día
Todos los issues de Blosc/python-blosc2
Issues similares
-
json_params_matcher fails on falsy top-level JSON primitives (0, False, "")Posiblemente ocupada @mayureshsonawane17 la tomó hoy. AbiertoWaiting for: Product Owner
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 5 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Los mantenedores suelen responder en 1 día
-
第二章思考题 8:Skill 追加到末尾不必每轮重新计算 KVAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
bojieli/ai-agent-book#1169 ·
Los mantenedores suelen responder en 1 día
-
priority:low ready-for-dev
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
OpenHands/extensions#738 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
micronaut-projects/micronaut-core#13677 ·
Los mantenedores suelen responder en 1 día