Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

server: /load with backend=polars pages a random 1M-row sample, shuffled, when the file has more rows

Abierto
#992 1 comentario 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

  • #1000 de @paddymul — cerrado sin fusionar
  • #1007 de @paddymul — cerrado sin fusionar

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
25/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Estancado
Stack tecnológico
python
Área
backend, data

Línea de trabajo

Start with buckaroo/server/data_loading_polars.py:40, buckaroo/dataflow/dataflow.py:420, and buckaroo/dataflow/dataflow_extras.py:68; trace how pre_stats_sample affects the dataflow's raw frame. Reproduce with the issue's 1.2M-row example and check the /load response. Done means the grid pages through the full frame in file order, while any stats sampling does not replace that frame.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

bug polars

Problem

/load with backend: "polars" pages through a random 1M-row sample, in shuffled order, when the file has more than 1M rows. Every row request, sort and search runs against the sample.

PolarsServerSampling.pre_limit = 1_000_000 (buckaroo/server/data_loading_polars.py:40) reads as a cap on the stats input, but CustomizableDataflow.__init__ passes pre_stats_sample(orig_df) in as the dataflow's raw frame (buckaroo/dataflow/dataflow.py:420), and pre_stats_sample calls df.sample(pre_limit) (buckaroo/dataflow/dataflow_extras.py:68). The sample becomes the table.

import polars as pl
from buckaroo.server.data_loading_polars import create_polars_dataflow

flow = create_polars_dataflow(pl.DataFrame({"row": range(1_200_000)}))
_, processed, _ = flow.widget_args_tuple
print(len(processed))                      # 1000000
print(processed["row"].head(5).to_list())  # e.g. [297452, 366261, 578517, 947309, 607234]

Over the server, POST /load on a 1.2M-row parquet answers rows: 1200000, the initial_state carries df_meta: {filtered_rows: 1000000, total_rows: 1200000}, and the first infinite_resp has length: 1000000 with the rows out of file order.

Impact

tallyman would like to hand its grid a parquet path through /load instead of a xorq build through /load_expr. Its entries routinely pass 1M rows (the largest is 78,024,179). Each would show a random slice of its rows, out of order, and the summary stats would describe that slice.

Suggested fix

  • Keep the frame the grid pages through whole. pre_limit = False on PolarsServerSampling does that, and is what XorqInfiniteSampling already uses.
  • If stats on very large frames need a cap, sample only what DFStatsClass reads and mark the stats as sampled, rather than replacing raw_df.

Context

Found while checking whether tallyman's grid can move from /load_expr to /load with backend: "polars". Reproduced on buckaroo 0.15.8; the same code is on main at 992fdb3.

The other gaps from the same check: #993, #994, #995, #996.

Lenguaje dominante
Python
Estrellas
685
Forks
17
Merge medio
1 d 4 h
PR fusionados (30 d)
29

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de buckaroo-data/buckaroo

Todos los issues de buckaroo-data/buckaroo

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.