Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

server: /load with backend=polars pages a random 1M-row sample, shuffled, when the file has more rows

Aperta
#992 1 commento 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 1 giorno

Nessuno ha ancora preso questa issue.

  • #1000 di @paddymul — chiusa senza merge
  • #1007 di @paddymul — chiusa senza merge

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
25/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Ferma
Stack tecnologico
python
Ambito
backend, data

Direzione di ricerca

Start with buckaroo/server/data_loading_polars.py:40, buckaroo/dataflow/dataflow.py:420, and buckaroo/dataflow/dataflow_extras.py:68; trace how pre_stats_sample affects the dataflow's raw frame. Reproduce with the issue's 1.2M-row example and check the /load response. Done means the grid pages through the full frame in file order, while any stats sampling does not replace that frame.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

bug polars

Problem

/load with backend: "polars" pages through a random 1M-row sample, in shuffled order, when the file has more than 1M rows. Every row request, sort and search runs against the sample.

PolarsServerSampling.pre_limit = 1_000_000 (buckaroo/server/data_loading_polars.py:40) reads as a cap on the stats input, but CustomizableDataflow.__init__ passes pre_stats_sample(orig_df) in as the dataflow's raw frame (buckaroo/dataflow/dataflow.py:420), and pre_stats_sample calls df.sample(pre_limit) (buckaroo/dataflow/dataflow_extras.py:68). The sample becomes the table.

import polars as pl
from buckaroo.server.data_loading_polars import create_polars_dataflow

flow = create_polars_dataflow(pl.DataFrame({"row": range(1_200_000)}))
_, processed, _ = flow.widget_args_tuple
print(len(processed))                      # 1000000
print(processed["row"].head(5).to_list())  # e.g. [297452, 366261, 578517, 947309, 607234]

Over the server, POST /load on a 1.2M-row parquet answers rows: 1200000, the initial_state carries df_meta: {filtered_rows: 1000000, total_rows: 1200000}, and the first infinite_resp has length: 1000000 with the rows out of file order.

Impact

tallyman would like to hand its grid a parquet path through /load instead of a xorq build through /load_expr. Its entries routinely pass 1M rows (the largest is 78,024,179). Each would show a random slice of its rows, out of order, and the summary stats would describe that slice.

Suggested fix

  • Keep the frame the grid pages through whole. pre_limit = False on PolarsServerSampling does that, and is what XorqInfiniteSampling already uses.
  • If stats on very large frames need a cap, sample only what DFStatsClass reads and mark the stats as sampled, rather than replacing raw_df.

Context

Found while checking whether tallyman's grid can move from /load_expr to /load with backend: "polars". Reproduced on buckaroo 0.15.8; the same code is on main at 992fdb3.

The other gaps from the same check: #993, #994, #995, #996.

Lingua principale
Python
Stelle
685
Fork
17
Merge medio
1g 4h
PR unite (30g)
29

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di buckaroo-data/buckaroo

Tutte le issue di buckaroo-data/buckaroo

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.