Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

server: /load with backend=polars pages a random 1M-row sample, shuffled, when the file has more rows

オープン
#992 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

  • #1000 @paddymul による — マージされずにクローズ
  • #1007 @paddymul による — マージされずにクローズ

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
25/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
停滞
技術スタック
python
領域
backend, data

調査の方向性

Start with buckaroo/server/data_loading_polars.py:40, buckaroo/dataflow/dataflow.py:420, and buckaroo/dataflow/dataflow_extras.py:68; trace how pre_stats_sample affects the dataflow's raw frame. Reproduce with the issue's 1.2M-row example and check the /load response. Done means the grid pages through the full frame in file order, while any stats sampling does not replace that frame.

索引モデルが issue の本文から書いたものです。

説明

bug polars

Problem

/load with backend: "polars" pages through a random 1M-row sample, in shuffled order, when the file has more than 1M rows. Every row request, sort and search runs against the sample.

PolarsServerSampling.pre_limit = 1_000_000 (buckaroo/server/data_loading_polars.py:40) reads as a cap on the stats input, but CustomizableDataflow.__init__ passes pre_stats_sample(orig_df) in as the dataflow's raw frame (buckaroo/dataflow/dataflow.py:420), and pre_stats_sample calls df.sample(pre_limit) (buckaroo/dataflow/dataflow_extras.py:68). The sample becomes the table.

import polars as pl
from buckaroo.server.data_loading_polars import create_polars_dataflow

flow = create_polars_dataflow(pl.DataFrame({"row": range(1_200_000)}))
_, processed, _ = flow.widget_args_tuple
print(len(processed))                      # 1000000
print(processed["row"].head(5).to_list())  # e.g. [297452, 366261, 578517, 947309, 607234]

Over the server, POST /load on a 1.2M-row parquet answers rows: 1200000, the initial_state carries df_meta: {filtered_rows: 1000000, total_rows: 1200000}, and the first infinite_resp has length: 1000000 with the rows out of file order.

Impact

tallyman would like to hand its grid a parquet path through /load instead of a xorq build through /load_expr. Its entries routinely pass 1M rows (the largest is 78,024,179). Each would show a random slice of its rows, out of order, and the summary stats would describe that slice.

Suggested fix

  • Keep the frame the grid pages through whole. pre_limit = False on PolarsServerSampling does that, and is what XorqInfiniteSampling already uses.
  • If stats on very large frames need a cap, sample only what DFStatsClass reads and mark the stats as sampled, rather than replacing raw_df.

Context

Found while checking whether tallyman's grid can move from /load_expr to /load with backend: "polars". Reproduced on buckaroo 0.15.8; the same code is on main at 992fdb3.

The other gaps from the same check: #993, #994, #995, #996.

主要言語
Python
スター
685
フォーク
17
平均マージ
1日 4時間
マージ済み PR(30日)
29

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

buckaroo-data/buckaroo のほかの issue

buckaroo-data/buckaroo の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。