server: /load with backend=polars pages a random 1M-row sample, shuffled, when the file has more rows
メンテナーはふだん 1 日以内に返信
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 25/100
調査の方向性
Start with buckaroo/server/data_loading_polars.py:40, buckaroo/dataflow/dataflow.py:420, and buckaroo/dataflow/dataflow_extras.py:68; trace how pre_stats_sample affects the dataflow's raw frame. Reproduce with the issue's 1.2M-row example and check the /load response. Done means the grid pages through the full frame in file order, while any stats sampling does not replace that frame.
索引モデルが issue の本文から書いたものです。
説明
Problem
/load with backend: "polars" pages through a random 1M-row sample, in shuffled order, when the file has more than 1M rows. Every row request, sort and search runs against the sample.
PolarsServerSampling.pre_limit = 1_000_000 (buckaroo/server/data_loading_polars.py:40) reads as a cap on the stats input, but CustomizableDataflow.__init__ passes pre_stats_sample(orig_df) in as the dataflow's raw frame (buckaroo/dataflow/dataflow.py:420), and pre_stats_sample calls df.sample(pre_limit) (buckaroo/dataflow/dataflow_extras.py:68). The sample becomes the table.
import polars as pl
from buckaroo.server.data_loading_polars import create_polars_dataflow
flow = create_polars_dataflow(pl.DataFrame({"row": range(1_200_000)}))
_, processed, _ = flow.widget_args_tuple
print(len(processed)) # 1000000
print(processed["row"].head(5).to_list()) # e.g. [297452, 366261, 578517, 947309, 607234]
Over the server, POST /load on a 1.2M-row parquet answers rows: 1200000, the initial_state carries df_meta: {filtered_rows: 1000000, total_rows: 1200000}, and the first infinite_resp has length: 1000000 with the rows out of file order.
Impact
tallyman would like to hand its grid a parquet path through /load instead of a xorq build through /load_expr. Its entries routinely pass 1M rows (the largest is 78,024,179). Each would show a random slice of its rows, out of order, and the summary stats would describe that slice.
Suggested fix
- Keep the frame the grid pages through whole.
pre_limit = FalseonPolarsServerSamplingdoes that, and is whatXorqInfiniteSamplingalready uses. - If stats on very large frames need a cap, sample only what
DFStatsClassreads and mark the stats as sampled, rather than replacingraw_df.
Context
Found while checking whether tallyman's grid can move from /load_expr to /load with backend: "polars". Reproduced on buckaroo 0.15.8; the same code is on main at 992fdb3.
The other gaps from the same check: #993, #994, #995, #996.
- 主要言語
- Python
- スター
- 685
- フォーク
- 17
- 平均マージ
- 1日 4時間
- マージ済み PR(30日)
- 29
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
buckaroo-data/buckaroo のほかの issue
-
Histograms JS
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
buckaroo-data/buckaroo#1089 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Histograms
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
buckaroo-data/buckaroo#1087 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Histograms
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
buckaroo-data/buckaroo#1086 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Histograms
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
buckaroo-data/buckaroo#1083 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
Histograms
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
buckaroo-data/buckaroo#1082 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
buckaroo-data/buckaroo の issue をすべて見る
似ている issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
rpm-software-management/mock#1824 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
jpata/particleflow#520 ·
メンテナーはふだん 1 日以内に返信
-
bug good first issue hacktoberfest
難易度 1/5 1時間未満 初心者へのやさしさ 78/100
gridhead/gi-loadouts#699 ·
メンテナーはふだん 13 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 86/100
FinanceFlash/unvibecode#206 ·
メンテナーはふだん 1 日以内に返信