Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

polars stats: value_counts and mode run as two group-bys per column, one column at a time

オープン
#997 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
48/100
issue の種類
リファクタリング
明瞭さ
おおむね明確
活発さ
活発
技術スタック
python
領域
data, performance

調査の方向性

Start with buckaroo/customizations/pl_stats_v2.py and buckaroo/pluggable_analysis_framework/df_stats_v2.py, then trace how StatPipeline passes each stat a pl.Series. First verify that taking mode from the existing value-counts result preserves behavior, including tied values; then investigate batching the per-column group-bys with pl.collect_all or a single select. Done means the redundant mode group-by is removed and batched results are handed to each stat correctly, with the reported performance improvement checked against equivalent stats.

索引モデルが issue の本文から書いたものです。

説明

performance polars

Problem

With the 50,000-row sample in PlDfStatsV2.get_operating_df (buckaroo/pluggable_analysis_framework/df_stats_v2.py:93) bypassed, PlDfStatsV2 takes 12.4 s on a 10,803,012-row, 43-column parquet (28 String columns). Most of that is two group-bys per column in pl_base_summary_stats (buckaroo/customizations/pl_stats_v2.py:77), run one column at a time:

  • _pl_vc_to_pd calls ser.drop_nulls().value_counts(sort=True) (pl_stats_v2.py:69): 6.8 s over the 43 columns. The sort isn't the cost (sort=False is 6.4 s), and the conversion to a pandas Series is 0.14 s in total. The widest column, 3,059,039 distinct strings, takes 1.6 s on its own.
  • ser.drop_nulls().mode().item(0) (pl_stats_v2.py:86): 2.5 s. It is a second group-by on the same column, and its answer is the first row of the value_counts computed just before it.

StatPipeline hands each stat one pl.Series at a time, so polars never runs these across columns in parallel. The same 43 value_counts collected together with pl.collect_all over one lazy select per column take 2.8 s instead of 6.9 s.

Impact

This is the cost of exact stats on the polars path. On the same file, XorqServerDataflow over deferred_read_parquet computes its stats in 3.4 s. It matters for /load with backend: "polars" once #992 stops sampling the frame: today stats on large frames come from the 50,000-row sample, so distinct_count on a column of unique IDs reads 50,000.

Suggested fix

  • Take mode from the first row of the value_counts that pl_base_summary_stats already computes, instead of calling mode(). Saves about 2.5 s here. Both leave the choice among tied values unspecified today.
  • Compute the per-column group-bys for all columns in one parallel collect (pl.collect_all, or a single select) and hand each stat its column's result. This needs the polars stat pipeline to batch across columns, so it is the larger change.

Context

Found while measuring /load with backend: "polars" for tallyman (#992, #993). Measured on main at 992fdb3 with polars 1.35.2 on an Apple M4 Pro (14 cores), frame already in memory before timing.

主要言語
Python
スター
685
フォーク
17
平均マージ
1日 4時間
マージ済み PR(30日)
29

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

buckaroo-data/buckaroo のほかの issue

buckaroo-data/buckaroo の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。