Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Fail closed: aggregation and construction paths that silently return unweighted results

オープン
#333 コメント 0 件 リアクション 0 件 担当者 1 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

@juaristi22 がすでに取り組んでいます。

2026年9月22日 から。

評価

この issue はまだ評価されていません。

説明

bug

Several ordinary pandas spellings return an unweighted number, or a Micro object whose weights were reset to 1, with no warning. All 852 tests pass, so none of these paths is covered. Reproduced on main at c477bae (microdf 1.5.9) with pandas 3.0.6; the review that found them also reproduced them on pandas 2.3.3.

import numpy as np, pandas as pd, microdf as mdf
df = mdf.MicroDataFrame({"x": [10., 100.], "g": ["a", "a"]}, weights=[9, 1])
# weighted mean is 19; unweighted is 55

df.groupby("g").x.mean()                                          # 19.0  correct
df.groupby("g").agg({"x": "mean"})                                # 55.0  no warning
df.groupby("g").agg(m=("x", "mean"))                              # 55.0
df.groupby("g").x.agg(lambda z: z.mean())                         # 55.0
df.pivot_table(index="g", values="x", aggfunc=lambda z: z.mean()) # 55.0
df.apply(lambda r: r.x, axis=1).mean()                            # 55.0, plain Series
np.average(df.x); np.median(df.x)                                 # 55.0
df.x.value_counts()                                               # {10: 1, 100: 1}
df.x.rolling(2).mean()                                            # [nan, 55.0]
df.mean(numeric_only=True)                                        # empty Series
df.groupby("g").mean(numeric_only=True)                           # empty DataFrame
df.median(numeric_only=True)                                      # empty

mdf.MicroDataFrame(df).weights                                    # [1.0, 1.0]  reset, still looks weighted
pd.cut(df.x, 2).weights                                           # [1.0, 1.0]

a = mdf.MicroSeries([10., 100.], weights=[9, 1]); b = mdf.MicroSeries([1., 2.], weights=[1, 9])
(a + b).mean(), (b + a).mean()                                    # 20.1, 92.9  left operand's weights win
pd.concat([pd.DataFrame({"x": [1.]}), df])                        # plain DataFrame (weighted-first raises)

Also inherited and unweighted: sem, skew, kurt, prod, mode, idxmax/idxmin. Separately, df.dropna() raises Cannot transpose row weights onto columns on any all-numeric frame (it works when a non-numeric column is present).

#264 reported the agg case in November 2025 and was closed; its repro still fails.

What "fixed" means

For each path above, either return the weighted result or raise. Nothing may return an unweighted number, or a Micro object with silently reset weights, from a weighted input.

  • groupby(...).agg with dict, named and callable forms; pivot_table with a callable; SeriesGroupBy.agg(callable): weighted, routed through the same code as the named reductions.
  • Row-wise DataFrame.apply(axis=1): return a MicroSeries with the frame's weights, or raise.
  • numeric_only=True on frame and groupby reductions: weighted result over the numeric columns.
  • MicroDataFrame(mdf) / MicroSeries(ms) without weights=: inherit the source's weights.
  • pd.cut, pd.qcut, pd.to_numeric, explode, rolling/expanding/ewm on a MicroSeries: carry weights or raise. Silently resetting to 1 is the worst of the three outcomes.
  • Arithmetic between Micro objects with different weights: raise unless the weights are equal. Add MicroSeries.mean etc. to __array_function__ so np.average and np.median dispatch, or raise.
  • value_counts, mode, sem, skew, kurt, prod, idxmax, idxmin: weighted where the definition is standard, otherwise raise with a message naming the plain-pandas escape (pd.Series(s)).
  • dropna() on all-numeric frames: works.
  • Every case above becomes a regression test with the numerical weighted answer asserted, run on both pandas 2 and 3.
  • docs/ gains a support matrix generated from the tests, and the README stops implying weights survive every operation.

Credit: @baogorek filed #264 and #265, which first described this class of failure.

主要言語
Python
スター
16
フォーク
10
平均マージ
5日 3時間
マージ済み PR(30日)
21

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

PolicyEngine/microdf のほかの issue

PolicyEngine/microdf の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。