Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Fail closed: aggregation and construction paths that silently return unweighted results

Đang mở
#333 0 bình luận 0 reaction 1 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@juaristi22 đang làm issue này rồi.

Từ ngày 22/9/2026.

Đánh giá

Issue này chưa được đánh giá.

Mô tả

bug

Several ordinary pandas spellings return an unweighted number, or a Micro object whose weights were reset to 1, with no warning. All 852 tests pass, so none of these paths is covered. Reproduced on main at c477bae (microdf 1.5.9) with pandas 3.0.6; the review that found them also reproduced them on pandas 2.3.3.

import numpy as np, pandas as pd, microdf as mdf
df = mdf.MicroDataFrame({"x": [10., 100.], "g": ["a", "a"]}, weights=[9, 1])
# weighted mean is 19; unweighted is 55

df.groupby("g").x.mean()                                          # 19.0  correct
df.groupby("g").agg({"x": "mean"})                                # 55.0  no warning
df.groupby("g").agg(m=("x", "mean"))                              # 55.0
df.groupby("g").x.agg(lambda z: z.mean())                         # 55.0
df.pivot_table(index="g", values="x", aggfunc=lambda z: z.mean()) # 55.0
df.apply(lambda r: r.x, axis=1).mean()                            # 55.0, plain Series
np.average(df.x); np.median(df.x)                                 # 55.0
df.x.value_counts()                                               # {10: 1, 100: 1}
df.x.rolling(2).mean()                                            # [nan, 55.0]
df.mean(numeric_only=True)                                        # empty Series
df.groupby("g").mean(numeric_only=True)                           # empty DataFrame
df.median(numeric_only=True)                                      # empty

mdf.MicroDataFrame(df).weights                                    # [1.0, 1.0]  reset, still looks weighted
pd.cut(df.x, 2).weights                                           # [1.0, 1.0]

a = mdf.MicroSeries([10., 100.], weights=[9, 1]); b = mdf.MicroSeries([1., 2.], weights=[1, 9])
(a + b).mean(), (b + a).mean()                                    # 20.1, 92.9  left operand's weights win
pd.concat([pd.DataFrame({"x": [1.]}), df])                        # plain DataFrame (weighted-first raises)

Also inherited and unweighted: sem, skew, kurt, prod, mode, idxmax/idxmin. Separately, df.dropna() raises Cannot transpose row weights onto columns on any all-numeric frame (it works when a non-numeric column is present).

#264 reported the agg case in November 2025 and was closed; its repro still fails.

What "fixed" means

For each path above, either return the weighted result or raise. Nothing may return an unweighted number, or a Micro object with silently reset weights, from a weighted input.

  • groupby(...).agg with dict, named and callable forms; pivot_table with a callable; SeriesGroupBy.agg(callable): weighted, routed through the same code as the named reductions.
  • Row-wise DataFrame.apply(axis=1): return a MicroSeries with the frame's weights, or raise.
  • numeric_only=True on frame and groupby reductions: weighted result over the numeric columns.
  • MicroDataFrame(mdf) / MicroSeries(ms) without weights=: inherit the source's weights.
  • pd.cut, pd.qcut, pd.to_numeric, explode, rolling/expanding/ewm on a MicroSeries: carry weights or raise. Silently resetting to 1 is the worst of the three outcomes.
  • Arithmetic between Micro objects with different weights: raise unless the weights are equal. Add MicroSeries.mean etc. to __array_function__ so np.average and np.median dispatch, or raise.
  • value_counts, mode, sem, skew, kurt, prod, idxmax, idxmin: weighted where the definition is standard, otherwise raise with a message naming the plain-pandas escape (pd.Series(s)).
  • dropna() on all-numeric frames: works.
  • Every case above becomes a regression test with the numerical weighted answer asserted, run on both pandas 2 and 3.
  • docs/ gains a support matrix generated from the tests, and the README stops implying weights survive every operation.

Credit: @baogorek filed #264 and #265, which first described this class of failure.

Ngôn ngữ chính
Python
Star
16
Fork
10
Merge trung bình
13 giờ 43 phút
Pull request đã merge (30 ngày)
15

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của PolicyEngine/microdf

Tất cả issue của PolicyEngine/microdf

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.