Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Fail closed: aggregation and construction paths that silently return unweighted results

Open
#333 0 comments 0 reactions 1 assignee View on GitHub

Maintainers usually reply within 1 day

@juaristi22 is already working on this.

Since Sep 22, 2026.

Assessment

This issue has not been assessed yet.

Description

bug

Several ordinary pandas spellings return an unweighted number, or a Micro object whose weights were reset to 1, with no warning. All 852 tests pass, so none of these paths is covered. Reproduced on main at c477bae (microdf 1.5.9) with pandas 3.0.6; the review that found them also reproduced them on pandas 2.3.3.

import numpy as np, pandas as pd, microdf as mdf
df = mdf.MicroDataFrame({"x": [10., 100.], "g": ["a", "a"]}, weights=[9, 1])
# weighted mean is 19; unweighted is 55

df.groupby("g").x.mean()                                          # 19.0  correct
df.groupby("g").agg({"x": "mean"})                                # 55.0  no warning
df.groupby("g").agg(m=("x", "mean"))                              # 55.0
df.groupby("g").x.agg(lambda z: z.mean())                         # 55.0
df.pivot_table(index="g", values="x", aggfunc=lambda z: z.mean()) # 55.0
df.apply(lambda r: r.x, axis=1).mean()                            # 55.0, plain Series
np.average(df.x); np.median(df.x)                                 # 55.0
df.x.value_counts()                                               # {10: 1, 100: 1}
df.x.rolling(2).mean()                                            # [nan, 55.0]
df.mean(numeric_only=True)                                        # empty Series
df.groupby("g").mean(numeric_only=True)                           # empty DataFrame
df.median(numeric_only=True)                                      # empty

mdf.MicroDataFrame(df).weights                                    # [1.0, 1.0]  reset, still looks weighted
pd.cut(df.x, 2).weights                                           # [1.0, 1.0]

a = mdf.MicroSeries([10., 100.], weights=[9, 1]); b = mdf.MicroSeries([1., 2.], weights=[1, 9])
(a + b).mean(), (b + a).mean()                                    # 20.1, 92.9  left operand's weights win
pd.concat([pd.DataFrame({"x": [1.]}), df])                        # plain DataFrame (weighted-first raises)

Also inherited and unweighted: sem, skew, kurt, prod, mode, idxmax/idxmin. Separately, df.dropna() raises Cannot transpose row weights onto columns on any all-numeric frame (it works when a non-numeric column is present).

#264 reported the agg case in November 2025 and was closed; its repro still fails.

What "fixed" means

For each path above, either return the weighted result or raise. Nothing may return an unweighted number, or a Micro object with silently reset weights, from a weighted input.

  • groupby(...).agg with dict, named and callable forms; pivot_table with a callable; SeriesGroupBy.agg(callable): weighted, routed through the same code as the named reductions.
  • Row-wise DataFrame.apply(axis=1): return a MicroSeries with the frame's weights, or raise.
  • numeric_only=True on frame and groupby reductions: weighted result over the numeric columns.
  • MicroDataFrame(mdf) / MicroSeries(ms) without weights=: inherit the source's weights.
  • pd.cut, pd.qcut, pd.to_numeric, explode, rolling/expanding/ewm on a MicroSeries: carry weights or raise. Silently resetting to 1 is the worst of the three outcomes.
  • Arithmetic between Micro objects with different weights: raise unless the weights are equal. Add MicroSeries.mean etc. to __array_function__ so np.average and np.median dispatch, or raise.
  • value_counts, mode, sem, skew, kurt, prod, idxmax, idxmin: weighted where the definition is standard, otherwise raise with a message naming the plain-pandas escape (pd.Series(s)).
  • dropna() on all-numeric frames: works.
  • Every case above becomes a regression test with the numerical weighted answer asserted, run on both pandas 2 and 3.
  • docs/ gains a support matrix generated from the tests, and the README stops implying weights survive every operation.

Credit: @baogorek filed #264 and #265, which first described this class of failure.

Dominant language
Python
Stars
16
Forks
10
Avg merge
13h 43m
Merged PRs (30d)
15

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from PolicyEngine/microdf

All issues in PolicyEngine/microdf

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.