Hacktoberfest 2026: die Issues, die Maintainer für den Oktober markiert haben – offen und einsteigerfreundlich. Hacktoberfest-Issues durchsuchen

groupby().agg() silently ignores weights, producing incorrect results

Offen
#264 2 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Maintainer antworten meist innerhalb von 1 Tag

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
3/5
Geschätzter Aufwand
1-2 Tage
Anfängerfreundlichkeit
65/100
Issue-Typ
Bug
Klarheit
Klar beschrieben
Aktivitätsstatus
Aktiv
Tech-Stack
pandas, python

Rechercherichtung

The issue is in microdf/microdataframe.py lines 644-680, where the MicroDataFrameGroupBy class does not override .agg(). Start by reading the existing weighted method overrides (like sum). Determine how to parse aggregation specs and apply weights, or implement a clear error/warning. Run the provided test case to verify the fix.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

Problem

MicroDataFrame's groupby operations silently ignore weights when using .agg(), producing incorrect (unweighted) results without any warning or error.

Example
import microdf as mdf
import numpy as np

df = mdf.MicroDataFrame(
    {"group": ["A", "A", "B", "B"], "value": [10, 20, 30, 40]},
    weights=np.array([2, 3, 1, 4])
)

# CORRECT (weighted):
df.groupby("group").value.sum()
# A: 10*2 + 20*3 = 80.0
# B: 30*1 + 40*4 = 190.0

# INCORRECT (unweighted) - no warning!
df.groupby("group").agg({'value': 'sum'})
# A: 10 + 20 = 30  ❌
# B: 30 + 40 = 70  ❌
Impact

This is a critical data correctness issue because:

  1. Silent failure: No error or warning - just wrong numbers
  2. Natural usage pattern: .agg() is a standard pandas idiom for multi-column aggregation
  3. Plausible results: The unweighted numbers look reasonable, making bugs hard to detect
  4. Real-world consequences: Users analyzing survey data (CPS, ACS, etc.) will get wildly incorrect population estimates
Root Cause

Looking at microdf/microdataframe.py:644-680, the MicroDataFrameGroupBy class:

  • ✓ Overrides specific methods like sum(), mean(), etc. to apply weights
  • ✗ Does NOT override .agg() or .aggregate(), so they fall back to pandas' unweighted implementation
Proposed Solutions

Option 1: Override .agg() to apply weights (Best)

  • Implement MicroDataFrameGroupBy.agg() to properly handle weights
  • Parse the aggregation specifications and route to weighted methods

Option 2: Raise an error (Safer than current behavior)

def agg(self, *args, **kwargs):
    raise NotImplementedError(
        "MicroDataFrameGroupBy.agg() does not support weights. "
        "Use df.groupby(col).column.sum() instead."
    )

Option 3: Emit a loud warning

def agg(self, *args, **kwargs):
    warnings.warn(
        "MicroDataFrameGroupBy.agg() ignores weights! Results will be unweighted.",
        UserWarning,
        stacklevel=2
    )
    return super().agg(*args, **kwargs)
Related Issues

This extends #193, which identified similar problems with .groupby()[[cols]].sum() but didn't specifically address .agg().

Additional Test Cases Needed
def test_agg_with_weights():
    """Test that .agg() applies weights correctly or raises an error"""
    df = mdf.MicroDataFrame(
        {"group": ["A", "A", "B"], "value": [10, 20, 30]},
        weights=np.array([2, 3, 4])
    )
    
    # These should either work correctly or raise NotImplementedError
    result = df.groupby("group").agg({'value': 'sum'})
    
    # If implemented, should equal weighted sums
    # A: 10*2 + 20*3 = 80
    # B: 30*4 = 120
    expected = pd.DataFrame({'value': [80.0, 120.0]}, index=['A', 'B'])
    expected.index.name = 'group'
    
    # Should NOT be unweighted sums (30, 30)
    assert not result.equals(pd.DataFrame({'value': [30, 30]}))
Priority

HIGH - This is a data correctness bug that produces silently wrong results in a library designed for weighted survey analysis.

Vorherrschende Sprache
Python
Sterne
16
Forks
10
Ø Merge
5 T. 3 Std.
Gemergte PRs (30 T.)
21

Entwicklungsumgebung

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus PolicyEngine/microdf

Alle Issues in PolicyEngine/microdf

Ähnliche Issues

Weitere Issues zu Python

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.