Hacktoberfest 2026: as issues que os mantenedores marcaram para outubro, abertas e boas para iniciantes. Ver issues do Hacktoberfest

groupby().agg() silently ignores weights, producing incorrect results

Aberta
#264 2 comentários 0 reações 0 responsáveis Ver no GitHub

Mantenedores costumam responder em até 1 dia

Ninguém assumiu esta issue ainda.

Avaliação

Dificuldade
3/5
Tempo estimado
1-2 dias
Facilidade para iniciantes
65/100
Tipo de issue
Bug
Clareza
Claramente especificada
Status de atividade
Ativa
Stack de tecnologia
pandas, python

Direção de pesquisa

The issue is in microdf/microdataframe.py lines 644-680, where the MicroDataFrameGroupBy class does not override .agg(). Start by reading the existing weighted method overrides (like sum). Determine how to parse aggregation specs and apply weights, or implement a clear error/warning. Run the provided test case to verify the fix.

Escrita pelo modelo de indexação a partir do texto da issue.

Descrição

Problem

MicroDataFrame's groupby operations silently ignore weights when using .agg(), producing incorrect (unweighted) results without any warning or error.

Example
import microdf as mdf
import numpy as np

df = mdf.MicroDataFrame(
    {"group": ["A", "A", "B", "B"], "value": [10, 20, 30, 40]},
    weights=np.array([2, 3, 1, 4])
)

# CORRECT (weighted):
df.groupby("group").value.sum()
# A: 10*2 + 20*3 = 80.0
# B: 30*1 + 40*4 = 190.0

# INCORRECT (unweighted) - no warning!
df.groupby("group").agg({'value': 'sum'})
# A: 10 + 20 = 30  ❌
# B: 30 + 40 = 70  ❌
Impact

This is a critical data correctness issue because:

  1. Silent failure: No error or warning - just wrong numbers
  2. Natural usage pattern: .agg() is a standard pandas idiom for multi-column aggregation
  3. Plausible results: The unweighted numbers look reasonable, making bugs hard to detect
  4. Real-world consequences: Users analyzing survey data (CPS, ACS, etc.) will get wildly incorrect population estimates
Root Cause

Looking at microdf/microdataframe.py:644-680, the MicroDataFrameGroupBy class:

  • ✓ Overrides specific methods like sum(), mean(), etc. to apply weights
  • ✗ Does NOT override .agg() or .aggregate(), so they fall back to pandas' unweighted implementation
Proposed Solutions

Option 1: Override .agg() to apply weights (Best)

  • Implement MicroDataFrameGroupBy.agg() to properly handle weights
  • Parse the aggregation specifications and route to weighted methods

Option 2: Raise an error (Safer than current behavior)

def agg(self, *args, **kwargs):
    raise NotImplementedError(
        "MicroDataFrameGroupBy.agg() does not support weights. "
        "Use df.groupby(col).column.sum() instead."
    )

Option 3: Emit a loud warning

def agg(self, *args, **kwargs):
    warnings.warn(
        "MicroDataFrameGroupBy.agg() ignores weights! Results will be unweighted.",
        UserWarning,
        stacklevel=2
    )
    return super().agg(*args, **kwargs)
Related Issues

This extends #193, which identified similar problems with .groupby()[[cols]].sum() but didn't specifically address .agg().

Additional Test Cases Needed
def test_agg_with_weights():
    """Test that .agg() applies weights correctly or raises an error"""
    df = mdf.MicroDataFrame(
        {"group": ["A", "A", "B"], "value": [10, 20, 30]},
        weights=np.array([2, 3, 4])
    )
    
    # These should either work correctly or raise NotImplementedError
    result = df.groupby("group").agg({'value': 'sum'})
    
    # If implemented, should equal weighted sums
    # A: 10*2 + 20*3 = 80
    # B: 30*4 = 120
    expected = pd.DataFrame({'value': [80.0, 120.0]}, index=['A', 'B'])
    expected.index.name = 'group'
    
    # Should NOT be unweighted sums (30, 30)
    assert not result.equals(pd.DataFrame({'value': [30, 30]}))
Priority

HIGH - This is a data correctness bug that produces silently wrong results in a library designed for weighted survey analysis.

Linguagem predominante
Python
Estrelas
16
Forks
10
Merge médio
5d 3h
PRs com merge (30d)
21

Preparar o ambiente

Primeiros passos

  1. Leia a issue inteira e depois o guia de contribuição do projeto.
  2. Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
  3. Faça um fork do repositório e trabalhe em uma branch.
  4. Abra um pull request que referencie o número da issue.

Mais de PolicyEngine/microdf

Todas as issues de PolicyEngine/microdf

Issues semelhantes

Mais issues de Python

Receba novas issues na sua caixa de entrada

Um resumo curto de issues do GitHub para quem está começando.