groupby().agg() silently ignores weights, producing incorrect results
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 3/5
- Tempo estimado
- 1-2 dias
- Facilidade para iniciantes
- 65/100
- Tipo de issue
- Bug
- Clareza
- Claramente especificada
- Status de atividade
- Ativa
- Domínio
- backend-api-design, data, data-engineering
Direção de pesquisa
The issue is in microdf/microdataframe.py lines 644-680, where the MicroDataFrameGroupBy class does not override .agg(). Start by reading the existing weighted method overrides (like sum). Determine how to parse aggregation specs and apply weights, or implement a clear error/warning. Run the provided test case to verify the fix.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Problem
MicroDataFrame's groupby operations silently ignore weights when using .agg(), producing incorrect (unweighted) results without any warning or error.
Example
import microdf as mdf
import numpy as np
df = mdf.MicroDataFrame(
{"group": ["A", "A", "B", "B"], "value": [10, 20, 30, 40]},
weights=np.array([2, 3, 1, 4])
)
# CORRECT (weighted):
df.groupby("group").value.sum()
# A: 10*2 + 20*3 = 80.0
# B: 30*1 + 40*4 = 190.0
# INCORRECT (unweighted) - no warning!
df.groupby("group").agg({'value': 'sum'})
# A: 10 + 20 = 30 ❌
# B: 30 + 40 = 70 ❌
Impact
This is a critical data correctness issue because:
- Silent failure: No error or warning - just wrong numbers
- Natural usage pattern:
.agg()is a standard pandas idiom for multi-column aggregation - Plausible results: The unweighted numbers look reasonable, making bugs hard to detect
- Real-world consequences: Users analyzing survey data (CPS, ACS, etc.) will get wildly incorrect population estimates
Root Cause
Looking at microdf/microdataframe.py:644-680, the MicroDataFrameGroupBy class:
- ✓ Overrides specific methods like
sum(),mean(), etc. to apply weights - ✗ Does NOT override
.agg()or.aggregate(), so they fall back to pandas' unweighted implementation
Proposed Solutions
Option 1: Override .agg() to apply weights (Best)
- Implement
MicroDataFrameGroupBy.agg()to properly handle weights - Parse the aggregation specifications and route to weighted methods
Option 2: Raise an error (Safer than current behavior)
def agg(self, *args, **kwargs):
raise NotImplementedError(
"MicroDataFrameGroupBy.agg() does not support weights. "
"Use df.groupby(col).column.sum() instead."
)
Option 3: Emit a loud warning
def agg(self, *args, **kwargs):
warnings.warn(
"MicroDataFrameGroupBy.agg() ignores weights! Results will be unweighted.",
UserWarning,
stacklevel=2
)
return super().agg(*args, **kwargs)
Related Issues
This extends #193, which identified similar problems with .groupby()[[cols]].sum() but didn't specifically address .agg().
Additional Test Cases Needed
def test_agg_with_weights():
"""Test that .agg() applies weights correctly or raises an error"""
df = mdf.MicroDataFrame(
{"group": ["A", "A", "B"], "value": [10, 20, 30]},
weights=np.array([2, 3, 4])
)
# These should either work correctly or raise NotImplementedError
result = df.groupby("group").agg({'value': 'sum'})
# If implemented, should equal weighted sums
# A: 10*2 + 20*3 = 80
# B: 30*4 = 120
expected = pd.DataFrame({'value': [80.0, 120.0]}, index=['A', 'B'])
expected.index.name = 'group'
# Should NOT be unweighted sums (30, 30)
assert not result.equals(pd.DataFrame({'value': [30, 30]}))
Priority
HIGH - This is a data correctness bug that produces silently wrong results in a library designed for weighted survey analysis.
- Linguagem predominante
- Python
- Estrelas
- 16
- Forks
- 10
- Merge médio
- 5d 3h
- PRs com merge (30d)
- 21
Preparar o ambiente
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de PolicyEngine/microdf
-
docs/examples.md still says MicroDataFrame.cov() and .corr() are unweightedTalvez já em andamento @juaristi22 assumiu há 7 dias. Aberta
PolicyEngine/microdf#335 · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
Poverty gap docstrings overclaim FGT indices, and the poverty estimators have no testsTalvez já em andamento @juaristi22 assumiu há 7 dias. Aberta
PolicyEngine/microdf#334 · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
Fail closed: aggregation and construction paths that silently return unweighted resultsTalvez já em andamento @juaristi22 assumiu há 7 dias. Abertabug
PolicyEngine/microdf#333 · 1 responsável ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 30/100
PolicyEngine/microdf#314 ·
Mantenedores costumam responder em até 1 dia
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 45/100
PolicyEngine/microdf#223 ·
Mantenedores costumam responder em até 1 dia
Todas as issues de PolicyEngine/microdf
Issues semelhantes
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
Mantenedores costumam responder em até 1 dia
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 90/100
Mantenedores costumam responder em até 1 dia
-
https://search.utilibre.orgAbertainstance instance add
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
searxng/searx-instances#941 · 1 comentário ·
-
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 92/100
FluidNumerics/fluid-walk-blocker#89 ·
Mantenedores costumam responder em até 1 dia
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 84/100
Mantenedores costumam responder em até 1 dia