Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Two-sample SMD variance multiplies by degrees of freedom instead of dividing

Aperta Adatta ai principianti
#143 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
86/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
data

Direzione di ricerca

Inizia dal punto di ingresso compute_measure in pymare.effectsize e riproduci la varianza SMD a due gruppi riportata per dimensioni del campione pari a 50 e 500. Aggiungi la copertura di regressione proposta per d = 0, d = 1 e il ridimensionamento della dimensione del campione, quindi esegui i test mirati e conferma che l’intera suite non-Stan continui a passare.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

PyMARE makes a nonzero two-group effect look less certain as the study gets larger. With equal SDs and Cohen’s d = 1, increasing each group from 50 to 500 raises the reported sampling variance from 49.04 to 499.004. It should fall from 0.045102 to 0.004501.

This variance sets study weights and can materially change a meta-analysis.

Cause

Release 0.0.12 and current master at 613c06ba432451ad536fb06877fe9fd653ceef85 encode:

d**2 / 2 * (n1 + n2 - 2)

The standard large-sample term is:

d**2 / (2 * (n1 + n2 - 2))

The missing parentheses multiply by the residual degrees of freedom instead of dividing by it. The same bad v_d flows into Hedges’ g (SMD).

Minimal reproduction

from pymare.effectsize import compute_measure

for n in (50, 500):
    d, variance = compute_measure(
        "D", m1=10, sd1=2, n1=n, m2=8, sd2=2, n2=n
    )
    print(n, d, variance)

Released output:

50  1.0  49.04
500 1.0 499.004

Expected:

50  1.0  0.04510204081632653
500 1.0 0.004501002004008016

A zero-effect control is unchanged at 0.04 for n1 = n2 = 50 because the disputed term vanishes.

Real-data consequence

I replayed the nine-study Normand (1999) stroke length-of-stay example published by metafor:

https://www.metafor-project.org/doku.php/tips:assembling_data_smd

PyMARE’s SMD point estimates match, but its variances are 8.35× to 8,448× too large. Feeding those outputs into PyMARE’s REML estimator changes:

Input variance Pooled SMD SE p tau²
PyMARE 0.0.12 -0.0567 0.6678 0.9323 0
Corrected formula -0.5369 0.3082 0.0815 0.7902
Published metafor result -0.5371 0.3087 0.0818 0.7908

This demonstrates a changed analysis result on real data. I have not identified a published analysis that used PyMARE’s affected converter, so no changed published conclusion is claimed.

Proposed fix and checks

The fix is the one-line parenthesis change above plus a regression test covering d = 0, d = 1, and sample-size scaling. On Python 3.12.14, the focused file passes 9 tests; the full non-Stan suite passes 770 tests with 27 skips.

Lingua principale
Python
Stelle
58
Fork
16
Metriche di merge delle PR
Nessuna PR unita negli ultimi 30g

Preparare l'ambiente

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di neurostuff/PyMARE

Tutte le issue di neurostuff/PyMARE

Issue simili

Altre issue su Python

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.