Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

heavy_tails.md identifies power laws only visually — add a goodness-of-fit test

Aperta
#807 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

Nessuno ha ancora preso questa issue.

Valutazione

Difficoltà
4/5
Tempo stimato
3-5 giorni
Idoneità per principianti
48/100
Tipo di issue
Documentazione
Chiarezza
Abbastanza chiara
Stato di attività
Tranquilla
Stack tecnologico
jupyter-notebook, python

Direzione di ricerca

Inizia da lectures/heavy_tails.md, in particolare dal materiale su CCDF empirica, Q-Q plot e ht_ex4. Studia il Hill estimator e la procedura basata su KS di Clauset–Shalizi–Newman, quindi aggiungi una breve sezione, qualifica la diagnostica visiva esistente e ricollega il metodo a ht_ex4, in modo che i lettori possano valutare dati Pareto rispetto a dati lognormali.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

heavy_tails.md defines power laws and Pareto tails formally, then identifies them by eye: "All plots are in log-log, so that a power law shows up as a linear log-log plot, at least in the upper tail." The lecture builds empirical CCDFs and Q-Q plots for firm size and city size, and stops there.

It has no goodness-of-fit test, no tail-index estimator, and no note that log-log linearity is a weak diagnostic. Searching the lecture for "goodness", "test for", "Hill estimator", "KS test", "Kolmogorov" or "Clauset" returns nothing.

Why this matters

Exercise ht_ex4 asks the reader to compare a Pareto distribution against a mean-and-median-matched lognormal, for the present discounted value of corporate tax revenue, and to observe the difference. The lecture therefore poses the Pareto-versus-lognormal question and gives the reader no way to settle it from data. The same comparison appears, also unresolved, as Exercise 2.2.10 of Economic Networks.

Eyeballing a log-log plot for straightness is precisely the practice the goodness-of-fit literature exists to caution against, so teaching only the visual method leaves readers with a diagnostic that looks more reliable than it is.

Suggested scope

A short section, not a new lecture:

  • Estimating the tail index — the Hill estimator, and its sensitivity to where the tail is deemed to start
  • Testing the hypothesis — the Clauset–Shalizi–Newman KS-based procedure is the standard reference and has a widely used implementation
  • A sentence of honesty in the existing visual material, noting that log-log linearity is suggestive rather than conclusive

Enough that a reader can answer "is this actually a power law?" instead of "does this look straight?". Wiring it back to ht_ex4 would close the loop on an exercise that currently ends in an observation rather than an answer.

Spun out of QuantEcon/meta#141, which collected it as the one concrete deliverable inside a broader proposal.

Lingua principale
Jupyter Notebook
Stelle
65
Fork
31
Merge medio
17h 18m
PR unite (30g)
5

Preparare l'ambiente

Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di QuantEcon/lecture-python-intro

Tutte le issue di QuantEcon/lecture-python-intro

Issue simili

Altre issue su Data Engineering

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.