heavy_tails.md identifies power laws only visually — add a goodness-of-fit test
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
- Tipo di issue
- Documentazione
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Tranquilla
- Stack tecnologico
- jupyter-notebook, python
- Ambito
- data, documentation
Direzione di ricerca
Inizia da lectures/heavy_tails.md, in particolare dal materiale su CCDF empirica, Q-Q plot e ht_ex4. Studia il Hill estimator e la procedura basata su KS di Clauset–Shalizi–Newman, quindi aggiungi una breve sezione, qualifica la diagnostica visiva esistente e ricollega il metodo a ht_ex4, in modo che i lettori possano valutare dati Pareto rispetto a dati lognormali.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
heavy_tails.md defines power laws and Pareto tails formally, then identifies them by eye: "All plots are in log-log, so that a power law shows up as a linear log-log plot, at least in the upper tail." The lecture builds empirical CCDFs and Q-Q plots for firm size and city size, and stops there.
It has no goodness-of-fit test, no tail-index estimator, and no note that log-log linearity is a weak diagnostic. Searching the lecture for "goodness", "test for", "Hill estimator", "KS test", "Kolmogorov" or "Clauset" returns nothing.
Why this matters
Exercise ht_ex4 asks the reader to compare a Pareto distribution against a mean-and-median-matched lognormal, for the present discounted value of corporate tax revenue, and to observe the difference. The lecture therefore poses the Pareto-versus-lognormal question and gives the reader no way to settle it from data. The same comparison appears, also unresolved, as Exercise 2.2.10 of Economic Networks.
Eyeballing a log-log plot for straightness is precisely the practice the goodness-of-fit literature exists to caution against, so teaching only the visual method leaves readers with a diagnostic that looks more reliable than it is.
Suggested scope
A short section, not a new lecture:
- Estimating the tail index — the Hill estimator, and its sensitivity to where the tail is deemed to start
- Testing the hypothesis — the Clauset–Shalizi–Newman KS-based procedure is the standard reference and has a widely used implementation
- A sentence of honesty in the existing visual material, noting that log-log linearity is suggestive rather than conclusive
Enough that a reader can answer "is this actually a power law?" instead of "does this look straight?". Wiring it back to ht_ex4 would close the loop on an exercise that currently ends in an observation rather than an answer.
Spun out of QuantEcon/meta#141, which collected it as the one concrete deliverable inside a broader proposal.
- Lingua principale
- Jupyter Notebook
- Stelle
- 65
- Fork
- 31
- Merge medio
- 17h 18m
- PR unite (30g)
- 5
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di QuantEcon/lecture-python-intro
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
-
enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
Tutte le issue di QuantEcon/lecture-python-intro
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
deepset-ai/haystack#13074 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
langflow-ai/langflow#15510 ·
I maintainer di solito rispondono entro 1 giorno
-
area:analytics bug effort:S level:L3
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
mczielinski/ob-analytics#314 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Community
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
I maintainer di solito rispondono entro 3 giorni
-
Add: BHOOMI 24x7Apertachannels:add check:passed
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 75/100
I maintainer di solito rispondono entro 4 giorni