heavy_tails.md identifies power laws only visually — add a goodness-of-fit test
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Documentation
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- jupyter-notebook, python
- Domain
- data, documentation
Research direction
Start with lectures/heavy_tails.md, especially the empirical CCDF, Q-Q plot, and ht_ex4 material. Research the Hill estimator and the Clauset–Shalizi–Newman KS-based procedure, then add a short section, qualify the existing visual diagnostic, and connect the method back to ht_ex4 so readers can assess Pareto versus lognormal data.
Written by the indexing model from the issue text.
Description
heavy_tails.md defines power laws and Pareto tails formally, then identifies them by eye: "All plots are in log-log, so that a power law shows up as a linear log-log plot, at least in the upper tail." The lecture builds empirical CCDFs and Q-Q plots for firm size and city size, and stops there.
It has no goodness-of-fit test, no tail-index estimator, and no note that log-log linearity is a weak diagnostic. Searching the lecture for "goodness", "test for", "Hill estimator", "KS test", "Kolmogorov" or "Clauset" returns nothing.
Why this matters
Exercise ht_ex4 asks the reader to compare a Pareto distribution against a mean-and-median-matched lognormal, for the present discounted value of corporate tax revenue, and to observe the difference. The lecture therefore poses the Pareto-versus-lognormal question and gives the reader no way to settle it from data. The same comparison appears, also unresolved, as Exercise 2.2.10 of Economic Networks.
Eyeballing a log-log plot for straightness is precisely the practice the goodness-of-fit literature exists to caution against, so teaching only the visual method leaves readers with a diagnostic that looks more reliable than it is.
Suggested scope
A short section, not a new lecture:
- Estimating the tail index — the Hill estimator, and its sensitivity to where the tail is deemed to start
- Testing the hypothesis — the Clauset–Shalizi–Newman KS-based procedure is the standard reference and has a widely used implementation
- A sentence of honesty in the existing visual material, noting that log-log linearity is suggestive rather than conclusive
Enough that a reader can answer "is this actually a power law?" instead of "does this look straight?". Wiring it back to ht_ex4 would close the loop on an exercise that currently ends in an observation rather than an answer.
Spun out of QuantEcon/meta#141, which collected it as the one concrete deliverable inside a broader proposal.
- Dominant language
- Jupyter Notebook
- Stars
- 65
- Forks
- 32
- Avg merge
- 4d 14h
- Merged PRs (30d)
- 6
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from QuantEcon/lecture-python-intro
-
Difficulty 1/5 Under an hour Newbie friendliness 92/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
Difficulty 1/5 Under an hour Newbie friendliness 90/100
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
-
enhancement
Difficulty 2/5 Half a day Newbie friendliness 76/100
QuantEcon/lecture-python-intro#764 · 4 comments ·
All issues in QuantEcon/lecture-python-intro
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
open-compass/VLMEvalKit#1705 ·
-
Feature Request units
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
spacetelescope/synphot_refactor#447 · 2 comments ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
TauricResearch/TradingAgents#1397 ·