MediaWiki API Project: Source-Type Composition and "Agenda Setting" Sources
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Ferma
- Stack tecnologico
- jupyter-notebook, plotly, python
- Ambito
- api, data-engineering, data-visualization
Direzione di ricerca
Inizia definendo la coorte di pagine in seed_pages.csv e le regole di classificazione degli URL in source_rules.yaml, quindi consulta la documentazione di MediaWiki Revisions and Query API. Implementa i passaggi di estrazione, normalizzazione, classificazione, metriche e dashboard descritti nell’issue. Il lavoro è completo quando sono disponibili gli artefatti parquet/CSV elencati, la dashboard, il README dei metodi e i test per il parsing degli URL, la copertura del classificatore e i controlli di coerenza di HHI.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Overview
Audit source-type composition and potential agenda setting on Wikipedia by building link graphs for sensitive topics and measuring whether pages disproportionately link to narrow sets of sources (e.g., state media, tabloids, academic journals, reputable newspapers). Deliver a reproducible dataset and dashboard that show concentration, source diversity, and cross-topic differences.
Action Items
If this is the beginning (research & design)
-
Define scope: 100–300 English Wikipedia pages across elections, policing, migration, public health, human rights, and climate disinformation (
seed_pages.csv). -
Build a source taxonomy:
source_rules.yamlmapping URL patterns → classes (e.g., peer-reviewed journal, mainstream news, state media, think tank, government site, company site, tabloid, blog). Include overrides and known aliases/redirects. -
Metrics & windows: current-state snapshot plus a 3–5 year trend (annual). Metrics include: unique sources per page, Herfindahl–Hirschman Index (HHI) of sources, share by class, top domains, and change over time.
-
Methods:
- Extract references/links from article content via revisions API (
prop=revisions&rvslots=main&rvprop=content) for selected waypoints (now and past years). - Parse citations and external links (template fields like
|url=and bare links), normalize URLs, and classify to taxonomy. - Build page→source bipartite graph and compute per-page concentration and per-topic distributions.
- Extract references/links from article content via revisions API (
-
Tooling (choose pairs and keep consistent):
requestsorhttpx;pandasorpolars;mwparserfromhellorwikitextparser; storageduckdborsqlite; vizaltairorplotly; graphnetworkxorigraph. -
Ethics: aggregate reporting; avoid naming individual editors; clarify that links ≠ endorsement and many links are citations.
If researched and ready (implementation steps)
-
Seed & resolve
- Ingest
seed_pages.csv; resolvepageidand record redirects.
- Ingest
-
Timepoints & pulls
- For each page, fetch content for
t0(e.g., Jan 1 three years ago),t1(Jan 1 two years ago),t2(Jan 1 last year), andt_now. Useprop=revisionsto list revids around those dates and then fetch content for selectedrevids.
- For each page, fetch content for
-
Extract & normalize
- From wikitext, extract: (a) citation templates’
|url=fields, (b) external links[http(s)://...], (c) archive URLs → expand to original when present. Normalize to registrable domain + path stem; drop tracking params; handledoi:andpmid:separately.
- From wikitext, extract: (a) citation templates’
-
Classify sources
- Apply
source_rules.yaml(domain regex + path hints) to map each URL to a class; add a small manual override list and an “unknown” bucket.
- Apply
-
Graph & metrics
- Build bipartite graph (page ↔ source domain). Compute per-page: unique source count, HHI, top source share, class shares. Aggregate per topic and over time.
-
Deliver
- Artifacts:
links_raw.parquet,sources_classified.parquet,page_metrics.parquet,topic_yearly.parquet,graph_edgelist.parquet,metrics.csv. - Dashboard: per-topic class shares, HHI distributions, top domains table, change-over-time charts.
- Methods README with parsing heuristics, taxonomy, and known edge cases (templates, archives).
- Artifacts:
-
Quality & Ops
- Caching and retries; persist raw JSON; version the taxonomy rules.
- Tests: URL parser precision, classifier coverage, HHI sanity checks.
- Optional: scheduled yearly refresh; diff reports.
Resources/Instructions
API docs to pin in repo
- Action API overview:
API:Action_API - Revisions (timestamps, content):
API:Revisions - Query continuation & etiquette:
API:Query - (Optional) Exturlusage (to spot-check present-day links to a given domain):
API:Exturlusage
Suggested libraries (choose pairs)
- HTTP:
requests|httpx - DataFrames:
pandas|polars - Parsing:
mwparserfromhell|wikitextparser - Storage:
duckdb|sqlite - Graph:
networkx|igraph - Viz:
altair|plotly
Sample queries
# Revisions near a given date (to pick a waypoint revid)
action=query&prop=revisions&rvprop=ids|timestamp&rvlimit=max&rvstart=2022-01-02T00:00:00Z&rvend=2021-12-31T00:00:00Z&titles=<TITLE>
# Fetch content for a specific revision id
action=query&prop=revisions&revids=<REVID>&rvslots=main&rvprop=content
# Present-day pages that link a domain (spot check)
action=query&list=exturlusage&euquery=example.com&eulimit=max&eunamespace=0
Data handling & ethics
-
Report at page/topic aggregates; avoid editor-level commentary.
-
Note that many links are citations; classify “archive.org” by its original URL when available.
-
Keep an “unknown/unclassified” class; document coverage %.
-
Error handling:
try/exceptwith helpful prints for file-not-found; terminate with a trace on dtype mismatches; warn on partial parsing. -
If this issue requires access to 311 data, please answer the following questions:
- Not applicable.
- N/A
- N/A
- N/A
Project Outline (detailed plan for this idea) in details:
Research question
Do sensitive-topic pages rely on a narrow set of source types (low diversity/high concentration), and how has the mix shifted over the last 3–5 years?
Data sources & modules
prop=revisions(content) at yearly waypoints.list=exturlusagefor current-state validation.- Local taxonomy (
source_rules.yaml) and override list.
Method
- Define page cohort and taxonomy rules.
- Pull wikitext for 3–4 timepoints per page; extract and normalize URLs (expand archived links to originals).
- Classify each URL into a source class; compute page-level and topic-level metrics; construct a bipartite graph.
- Analyze diversity (HHI), top-source share, and class shares over time; identify pages/topics with unusually high concentration.
Key metrics
- Unique sources per page; HHI; top-source share.
- Class distribution (% journals, % mainstream news, % state media, etc.).
- Change per year; pages entering/leaving high-concentration status.
- Coverage of classification rules (% URLs classified).
Deliverables
- Clean tables (
links_raw.parquet,sources_classified.parquet,page_metrics.parquet,topic_yearly.parquet,graph_edgelist.parquet). - Notebook +
reports/source_type_audit.md. - Streamlit/Altair dashboard with diversity plots and top domains.
Caveats & limitations
- Citation templates vary by article; some links are nested or parameterized.
- Archive links and URL shorteners require expansion; some originals are unreachable.
- Classification is heuristic—maintain review samples and report precision/coverage.
Implementation notes
- Keys:
(pageid, revid, url_norm); normalize domains (registrable) and paths (trim UTM). - Persist a query manifest and rule versions with artifacts.
- Add small labeled sets to validate parsing and classification; publish confusion examples.
- Lingua principale
- Jupyter Notebook
- Stelle
- 33
- Fork
- 23
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di hackforla/data-science
-
H2 — Pandemic disruption accelerated a decline that began earlier | Student decline: More than pandemic learning loss aloneForse di nuovo libera @MissBrandyLea l’ha presa 53 giorni fa e non c’è nessuna pull request aperta. Aperta
hackforla/data-science#283 · 8 commenti · 1 assegnatario ·
-
complexity: Large CoP: Data Science project duration: one time role: data science size: 5pt
Difficoltà 5/5 Più di una settimana Idoneità per principianti 20/100
hackforla/data-science#281 ·
-
Investigate America’s Declining HealthspanForse di nuovo libera @wesdufelmeier l’ha presa 60 giorni fa e non c’è nessuna pull request aperta. Apertacomplexity: Large CoP: Data Science project duration: one time role: data science size: 8pt
hackforla/data-science#280 · 8 commenti · 1 assegnatario ·
-
complexity: Large CoP: Data Science project duration: one time role: data analysis role: data science size: 8pt
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
hackforla/data-science#279 ·
-
complexity: Large CoP: Data Science project duration: one time role: data science size: 8pt
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
hackforla/data-science#278 ·
Tutte le issue di hackforla/data-science
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
I maintainer di solito rispondono entro 1 giorno
-
friction
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
kentcdodds/kody#3263 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
PolicyEngine/policyengine-us#10073 ·
I maintainer di solito rispondono entro 2 giorni
-
bug release:v5.56
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Jason-Vaughan/TangleClaw#2270 ·
I maintainer di solito rispondono entro 1 giorno
-
area:jobads-cv BE mvp P2
Difficoltà 2/5 1-3 ore Idoneità per principianti 62/100
klasolsson81/jobbliggaren#2099 ·
I maintainer di solito rispondono entro 1 giorno