US production: /us/metadata returns 500 and state-region simulations fail with connection reset at the simulation gateway
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- github-actions, python, sqlalchemy
- Ambito
- api, backend, cloud, observability
Direzione di ricerca
Inizia dai log di produzione per GET /us/metadata e dal percorso API → gateway → Modal executor, in particolare policyengine-simulation-py4-20-3, e confrontali con l'endpoint UK funzionante. Traccia il motivo per cui le richieste continuano a essere in elaborazione o restituiscono reset della connessione dopo la migrazione della persistenza e le modifiche a Cloud Run. Il lavoro è completato quando i metadati US restituiscono 200, le esecuzioni state-region terminano, le esecuzioni nazionali terminano normalmente e il monitoraggio esegue probe su questi percorsi.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The US v1 production stack is partially down, and monitoring can't see it because Better Stack only probes the API root (which returns 200):
GET /us/metadatareturns HTTP 500 — observed on 2026-08-06, 2026-08-15, and 2026-08-21 (every check)./uk/metadataand/return 200. The app's US parameter pages depend on this endpoint.- US state-region society-wide simulations fail —
GET /us/economy/{policy}/over/2?region=state/or&time_period=2027returns{"status": "error", "message": "[Errno 104] Connection reset by peer", "result": null}, consistently on retry. Previously cached state results now error too (the cached result is gone and the recompute fails). - National runs hang —
region=usrequests sit in"computing"for 10+ minutes (normal completion is a few minutes).
As of 2026-08-21 18:00 UTC, every executor app in the Modal main environment (policyengine-simulation-py4-18-*, py4-20-3, py4-22-0) and policyengine-simulation-gateway show deployed with 0 running tasks while the API reports runs computing — requests are not starting on Modal. The reset-by-peer surfaces at the API → gateway → executor hop for the route production uses (resolved_app_name: policyengine-simulation-py4-20-3 on completed runs from before the failure).
Timeline (all observations read-only, from outside)
- 2026-07-17: state runs healthy (e.g. policy 98001
state/or2027 → −$25.2M, stamped pe-us 1.764.6 /populace-us-2024-buildi-sparse…/policyengine-simulation-py4-20-3). - 2026-08-06:
/us/metadata500; state runs still complete with the identical stamp. - 2026-08-15:
/us/metadata500; state runs still complete. - 2026-08-21:
/us/metadata500; state runs fail with connection reset; national hangs.
Better Stack logged only two brief auto-resolved root incidents (8/7 09:14–09:19, 8/14 04:47–04:52) in the same window.
Repro
curl -s -o /dev/null -w "%{http_code}\n" https://api.policyengine.org/us/metadata # 500
curl -s -o /dev/null -w "%{http_code}\n" https://api.policyengine.org/uk/metadata # 200
curl -s "https://api.policyengine.org/us/economy/98001/over/2?region=state/or&time_period=2027"
# {"status": "error", "message": "[Errno 104] Connection reset by peer", "result": null}
Possibly related
- The metadata failure window starts after the v1 persistence migration to SQLAlchemy/Alembic (#3788, 8/12) and Cloud Run revision changes in sim-api (#657, 7/28) — correlation only; I have not looked at logs.
- #2430 (simulation-API 500s surface as
status: ok+result: null) describes the adjacent failure mode; this one at least surfaces asstatus: error. - Suggest adding
/us/metadataand a state-region simulation smoke to Better Stack — the root probe can't catch either.
User-facing impact: every state-level analysis on policyengine.org (US) currently fails; external partners sharing state reform links hit errors.
- Lingua principale
- Python
- Stelle
- 18
- Fork
- 33
- Merge medio
- 1g 7h
- PR unite (30g)
- 22
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di PolicyEngine/policyengine-api
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 55/100
PolicyEngine/policyengine-api#3865 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
PolicyEngine/policyengine-api#3864 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 42/100
PolicyEngine/policyengine-api#3862 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
PolicyEngine/policyengine-api#3860 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 38/100
PolicyEngine/policyengine-api#3823 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di PolicyEngine/policyengine-api
Issue simili
-
[Bug]: Non-vision image fallback calls vision_analyze with an empty source for oversized inline images and tells the model the image is corruptForse già presa @liuhao1024 l’ha presa oggi. Apertacomp/agent P2 tool/vision type/bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
NousResearch/hermes-agent#132605 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
pymc-labs/pymc-marketing#3102 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
vllm-project/compressed-tensors#920 ·
I maintainer di solito rispondono entro 1 giorno
-
Add category README filesApertadocumentation good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
-
accepted bug wg/router-models-inference-runtime
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
vllm-project/semantic-router#4524 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno