Best practice for persisting large accumulated LangGraph state without hitting ScheduleActivityTask payload size limits
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Documentazione
- Chiarezza
- Da chiarire
- Stato di attività
- Attiva
- Stack tecnologico
- python
- Ambito
- backend, distributed-systems
Direzione di ricerca
Start with the LangGraphPlugin activity boundary and the workflow.execute_activity call described in the issue, then review how the Postgres-backed checkpointer and Temporal activity payload limits interact. The issue is complete when maintainers document a supported persistence pattern and clarify whether payload codecs, chunking, or streaming are appropriate for large state.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
We're running LangGraph-based agents through temporalio.contrib.langgraph.LangGraphPlugin, where each graph node executes as a Temporal Activity. The workflows accumulate state across many sequential (and some parallel) node executions — investigation-style agents with 20-30+ nodes producing tool-call transcripts, messages, etc.
Before wrapping these graphs with Temporal, we ran them directly against LangGraph's own Postgres-backed checkpointer, which persists automatically after every node (every "superstep"). Once we moved execution behind Temporal Activities, that per-node persistence went away — the nodes now run as separate Activities without a shared in-process checkpointer, so nothing writes to Postgres until we explicitly do so ourselves.
Our first attempt reintroduced persistence as a single workflow.execute_activity(persist_result_activity, args=[...full accumulated final state...]) call after ainvoke() returns, so it survives even if no client ever reconnects to read handle.result(). That worked for small/medium runs, but a larger run hit:
BadScheduleActivityAttributes: ScheduleActivityTaskCommandAttributes.Input exceeds size limit.
...which terminated the entire workflow server-side (WORKFLOW_EXECUTION_TERMINATED) — a total loss, worse than the gap we were trying to close. Individual per-node Activity payloads never approached the limit (29/29 node-Activities completed fine); only the one bulk end-of-run payload did.
Our leading fix: restore the original per-node persistence behavior — write each node's own delta to Postgres from inside that node's own Activity execution (no extra execute_activity hop needed, since the node body already runs as an Activity) — rather than shipping the whole accumulated state through Temporal in one call at the end. This keeps every persistence write proportional to a single node's output, matching what the pre-Temporal execution already did natively.
Questions for the maintainers:
- Is per-node/per-Activity persistence the idiomatic pattern here, or is there Temporal-native support for this (e.g. something LangGraphPlugin already offers) that we're missing?
- We also considered gzip-compressing the final payload before passing it as Activity args — it only raises the ceiling rather than removing it, and doesn't restore true mid-run durability (a crash just before that one activity still loses the whole run's state). Are there better-supported options for cases where a single large Activity payload is genuinely unavoidable (custom
PayloadCodec, chunking, streaming activity results)? - Any general guidance on structuring Activities around something like LangGraph, where per-node output sizes vary widely and the framework's own native checkpointing gets bypassed by the Activity boundary?
- Lingua principale
- Python
- Stelle
- 1.2k
- Fork
- 241
- Merge medio
- 3g 2h
- PR unite (30g)
- 49
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di temporalio/sdk-python
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
temporalio/sdk-python#1517 · 10 commenti ·
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
temporalio/sdk-python#496 ·
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
temporalio/sdk-python#1890 ·
-
[Bug] Local activity resolutions regrouped on replay since 1.32.0, delivering the wrong payload Aperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
temporalio/sdk-python#1881 · 1 commento ·
-
bug
temporalio/sdk-python#1817 · 1 commento · 1 assegnatario ·
Tutte le issue di temporalio/sdk-python
Issue simili
-
area: harness bug status: needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Human-Agent-Society/reef#625 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 80/100
learningequality/kolibri#15351 · 2 commenti ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Name consistency Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
eellak/triplestore#65 · 1 commento ·