Skip to content

Best practice for persisting large accumulated LangGraph state without hitting ScheduleActivityTask payload size limits #1894

Description

@dorelad-upwind

We're running LangGraph-based agents through temporalio.contrib.langgraph.LangGraphPlugin, where each graph node executes as a Temporal Activity. The workflows accumulate state across many sequential (and some parallel) node executions — investigation-style agents with 20-30+ nodes producing tool-call transcripts, messages, etc.

Before wrapping these graphs with Temporal, we ran them directly against LangGraph's own Postgres-backed checkpointer, which persists automatically after every node (every "superstep"). Once we moved execution behind Temporal Activities, that per-node persistence went away — the nodes now run as separate Activities without a shared in-process checkpointer, so nothing writes to Postgres until we explicitly do so ourselves.

Our first attempt reintroduced persistence as a single workflow.execute_activity(persist_result_activity, args=[...full accumulated final state...]) call after ainvoke() returns, so it survives even if no client ever reconnects to read handle.result(). That worked for small/medium runs, but a larger run hit:

BadScheduleActivityAttributes: ScheduleActivityTaskCommandAttributes.Input exceeds size limit.

...which terminated the entire workflow server-side (WORKFLOW_EXECUTION_TERMINATED) — a total loss, worse than the gap we were trying to close. Individual per-node Activity payloads never approached the limit (29/29 node-Activities completed fine); only the one bulk end-of-run payload did.

Our leading fix: restore the original per-node persistence behavior — write each node's own delta to Postgres from inside that node's own Activity execution (no extra execute_activity hop needed, since the node body already runs as an Activity) — rather than shipping the whole accumulated state through Temporal in one call at the end. This keeps every persistence write proportional to a single node's output, matching what the pre-Temporal execution already did natively.

Questions for the maintainers:

  1. Is per-node/per-Activity persistence the idiomatic pattern here, or is there Temporal-native support for this (e.g. something LangGraphPlugin already offers) that we're missing?
  2. We also considered gzip-compressing the final payload before passing it as Activity args — it only raises the ceiling rather than removing it, and doesn't restore true mid-run durability (a crash just before that one activity still loses the whole run's state). Are there better-supported options for cases where a single large Activity payload is genuinely unavoidable (custom PayloadCodec, chunking, streaming activity results)?
  3. Any general guidance on structuring Activities around something like LangGraph, where per-node output sizes vary widely and the framework's own native checkpointing gets bypassed by the Activity boundary?

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions