Async persistence restored before attach() never rejoins an in-flight run
Maintainers usually reply within 1 day
@tombeckenham is already working on this.
Since Oct 6, 2026.
Assessment
This issue has not been assessed yet.
Description
TanStack AI version
@tanstack/ai-client 0.36.1 and 0.37.0 — reproduced against clean npm installations of both. Originally observed through @tanstack/ai-react 0.29.4.
Framework/Library version
Originally observed in React Native with an AsyncStorage-backed ChatClientPersistence adapter. The isolated reproduction below uses Node.js 24.14.1, with no React, native runtime, server, or model API required.
Describe the bug and the steps to reproduce it
An in-flight chat does not reconnect after a cold start when the asynchronous persistence read completes before the view calls client.attach().
The transcript restores successfully, but connection.joinRun() is never called. The partial assistant message therefore remains frozen even though the server still owns the generation and its replay log. Attaching before the same async read resolves works.
- Persist a combined
{ messages, resume: { resumeState: { threadId, runId } } }record for an in-flight run, with no pending interrupts. - Construct a fresh
ChatClientwith an asynchronousgetItemand a connection supportingjoinRun. - Let
getItemresolve before mounting/attaching the view. - Call
client.attach(). - Observe that the messages are restored, but no rejoin occurs.
Expected: construction/hydration performs no network I/O before attachment; after attachment, the client joins the restored run exactly once, regardless of whether storage or attachment happened first.
Actual: only the attachment-first ordering calls joinRun. Storage-first restores the transcript but permanently skips the rejoin for that mount.
Your Minimal, Reproducible Example - (Sandbox Highly Recommended)
Self-contained executable reproduction below. It does not require application code, API keys, or a running backend.
In an empty directory:
npm init -y
npm install @tanstack/[email protected]
# Save the following as repro.mjs
node repro.mjs
The same failure occurs with @tanstack/[email protected].
import assert from 'node:assert/strict';
import { setTimeout as delay } from 'node:timers/promises';
import { ChatClient } from '@tanstack/ai-client';
async function check(restoreBeforeAttach) {
const joined = [];
const snapshot = {
messages: [{ id: 'user-1', role: 'user', parts: [{ type: 'text', content: 'Hello' }] }],
resume: { resumeState: { threadId: 'thread-1', runId: 'run-1' } },
};
const client = new ChatClient({
threadId: 'thread-1',
persistence: {
getItem: async () => structuredClone(snapshot),
setItem: async () => {},
removeItem: async () => {},
},
connection: {
async *connect() { throw new Error('A restore must not start a new generation'); },
async *joinRun(runId) {
joined.push(runId);
yield { type: 'RUN_STARTED', threadId: 'thread-1', runId, timestamp: Date.now() };
yield { type: 'TEXT_MESSAGE_START', messageId: 'reply-1', role: 'assistant', timestamp: Date.now() };
yield { type: 'TEXT_MESSAGE_CONTENT', messageId: 'reply-1', delta: 'Resumed reply', timestamp: Date.now() };
yield { type: 'TEXT_MESSAGE_END', messageId: 'reply-1', timestamp: Date.now() };
yield { type: 'RUN_FINISHED', threadId: 'thread-1', runId, timestamp: Date.now() };
},
},
});
try {
if (restoreBeforeAttach) {
await delay(0); // Async storage resolves before the view's mount effect.
assert.equal(client.getMessages().length, 1); // Transcript restored successfully.
}
client.attach();
await delay(50);
assert.deepEqual(joined, ['run-1'], `restoreBeforeAttach=${restoreBeforeAttach}`);
console.log(`PASS restoreBeforeAttach=${restoreBeforeAttach}`);
} finally {
client.detach();
client.dispose();
}
}
await check(false); // Control: attach first, then storage resolves. Passes.
await check(true); // Bug: storage resolves first. joinRun is never called.
Observed output on both unpatched versions:
PASS restoreBeforeAttach=false
AssertionError [ERR_ASSERTION]: restoreBeforeAttach=true
+ actual - expected
+ []
- [ 'run-1' ]
Root cause
In packages/ai-client/src/chat-client.ts:
- Constructor-time
rejoinRunIdis populated from synchronously restored state (orinitialResumeSnapshot). An async persistence read leaves it unset. applyPersistedResume()receives the async snapshot and callsmaybeRejoinInFlight(runId).- If the view has not attached yet,
maybeRejoinInFlight()returns becausetailingis false. This guard is correct: an unmounted/discarded client must not open a connection. - The async run ID is not retained for a later attachment.
attach()checks the constructor-timerejoinRunId, so it has no pending rejoin to perform.
Suggested fix and validation
Retain a bare in-flight run ID when async hydration completes, so attach() can consume it later. The core change tested locally is:
- private readonly rejoinRunId: string | null | undefined
+ private rejoinRunId: string | null | undefined
private applyPersistedResume(snapshot: ChatResumeSnapshot): void {
this.applyResumeSnapshot(snapshot)
const hasInterrupts =
Array.isArray(snapshot.pendingInterrupts) &&
snapshot.pendingInterrupts.length > 0
const runId = snapshot.resumeState?.runId
+ this.rejoinRunId = hasInterrupts ? null : runId
if (!hasInterrupts && runId) {
this.maybeRejoinInFlight(runId)
}
}
I verified that adding the equivalent assignment to the published 0.37.0 runtime makes both cases in the standalone reproduction pass. A complete fix should also keep this remembered ID current and clear it when the run becomes terminal, rather than leaving a stale constructor-era ID for later reattachments. Interrupt-only snapshots should retain their existing behavior instead of being automatically tailed.
The application-level regression test also recreates the client from a frozen JSON storage snapshot and routes the real client through a durable NDJSON backend. It checks both storage/attachment orderings, replay without a second generation POST, and no duplicate assistant text. Model output is simulated.
Related reports checked
- #1429 concerns server-authoritative hydration choosing a parent interrupt over an active continuation. This report uses a client-side storage adapter and a bare run pointer, with no interrupts.
- #1058 concerns draining queued actions after a rejoined client-tool run. Here
joinRunis never called in the first place. - #1620 deduplicates server hydration requests in Strict Mode. This race is in async client-side persistence restoration, and the reproduction still fails on 0.37.0.
Screenshots or Videos (Optional)
The executable assertions above reproduce the failure without a UI.
Do you intend to try to help solve this bug with your own PR?
A tested local workaround and reproduction are included above; no PR is being opened with this report.
Terms & Code of Conduct
- I agree to follow this project's Code of Conduct.
- I understand that a bug without a reliable, debuggable reproduction may not be fixed and may be closed.
- Dominant language
- TypeScript
- Stars
- 3.2k
- Forks
- 361
- Avg merge
- 1d 22h
- Merged PRs (30d)
- 218
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from TanStack/ai
-
ai-openrouter: malformed Chat Completions tool arguments execute as an empty objectPossibly taken @tombeckenham claimed this 1 day ago. Openhas-pr waiting-on: maintainer
TanStack/ai#1689 · 1 assignee ·
Maintainers usually reply within 1 day
-
StreamProcessor discards metadata from REASONING_MESSAGE_STARTPossibly taken @tombeckenham claimed this 1 day ago. Openwaiting-on: maintainer
TanStack/ai#1667 · 1 assignee ·
Maintainers usually reply within 1 day
-
Ollama chat adapter does not forward the request abort signal to the ollama SDKPossibly taken @tombeckenham claimed this 3 days ago. Openwaiting-on: maintainer
TanStack/ai#1644 · 1 assignee ·
Maintainers usually reply within 1 day
-
`ai-opencode`: `RUN_FINISHED` never arrives when an adapter teardown step does not settlePossibly taken @tombeckenham claimed this 4 days ago. Openwaiting-on: maintainer
TanStack/ai#1638 · 1 comment · 1 assignee ·
Maintainers usually reply within 1 day
-
ai-mcp: support publishing MCP list changes to subscribed clientsPossibly taken @jherr claimed this 4 days ago. Openwaiting-on: maintainer
Difficulty 5/5 Over a week Newbie friendliness 28/100
TanStack/ai#1631 · 1 assignee ·
Maintainers usually reply within 1 day
Similar issues
-
effort:S priority:P2
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
cameri/nostream#811 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 62/100
dam-agents/dam#4562 ·
Maintainers usually reply within 1 day
-
bug p3 triaged
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
Maintainers usually reply within 1 day
-
bug javascript P2-medium python release:v3.1
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
adrirubio/claude-deck#546 ·
Maintainers usually reply within 1 day