One fixture for agent tests: run(), continue(), and loaded history
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 32/100
- Tipo di issue
- Funzionalità
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- postgres, redis, typescript
- Ambito
- testing
Direzione di ricerca
Start in packages/junior-evals/src/fixture/ and compare the existing evals in evals/integration/ with the run, agent, history, continuation, fork, and setup-data contract. Check packages/junior/tests/ to preserve the boundary for tests that do not run the agent. Done means the fixture supports the documented real-agent flows, history loading, progress handling, judges, and Conversation reporting without changing Guardian or router evals.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Tests that run the agent should use one small fixture. Every turn under test runs the real agent with the real model. History, not mocks, keeps the tests cheap. This issue defines the fixture contract so that implementation can start.
This replaces the deterministic fake-model layer from #1425. It also changes the harness rules in #1398, which fake the model stream.
Why
Tests that run the agent depend on runtime internals. Routine refactors rewrite them, so they stop catching regressions.
- In the last six months, 61% of commits that changed agent code also changed these tests. In
ref:commits, 40% changed existing assertions. One reshape ofAgentRun(c0fee35ff) changed 33 test files. - There are 8 ways to set up a run, 11 ways to fake the model, and 5 or more ways to load history.
- Tests that fake the model still send about 320 live AI Gateway requests per run, because nothing fakes the turn router, titles, or the reply policy. With a Vercel login, these are real billed calls of 1.7–7 seconds, and they caused most local timeouts.
- Eval
run()allows one scenario per test. Messages ineventscallhandleNewMentiondirectly, so later eval messages skip the mailbox path that production uses. preloadHistorywrites four stores by hand. It does not write turn lifecycle events or themessage:<id>:agentkey, so a fork of preloaded history fails.
Principles
- The agent is one unit. Tests do not mock the model or any other part of the agent.
- History, not mocks, keeps tests cheap. Only the turn under test runs.
- A test touches the product in three places only: inputs through app routes, mocked third-party APIs, and what people and the model see.
- A refactor that keeps behavior does not change test assertions.
Where tests live
- Every test that runs the agent is an eval in
packages/junior-evals. Must-pass tests are integration evals inevals/integration/. Scored tests are behavioral evals. - The fixture is in
packages/junior-evals/src/fixture/. packages/junior/tests/keeps the tests that do not run the agent, such as API routes, storage, and ingress parsing. These tests fail on any AI Gateway request.- Guardian and router evals do not change.
Contract
// Test context
run(input: Input | Input[], options?: Options): Promise<Conversation>;
agent(options: JuniorAppOptions): Promise<{ run: typeof run }>;
type Options = {
history?: HistoryItem[] | RecordedConversation;
criteria?: Rubric;
onProgress?: (
progress: TurnProgress,
actions: { send(input: Input): Promise<void> },
) => void | Promise<void>;
};
// Returned by every call. The fields describe that call only.
type Conversation = {
conversationId: string;
replies: Reply[]; // assistant messages that people saw
toolCalls: ToolCall[]; // tool calls of the agent, with their results
reactions: string[]; // Slack reactions that Junior added
turns: Turn[]; // each Turn in order: status, replies, toolCalls
evalRun: HarnessRun; // the vitest-evals run, for other judges
continue(input: Input | Input[], options?: Options): Promise<Conversation>;
fork(reply: Reply | HistoryReply): Promise<Conversation>;
};
type TurnProgress =
| { type: "model_request" }
| { type: "tool_request"; name: string; args: unknown }
| { type: "reply"; text: string } // Slack Conversations only
| { type: "paused" };
Rules
run()starts a new Conversation on the agent of the test and sends the input. Tworun()calls never share a Conversation.conversation.continue()sends the next input to the same Conversation. A returned value is not a snapshot.continue()on an earlier value continues the Conversation at its current state.conversation.fork(reply)calls the forks API route at that reply and returns the fork as a new Conversation. It does not run a turn. The reply comes fromconversation.repliesor from areply()history item.- A call returns when the agent is idle: the queue is drained, the replies are delivered, and the work that the turns started, such as titles, is finished. A call fails when the agent is not idle within 60 seconds.
- All Conversations of a test use the same agent. Data outside a Conversation, such as memories and automations, stays for the whole test.
turn.statususes the turn states of the reporting API:started,succeeded,no_reply, orfailed. A turn that waits for authorization staysstarted.agent(options)takes the same options ascreateApp()and returns the samerun. The default agent iscreateApp()with the default options.
Inputs
An input says what reached Junior. The fixture sends it through the app route that production uses.
| Input | App route |
|---|---|
slackMention(text), slackThreadMessage(text) |
Slack Events API webhook, with a valid signature |
webMessage(text) |
POST /api/conversations or POST /api/conversations/:id/messages |
heartbeat() |
the heartbeat route, which runs the due automations |
githubWebhook(...), event(...) |
the provider webhook route, or Event ingest |
completeAuth(provider) |
the OAuth or MCP OAuth callback route |
[first, second] |
both inputs arrive before the worker runs, as one mailbox batch |
run(slackMention(text))posts to a new Slack thread. Builder options set the author and the channel type, for example a direct message.continue(webMessage(text))on a Slack Conversation is a dashboard continue.run(heartbeat())returns the Conversation that the due automation started.- When a turn asks for authorization, the prompt is in the Slack mock.
continue(completeAuth(provider))completes the flow and returns the resumed turn.
onProgress
The test reacts to what the turn does. It does not choose a point in the turn by itself.
model_request: the agent sent a model request. The request waits at the AI Gateway until the handler finishes.tool_request: the model asked for a tool. The response reaches the agent after the handler finishes, so a sent input arrives before the tool runs.reply: Junior delivered a Slack reply. The Slack mock responds after the handler finishes. Web Conversations have no delivery, so they never reportreply.paused: a turn stopped before it finished, for example at its deadline, and Junior queued the rest of it. The rest waits in the queue until the handler finishes, so a sent input arrives before the turn continues.send(input)posts the input through its app route and returns when the input is in the mailbox. Unlikecontinue(), it does not wait for the agent to be idle.- The product decides what a sent input does: it steers the running turn, it waits as a follow-up turn, or it stops the turn. The test asserts the outcome.
- The fixture sees progress in the AI Gateway traffic, the Slack mock, and the queue that it replaces. It does not hook
AgentRun.onEventor other runtime internals.
History
historyonrun()orcontinue()loads earlier turns as stored data before the input. Loading never runs the agent or calls the model.- Items use the existing builders:
slackMention,slackThreadMessage,reply(with optional tool history), andwebMessage. - The loader writes rows with the product functions that turns use:
recordActivity, the turn lifecycle service, andcommitAcceptedReply. - For Slack, the loader adds the messages to the Slack mock. It subscribes the thread when the history contains a reply.
historyalso accepts a sanitized recorded conversation: JSON event rows exported from a real Conversation. The loader inserts them as they are, ascopyForkEventsdoes. The export drops the event types that forks do not copy, and it replaces user ids, emails, team ids, and names.- Loaded history is not part of a call result.
- One test compares loaded history with the event rows of a real turn, without ids, timestamps, sequence numbers, and model usage. One test forks at a loaded reply. One test runs a turn on every recorded conversation.
Setup data
Each kind of setup data has one insert function. It writes through the product store function for that kind.
| Function | Writes with |
|---|---|
insertScheduledAutomation({ destination, ... }) |
saveScheduledAutomation |
insertEventAutomation({ destination, ... }) |
createEventAutomation |
insertWatch({ conversation, ... }) |
createWatch |
insertMemory(...) |
the memory plugin store |
insertCredential(...) |
the credential token store |
Insert functions only write data. They never run turns, and they contain no assertions.
Judges and assertions
- With
criteria, a call scores its replies withRubricJudgeand fails the test below the threshold. The judge receives the user-visible text of the Conversation so far, and it scores the assistant messages of that call. expect(conversation.evalRun).toSatisfyJudge(...)runs any other judge.- Exact assertions check facts that do not depend on wording: reply counts, turn states, tool calls, Slack calls, and stored effects read through Junior APIs. A judge checks wording.
- Tests do not assert on stored rows, turn records,
AgentRunobjects, or idempotency keys. - After each call, the fixture records every Conversation of the test on
task.meta.harness, so eval reports show the full Conversation.
What is real and what is mocked
| Real | the agent, the model, Guardian, the turn router, titles, the reply policy, compaction, Postgres, Redis |
| Mocked with MSW | Slack, GitHub, MCP fixtures, OAuth providers, image generation |
| Replayed | webFetch, webSearch |
| Replaced in-process | the Vercel queue transport, waitUntil |
Examples
test("recalls the earlier ask", async ({ run }) => {
const conversation = await run(slackMention("what did i just ask?"), {
history: [
slackMention("I need the budget by Friday."),
reply("Got it: budget due Friday."),
],
criteria: rubric({ pass: ["Recalls the budget and the Friday deadline."] }),
});
expect(conversation.replies).toHaveLength(1);
const next = await conversation.continue(slackMention("now draft the email"));
expect(next.replies).toHaveLength(1);
});
test("stop interrupts the running turn", async ({ run }) => {
const conversation = await run(slackMention("summarize the whole channel"), {
onProgress: async (progress, { send }) => {
if (progress.type === "model_request") {
await send(slackThreadMessage("stop"));
}
},
});
expect(conversation.replies.at(-1)?.text).toContain("stay out of this thread");
});
test("posts the digest when the automation is due", async ({ run }) => {
await insertScheduledAutomation({
destination: slackChannel(),
task: "Post the weekly digest.",
});
const digest = await run(heartbeat());
expect(digest.replies).toHaveLength(1);
});
test("a fork continues from the chosen reply", async ({ run }) => {
const decision = reply("We picked the blue option.");
const source = await run(webMessage("which option did we pick?"), {
history: [webMessage("Pick an option for the launch."), decision],
});
const fork = await source.fork(decision);
const next = await fork.continue(webMessage("try green instead"));
expect(next.replies).toHaveLength(1);
});
test("compacts history that does not fit the context window", async ({ agent }) => {
const { run } = await agent({ limits: { contextWindowTokens: SMALL_CONTEXT_WINDOW } });
const conversation = await run(slackMention("what did we decide for the launch?"), {
history: launchPlanningHistory, // larger than SMALL_CONTEXT_WINDOW
criteria: rubric({ pass: ["Gives the launch decision from the history."] }),
});
expect(conversation.turns[0].status).toBe("succeeded");
});
Product changes
createApp()accepts alimitsoption for the turn timeout, tool-call limit, slice limit, context window, and consecutive automated-turn limit, and a Slack option for the cross-actor steering mode. Without an option, each value keeps its environment default.- Each
createApp()call sets every runtime config value from its options or the default. A later app in the same process cannot inherit the plugins, profiles, or limits of an earlier app. - The app uses the
conversationWorkqueue in every place. Today the spawn-agent binding and the heartbeat callgetVercelConversationWorkQueue()directly. createApp()accepts a plugin task queue, as it acceptsconversationWorkQueue. Today a completed turn always sends plugin tasks to the Vercel queue. The fixture cannot replace that queue, so passive memory extraction never runs in fixture tests. The runtime already has asendPluginTaskoption, butcreateApp()does not expose it. With the option, the fixture runs plugin tasks in process, and a call is idle only after they finish.- Title work gets an owner. The worker that persists messages gives the title promise to
waitUntil. Today the work starts withvoid, so tests poll, and some title writes land after teardown. - Remove code that exists only for tests: the
streamFnparameter ofexecuteAgentRunandcreateAgentRunner, and the*ForTestsreset functions.
Users see no change in agent behavior.
Removed
- The eval
run()input:initialEvents,events,steer(),overrides, andpreloadHistory. - The old fixtures and model fakes:
createAgent,createConversationWebHarness,createConversationWorkSlackHarness,createTestChatRuntime,createModelStream,streamReplies,streamScript, the agent-runner helpers,mockAnthropicStream,mockTitleModel, andmockTurnRouterModel. - The eight component suites that mock the Pi agent with
MockAgent.
Test rules
scripts/check-test-architecture.mjs scans packages/junior/tests/ and packages/junior-evals/. Each rule starts with a list of the files that break it today. Each list can only get shorter.
- Tests in
packages/junior/tests/do not run the agent. - Agent tests import only the fixture, public types, and test libraries.
- No model fakes: no
vi.mockof@earendil-works/pi-agent-coreor@/chat/pi/client, noclass MockAgent, nocreateFauxCoreorfauxAssistantMessage, no inlinecompleteObject:orcompleteText:fakes, and no AI Gateway MSW handlers. - No
setPlugins(,Object.assign(botConfig, orvi.resetModules(in tests. - No
processConversationQueueMessage(,createSlackRuntime(, orcreateConversationWork(outside the fixture. anti-slop/no-module-mockingis on for integration and component tests.
Implementation order
- Add the test rules with today's lists. Give title work an owner. Add
errorlisteners to the Postgres pools. - Change
createApp(): limit options, a full config reset on each call, and one queue source. - Build the fixture:
run(),continue(),fork(),agent(), the inputs,onProgress, the call results,criteria, and reporting. - Add history loading, recorded conversations, the insert functions, and the loader tests.
- First slice: move
fork.test.ts, the Slack steering tests,continuity.eval.ts, and the scheduler credentials eval. Add the missing fork cases. Measure model cost, run time, and flakes. - Decide the open questions with those numbers.
- Move the other tests that run the agent. Replace the
MockAgentsuites. Remove the old fixtures, the model fakes, andstreamFn. - Add the plugin task queue option to
createApp()and run plugin tasks in the fixture. Do this before the memory evals move, because they check memories that passive extraction writes. - Move the remaining conversation evals.
- Update
policies/testing.md,policies/evals.md, and the test READMEs.
Open questions
- When do integration evals run on pull requests? Today the eval workflows skip pull requests that change only product code. Recommendation: run integration evals on each pull request that changes agent code. Keep behavioral evals on labels and their current path filters.
- How do tests cover provider errors, gateway timeouts, and empty answers? A real model does not produce them on request. Recommendation: a provider error or timeout is a third-party failure, so a test can inject it at the AI Gateway, as
queueSlackApiErrordoes for Slack. A test never scripts a model answer. The empty-answer retry stays a unit test of its rule. Decided:rejectNextModelRequest()makes the AI Gateway reject one agent request with a provider error. One eval checks the failure reply and the next turn. The messages for each kind of provider error stay unit tests. - Do agent tests get a Vercel Sandbox? The real agent can choose sandbox tools, and the sandbox needs the Quick Tunnel. In the last 400 runs of each eval workflow, 37 shards failed in the Quick Tunnel setup. Recommendation: no sandbox by default. A test that needs sandbox tools asks for a sandbox in its agent options. First check that the agent offers no sandbox tools when no sandbox is configured.
- Should this fixture move into vitest-evals later? This plan builds it in Junior first.
Out of scope
- Removing the module-level config globals. Each
createApp()call resets them, which is enough while tests run one app at a time per file. - Product gaps found in the same audit. They have their own issues:
- #2002: event automation dispatches never count toward the automated-turn limit.
- #2003: the chat README still says three Guardian rejections interrupt the execution slice.
- #2004: stale test references in the task-execution and runtime READMEs.
- Lingua principale
- TypeScript
- Stelle
- 367
- Fork
- 41
- Merge medio
- 6h 9m
- PR unite (30g)
- 188
Preparare l'ambiente
- Include un Dockerfile o un file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di getsentry/junior
-
documentation
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 64/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
getsentry/junior#335 · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di getsentry/junior
Issue simili
-
Difficoltà 1/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 1 giorno
-
core
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
vectorize-io/hindsight#5457 ·
I maintainer di solito rispondono entro 1 giorno
-
beginner friendly community contributions-welcome good first issue hacktoberfest help wanted testing up-for-grabs
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
lukilabs/beautiful-mermaid#160 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 66/100
rescript-lang/rescript-lang.org#1420 ·
I maintainer di solito rispondono entro 2 giorni