Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

One fixture for agent tests: run(), continue(), and loaded history

Abierto
#2,001 2 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
32/100
Tipo de issue
Nueva funcionalidad
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
postgres, redis, typescript
Área
testing

Línea de trabajo

Start in packages/junior-evals/src/fixture/ and compare the existing evals in evals/integration/ with the run, agent, history, continuation, fork, and setup-data contract. Check packages/junior/tests/ to preserve the boundary for tests that do not run the agent. Done means the fixture supports the documented real-agent flows, history loading, progress handling, judges, and Conversation reporting without changing Guardian or router evals.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Tests that run the agent should use one small fixture. Every turn under test runs the real agent with the real model. History, not mocks, keeps the tests cheap. This issue defines the fixture contract so that implementation can start.

This replaces the deterministic fake-model layer from #1425. It also changes the harness rules in #1398, which fake the model stream.

Why

Tests that run the agent depend on runtime internals. Routine refactors rewrite them, so they stop catching regressions.

  • In the last six months, 61% of commits that changed agent code also changed these tests. In ref: commits, 40% changed existing assertions. One reshape of AgentRun (c0fee35ff) changed 33 test files.
  • There are 8 ways to set up a run, 11 ways to fake the model, and 5 or more ways to load history.
  • Tests that fake the model still send about 320 live AI Gateway requests per run, because nothing fakes the turn router, titles, or the reply policy. With a Vercel login, these are real billed calls of 1.7–7 seconds, and they caused most local timeouts.
  • Eval run() allows one scenario per test. Messages in events call handleNewMention directly, so later eval messages skip the mailbox path that production uses.
  • preloadHistory writes four stores by hand. It does not write turn lifecycle events or the message:<id>:agent key, so a fork of preloaded history fails.

Principles

  • The agent is one unit. Tests do not mock the model or any other part of the agent.
  • History, not mocks, keeps tests cheap. Only the turn under test runs.
  • A test touches the product in three places only: inputs through app routes, mocked third-party APIs, and what people and the model see.
  • A refactor that keeps behavior does not change test assertions.

Where tests live

  • Every test that runs the agent is an eval in packages/junior-evals. Must-pass tests are integration evals in evals/integration/. Scored tests are behavioral evals.
  • The fixture is in packages/junior-evals/src/fixture/.
  • packages/junior/tests/ keeps the tests that do not run the agent, such as API routes, storage, and ingress parsing. These tests fail on any AI Gateway request.
  • Guardian and router evals do not change.

Contract

// Test context
run(input: Input | Input[], options?: Options): Promise<Conversation>;
agent(options: JuniorAppOptions): Promise<{ run: typeof run }>;

type Options = {
  history?: HistoryItem[] | RecordedConversation;
  criteria?: Rubric;
  onProgress?: (
    progress: TurnProgress,
    actions: { send(input: Input): Promise<void> },
  ) => void | Promise<void>;
};

// Returned by every call. The fields describe that call only.
type Conversation = {
  conversationId: string;
  replies: Reply[]; // assistant messages that people saw
  toolCalls: ToolCall[]; // tool calls of the agent, with their results
  reactions: string[]; // Slack reactions that Junior added
  turns: Turn[]; // each Turn in order: status, replies, toolCalls
  evalRun: HarnessRun; // the vitest-evals run, for other judges
  continue(input: Input | Input[], options?: Options): Promise<Conversation>;
  fork(reply: Reply | HistoryReply): Promise<Conversation>;
};

type TurnProgress =
  | { type: "model_request" }
  | { type: "tool_request"; name: string; args: unknown }
  | { type: "reply"; text: string } // Slack Conversations only
  | { type: "paused" };
Rules
  • run() starts a new Conversation on the agent of the test and sends the input. Two run() calls never share a Conversation.
  • conversation.continue() sends the next input to the same Conversation. A returned value is not a snapshot. continue() on an earlier value continues the Conversation at its current state.
  • conversation.fork(reply) calls the forks API route at that reply and returns the fork as a new Conversation. It does not run a turn. The reply comes from conversation.replies or from a reply() history item.
  • A call returns when the agent is idle: the queue is drained, the replies are delivered, and the work that the turns started, such as titles, is finished. A call fails when the agent is not idle within 60 seconds.
  • All Conversations of a test use the same agent. Data outside a Conversation, such as memories and automations, stays for the whole test.
  • turn.status uses the turn states of the reporting API: started, succeeded, no_reply, or failed. A turn that waits for authorization stays started.
  • agent(options) takes the same options as createApp() and returns the same run. The default agent is createApp() with the default options.
Inputs

An input says what reached Junior. The fixture sends it through the app route that production uses.

Input App route
slackMention(text), slackThreadMessage(text) Slack Events API webhook, with a valid signature
webMessage(text) POST /api/conversations or POST /api/conversations/:id/messages
heartbeat() the heartbeat route, which runs the due automations
githubWebhook(...), event(...) the provider webhook route, or Event ingest
completeAuth(provider) the OAuth or MCP OAuth callback route
[first, second] both inputs arrive before the worker runs, as one mailbox batch
  • run(slackMention(text)) posts to a new Slack thread. Builder options set the author and the channel type, for example a direct message.
  • continue(webMessage(text)) on a Slack Conversation is a dashboard continue.
  • run(heartbeat()) returns the Conversation that the due automation started.
  • When a turn asks for authorization, the prompt is in the Slack mock. continue(completeAuth(provider)) completes the flow and returns the resumed turn.
onProgress

The test reacts to what the turn does. It does not choose a point in the turn by itself.

  • model_request: the agent sent a model request. The request waits at the AI Gateway until the handler finishes.
  • tool_request: the model asked for a tool. The response reaches the agent after the handler finishes, so a sent input arrives before the tool runs.
  • reply: Junior delivered a Slack reply. The Slack mock responds after the handler finishes. Web Conversations have no delivery, so they never report reply.
  • paused: a turn stopped before it finished, for example at its deadline, and Junior queued the rest of it. The rest waits in the queue until the handler finishes, so a sent input arrives before the turn continues.
  • send(input) posts the input through its app route and returns when the input is in the mailbox. Unlike continue(), it does not wait for the agent to be idle.
  • The product decides what a sent input does: it steers the running turn, it waits as a follow-up turn, or it stops the turn. The test asserts the outcome.
  • The fixture sees progress in the AI Gateway traffic, the Slack mock, and the queue that it replaces. It does not hook AgentRun.onEvent or other runtime internals.
History
  • history on run() or continue() loads earlier turns as stored data before the input. Loading never runs the agent or calls the model.
  • Items use the existing builders: slackMention, slackThreadMessage, reply (with optional tool history), and webMessage.
  • The loader writes rows with the product functions that turns use: recordActivity, the turn lifecycle service, and commitAcceptedReply.
  • For Slack, the loader adds the messages to the Slack mock. It subscribes the thread when the history contains a reply.
  • history also accepts a sanitized recorded conversation: JSON event rows exported from a real Conversation. The loader inserts them as they are, as copyForkEvents does. The export drops the event types that forks do not copy, and it replaces user ids, emails, team ids, and names.
  • Loaded history is not part of a call result.
  • One test compares loaded history with the event rows of a real turn, without ids, timestamps, sequence numbers, and model usage. One test forks at a loaded reply. One test runs a turn on every recorded conversation.
Setup data

Each kind of setup data has one insert function. It writes through the product store function for that kind.

Function Writes with
insertScheduledAutomation({ destination, ... }) saveScheduledAutomation
insertEventAutomation({ destination, ... }) createEventAutomation
insertWatch({ conversation, ... }) createWatch
insertMemory(...) the memory plugin store
insertCredential(...) the credential token store

Insert functions only write data. They never run turns, and they contain no assertions.

Judges and assertions
  • With criteria, a call scores its replies with RubricJudge and fails the test below the threshold. The judge receives the user-visible text of the Conversation so far, and it scores the assistant messages of that call.
  • expect(conversation.evalRun).toSatisfyJudge(...) runs any other judge.
  • Exact assertions check facts that do not depend on wording: reply counts, turn states, tool calls, Slack calls, and stored effects read through Junior APIs. A judge checks wording.
  • Tests do not assert on stored rows, turn records, AgentRun objects, or idempotency keys.
  • After each call, the fixture records every Conversation of the test on task.meta.harness, so eval reports show the full Conversation.
What is real and what is mocked
Real the agent, the model, Guardian, the turn router, titles, the reply policy, compaction, Postgres, Redis
Mocked with MSW Slack, GitHub, MCP fixtures, OAuth providers, image generation
Replayed webFetch, webSearch
Replaced in-process the Vercel queue transport, waitUntil

Examples

test("recalls the earlier ask", async ({ run }) => {
  const conversation = await run(slackMention("what did i just ask?"), {
    history: [
      slackMention("I need the budget by Friday."),
      reply("Got it: budget due Friday."),
    ],
    criteria: rubric({ pass: ["Recalls the budget and the Friday deadline."] }),
  });
  expect(conversation.replies).toHaveLength(1);

  const next = await conversation.continue(slackMention("now draft the email"));
  expect(next.replies).toHaveLength(1);
});

test("stop interrupts the running turn", async ({ run }) => {
  const conversation = await run(slackMention("summarize the whole channel"), {
    onProgress: async (progress, { send }) => {
      if (progress.type === "model_request") {
        await send(slackThreadMessage("stop"));
      }
    },
  });
  expect(conversation.replies.at(-1)?.text).toContain("stay out of this thread");
});

test("posts the digest when the automation is due", async ({ run }) => {
  await insertScheduledAutomation({
    destination: slackChannel(),
    task: "Post the weekly digest.",
  });
  const digest = await run(heartbeat());
  expect(digest.replies).toHaveLength(1);
});

test("a fork continues from the chosen reply", async ({ run }) => {
  const decision = reply("We picked the blue option.");
  const source = await run(webMessage("which option did we pick?"), {
    history: [webMessage("Pick an option for the launch."), decision],
  });
  const fork = await source.fork(decision);
  const next = await fork.continue(webMessage("try green instead"));
  expect(next.replies).toHaveLength(1);
});

test("compacts history that does not fit the context window", async ({ agent }) => {
  const { run } = await agent({ limits: { contextWindowTokens: SMALL_CONTEXT_WINDOW } });
  const conversation = await run(slackMention("what did we decide for the launch?"), {
    history: launchPlanningHistory, // larger than SMALL_CONTEXT_WINDOW
    criteria: rubric({ pass: ["Gives the launch decision from the history."] }),
  });
  expect(conversation.turns[0].status).toBe("succeeded");
});

Product changes

  • createApp() accepts a limits option for the turn timeout, tool-call limit, slice limit, context window, and consecutive automated-turn limit, and a Slack option for the cross-actor steering mode. Without an option, each value keeps its environment default.
  • Each createApp() call sets every runtime config value from its options or the default. A later app in the same process cannot inherit the plugins, profiles, or limits of an earlier app.
  • The app uses the conversationWork queue in every place. Today the spawn-agent binding and the heartbeat call getVercelConversationWorkQueue() directly.
  • createApp() accepts a plugin task queue, as it accepts conversationWorkQueue. Today a completed turn always sends plugin tasks to the Vercel queue. The fixture cannot replace that queue, so passive memory extraction never runs in fixture tests. The runtime already has a sendPluginTask option, but createApp() does not expose it. With the option, the fixture runs plugin tasks in process, and a call is idle only after they finish.
  • Title work gets an owner. The worker that persists messages gives the title promise to waitUntil. Today the work starts with void, so tests poll, and some title writes land after teardown.
  • Remove code that exists only for tests: the streamFn parameter of executeAgentRun and createAgentRunner, and the *ForTests reset functions.

Users see no change in agent behavior.

Removed

  • The eval run() input: initialEvents, events, steer(), overrides, and preloadHistory.
  • The old fixtures and model fakes: createAgent, createConversationWebHarness, createConversationWorkSlackHarness, createTestChatRuntime, createModelStream, streamReplies, streamScript, the agent-runner helpers, mockAnthropicStream, mockTitleModel, and mockTurnRouterModel.
  • The eight component suites that mock the Pi agent with MockAgent.

Test rules

scripts/check-test-architecture.mjs scans packages/junior/tests/ and packages/junior-evals/. Each rule starts with a list of the files that break it today. Each list can only get shorter.

  • Tests in packages/junior/tests/ do not run the agent.
  • Agent tests import only the fixture, public types, and test libraries.
  • No model fakes: no vi.mock of @earendil-works/pi-agent-core or @/chat/pi/client, no class MockAgent, no createFauxCore or fauxAssistantMessage, no inline completeObject: or completeText: fakes, and no AI Gateway MSW handlers.
  • No setPlugins(, Object.assign(botConfig, or vi.resetModules( in tests.
  • No processConversationQueueMessage(, createSlackRuntime(, or createConversationWork( outside the fixture.
  • anti-slop/no-module-mocking is on for integration and component tests.

Implementation order

  • Add the test rules with today's lists. Give title work an owner. Add error listeners to the Postgres pools.
  • Change createApp(): limit options, a full config reset on each call, and one queue source.
  • Build the fixture: run(), continue(), fork(), agent(), the inputs, onProgress, the call results, criteria, and reporting.
  • Add history loading, recorded conversations, the insert functions, and the loader tests.
  • First slice: move fork.test.ts, the Slack steering tests, continuity.eval.ts, and the scheduler credentials eval. Add the missing fork cases. Measure model cost, run time, and flakes.
  • Decide the open questions with those numbers.
  • Move the other tests that run the agent. Replace the MockAgent suites. Remove the old fixtures, the model fakes, and streamFn.
  • Add the plugin task queue option to createApp() and run plugin tasks in the fixture. Do this before the memory evals move, because they check memories that passive extraction writes.
  • Move the remaining conversation evals.
  • Update policies/testing.md, policies/evals.md, and the test READMEs.

Open questions

  1. When do integration evals run on pull requests? Today the eval workflows skip pull requests that change only product code. Recommendation: run integration evals on each pull request that changes agent code. Keep behavioral evals on labels and their current path filters.
  2. How do tests cover provider errors, gateway timeouts, and empty answers? A real model does not produce them on request. Recommendation: a provider error or timeout is a third-party failure, so a test can inject it at the AI Gateway, as queueSlackApiError does for Slack. A test never scripts a model answer. The empty-answer retry stays a unit test of its rule. Decided: rejectNextModelRequest() makes the AI Gateway reject one agent request with a provider error. One eval checks the failure reply and the next turn. The messages for each kind of provider error stay unit tests.
  3. Do agent tests get a Vercel Sandbox? The real agent can choose sandbox tools, and the sandbox needs the Quick Tunnel. In the last 400 runs of each eval workflow, 37 shards failed in the Quick Tunnel setup. Recommendation: no sandbox by default. A test that needs sandbox tools asks for a sandbox in its agent options. First check that the agent offers no sandbox tools when no sandbox is configured.
  4. Should this fixture move into vitest-evals later? This plan builds it in Junior first.

Out of scope

  • Removing the module-level config globals. Each createApp() call resets them, which is enough while tests run one app at a time per file.
  • Product gaps found in the same audit. They have their own issues:
    • #2002: event automation dispatches never count toward the automated-turn limit.
    • #2003: the chat README still says three Guardian rejections interrupt the execution slice.
    • #2004: stale test references in the task-execution and runtime READMEs.
Lenguaje dominante
TypeScript
Estrellas
367
Forks
41
Merge medio
6 h 9 min
PR fusionados (30 d)
188

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de getsentry/junior

Todos los issues de getsentry/junior

Issues similares

Más issues de TypeScript

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.