Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

TaskOutput corrupts valid UTF-8 when the tail preview starts inside a multibyte character

Open Beginner friendly
#4,140 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
85/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
node.js, typescript
Domain
backend

Research direction

Start from readTaskOutputSnapshot in packages/agent-core-v2/src/agent/task/persist.ts (L130-149) and the in-memory fallback in taskService.ts (L550-582): both slice at a raw byte offset before decoding, as does taskOutputTool.ts's 32 KiB request. Run the reporter's repro script (.tmp/repro-task-output-utf8.ts) to observe the replacement characters. Done means the preview starts on a code-point boundary, snapshot byte metadata matches the retained source bytes, and the persistence/output-access tests gain CJK, emoji, aligned-boundary, small-budget and ASCII cases.

Written by the indexing model from the issue text.

Description

What version of Kimi Code is running?

Source checkout of current main: 21406fb4c805cc8c715e6d1f16ad3fb5f25f4fe3. apps/kimi-code/package.json reports 2.1.1.

This is a source-level reproduction using the actual persistence implementation and repository test harness, not an installed-CLI/model end-to-end run. I rechecked upstream main before reporting.

Which open platform/subscription were you using?

Not applicable. No login, provider requests, or paid services are needed.

Which model were you using?

Not applicable. The defect occurs while constructing the task-output preview, before model inference.

What platform is your computer?

Windows x64 (Microsoft Windows NT 10.0.26200.0), Node.js v24.15.0, pnpm 10.33.0.

What issue are you seeing?

TaskOutput can insert replacement characters into a preview of a valid UTF-8 task log. Its 32 KiB tail window is decoded starting at an arbitrary byte offset, which can fall inside a multibyte character.

For example, a task output containing "中".repeat(10923) is 32,769 bytes. The 32,768-byte preview starts at byte 1 of the first three-byte character, and its text starts with ��中 instead of an intact suffix of the original output. The persisted output.log remains correct; corruption is introduced by preview construction.

This affects the text exposed through TaskOutput to the agent. It also reproduces in the in-memory preview fallback and after reopening persisted output. The full log can still be read separately, so this report is about preview corruption, not permanent loss of the log.

What steps can reproduce the bug?
  1. Check out the commit above and run pnpm install --frozen-lockfile with the supported Node/pnpm versions.
  2. Save the following as .tmp/repro-task-output-utf8.ts in the repository root.
  3. Run pnpm exec tsx .tmp/repro-task-output-utf8.ts.
import assert from 'node:assert/strict';
import { randomUUID } from 'node:crypto';
import { readFile } from 'node:fs/promises';
import { join } from 'node:path';
import { AgentTaskPersistence } from '../packages/agent-core-v2/src/agent/task/persist';
import { JsonAtomicDocumentStore } from '../packages/agent-core-v2/src/persistence/backends/node-fs/atomicDocumentStore';
import { FileStorageService } from '../packages/agent-core-v2/src/persistence/backends/node-fs/fileStorageService';

const root = join(process.cwd(), '.tmp', 'utf8-repro', randomUUID());
const scope = 'session/agents/main';
const storage = new FileStorageService(root);
const makePersistence = () => new AgentTaskPersistence(
  join(root, scope), scope, new JsonAtomicDocumentStore(storage), storage,
);
const persistence = makePersistence();
const taskId = 'bash-11111111';
const original = '中'.repeat(10923);
await persistence.appendTaskOutput(taskId, original);
assert.equal(await readFile(persistence.taskOutputFile(taskId), 'utf8'), original);
const snapshot = await persistence.readTaskOutputSnapshot(taskId, 32 * 1024);
assert.ok(snapshot);
assert.deepEqual(await makePersistence().readTaskOutputSnapshot(taskId, 32 * 1024), snapshot);
console.log(JSON.stringify({
  outputSizeBytes: snapshot.outputSizeBytes,
  previewBytes: snapshot.previewBytes,
  previewEncodedBytes: Buffer.byteLength(snapshot.preview),
  prefix: [...snapshot.preview].slice(0, 3).join(''),
  fullLogIntact: true,
  sameAfterReopen: true,
}));
assert.equal(snapshot.preview.startsWith('\uFFFD\uFFFD'), true);

Actual output (the assertion intentionally confirms the current defect):

{"outputSizeBytes":32769,"previewBytes":32768,"previewEncodedBytes":32772,"prefix":"��中","fullLogIntact":true,"sameAfterReopen":true}

Additional local observations using the same real persistence implementation:

Input Output bytes Replacement characters in preview Full log intact
"中".repeat(10923) 32769 2 Yes
"😀".repeat(8192) + "x" 32769 3 Yes
"é".repeat(16384) + "x" 32769 1 Yes
"abc" + "中".repeat(10922) + "xy" (aligned boundary) 32771 0 Yes
"a".repeat(32769) 32769 0 Yes
"中😀é" (no truncation) 9 0 Yes

I also ran an observation test using the existing createTestAgent/task-service harness: a foreground ProcessTask supplied the same valid text, the live fallback preview began with ��, a fresh manager loaded the completed task and its persisted log, and the real TaskOutputTool result contained [output]\n��中. That observation test passed on unmodified main. It uses controlled Node streams, no provider, and no arbitrary sleeps. This is harness coverage, not a claim of a full CLI session reproduction.

What is the expected behavior?

A tail preview of valid UTF-8 should contain only complete code points from the original text. When the byte budget lands inside a character, omit that partial leading character rather than manufacturing �.

Keep the existing byte budget, full-log contents, output path, and ASCII behavior. Preview byte metadata should describe the actual retained source-byte range.

Additional information

The observed cause is the byte slice followed by independent decoding:

Existing snapshot tests cover truncation with ASCII, which cannot expose this boundary condition.

Related but different: #3756 concerns UTF-16 surrogate-pair slicing in generic tool-result truncation and the output accumulator. This report concerns UTF-8 byte slicing in task snapshots; the CJK reproduction contains no surrogate pairs and bypasses generic truncation entirely. #733 concerned v1 ring-buffer byte accounting, rather than this v2 preview-decoding path. I searched open/closed issues and PRs for TaskOutput, UTF-8, Unicode, multibyte, and preview truncation; I did not find an active fix for this specific path.

Proposed scope, subject to maintainer approval: adjust UTF-8 tail boundaries in persisted and in-memory task snapshots, keep byte metadata consistent, and extend the existing persistence/output-access tests with CJK, emoji, aligned-boundary, small-budget, and ASCII controls. No new feature, public API, storage format, or grapheme-cluster handling is proposed. Adjacent terminal-event/ring-buffer truncation sites can be discussed separately if needed; I have not claimed to reproduce every truncation path.

I am willing to implement this narrowly scoped fix after a maintainer's /approve. No production fix or PR has been prepared.

Contribution
  • I am willing to submit a PR for this bug fix myself (please wait for maintainer approval in this issue first)
Dominant language
TypeScript
Stars
7.8k
Forks
1.3k
Avg merge
17h 23m
Merged PRs (30d)
157

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from MoonshotAI/kimi-code

All issues in MoonshotAI/kimi-code

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.