Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug]: Logs mixes legacy and generation-window decode estimates across historical rows without identifying the timing basis

Open
#6,567 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
52/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
typescript

Research direction

Start with src/server/management/shared.ts and the generation timestamp recording in src/server/request-log.ts; inspect how persisted timing fields are used for individual Logs rows and compare that with the legacy and generation-window formulas in the report. Use the supplied scalar fixture to verify the displayed estimates. Done means rows identify their timing basis and distinguish end-to-end, generation-window, and legacy estimates without fabricating timestamps or rewriting historical data.

Written by the indexing model from the issue text.

Description

bug gui
Client or integration

OpenCodex dashboard

Area

Dashboard

Summary

Historical and new requests viewed together after upgrading from 2.76.0 to 2.77.0 can show materially different decode-rate estimates in the same Logs tok/s display position without identifying the timing basis. This makes a measurement change look like a severe generation-speed regression.

Expected: identify the timing basis per row and distinguish end-to-end throughput from both generation-window and legacy post-visible-output estimates. Do not fabricate missing timestamps, rewrite historical data, or restore inflated numbers.

This report does not claim that a real latency problem is merely cosmetic. Actual pre-output latency also worsened in this installation and is tracked separately in #6568. Correcting the display must not be counted as repairing that latency.

Reproduction
  1. Retain successful 2.76.0 usage rows without genStartMs and lastOutputMs.
  2. Upgrade to 2.77.0 and record successful rows with those generation-window timestamps.
  3. Display both periods in Logs.
  4. Compare the estimated tok/s value following the end-to-end rate.

The legacy estimate uses outputTokens / ((durationMs - firstOutputMs) / 1000). Where generation timestamps exist, the new estimate uses outputTokens / ((lastOutputMs - genStartMs) / 1000). The former window begins at visible output but its numerator includes reported reasoning output tokens.

The scalar fixture below reproduces the arithmetic on one request: 133.844 tok/s with the legacy formula, 21.804 tok/s with the generation-window formula, and 6.157 tok/s end-to-end. Therefore an old displayed 100+ tok/s and a new 20–35 tok/s are not automatically comparable decoder-throughput measurements.

Version

2.77.0, compared with retained historical usage from 2.76.0. Bun 1.4.0 and official Codex CLI 0.160.0.

Operating system

Windows 11 x64; exact Windows feature-update version was not included in this report.

Provider and model

Native OpenAI Codex Responses route, GPT-6.1 Sol, medium reasoning, default service tier. Existing account selection unchanged.

Logs or error output
{
  "durationMs": 36871,
  "firstOutputMs": 35175,
  "genStartMs": 26115,
  "lastOutputMs": 36526,
  "outputTokens": 227,
  "reasoningOutputTokens": 168
}

Applying the same visible-token/window calculation to both periods produced medians near 33–34 tok/s while first-output waiting increased. Those periods are observational, and the visible-token calculation is itself an estimate, not a provider-signed decode measurement.

Screenshots and supporting files

The exact scalar fixture above is sufficient for the arithmetic reproduction. Raw screenshots and usage records are omitted because they contain unrelated user information. Source paths inspected include src/server/management/shared.ts and generation timestamp recording in src/server/request-log.ts.

A separate version-only 2.77 → 2.76 → 2.77 test used the official ephemeral CLI in a read-only sandbox. It showed pre-output waits in both versions; the sample is too small to establish or exclude a version-level latency regression. No unknown-result business operation was replayed, and no model/effort downgrade, account rotation, retry-safety override or global Fast change was made.

Redacted configuration

The display issue concerns persisted timing fields, not a provider secret or custom routing rule. All compared requests used the same native model/effort and default service tier. No authentication values or user configuration dump are attached.

Related: #6309 concerns aggregated throughput reporting; this report concerns incompatible timing bases on individual historical Logs rows. Suggested remediation is a clear per-row timing-basis label, consistent exported/aggregated semantics, and separate visibility for first-output latency and end-to-end throughput.

Checks
  • I searched existing issues and documentation.
  • I removed secrets, tokens, account details, request credentials, and personal data.
Dominant language
TypeScript
Stars
16.9k
Forks
1.3k
Avg merge
4h 24m
Merged PRs (30d)
574

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from lidge-jun/opencodex

All issues in lidge-jun/opencodex

Similar issues

More TypeScript issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.