Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Async engine drops rows on stale keep-alive connection errors (`RemoteProtocolError` / `ReadError` classified as non-retryable `API_ERROR`)

Open
#1,003 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

@mikeknep is already working on this.

Since Oct 9, 2026.

  • #1004 by @mikeknep — open

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
22/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
python

Research direction

Start with infer_error_kind_from_exception in data_designer/engine/models/clients/errors.py, then ModelRequestExecutor._should_retry in model_request_executor.py, and _NO_TRANSPORT_RETRY_CONFIG in factory.py. Done means httpx.RemoteProtocolError and ReadError/WriteError lead to a retry on a new connection. Pull request #1004 is already open against this issue, so check it before starting.

Written by the indexing model from the issue text.

Description

bug
Priority Level

Medium (Annoying but has workaround)

Describe the bug

On the async task-queue path, an HTTP request that fails because the client reused a pooled keep-alive connection the server had just closed is treated as a non-retryable error, and the row is dropped.

This happens when the server closes an idle keep-alive connection right as the client sends a new request on it. httpx then raises httpx.RemoteProtocolError("Server disconnected without sending a response.") or httpx.ReadError. The request never reaches the application, so a retry on a fresh connection always succeeds.

Three pieces of data_designer.engine.models.clients combine to make this fatal:

  1. factory.py: when request_admission is set (async engine), the adapter gets _NO_TRANSPORT_RETRY_CONFIG (max_retries=0). That turns off httpx_retries.RetryTransport, whose default RETRYABLE_EXCEPTIONS already includes httpx.RemoteProtocolError and httpx.NetworkError, and which allows POST in this config. Before this change, these errors were retried at the transport layer.
  2. errors.py::infer_error_kind_from_exception classifies by exception type name. RemoteProtocolError, ReadError, and WriteError contain neither "timeout" nor "connect", so they fall through to ProviderErrorKind.API_ERROR.
  3. model_request_executor.py::ModelRequestExecutor._should_retry: for errors without an HTTP status, it only retries ProviderErrorKind.API_CONNECTION.

As a result, the request surfaces as ModelAPIError ("An unexpected API error occurred…"). The async scheduler logs Non-retryable failure on <column>[rg=0, row=N] and the row is missing from the output.

The window is easiest to hit when a model client sits idle for about the server's keep-alive timeout between requests. One case is health checks followed by a slow sampler, such as PersonSamplerParams loading a Nemotron Personas parquet. httpx's default keepalive_expiry is 5s, and so is uvicorn's default timeout_keep_alive. With an idle gap near 5s, it's a coin flip which side closes the connection first.

Observed with data-designer-engine==0.9.3, httpx==0.28.1. The code paths above are unchanged on main (88ecb598cd) and in v0.9.4.

Steps/Code to reproduce bug

1. Classification (deterministic):

import httpx
from data_designer.engine.models.clients.errors import infer_error_kind_from_exception

for exc in [
    httpx.RemoteProtocolError("Server disconnected without sending a response."),
    httpx.ReadError(""),
    httpx.WriteError(""),
    httpx.ConnectError(""),
]:
    print(type(exc).__name__, "->", infer_error_kind_from_exception(exc))

Output:

RemoteProtocolError -> ProviderErrorKind.API_ERROR
ReadError -> ProviderErrorKind.API_ERROR
WriteError -> ProviderErrorKind.API_ERROR
ConnectError -> ProviderErrorKind.API_CONNECTION

Only API_CONNECTION is retried by ModelRequestExecutor._should_retry.

2. These exceptions come from ordinary keep-alive expiry (timing-dependent):

This shows that a plain uvicorn server and an httpx client, idling for about the server's keep-alive timeout, produce exactly these exception types:

import asyncio, collections, threading
import httpx, uvicorn

async def app(scope, receive, send):
    if scope["type"] != "http":
        return
    await receive()
    await send({"type": "http.response.start", "status": 200, "headers": [(b"content-type", b"application/json")]})
    await send({"type": "http.response.body", "body": b"{}"})

async def main():
    server = uvicorn.Server(uvicorn.Config(app, port=18765, timeout_keep_alive=1, log_level="error"))
    threading.Thread(target=server.run, daemon=True).start()
    while not server.started:
        await asyncio.sleep(0.05)
    errs = collections.Counter()
    async with httpx.AsyncClient() as c:
        for i in range(200):
            await c.post("http://127.0.0.1:18765/")
            await asyncio.sleep(1.0 + (i % 40) * 0.0005 - 0.005)  # idle ~= server keep-alive
            try:
                await c.post("http://127.0.0.1:18765/")
            except Exception as e:
                errs[f"{type(e).__name__}: {e}"] += 1
    print("failures out of 200:", dict(errs))
    server.should_exit = True

asyncio.run(main())

Output from one run:

failures out of 200: {'ReadError: ': 12, 'RemoteProtocolError: Server disconnected without sending a response.': 1}

3. End-to-end symptom:

Our setup is an OpenAI-compatible provider served by uvicorn with the default 5s keep-alive. It has two LLMTextColumnConfig columns on separate models, plus a SamplerColumnConfig with SamplerType.PERSON. Preview runs with 10 records on the async engine. Model health checks complete, the person sampler takes about 5s to load its dataset, and then the first LLM request on one model's pooled connection fails:

Non-retryable failure on response_from_b[rg=0, row=0]:   |----------
  | Cause: An unexpected API error occurred with model 'model-b' while running generation for column 'response_from_b'.
  | Solution: Try again in a few moments. Check with your model provider 'igw-mock-test-provider' if the issue persists.
  |----------
✅ Async generation complete [5.1s]: 19 ok, 1 failed across 2 column(s)
  |-- model: model-b
  |-- requests: success=9, failed=1, total=10, rpm=117

Server access logs contain all 19 successful requests and no entry for the failed one. The request never reached the application. In both occurrences we captured, the gap between the model's last health-check request and the failure was about 5.0s. Runs where the sampler took longer, for example 11.4s, always pass: by then the pool has already discarded the expired connection.

Expected behavior

A request that fails because a pooled connection was closed by the peer before any response bytes arrived should be retried on a new connection, as it was before the async path disabled transport-level retries. The row should not be dropped.

Possible fixes, in rough order of preference:

  • Classify httpx.RemoteProtocolError and httpx.NetworkError (ReadError, WriteError) as ProviderErrorKind.API_CONNECTION in infer_error_kind_from_exception. A type-based check (isinstance(exc, (httpx.NetworkError, httpx.RemoteProtocolError))) would be more robust than the name-substring heuristic. That would let ModelRequestExecutor retry them under the existing RetryConfig budget.
  • Or have _should_retry treat these transport errors as retryable directly.
  • Optionally, use a client keepalive_expiry shorter than common server keep-alive defaults. That would make the race much less likely in the first place, though it can't remove it.

One caveat: a ReadError can in theory happen after the server has started processing a POST, so a retry could duplicate a generation. The previous transport-level retry policy already accepted that tradeoff (allowed_methods includes POST), and for LLM completions it seems far better than silently dropping the row.

Agent Diagnostic / Prior Investigation

An agent investigated this from NeMo Helix CI logs. A Data Designer e2e test uses a mock model provider with fixed responses, and it flaked twice with assert 9 == 10 on preview row count. Findings:

  • In both failures the failed request has no matching entry in the server's request log, while all other completions returned 200. The failure is client-side, before the request was handled.
  • The failing request was the first one sent on model-b's client after an idle gap of about 5.0s. That gap was the person sampler's dataset load, measured from the health-check request. Server keep-alive was uvicorn's default 5s; httpx's default keepalive_expiry is also 5s.
  • I traced the classification and retry path through infer_error_kind_from_exception → ProviderErrorKind.API_ERROR → ModelRequestExecutor._should_retry (returns False) → ModelAPIError → the async scheduler treats it as non-retryable (it isn't in PRESERVED_RETRYABLE_ERRORS).
  • I confirmed that create_model_client passes _NO_TRANSPORT_RETRY_CONFIG when request_admission is set. That disables httpx_retries, which would otherwise retry RemoteProtocolError / NetworkError on POST.
  • I reproduced the exception types with the standalone uvicorn/httpx script above: 13 of 200 failures with the idle time near the keep-alive timeout.
  • I searched existing issues for "RemoteProtocolError", "keep-alive retry", and "Non-retryable". Nothing matched. #948 is related: unmapped 4xx statuses are also reported as "unexpected API error".

On our side, we're working around it by raising the server keep-alive above httpx's 5s client expiry. The client then always retires the connection first.

Additional context
  • data-designer-engine 0.9.3 (code unchanged on main @ 88ecb598cd / v0.9.4)
  • httpx 0.28.1, httpx-retries (default RETRYABLE_EXCEPTIONS = TimeoutException, NetworkError, RemoteProtocolError), uvicorn 0.52.1
  • Python 3.13, Linux (GitHub Actions runner); reproduced the exception types locally on macOS
  • Async engine (⚡ Using async task-queue preview)
Checklist
  • I reproduced this issue or provided a minimal example
  • I searched the docs/issues myself, or had my agent do so
  • If I used an agent, I included its diagnostics above
Dominant language
Python
Stars
2.3k
Forks
219
Avg merge
1d 21h
Merged PRs (30d)
38

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from NVIDIA-NeMo/DataDesigner

All issues in NVIDIA-NeMo/DataDesigner

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.