Async engine drops rows on stale keep-alive connection errors (`RemoteProtocolError` / `ReadError` classified as non-retryable `API_ERROR`)
Maintainers usually reply within 1 day
Assessment
- Difficulty
- 3/5
- Estimated time
- 1-2 days
- Newbie friendliness
- 22/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- python
- Domain
- backend, networking
Research direction
Start with infer_error_kind_from_exception in data_designer/engine/models/clients/errors.py, then ModelRequestExecutor._should_retry in model_request_executor.py, and _NO_TRANSPORT_RETRY_CONFIG in factory.py. Done means httpx.RemoteProtocolError and ReadError/WriteError lead to a retry on a new connection. Pull request #1004 is already open against this issue, so check it before starting.
Written by the indexing model from the issue text.
Description
Priority Level
Medium (Annoying but has workaround)
Describe the bug
On the async task-queue path, an HTTP request that fails because the client reused a pooled keep-alive connection the server had just closed is treated as a non-retryable error, and the row is dropped.
This happens when the server closes an idle keep-alive connection right as the client sends a new request on it. httpx then raises httpx.RemoteProtocolError("Server disconnected without sending a response.") or httpx.ReadError. The request never reaches the application, so a retry on a fresh connection always succeeds.
Three pieces of data_designer.engine.models.clients combine to make this fatal:
factory.py: whenrequest_admissionis set (async engine), the adapter gets_NO_TRANSPORT_RETRY_CONFIG(max_retries=0). That turns offhttpx_retries.RetryTransport, whose defaultRETRYABLE_EXCEPTIONSalready includeshttpx.RemoteProtocolErrorandhttpx.NetworkError, and which allows POST in this config. Before this change, these errors were retried at the transport layer.errors.py::infer_error_kind_from_exceptionclassifies by exception type name.RemoteProtocolError,ReadError, andWriteErrorcontain neither"timeout"nor"connect", so they fall through toProviderErrorKind.API_ERROR.model_request_executor.py::ModelRequestExecutor._should_retry: for errors without an HTTP status, it only retriesProviderErrorKind.API_CONNECTION.
As a result, the request surfaces as ModelAPIError ("An unexpected API error occurred…"). The async scheduler logs Non-retryable failure on <column>[rg=0, row=N] and the row is missing from the output.
The window is easiest to hit when a model client sits idle for about the server's keep-alive timeout between requests. One case is health checks followed by a slow sampler, such as PersonSamplerParams loading a Nemotron Personas parquet. httpx's default keepalive_expiry is 5s, and so is uvicorn's default timeout_keep_alive. With an idle gap near 5s, it's a coin flip which side closes the connection first.
Observed with data-designer-engine==0.9.3, httpx==0.28.1. The code paths above are unchanged on main (88ecb598cd) and in v0.9.4.
Steps/Code to reproduce bug
1. Classification (deterministic):
import httpx
from data_designer.engine.models.clients.errors import infer_error_kind_from_exception
for exc in [
httpx.RemoteProtocolError("Server disconnected without sending a response."),
httpx.ReadError(""),
httpx.WriteError(""),
httpx.ConnectError(""),
]:
print(type(exc).__name__, "->", infer_error_kind_from_exception(exc))
Output:
RemoteProtocolError -> ProviderErrorKind.API_ERROR
ReadError -> ProviderErrorKind.API_ERROR
WriteError -> ProviderErrorKind.API_ERROR
ConnectError -> ProviderErrorKind.API_CONNECTION
Only API_CONNECTION is retried by ModelRequestExecutor._should_retry.
2. These exceptions come from ordinary keep-alive expiry (timing-dependent):
This shows that a plain uvicorn server and an httpx client, idling for about the server's keep-alive timeout, produce exactly these exception types:
import asyncio, collections, threading
import httpx, uvicorn
async def app(scope, receive, send):
if scope["type"] != "http":
return
await receive()
await send({"type": "http.response.start", "status": 200, "headers": [(b"content-type", b"application/json")]})
await send({"type": "http.response.body", "body": b"{}"})
async def main():
server = uvicorn.Server(uvicorn.Config(app, port=18765, timeout_keep_alive=1, log_level="error"))
threading.Thread(target=server.run, daemon=True).start()
while not server.started:
await asyncio.sleep(0.05)
errs = collections.Counter()
async with httpx.AsyncClient() as c:
for i in range(200):
await c.post("http://127.0.0.1:18765/")
await asyncio.sleep(1.0 + (i % 40) * 0.0005 - 0.005) # idle ~= server keep-alive
try:
await c.post("http://127.0.0.1:18765/")
except Exception as e:
errs[f"{type(e).__name__}: {e}"] += 1
print("failures out of 200:", dict(errs))
server.should_exit = True
asyncio.run(main())
Output from one run:
failures out of 200: {'ReadError: ': 12, 'RemoteProtocolError: Server disconnected without sending a response.': 1}
3. End-to-end symptom:
Our setup is an OpenAI-compatible provider served by uvicorn with the default 5s keep-alive. It has two LLMTextColumnConfig columns on separate models, plus a SamplerColumnConfig with SamplerType.PERSON. Preview runs with 10 records on the async engine. Model health checks complete, the person sampler takes about 5s to load its dataset, and then the first LLM request on one model's pooled connection fails:
Non-retryable failure on response_from_b[rg=0, row=0]: |----------
| Cause: An unexpected API error occurred with model 'model-b' while running generation for column 'response_from_b'.
| Solution: Try again in a few moments. Check with your model provider 'igw-mock-test-provider' if the issue persists.
|----------
✅ Async generation complete [5.1s]: 19 ok, 1 failed across 2 column(s)
|-- model: model-b
|-- requests: success=9, failed=1, total=10, rpm=117
Server access logs contain all 19 successful requests and no entry for the failed one. The request never reached the application. In both occurrences we captured, the gap between the model's last health-check request and the failure was about 5.0s. Runs where the sampler took longer, for example 11.4s, always pass: by then the pool has already discarded the expired connection.
Expected behavior
A request that fails because a pooled connection was closed by the peer before any response bytes arrived should be retried on a new connection, as it was before the async path disabled transport-level retries. The row should not be dropped.
Possible fixes, in rough order of preference:
- Classify
httpx.RemoteProtocolErrorandhttpx.NetworkError(ReadError,WriteError) asProviderErrorKind.API_CONNECTIONininfer_error_kind_from_exception. A type-based check (isinstance(exc, (httpx.NetworkError, httpx.RemoteProtocolError))) would be more robust than the name-substring heuristic. That would letModelRequestExecutorretry them under the existingRetryConfigbudget. - Or have
_should_retrytreat these transport errors as retryable directly. - Optionally, use a client
keepalive_expiryshorter than common server keep-alive defaults. That would make the race much less likely in the first place, though it can't remove it.
One caveat: a ReadError can in theory happen after the server has started processing a POST, so a retry could duplicate a generation. The previous transport-level retry policy already accepted that tradeoff (allowed_methods includes POST), and for LLM completions it seems far better than silently dropping the row.
Agent Diagnostic / Prior Investigation
An agent investigated this from NeMo Helix CI logs. A Data Designer e2e test uses a mock model provider with fixed responses, and it flaked twice with assert 9 == 10 on preview row count. Findings:
- In both failures the failed request has no matching entry in the server's request log, while all other completions returned 200. The failure is client-side, before the request was handled.
- The failing request was the first one sent on
model-b's client after an idle gap of about 5.0s. That gap was the person sampler's dataset load, measured from the health-check request. Server keep-alive was uvicorn's default 5s; httpx's defaultkeepalive_expiryis also 5s. - I traced the classification and retry path through
infer_error_kind_from_exception→ProviderErrorKind.API_ERROR→ModelRequestExecutor._should_retry(returns False) →ModelAPIError→ the async scheduler treats it as non-retryable (it isn't inPRESERVED_RETRYABLE_ERRORS). - I confirmed that
create_model_clientpasses_NO_TRANSPORT_RETRY_CONFIGwhenrequest_admissionis set. That disableshttpx_retries, which would otherwise retryRemoteProtocolError/NetworkErroron POST. - I reproduced the exception types with the standalone uvicorn/httpx script above: 13 of 200 failures with the idle time near the keep-alive timeout.
- I searched existing issues for "RemoteProtocolError", "keep-alive retry", and "Non-retryable". Nothing matched. #948 is related: unmapped 4xx statuses are also reported as "unexpected API error".
On our side, we're working around it by raising the server keep-alive above httpx's 5s client expiry. The client then always retires the connection first.
Additional context
data-designer-engine0.9.3 (code unchanged onmain@ 88ecb598cd / v0.9.4)httpx0.28.1,httpx-retries(defaultRETRYABLE_EXCEPTIONS=TimeoutException,NetworkError,RemoteProtocolError),uvicorn0.52.1- Python 3.13, Linux (GitHub Actions runner); reproduced the exception types locally on macOS
- Async engine (
⚡ Using async task-queue preview)
Checklist
- I reproduced this issue or provided a minimal example
- I searched the docs/issues myself, or had my agent do so
- If I used an agent, I included its diagnostics above
- Dominant language
- Python
- Stars
- 2.3k
- Forks
- 219
- Avg merge
- 1d 21h
- Merged PRs (30d)
- 38
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from NVIDIA-NeMo/DataDesigner
-
docs: required_columns description is incomplete for LLM and multimodal columnsPossibly taken @nightcityblade claimed this 2 days ago. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
NVIDIA-NeMo/DataDesigner#1002 · 3 comments ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
NVIDIA-NeMo/DataDesigner#995 ·
Maintainers usually reply within 1 day
-
enforce \from future import annotations` via ruff FA102 rule`Possibly taken @chethanuk claimed this 30 days ago. Opentask
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
NVIDIA-NeMo/DataDesigner#760 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 63/100
NVIDIA-NeMo/DataDesigner#1001 ·
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 53/100
NVIDIA-NeMo/DataDesigner#996 ·
Maintainers usually reply within 1 day
All issues in NVIDIA-NeMo/DataDesigner
Similar issues
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitOpenneeds-triage
Difficulty 2/5 1-3 hours Newbie friendliness 77/100
krkn-chaos/krkn#1627 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
NousResearch/hermes-agent#136483 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
[BUG] LazyStackedTensorDictStore zeroes the last byte of a new key set on the last elementPossibly taken @peterdsharpe claimed this today. Openbug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
pytorch/tensordict#2307 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Maintainers usually reply within 1 day