Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Async provider can strand analysis after reconnect when reusing streamed HTTP connections

Open
#45 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
78/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python

Research direction

Start with example-provider.py at the async provider rewrite from d0eeb24, focusing on the separate streamed submission session and the pooled /api/external-engine/work acquisition loop. Reproduce the reconnect-and-step sequence, then compare it with the submission-only force-close control described in the issue. Done means later positions stream analysis after reconnect, while acquisition remains pooled and normal stop requests still produce bestmove.

Written by the indexing model from the issue text.

Description

Summary

Since the async provider rewrite in d0eeb24, stepping to a new analysis
position after a browser/Lichess reconnect can immediately close the provider's
new streamed submission. The browser eventually reports:

Your external engine does not appear to be connected.

The provider and UCI engine are still running. The provider returns to polling
/api/external-engine/work, but receives only 204 No Content, so analysis does
not recover automatically.

A controlled comparison found:

Provider Streamed submission connection Result
a93f733 fresh requests.post() connection passes
d0eeb24 pooled aiohttp.ClientSession connection fails
d0eeb24 + TCPConnector(force_close=True) for submissions fresh aiohttp connection passes

Only the submission connector was changed in the final A/B test. Work
acquisition, async UCI handling, keepalives, engine configuration, position,
and interaction sequence were unchanged.

Regression range

d0eeb24 merged the async/aiohttp long-search keepalive implementation in
#44.

Environment

  • lichess.org production and https://engine.lichess.ovh
  • Debian forky/sid
  • Python 3.14.7
  • aiohttp 3.14.3
  • UCI engine: CrazyAra 1.0.5, Crazyhouse, MultiPV 5

The engine appears incidental: it continuously emits valid UCI info records,
and every received stop in these tests produced a terminal bestmove.

Reproduction

  1. Start example-provider.py from d0eeb24 with a Crazyhouse-capable UCI
    engine and debug logging.
  2. Open https://lichess.org/mTjhThgB/white#11 and enable the external engine.
  3. Leave analysis running on ply 11 for approximately 20 seconds.
  4. Advance once to ply 12 (6...d5).
  5. Observe that the new job is acquired and starts, but its streamed submission
    closes immediately. The GUI later displays the disconnected-provider dialog
    and does not recover automatically.

This also occurred while navigating other completed Crazyhouse games forward
and backward, but the sequence above was used for the controlled comparison.

Actual behavior

The failing run acquired the ply-12 job and started the engine normally, then
reported a closed stream one millisecond later:

2026-09-25 12:31:25,375 DEBUG:root:Acquire response: 200
2026-09-25 12:31:25,375 DEBUG:root:Acquired job d59DUk0XDyatI8iV
2026-09-25 12:31:25,392 INFO:root:Handling job d59DUk0XDyatI8iV
2026-09-25 12:31:25,392 DEBUG:root:2589856 << position fen ... moves ... f2f3 d7d5
2026-09-25 12:31:25,392 DEBUG:root:2589856 << go movetime 86400000
2026-09-25 12:31:25,392 DEBUG:root:Waiting for work
2026-09-25 12:31:25,393 INFO:root:Connection closed while streaming analysis
2026-09-25 12:31:25,393 DEBUG:root:2589856 << stop
2026-09-25 12:31:26,317 DEBUG:root:2589856 >> bestmove e1g1

The provider then remained connected and polled successfully, but the broker
returned only 204 responses while the GUI showed the error:

2026-09-25 12:31:35,564 DEBUG:root:Acquire response: 204
2026-09-25 12:31:45,735 DEBUG:root:Acquire response: 204
2026-09-25 12:31:55,907 DEBUG:root:Acquire response: 204
2026-09-25 12:32:06,078 DEBUG:root:Acquire response: 204
2026-09-25 12:32:16,249 DEBUG:root:Acquire response: 204

The UCI engine did not crash or hang. It honored stop and returned
bestmove in under one second.

Expected behavior

The newly acquired position should remain attached and stream analysis until it
is replaced or the user stops analysis. A transient browser reconnect should
not poison the HTTP transport used by a later job.

Controls

Synchronous control: a93f733

The exact pause-and-step sequence passed on a clean a93f733 worktree. Analysis
continued through 6...d5 and several later positions. Stopping analysis in
the GUI completed the current submission with HTTP 200 and the engine returned
bestmove. No disconnected-provider dialog appeared.

Bare a6ef15a was not tested separately because its relevant synchronous HTTP
lifecycle is the same; a93f733 additionally forwards the required terminal
bestmove.

Async control with fresh submission connections

The test was repeated on d0eeb24 with only this change:

submit_connector = aiohttp.TCPConnector(force_close=True)

async with (
    aiohttp.ClientSession(headers=auth_headers) as http,
    aiohttp.ClientSession(
        timeout=stream_timeout,
        connector=submit_connector,
    ) as submit_http,
):
    ...

The reproduction then passed:

2026-09-25 13:00:43,499 DEBUG:root:Acquired job 6dkeH6UrUh1s6nq3
2026-09-25 13:00:43,511 INFO:root:Handling job 6dkeH6UrUh1s6nq3
2026-09-25 13:00:43,855 DEBUG:root:2619837 << position fen ... moves ... f2f3 d7d5
2026-09-25 13:00:43,855 DEBUG:root:2619837 << go movetime 86400000
2026-09-25 13:00:54,920 DEBUG:root:2619837 >> info depth 32 ...
2026-09-25 13:00:58,922 DEBUG:root:Acquired job h4yqk5fKR5cIr6yT
2026-09-25 13:00:58,923 DEBUG:root:2619837 << stop
2026-09-25 13:00:58,932 DEBUG:root:2619837 >> bestmove e1g1

Further navigation also worked. The final GUI stop was clean:

2026-09-25 13:01:52,319 DEBUG:root:2619837 << stop
2026-09-25 13:01:52,329 DEBUG:root:2619837 >> bestmove f3e4

There were no Connection closed while streaming analysis records in the
force-close run.

Analysis

The async provider uses one long-lived aiohttp.ClientSession for all streamed
work submissions. These requests have indefinite, chunked bodies and can be
ended early when the browser disconnects or reconnects. Reusing the associated
HTTP connection for a later streamed job triggers the immediate-close behavior
in this reproduction.

The synchronous provider uses the top-level requests.post() helper for each
submission, which creates a fresh session/connection. Scoping
force_close=True to the async submission connector restores that behavior
without disabling pooling for the independent /work acquisition loop.

Proposed fix

Use a non-reusing connector for the streamed submission session:

 stream_timeout = aiohttp.ClientTimeout(total=None)
+submit_connector = aiohttp.TCPConnector(force_close=True)

 async with (
     aiohttp.ClientSession(headers=auth_headers) as http,
-    aiohttp.ClientSession(timeout=stream_timeout) as submit_http,
+    aiohttp.ClientSession(
+        timeout=stream_timeout,
+        connector=submit_connector,
+    ) as submit_http,
 ):

The acquisition session should remain pooled.

Validation performed

  • Reproduced repeatedly with pooled async submissions.
  • Passed the identical sequence with clean a93f733.
  • Passed the identical sequence with d0eeb24 plus submission-only
    force_close=True.
  • Continued through multiple later positions in both passing controls.
  • Verified clean GUI stop (stop followed by bestmove) in both controls.
  • Independently stress-tested the UCI engine with 25 immediate duplicate-stop
    cycles; all 25 returned bestmove.

Related issue

#35 reported random
503 errors in 2023, but had no reliable reproduction and predates the async
provider rewrite. This report has a deterministic sequence, a commit-level
regression boundary, timestamped provider/UCI evidence, and a single-variable
working control.

Proposed fix

(nssy/external-engine f6feea2).

Dominant language
Python
Stars
94
Forks
29
Avg merge
6d 3h
Merged PRs (30d)
2

Getting set up

We have not checked this project's setup files yet. Start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from lichess-org/external-engine

All issues in lichess-org/external-engine

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.