Connection.close(timeout=) waits forever on a pending cancel when the server never acknowledges it
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 68/100
Línea de trabajo
Comienza en asyncpg/protocol/protocol.pyx leyendo close(self, timeout), _request_cancel(), _handle_waiter_on_connection_lost() y _on_connection_lost(). Reproduce el caso de servidor congelado descrito en el issue y, después, verifica que close(timeout=2) retorna y que la pérdida del transporte no deja cancel_waiter en estado pending, mientras la conexión vuelve a abortar como se espera.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
When a statement times out (command_timeout) while the server, or a pooler in front of it, is frozen, asyncpg requests a cancel and then every later operation on that connection, including close(timeout=...), awaits the cancel acknowledgement with no bound. The timeout argument of close() does not cover that wait, and a subsequent transport loss does not resolve it either, so the connection can never be closed gracefully and any caller that awaits close() hangs indefinitely.
Versions
- asyncpg 0.30.0 and 0.31.0 (same code shape in both)
- Python 3.12.3, Linux
- Observed through SQLAlchemy 2.0.52's asyncpg dialect, which calls
Connection.close(timeout=2)when invalidating a connection after aTimeoutError, but the behaviour is asyncpg's.
Where in the source (0.31.0)
asyncpg/protocol/protocol.pyx,close(self, timeout): awaitsself.cancel_sent_waiterand thenif self.cancel_waiter is not None: await self.cancel_waiterbefore the part that is guarded bytimeout._request_cancel()(called from_on_timeout()) createscancel_waiter; it is resolved only by a ReadyForQuery arriving on the original socket._handle_waiter_on_connection_lost()and_on_connection_lost()resolveself.waiteronly;cancel_waiteris left pending when the transport is lost.abort()returns early whenself.closingis already set, so cancelling a stuckclose()from outside and then callingConnection._abort()does not close the transport.
Reproduction
- Run PostgreSQL behind pgbouncer (transaction pooling), or plain PostgreSQL.
- Open a connection with
command_timeout=5, runSELECT pg_sleep(40). - While it runs, freeze the server process (
podman pause/kill -STOPon postgres, or on pgbouncer). - The statement raises
asyncio.TimeoutErrorafter 5 s and asyncpg starts a cancel task. - Now
await conn.close(timeout=2): it never returns while the freeze lasts. If the frozen side is later closed by a pooler timeout (pgbouncerquery_timeoutcloses the client socket),close()still never returns because the transport loss resolves only the query waiter.
Observed with an asyncio task-stack watchdog: the caller sits in Connection.close → protocol.close → await self.cancel_waiter, and the _cancel task sits in connect_utils awaiting the cancel connection's on_disconnect, for as long as the server stays frozen (minutes; unbounded).
Expected
close(timeout=t)should bound the wait for the cancel acknowledgement byt(or by the connection'scommand_timeout) and fall back to aborting the transport._on_connection_lost()should resolvecancel_waiter(with the same connection-lost exception it uses forwaiter) so that a lost transport cannot leave a permanently pending cancel.
Workaround we use
An application-level guard that, on TimeoutError/CancelledError, runs asyncio.wait_for(conn.close(timeout=g), g) and on expiry calls conn.terminate() followed by an explicit conn._transport.abort(), because terminate() alone leaves the socket open once close() has marked the protocol as closing.
- Lenguaje dominante
- Python
- Estrellas
- 8.1k
- Forks
- 469
- Merge medio
- 4 h 42 min
- PR fusionados (30 d)
- 6
Guía de contribución
No hay ninguna guía de contribución indexada para este repositorio
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de MagicStack/asyncpg
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
MagicStack/asyncpg#1357 ·
-
tests fail in 2032 Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
MagicStack/asyncpg#997 ·
-
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
MagicStack/asyncpg#1354 ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 72/100
MagicStack/asyncpg#1342 ·
-
Dificultad 3/5 1-2 días Aptitud para principiantes 56/100
MagicStack/asyncpg#1340 · 1 comentario ·
Todos los issues de MagicStack/asyncpg
Issues similares
-
sponsored
Dificultad 2/5 1-3 horas Aptitud para principiantes 65/100
-
opensubtitlescom: moviehash never sent when opensubtitles (.org) is not in the provider list Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Diaoul/subliminal#1382 ·
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
-
triage/confirmed
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
agentscope-ai/agentscope#2775 ·
-
worker.gpuVendors silently accepts unsupported/misspelled vendor names — no validation guard Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100