Connection.close(timeout=) waits forever on a pending cancel when the server never acknowledges it

オープン
#1,356 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
68/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
活発
技術スタック
postgresql, python
領域
databases

調査の方向性

asyncpg/protocol/protocol.pyx で close(self, timeout)、_request_cancel()、_handle_waiter_on_connection_lost()、_on_connection_lost() を読み始めます。issue に記載されているサーバーがフリーズしたケースを再現し、その後 close(timeout=2) が戻り、transport の喪失によって cancel_waiter が pending のまま残らないことを確認します。その間、接続は期待どおり abort にフォールバックします。

索引モデルが issue の本文から書いたものです。

説明

Summary

When a statement times out (command_timeout) while the server, or a pooler in front of it, is frozen, asyncpg requests a cancel and then every later operation on that connection, including close(timeout=...), awaits the cancel acknowledgement with no bound. The timeout argument of close() does not cover that wait, and a subsequent transport loss does not resolve it either, so the connection can never be closed gracefully and any caller that awaits close() hangs indefinitely.

Versions

  • asyncpg 0.30.0 and 0.31.0 (same code shape in both)
  • Python 3.12.3, Linux
  • Observed through SQLAlchemy 2.0.52's asyncpg dialect, which calls Connection.close(timeout=2) when invalidating a connection after a TimeoutError, but the behaviour is asyncpg's.

Where in the source (0.31.0)

  • asyncpg/protocol/protocol.pyx, close(self, timeout): awaits self.cancel_sent_waiter and then if self.cancel_waiter is not None: await self.cancel_waiter before the part that is guarded by timeout.
  • _request_cancel() (called from _on_timeout()) creates cancel_waiter; it is resolved only by a ReadyForQuery arriving on the original socket.
  • _handle_waiter_on_connection_lost() and _on_connection_lost() resolve self.waiter only; cancel_waiter is left pending when the transport is lost.
  • abort() returns early when self.closing is already set, so cancelling a stuck close() from outside and then calling Connection._abort() does not close the transport.

Reproduction

  1. Run PostgreSQL behind pgbouncer (transaction pooling), or plain PostgreSQL.
  2. Open a connection with command_timeout=5, run SELECT pg_sleep(40).
  3. While it runs, freeze the server process (podman pause / kill -STOP on postgres, or on pgbouncer).
  4. The statement raises asyncio.TimeoutError after 5 s and asyncpg starts a cancel task.
  5. Now await conn.close(timeout=2): it never returns while the freeze lasts. If the frozen side is later closed by a pooler timeout (pgbouncer query_timeout closes the client socket), close() still never returns because the transport loss resolves only the query waiter.

Observed with an asyncio task-stack watchdog: the caller sits in Connection.closeprotocol.closeawait self.cancel_waiter, and the _cancel task sits in connect_utils awaiting the cancel connection's on_disconnect, for as long as the server stays frozen (minutes; unbounded).

Expected

  • close(timeout=t) should bound the wait for the cancel acknowledgement by t (or by the connection's command_timeout) and fall back to aborting the transport.
  • _on_connection_lost() should resolve cancel_waiter (with the same connection-lost exception it uses for waiter) so that a lost transport cannot leave a permanently pending cancel.

Workaround we use

An application-level guard that, on TimeoutError/CancelledError, runs asyncio.wait_for(conn.close(timeout=g), g) and on expiry calls conn.terminate() followed by an explicit conn._transport.abort(), because terminate() alone leaves the socket open once close() has marked the protocol as closing.

主要言語
Python
スター
8.1k
フォーク
469
平均マージ
4時間 42分
マージ済み PR(30日)
6

コントリビューションガイド

このリポジトリのコントリビューションガイドは索引されていません

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

MagicStack/asyncpg のほかの issue

MagicStack/asyncpg の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。