socket_timeout includes queueing time in a shared, unconfigurable, process-wide thread pool, causing misleading QueryTimeoutErrors on unrelated queries
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- postgresql, python
- Lĩnh vực
- backend, databases, performance
Hướng nghiên cứu
Start with utils/decorators.py, especially timeout() and preserve_transaction_status_with_timeout(), then trace services_container.get_thread_pool("DriverDialectExecutor"). Reproduce the reported contention with concurrent pg_sleep(8) calls and SELECT 1. Done should establish and test an agreed fix for queueing time and executor concurrency, rather than merely documenting the behavior.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
socket_timeout is enforced by submitting the underlying driver call to a shared ThreadPoolExecutor and calling future.result(timeout=socket_timeout) (utils/decorators.py, timeout() / preserve_transaction_status_with_timeout()). That executor is a single, process-wide singleton — services_container.get_thread_pool("DriverDialectExecutor") — shared by every connection and every cursor operation in the process, not scoped per-connection and not related to the size of the caller's own DB connection pool. Its size is never set explicitly anywhere in the wrapper (get_thread_pool(name) is always called with no max_workers), so it falls back to Python's generic ThreadPoolExecutor default: min(32, os.cpu_count() + 4).
Because future.result(timeout=...) starts its clock at executor.submit(), not at the moment the submitted call actually starts running, any time a submission spends queued behind other in-flight driver calls counts fully against socket_timeout. If enough concurrent driver calls are in flight to fill that small, fixed-size pool (very plausible for any app that pools more DB connections than cpu_count + 4, which is a common and reasonable configuration), unrelated and otherwise-trivially-fast queries start timing out — not because they're slow, but because they never got a worker thread in time. The resulting QueryTimeoutError is indistinguishable from a genuine slow query/lock wait, which makes this very hard to diagnose: the error attributes to whatever SQL happened to be queued, with no indication the real cause is thread-pool contention from unrelated concurrent calls elsewhere in the process.
Expected Behavior
socket_timeout should bound the time the actual driver/socket operation takes to execute, not scheduling delay inside an internal, undocumented thread pool. At minimum, the pool used to enforce this timeout should either scale with the caller's configured connection concurrency (e.g. pool size) or be explicitly configurable, so that setting a connection pool size larger than cpu_count + 4 doesn't silently create a hidden concurrency ceiling that's lower than the connection pool itself.
What plugins are used? What other connection properties were set?
plugins="iam", wrapper_dialect="rds-pg", wrapper_driver_dialect="psycopg"
Current Behavior
Under a burst of concurrent database contention (in our case: multiple concurrent transactions blocked on a Postgres row lock, each occupying a worker thread in DriverDialectExecutor for up to their full socket_timeout while waiting on Cursor.execute), a completely unrelated query — targeting a different table with no lock contention of its own — also raised QueryTimeoutError on Cursor.execute, having exceeded its 5 second socket_timeout.
We confirmed that this specific statement showed negligible database-side load and an average execution latency of 0.84ms across the whole incident window — i.e., the statement itself was never slow at the database. This is consistent with the delay happening entirely client-side: the call was queued in the shared DriverDialectExecutor pool behind the other concurrently-blocked calls, and the queueing time alone exceeded the 5 second timeout before the query was ever dispatched to Postgres.
Reproduction Steps
This can be reproduced by saturating the shared executor with concurrent slow calls:
- On a host/container where
os.cpu_count()is small (e.g. 2, giving a default pool size ofmin(32, 2+4) = 6), openN > 6connections viaAwsWrapperConnection.connect(psycopg.connect, ..., socket_timeout=5). - On
N - 1of those connections, concurrently (e.g. one thread per connection) execute a query that runs longer thansocket_timeoutbut well under the connection's own network timeout — e.g.SELECT pg_sleep(8)— so each of those calls occupies aDriverDialectExecutorworker thread for ~8 seconds. - On one additional, otherwise-idle connection, concurrently execute a trivial, instantaneous query:
SELECT 1. - Observe: the
SELECT 1call raisesaws_advanced_python_wrapper.QueryTimeoutError(Cursor.executetimeout) even thoughSELECT 1never reaches the database in a way that would take anywhere near 5 seconds — it's queued behind the 6pg_sleep(8)calls in the shared pool and itsfuture.result(timeout=5)clock (started at submission) expires first.
Possible Solution
A few options, roughly in order of how much they'd change existing behavior:
- Expose the
DriverDialectExecutor(and equivalent per-dialect) pool size as a configurable wrapper property (e.g.DRIVER_DIALECT_EXECUTOR_MAX_WORKERS), defaulting to something scaled to expected DB concurrency rather thancpu_count + 4. - Size the pool based on the connection provider's own pool configuration (e.g.
pool_size + max_overflow) when a pooled connection provider is in use, so it can't become a stricter bottleneck than the connection pool it's serving. - At minimum, document clearly that
socket_timeoutincludes time spent queued in a shared, process-wide thread pool of this default size, so callers can reason about whether their expected concurrent query volume can exceed it.
Additional Information/Context
No response
The AWS Advanced Python Wrapper version used
3.0.0
python version used
3.13.15
Operating System and version
Debian 13 (Trixie)
- Ngôn ngữ chính
- Python
- Star
- 99
- Fork
- 22
- Merge trung bình
- 1 ngày 7 giờ
- Pull request đã merge (30 ngày)
- 4
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của aws/aws-advanced-python-wrapper
-
[aio] host_monitoring_v2 without a topology plugin fails the first statement on cluster-endpoint connectionsCó thể đã có người làm @AhmadMasry đã nhận 1 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 49/100
aws/aws-advanced-python-wrapper#1288 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statementCó thể đã có người làm @AhmadMasry đã nhận 1 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
aws/aws-advanced-python-wrapper#1287 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Async: every connection opens its own topology monitor connection (pool of N holds 2N connections)Có thể đã có người làm @AhmadMasry đã nhận 1 ngày trước. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 25/100
aws/aws-advanced-python-wrapper#1284 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 72/100
aws/aws-advanced-python-wrapper#1276 · 2 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
aws/aws-advanced-python-wrapper#1275 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của aws/aws-advanced-python-wrapper
Issue tương tự
-
Link Checker ReportĐang mởautomated issue report
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 85/100
RapidAI/RapidOCRDocs#119 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
btclib-org/btclib-node#1833 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
IRIS reader: no-data velocity bins (DB_VEL, DB_VELC) returned as 0.0 m/s instead of NaNCó thể đã có người làm @syedhamidali đã nhận hôm nay. Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
elodin-sys/elodin#890 ·
Maintainer thường phản hồi trong vòng 1 ngày