Async: every connection opens its own topology monitor connection (pool of N holds 2N connections)
维护者通常 1 天内回复
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 25/100
调研方向
Start with aio/wrapper.py, aio/host_list_provider.py, and aio/host_monitoring_plugin.py to compare provider construction, monitor sharing, and resource release; inspect the linked pull requests first, since work is already underway. Reproduce the issue using the script in the report and check that connections to one cluster share a monitor and that it stops after the last connection releases it.
由索引模型根据 Issue 内容生成。
描述
Describe the bug
With the async API (aws_advanced_python_wrapper.aio), every application connection that uses a topology-aware plugin (failover, failover_v2, read_write_splitting, custom_endpoint, ...) starts its own AsyncClusterTopologyMonitor, and each monitor opens its own dedicated monitoring connection. A pool of N connections to one cluster therefore holds 2N server connections.
The sync wrapper shares one topology monitor per cluster id, so the same pool holds N + 1 connections.
Cause:
aio/wrapper.py(AsyncAwsWrapperConnection.connect) builds a new host list provider via_build_host_list_provider(...)on every connect.AsyncAuroraHostListProvider._get_or_create_monitor()(aio/host_list_provider.py) stores the monitor on the provider instance (self._monitor). It isn't shared across providers with the same cluster id.- Sync
RdsHostListProvider._get_or_create_monitor()(host_list_provider.py) usesmonitor_service.run_if_absent(ClusterTopologyMonitorImpl, self.get_cluster_id(), ...), so it is shared.
There's a second problem: closing the connection doesn't stop its monitor. AsyncPluginServiceImpl.release_resources() only releases the host list provider when it is an AsyncCanReleaseResources, and AsyncAuroraHostListProvider isn't one. The monitor is only stopped by release_resources_async() at shutdown. Under a pool that replaces connections (failover, pool_pre_ping invalidation, pool_recycle), monitoring connections build up over time.
Expected Behavior
As in the sync wrapper, connections to the same cluster share one topology monitor and one monitoring connection, and the monitor stops once no connection uses it.
What plugins are used? What other connection properties were set?
wrapper_dialect=aurora-pg&wrapper_plugins=failover,host_monitoring_v2 via create_async_engine("postgresql+aws_wrapper_psycopg://..."). Any topology-aware plugin triggers it. host_monitoring_v2 alone does not, because its monitors are already shared per host.
Current Behavior
A user reported that with an async SQLAlchemy engine and a pool of 10 connections, Aurora PostgreSQL showed 20 connections from the app. The sync wrapper showed 11 for the same setup.
The script below reproduces it without a database:
app connections: 10
distinct cluster ids: 1
running topology monitors: 10
monitor connections opened: 10
Reproduction Steps
import asyncio
from unittest.mock import AsyncMock, MagicMock
from aws_advanced_python_wrapper.aio import release_resources_async
from aws_advanced_python_wrapper.aio.host_list_provider import \
AsyncAuroraHostListProvider
from aws_advanced_python_wrapper.utils.properties import Properties
POOL_SIZE = 10
monitor_conns_opened = 0
async def monitor_conn_factory():
# Stands in for aio/wrapper.py's _monitor_conn_factory.
global monitor_conns_opened
monitor_conns_opened += 1
return MagicMock()
async def main():
props = Properties({"host": "db.cluster-xyz.us-east-1.rds.amazonaws.com",
"port": "5432", "plugins": "failover"})
providers = []
for i in range(POOL_SIZE):
# aio/wrapper.py builds a new provider on every connect.
driver_dialect = MagicMock()
driver_dialect.is_closed = AsyncMock(return_value=False)
p = AsyncAuroraHostListProvider(
props, driver_dialect, monitor_connection_factory=monitor_conn_factory)
try:
await p.refresh(MagicMock()) # what connect does
except Exception:
pass # the fake connection can't run the topology query
providers.append(p)
await asyncio.sleep(0.5)
running = sum(1 for p in providers if p._monitor is not None and p._monitor.is_running())
print(f"app connections: {POOL_SIZE}")
print(f"distinct cluster ids: {len({p.get_cluster_id() for p in providers})}")
print(f"running topology monitors: {running}")
print(f"monitor connections opened: {monitor_conns_opened}")
await release_resources_async()
asyncio.run(main())
Possible Solution
Share AsyncClusterTopologyMonitor per cluster id in a module-level registry, the way aio/host_monitoring_plugin.py already shares EFM monitors per host. Add release_resources() to the async topology providers so closing a connection releases its reference, and stop the monitor when the last reference is released.
Additional Information/Context
Found while helping a user who moved from psycopg's AsyncConnectionPool to an async SQLAlchemy engine. I'm working on a fix and will open a PR that references this issue.
The AWS Advanced Python Wrapper version used
main at c2cd657 (async API, not yet released)
python version used
Python 3.14
Operating System and version
macOS (reproduced locally); the user's environment wasn't stated
- 主要语言
- Python
- 星标
- 99
- 派生
- 22
- 平均合并
- 1 天 7 小时
- 30 天内合并 PR
- 4
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/aws-advanced-python-wrapper 的其他 Issue
-
[aio] host_monitoring_v2 without a topology plugin fails the first statement on cluster-endpoint connections可能已有人在做 @AhmadMasry 于 1 天前认领。 未关闭
难度 4/5 3-5 天 新手友好度 49/100
aws/aws-advanced-python-wrapper#1288 ·
维护者通常 1 天内回复
-
[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statement可能已有人在做 @AhmadMasry 于 1 天前认领。 未关闭
难度 4/5 3-5 天 新手友好度 35/100
aws/aws-advanced-python-wrapper#1287 ·
维护者通常 1 天内回复
-
bug
难度 5/5 一周以上 新手友好度 35/100
aws/aws-advanced-python-wrapper#1278 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 72/100
aws/aws-advanced-python-wrapper#1276 · 2 条评论 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 48/100
aws/aws-advanced-python-wrapper#1275 · 1 条评论 ·
维护者通常 1 天内回复
查看 aws/aws-advanced-python-wrapper 的全部 Issue
相似的 Issue
-
automated issue report
难度 1/5 1 小时以内 新手友好度 85/100
RapidAI/RapidOCRDocs#119 ·
-
难度 2/5 1-3 小时 新手友好度 70/100
-
难度 2/5 1-3 小时 新手友好度 85/100
btclib-org/btclib-node#1833 ·
维护者通常 1 天内回复
-
IRIS reader: no-data velocity bins (DB_VEL, DB_VELC) returned as 0.0 m/s instead of NaN可能已有人在做 @syedhamidali 今天认领。 未关闭
难度 2/5 1-3 小时 新手友好度 72/100
维护者通常 2 天内回复
-
难度 1/5 1 小时以内 新手友好度 80/100
elodin-sys/elodin#890 ·
维护者通常 1 天内回复