[aio] host_monitoring_v2 without a topology plugin fails the first statement on cluster-endpoint connections
维护者通常 1 天内回复
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 49/100
调研方向
Start with aio/wrapper.py and _build_host_list_provider, then compare the Aurora dialect provider selection in database_dialect.py with the async behavior. Trace how AsyncHostMonitoringPlugin._get_monitoring_host_info identifies and caches a host, using the described mocked reproduction as a starting point. Done when the first statement succeeds through a cluster endpoint and monitoring targets the connected instance, with regression coverage for the failure and cache behavior.
由索引模型根据 Issue 内容生成。
描述
Describe the bug
With the async API, host_monitoring_v2 used without a topology plugin (no failover, failover_v2, aurora_connection_tracker, ...) fails the first statement on every new connection made through an Aurora cluster endpoint, with AwsWrapperError: [HostMonitoringV2Plugin] Unable to identify the connected database instance. Later statements on that connection succeed, but the monitor watches the cluster endpoint's DNS name instead of the instance the connection is on.
Cause:
aio/wrapper.py_build_host_list_providerreturnsAsyncStaticHostListProviderwhenpluginscontains none of_TOPOLOGY_REQUIRING_PLUGINS, whatever the database dialect.host_monitoring_v2isn't in that set.- For a cluster endpoint,
AsyncHostMonitoringPlugin._get_monitoring_host_infocallsplugin_service.identify_connection. With the static provider the topology has nohost_id, so it returnsNoneand the plugin raises before running the statement. _get_monitoring_host_infoassignsself._monitoring_host_info = current_host_infobefore identifying, so on the next statement it returns the cached cluster endpoint and monitors that.
The sync wrapper doesn't hit this: Aurora dialects always use RdsHostListProvider (database_dialect.py, AuroraPgDialect.get_host_list_provider_supplier), regardless of plugins, so identify_connection resolves the instance.
Expected Behavior
As in sync, host_monitoring_v2 alone on a cluster endpoint identifies the instance the connection landed on and monitors it, and the first statement runs normally.
What plugins are used? What other connection properties were set?
wrapper_dialect=aurora-pg&wrapper_plugins=host_monitoring_v2, connecting through the cluster (writer) endpoint. With failover added, or through an instance endpoint, it works.
Current Behavior
Checked without a database by building the provider the way connect does and calling _get_monitoring_host_info with get_instance_id patched to return an instance id:
AsyncStaticHostListProvider mydb.cluster-xyz... -> 1st call: AwsWrapperError: Unable to identify the connected database instance
2nd call: monitors mydb.cluster-xyz.us-east-1.rds.amazonaws.com
AsyncStaticHostListProvider inst-1.xyz... -> monitors inst-1.xyz.us-east-1.rds.amazonaws.com
Reproduction Steps
import asyncio
from sqlalchemy import text
from sqlalchemy.ext.asyncio import create_async_engine
from aws_advanced_python_wrapper.aio import release_resources_async
async def main():
engine = create_async_engine(
"postgresql+aws_wrapper_psycopg://USER:PASSWORD@<name>.cluster-<id>.<region>.rds.amazonaws.com:5432/DB"
"?wrapper_dialect=aurora-pg&wrapper_plugins=host_monitoring_v2")
try:
async with engine.connect() as conn:
await conn.execute(text("SELECT 1")) # raises AwsWrapperError
finally:
await engine.dispose()
await release_resources_async()
asyncio.run(main())
Possible Solution
Match sync: pick the topology-aware provider from the Aurora/Multi-AZ/Global dialect even when no topology plugin is configured, or add host_monitoring_v2 (and host_monitoring) to _TOPOLOGY_REQUIRING_PLUGINS. Either way, resolve the instance (the topology monitor is shared per cluster after #1285, so this costs one monitoring connection per cluster). Separately, _get_monitoring_host_info shouldn't cache the cluster endpoint when identification fails.
Additional Information/Context
Found while answering a user question about using host_monitoring_v2 without failover. Related: #1284.
The AWS Advanced Python Wrapper version used
main at 8feea85 (3.1.0)
python version used
Python 3.14
Operating System and version
macOS (checked locally, not against a live cluster)
- 主要语言
- Python
- 星标
- 99
- 派生
- 22
- 平均合并
- 1 天 7 小时
- 30 天内合并 PR
- 4
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
aws/aws-advanced-python-wrapper 的其他 Issue
-
[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statement可能已有人在做 @AhmadMasry 于 1 天前认领。 未关闭
难度 4/5 3-5 天 新手友好度 35/100
aws/aws-advanced-python-wrapper#1287 ·
维护者通常 1 天内回复
-
Async: every connection opens its own topology monitor connection (pool of N holds 2N connections)可能已有人在做 @AhmadMasry 于 1 天前认领。 未关闭
难度 4/5 3-5 天 新手友好度 25/100
aws/aws-advanced-python-wrapper#1284 · 1 条评论 ·
维护者通常 1 天内回复
-
bug
难度 5/5 一周以上 新手友好度 35/100
aws/aws-advanced-python-wrapper#1278 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 72/100
aws/aws-advanced-python-wrapper#1276 · 2 条评论 ·
维护者通常 1 天内回复
-
bug
难度 4/5 3-5 天 新手友好度 48/100
aws/aws-advanced-python-wrapper#1275 · 1 条评论 ·
维护者通常 1 天内回复
查看 aws/aws-advanced-python-wrapper 的全部 Issue
相似的 Issue
-
changelog investigate
难度 2/5 1-3 小时 新手友好度 62/100
ramnes/notion-sdk-py#409 ·
-
good first issue help wanted
难度 2/5 1-3 小时 新手友好度 72/100
lindicaphxag-tech/kaggle#28 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 62/100
BSData/horus-heresy-3rd-edition#3211 ·
维护者通常 1 天内回复
-
bug needs-triage
难度 2/5 1-3 小时 新手友好度 70/100
维护者通常 1 天内回复
-
Unreachable-proxy mount test depends on fixed port 9999可能已有人在做 关联的 PR 仍在进行中或已合并。 未关闭bug tests
难度 2/5 1-3 小时 新手友好度 76/100
维护者通常 1 天内回复