[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statement
Les mainteneurs répondent en général sous 1 jour
@AhmadMasry y travaille déjà.
Depuis le 6/10/2026.
Évaluation
- Difficulté
- 4/5
- Temps estimé
- 3-5 jours
- Accessibilité débutants
- 35/100
Piste de recherche
Start with aws_advanced_python_wrapper/aio/host_monitoring_plugin.py, especially _AsyncMonitorServiceV2.start_monitoring(), _get_or_create_monitor(), _cleanup_idle_monitors(), and AsyncHostMonitorV2.is_usable() at the cited lines. Compare the async behavior with the sync sharing and locking in host_monitoring_v2_plugin.py:511-535, then add or update tests for multiple event loops. Done means live loops retain their own monitors, idle or ended monitors are cleaned up, and repeated statements do not continually create monitoring connections.
Rédigé par le modèle d'indexation à partir du texte de l'issue.
Description
Describe the bug
With the async API, host_monitoring_v2 (EFM v2) misbehaves when an application runs more than one asyncio event loop, for example one loop per worker thread. Each loop stops the other loops' failure-detection monitors whenever it asks for one of its own. Monitors are destroyed and recreated continuously, each recreation opens a new monitoring connection, and in-flight statements on the other loop silently lose failure detection.
Cause, in aws_advanced_python_wrapper/aio/host_monitoring_plugin.py:
_AsyncMonitorServiceV2.start_monitoring()calls_get_or_create_monitor()for every monitored statement (:537)._get_or_create_monitor()first calls_cleanup_idle_monitors()(:160), which stops and removes every monitor that isn'tis_usable()(:140-149).AsyncHostMonitorV2.is_usable()returnsFalsewhenever the monitor was started on a different event loop from the caller's (:320-326).- The registry key is the detection settings plus the host URL, without the event loop (:108, :161-163).
So every statement on loop B stops all of loop A's monitors, whatever their host or settings, and the next statement on loop A does the same to loop B. Different detection settings per loop don't help, because the cleanup step looks at every monitor regardless of key. The module-level _monitors dict is also read and written from several threads without a lock.
The sync plugin shares one monitor per settings and host through monitor_service.run_if_absent (host_monitoring_v2_plugin.py:511-535), under the monitor service's RLock, and only disposes of a monitor when it has expired with no active contexts or is stuck.
Expected Behavior
Each event loop keeps its own monitor per host and settings (an asyncio task, and the connections it aborts, belong to one loop), and a loop never stops another live loop's monitors. Monitors are only disposed of when idle past expiry, when their task has ended, or when their loop is closed.
What plugins are used? What other connection properties were set?
wrapper_plugins=host_monitoring_v2 (alone or with failover), any dialect, with two or more threads that each run their own event loop.
Current Behavior
The script below runs two threads with one event loop each, taking turns to start a statement against the same host while the other loop's statement is still running:
statements: 40 on 2 loops, 1 host, same settings
monitors created: 40
monitoring connections opened: 40
A monitor still running mid-statement: False
With different detection settings per loop: 40 monitors, 39 monitoring connections, same result for loop A.
Reproduction Steps
import asyncio, threading
from unittest.mock import AsyncMock, MagicMock
from aws_advanced_python_wrapper.aio import host_monitoring_plugin as hm
from aws_advanced_python_wrapper.hostinfo import HostInfo
from aws_advanced_python_wrapper.utils.properties import Properties
opened = created = 0
lock = threading.Lock()
_orig_start = hm.AsyncHostMonitorV2.start
def counting_start(self):
global created
with lock:
created += 1
_orig_start(self)
hm.AsyncHostMonitorV2.start = counting_start
def make_service():
svc = MagicMock()
async def force_connect(*a, **k):
global opened
with lock:
opened += 1
return MagicMock()
svc.force_connect = force_connect
dd = MagicMock()
dd.is_closed = AsyncMock(return_value=False)
dd.ping = AsyncMock(return_value=True)
dd.abort_connection = AsyncMock()
svc.driver_dialect = dd
svc.get_telemetry_factory.return_value = MagicMock(
open_telemetry_context=MagicMock(return_value=None))
return svc
HOST = HostInfo("inst-1.xyz.us-east-1.rds.amazonaws.com", 5432)
STEPS = 20
gate = [threading.Event() for _ in range(2 * STEPS + 1)]
gate[0].set()
result = {}
async def worker(idx):
service = hm._AsyncMonitorServiceV2(make_service())
props = Properties({"host": HOST.host})
for step in range(STEPS):
turn = 2 * step + idx
await asyncio.to_thread(gate[turn].wait)
ctx = await service.start_monitoring(MagicMock(), HOST, props, 0, 100, 3)
monitor = hm._monitors.get(hm._monitor_key(0, 100, 3, HOST.url))
await asyncio.sleep(0.05)
gate[turn + 1].set() # the other loop starts a statement
await asyncio.sleep(0.05)
if idx == 0 and step == STEPS - 1:
result["alive"] = monitor is not None and not monitor._stopped
await service.stop_monitoring(ctx, None)
threads = [threading.Thread(target=lambda i=i: asyncio.run(worker(i))) for i in range(2)]
for t in threads:
t.start()
for t in threads:
t.join()
print(f"statements: {2 * STEPS} on 2 loops, 1 host, same settings")
print(f"monitors created: {created}")
print(f"monitoring connections opened: {opened}")
print(f"A monitor still running mid-statement: {result['alive']}")
Possible Solution
- Add the event loop to the registry key, so each loop gets its own monitor per host and settings.
- In
_cleanup_idle_monitors, treat a monitor as unusable only when it is stopped, its task has finished, or its loop is closed, not when its loop differs from the caller's. Keep the idle-expiry rule. - Protect
_monitorswith a module-level lock, as sync's monitor service does.
I'm working on a fix and will open a PR that references this issue.
The AWS Advanced Python Wrapper version used
main at 8feea85 (3.1.0)
python version used
Python 3.14
Operating System and version
macOS (reproduced locally)
- Langage dominant
- Python
- Étoiles
- 99
- Forks
- 22
- Merge moyen
- 1 j 7 h
- PR mergées (30 j)
- 4
Préparer son environnement
- Aucun Dockerfile ni fichier Docker Compose
- Propose un modèle de pull request
- Lire le guide de contribution
Par où commencer
- Lisez l'issue en entier, puis le guide de contribution du projet.
- Signalez en commentaire que vous la prenez — cela évite que deux personnes fassent le même travail.
- Forkez le dépôt et travaillez sur une branche.
- Ouvrez une pull request qui référence le numéro de l'issue.
Autres issues de aws/aws-advanced-python-wrapper
-
[aio] host_monitoring_v2 without a topology plugin fails the first statement on cluster-endpoint connectionsPeut-être pris @AhmadMasry l’a pris il y a 1 jour. Ouverte
Difficulté 4/5 3-5 jours Accessibilité débutants 49/100
aws/aws-advanced-python-wrapper#1288 ·
Les mainteneurs répondent en général sous 1 jour
-
Async: every connection opens its own topology monitor connection (pool of N holds 2N connections)Peut-être pris @AhmadMasry l’a pris il y a 1 jour. Ouverte
Difficulté 4/5 3-5 jours Accessibilité débutants 25/100
aws/aws-advanced-python-wrapper#1284 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
-
bug
Difficulté 5/5 Plus d'une semaine Accessibilité débutants 35/100
aws/aws-advanced-python-wrapper#1278 ·
Les mainteneurs répondent en général sous 1 jour
-
bug
Difficulté 4/5 3-5 jours Accessibilité débutants 72/100
aws/aws-advanced-python-wrapper#1276 · 2 commentaires ·
Les mainteneurs répondent en général sous 1 jour
-
bug
Difficulté 4/5 3-5 jours Accessibilité débutants 48/100
aws/aws-advanced-python-wrapper#1275 · 1 commentaire ·
Les mainteneurs répondent en général sous 1 jour
Toutes les issues de aws/aws-advanced-python-wrapper
Issues similaires
-
enhancement
Difficulté 2/5 1-3 heures Accessibilité débutants 85/100
epam/ai-dial-quickapps-backend#628 ·
Les mainteneurs répondent en général sous 2 jours
-
0xlau.dev 已失效,切换成 timlau.meOuverte
Difficulté 1/5 Moins d'une heure Accessibilité débutants 69/100
timqian/chinese-independent-blogs#2235 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 68/100
eclipse-score/coverage_tool#27 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 65/100
bojieli/ai-agent-book#1174 ·
Les mainteneurs répondent en général sous 1 jour
-
Difficulté 2/5 1-3 heures Accessibilité débutants 82/100
RedHatQE/mtv-api-tests#721 ·
Les mainteneurs répondent en général sous 1 jour