Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statement

Abierto
#1,287 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

@AhmadMasry ya está trabajando en esto.

Desde el 6/10/2026.

  • #1289 de @AhmadMasry — cerrado sin fusionar
  • #1290 de @AhmadMasry — abierto

Evaluación

Dificultad
4/5
Tiempo estimado
3-5 días
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
aws, python

Línea de trabajo

Start with aws_advanced_python_wrapper/aio/host_monitoring_plugin.py, especially _AsyncMonitorServiceV2.start_monitoring(), _get_or_create_monitor(), _cleanup_idle_monitors(), and AsyncHostMonitorV2.is_usable() at the cited lines. Compare the async behavior with the sync sharing and locking in host_monitoring_v2_plugin.py:511-535, then add or update tests for multiple event loops. Done means live loops retain their own monitors, idle or ended monitors are cleaned up, and repeated statements do not continually create monitoring connections.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Describe the bug

With the async API, host_monitoring_v2 (EFM v2) misbehaves when an application runs more than one asyncio event loop, for example one loop per worker thread. Each loop stops the other loops' failure-detection monitors whenever it asks for one of its own. Monitors are destroyed and recreated continuously, each recreation opens a new monitoring connection, and in-flight statements on the other loop silently lose failure detection.

Cause, in aws_advanced_python_wrapper/aio/host_monitoring_plugin.py:

  1. _AsyncMonitorServiceV2.start_monitoring() calls _get_or_create_monitor() for every monitored statement (:537).
  2. _get_or_create_monitor() first calls _cleanup_idle_monitors() (:160), which stops and removes every monitor that isn't is_usable() (:140-149).
  3. AsyncHostMonitorV2.is_usable() returns False whenever the monitor was started on a different event loop from the caller's (:320-326).
  4. The registry key is the detection settings plus the host URL, without the event loop (:108, :161-163).

So every statement on loop B stops all of loop A's monitors, whatever their host or settings, and the next statement on loop A does the same to loop B. Different detection settings per loop don't help, because the cleanup step looks at every monitor regardless of key. The module-level _monitors dict is also read and written from several threads without a lock.

The sync plugin shares one monitor per settings and host through monitor_service.run_if_absent (host_monitoring_v2_plugin.py:511-535), under the monitor service's RLock, and only disposes of a monitor when it has expired with no active contexts or is stuck.

Expected Behavior

Each event loop keeps its own monitor per host and settings (an asyncio task, and the connections it aborts, belong to one loop), and a loop never stops another live loop's monitors. Monitors are only disposed of when idle past expiry, when their task has ended, or when their loop is closed.

What plugins are used? What other connection properties were set?

wrapper_plugins=host_monitoring_v2 (alone or with failover), any dialect, with two or more threads that each run their own event loop.

Current Behavior

The script below runs two threads with one event loop each, taking turns to start a statement against the same host while the other loop's statement is still running:

statements: 40 on 2 loops, 1 host, same settings
monitors created:              40
monitoring connections opened: 40
A monitor still running mid-statement: False

With different detection settings per loop: 40 monitors, 39 monitoring connections, same result for loop A.

Reproduction Steps
import asyncio, threading
from unittest.mock import AsyncMock, MagicMock
from aws_advanced_python_wrapper.aio import host_monitoring_plugin as hm
from aws_advanced_python_wrapper.hostinfo import HostInfo
from aws_advanced_python_wrapper.utils.properties import Properties

opened = created = 0
lock = threading.Lock()
_orig_start = hm.AsyncHostMonitorV2.start


def counting_start(self):
    global created
    with lock:
        created += 1
    _orig_start(self)


hm.AsyncHostMonitorV2.start = counting_start


def make_service():
    svc = MagicMock()

    async def force_connect(*a, **k):
        global opened
        with lock:
            opened += 1
        return MagicMock()
    svc.force_connect = force_connect
    dd = MagicMock()
    dd.is_closed = AsyncMock(return_value=False)
    dd.ping = AsyncMock(return_value=True)
    dd.abort_connection = AsyncMock()
    svc.driver_dialect = dd
    svc.get_telemetry_factory.return_value = MagicMock(
        open_telemetry_context=MagicMock(return_value=None))
    return svc


HOST = HostInfo("inst-1.xyz.us-east-1.rds.amazonaws.com", 5432)
STEPS = 20
gate = [threading.Event() for _ in range(2 * STEPS + 1)]
gate[0].set()
result = {}


async def worker(idx):
    service = hm._AsyncMonitorServiceV2(make_service())
    props = Properties({"host": HOST.host})
    for step in range(STEPS):
        turn = 2 * step + idx
        await asyncio.to_thread(gate[turn].wait)
        ctx = await service.start_monitoring(MagicMock(), HOST, props, 0, 100, 3)
        monitor = hm._monitors.get(hm._monitor_key(0, 100, 3, HOST.url))
        await asyncio.sleep(0.05)
        gate[turn + 1].set()          # the other loop starts a statement
        await asyncio.sleep(0.05)
        if idx == 0 and step == STEPS - 1:
            result["alive"] = monitor is not None and not monitor._stopped
        await service.stop_monitoring(ctx, None)


threads = [threading.Thread(target=lambda i=i: asyncio.run(worker(i))) for i in range(2)]
for t in threads:
    t.start()
for t in threads:
    t.join()
print(f"statements: {2 * STEPS} on 2 loops, 1 host, same settings")
print(f"monitors created:              {created}")
print(f"monitoring connections opened: {opened}")
print(f"A monitor still running mid-statement: {result['alive']}")
Possible Solution
  • Add the event loop to the registry key, so each loop gets its own monitor per host and settings.
  • In _cleanup_idle_monitors, treat a monitor as unusable only when it is stopped, its task has finished, or its loop is closed, not when its loop differs from the caller's. Keep the idle-expiry rule.
  • Protect _monitors with a module-level lock, as sync's monitor service does.

I'm working on a fix and will open a PR that references this issue.

The AWS Advanced Python Wrapper version used

main at 8feea85 (3.1.0)

python version used

Python 3.14

Operating System and version

macOS (reproduced locally)

Lenguaje dominante
Python
Estrellas
99
Forks
22
Merge medio
1 d 7 h
PR fusionados (30 d)
4

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de aws/aws-advanced-python-wrapper

Todos los issues de aws/aws-advanced-python-wrapper

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.