Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

[aio] host_monitoring_v2: event loops stop each other's monitors, recreating them on every statement

Đang mở
#1,287 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 1 ngày

@AhmadMasry đang làm issue này rồi.

Từ ngày 6/10/2026.

  • #1289 của @AhmadMasry — đã đóng, không merge
  • #1290 của @AhmadMasry — đang mở

Đánh giá

Độ khó
4/5
Thời gian dự kiến
3-5 ngày
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
aws, python
Lĩnh vực
backend, databases

Hướng nghiên cứu

Start with aws_advanced_python_wrapper/aio/host_monitoring_plugin.py, especially _AsyncMonitorServiceV2.start_monitoring(), _get_or_create_monitor(), _cleanup_idle_monitors(), and AsyncHostMonitorV2.is_usable() at the cited lines. Compare the async behavior with the sync sharing and locking in host_monitoring_v2_plugin.py:511-535, then add or update tests for multiple event loops. Done means live loops retain their own monitors, idle or ended monitors are cleaned up, and repeated statements do not continually create monitoring connections.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

Describe the bug

With the async API, host_monitoring_v2 (EFM v2) misbehaves when an application runs more than one asyncio event loop, for example one loop per worker thread. Each loop stops the other loops' failure-detection monitors whenever it asks for one of its own. Monitors are destroyed and recreated continuously, each recreation opens a new monitoring connection, and in-flight statements on the other loop silently lose failure detection.

Cause, in aws_advanced_python_wrapper/aio/host_monitoring_plugin.py:

  1. _AsyncMonitorServiceV2.start_monitoring() calls _get_or_create_monitor() for every monitored statement (:537).
  2. _get_or_create_monitor() first calls _cleanup_idle_monitors() (:160), which stops and removes every monitor that isn't is_usable() (:140-149).
  3. AsyncHostMonitorV2.is_usable() returns False whenever the monitor was started on a different event loop from the caller's (:320-326).
  4. The registry key is the detection settings plus the host URL, without the event loop (:108, :161-163).

So every statement on loop B stops all of loop A's monitors, whatever their host or settings, and the next statement on loop A does the same to loop B. Different detection settings per loop don't help, because the cleanup step looks at every monitor regardless of key. The module-level _monitors dict is also read and written from several threads without a lock.

The sync plugin shares one monitor per settings and host through monitor_service.run_if_absent (host_monitoring_v2_plugin.py:511-535), under the monitor service's RLock, and only disposes of a monitor when it has expired with no active contexts or is stuck.

Expected Behavior

Each event loop keeps its own monitor per host and settings (an asyncio task, and the connections it aborts, belong to one loop), and a loop never stops another live loop's monitors. Monitors are only disposed of when idle past expiry, when their task has ended, or when their loop is closed.

What plugins are used? What other connection properties were set?

wrapper_plugins=host_monitoring_v2 (alone or with failover), any dialect, with two or more threads that each run their own event loop.

Current Behavior

The script below runs two threads with one event loop each, taking turns to start a statement against the same host while the other loop's statement is still running:

statements: 40 on 2 loops, 1 host, same settings
monitors created:              40
monitoring connections opened: 40
A monitor still running mid-statement: False

With different detection settings per loop: 40 monitors, 39 monitoring connections, same result for loop A.

Reproduction Steps
import asyncio, threading
from unittest.mock import AsyncMock, MagicMock
from aws_advanced_python_wrapper.aio import host_monitoring_plugin as hm
from aws_advanced_python_wrapper.hostinfo import HostInfo
from aws_advanced_python_wrapper.utils.properties import Properties

opened = created = 0
lock = threading.Lock()
_orig_start = hm.AsyncHostMonitorV2.start


def counting_start(self):
    global created
    with lock:
        created += 1
    _orig_start(self)


hm.AsyncHostMonitorV2.start = counting_start


def make_service():
    svc = MagicMock()

    async def force_connect(*a, **k):
        global opened
        with lock:
            opened += 1
        return MagicMock()
    svc.force_connect = force_connect
    dd = MagicMock()
    dd.is_closed = AsyncMock(return_value=False)
    dd.ping = AsyncMock(return_value=True)
    dd.abort_connection = AsyncMock()
    svc.driver_dialect = dd
    svc.get_telemetry_factory.return_value = MagicMock(
        open_telemetry_context=MagicMock(return_value=None))
    return svc


HOST = HostInfo("inst-1.xyz.us-east-1.rds.amazonaws.com", 5432)
STEPS = 20
gate = [threading.Event() for _ in range(2 * STEPS + 1)]
gate[0].set()
result = {}


async def worker(idx):
    service = hm._AsyncMonitorServiceV2(make_service())
    props = Properties({"host": HOST.host})
    for step in range(STEPS):
        turn = 2 * step + idx
        await asyncio.to_thread(gate[turn].wait)
        ctx = await service.start_monitoring(MagicMock(), HOST, props, 0, 100, 3)
        monitor = hm._monitors.get(hm._monitor_key(0, 100, 3, HOST.url))
        await asyncio.sleep(0.05)
        gate[turn + 1].set()          # the other loop starts a statement
        await asyncio.sleep(0.05)
        if idx == 0 and step == STEPS - 1:
            result["alive"] = monitor is not None and not monitor._stopped
        await service.stop_monitoring(ctx, None)


threads = [threading.Thread(target=lambda i=i: asyncio.run(worker(i))) for i in range(2)]
for t in threads:
    t.start()
for t in threads:
    t.join()
print(f"statements: {2 * STEPS} on 2 loops, 1 host, same settings")
print(f"monitors created:              {created}")
print(f"monitoring connections opened: {opened}")
print(f"A monitor still running mid-statement: {result['alive']}")
Possible Solution
  • Add the event loop to the registry key, so each loop gets its own monitor per host and settings.
  • In _cleanup_idle_monitors, treat a monitor as unusable only when it is stopped, its task has finished, or its loop is closed, not when its loop differs from the caller's. Keep the idle-expiry rule.
  • Protect _monitors with a module-level lock, as sync's monitor service does.

I'm working on a fix and will open a PR that references this issue.

The AWS Advanced Python Wrapper version used

main at 8feea85 (3.1.0)

python version used

Python 3.14

Operating System and version

macOS (reproduced locally)

Ngôn ngữ chính
Python
Star
99
Fork
22
Merge trung bình
1 ngày 7 giờ
Pull request đã merge (30 ngày)
4

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của aws/aws-advanced-python-wrapper

Tất cả issue của aws/aws-advanced-python-wrapper

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.