[Bug]: Exception raised by a signal handler kills the sync dispatcher fiber, every later sync call spins at 100% CPU
I maintainer di solito rispondono entro 7 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 48/100
Direzione di ricerca
Start in playwright/_impl/_sync_base.py at _sync() and playwright/sync_api/_context_manager.py at greenlet_main; reproduce the signal-handler case from the issue. Trace how an exception raised during run_until_complete affects the dispatcher fiber and the pending task. Done means the exception reaches the calling sync code without leaving later sync calls spinning, while cleanup and subsequent pages remain usable.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Version
1.63.0 (also 1.52.0)
Steps to reproduce
import signal
from playwright.sync_api import sync_playwright
class TestTimeout(Exception):
pass
def on_alarm(signum, frame):
raise TestTimeout() # what pytest-timeout does with timeout_method=signal
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
signal.signal(signal.SIGALRM, on_alarm)
signal.alarm(1)
try:
page.wait_for_timeout(5000) # any call that is still pending when the signal arrives
except TestTimeout:
print("timeout delivered; dispatcher dead:", page._dispatcher_fiber.dead)
print("calling Playwright again...")
print(page.title()) # never returns: spins at 100% CPU
print("unreachable")
The same thing with pytest-timeout, which is how we hit it:
import pytest
from playwright.sync_api import sync_playwright
@pytest.fixture(scope="session")
def browser():
with sync_playwright() as p:
browser = p.chromium.launch()
yield browser
browser.close()
@pytest.mark.timeout(1)
def test_slow(browser):
page = browser.new_page()
page.wait_for_timeout(5000)
def test_next(browser):
page = browser.new_page() # hangs forever, the session never finishes
assert page.title() == ""
pytest test_repro.py -o timeout_method=signal
Expected behavior
The exception from the signal handler reaches the code that made the sync call (it does), and Playwright stays usable: page.title() returns, test_next passes.
Actual behavior
timeout delivered; dispatcher dead: True
calling Playwright again...
and then the process spins at 99% CPU forever. With pytest, test_slow fails with Failed: Timeout, then test_next (and the session fixture teardown) hang the same way.
A sync call blocks inside SyncBase._sync() in self._dispatcher_fiber.switch(), so the interpreter is running the dispatcher fiber (loop.run_until_complete(...) inside selector.select()) when the signal arrives. Python runs the handler in that frame, so the exception is raised inside the dispatcher fiber. It unwinds run_until_complete and greenlet_main, the fiber dies, and greenlet passes the exception to its parent, which is the greenlet waiting in _sync(). That part looks right. But from then on while not task.done(): self._dispatcher_fiber.switch() returns immediately on every iteration, because the switch goes to a dead greenlet.
Traceback of the timeout from our CI (1.52.0, Linux) shows where the signal lands:
File ".../playwright/sync_api/_generated.py", line 16713, in count
return mapping.from_maybe_impl(self._sync(self._impl_obj.count()))
File ".../playwright/_impl/_sync_base.py", line 113, in _sync
self._dispatcher_fiber.switch()
File ".../playwright/sync_api/_context_manager.py", line 56, in greenlet_main
self._loop.run_until_complete(self._connection.run_as_sync())
File ".../asyncio/base_events.py", line 1898, in _run_once
event_list = self._selector.select(timeout)
File ".../selectors.py", line 468, in select
fd_event_list = self._selector.poll(timeout, max_ev)
File ".../pytest_timeout.py", line 317, in handler
timeout_sigalrm(item, settings)
File ".../pytest_timeout.py", line 502, in timeout_sigalrm
pytest.fail(PYTEST_FAILURE_MESSAGE % settings.timeout)
File ".../_pytest/outcomes.py", line 163, in __call__
raise Failed(msg=reason, pytrace=pytrace)
In our CI the next call was the failure screenshot in a pytest_runtest_makereport hook. It spun until the job-level SIGINT, and the test ended up reported as passed in Allure, although its step had failed with the timeout.
Additional context
Related: #3187 adds a dispatcher_fiber.dead check to _sync(), so the next call raises TargetClosedError instead of spinning. That would turn the hang into an error, but the Playwright instance would still be unusable after a timeout, so every following test in the worker would fail. That PR covers a different trigger (the transport ends without Connection.cleanup()); a signal handler is a second way to reach the same loop.
What works for us as a workaround: a custom pytest_timeout_set_timer whose handler checks where it is running. When the current greenlet has a parent (that is, we are inside the dispatcher or an event listener fiber), it wakes the event loop and throws the exception into the root greenlet instead of raising it in place:
def handler(signum, frame):
try:
timeout_sigalrm(item, settings) # pytest-timeout: raises Failed
except pytest.fail.Exception as exc:
current = greenlet.getcurrent()
root = current
while root.parent is not None:
root = root.parent
if root is current:
raise
# select() was entered with a timeout computed before the next call's task
# existed, so without a wakeup the next call is never sent to the driver
asyncio.get_running_loop().call_soon_threadsafe(lambda: None)
root.throw(exc)
The dispatcher stays alive, the pending task is just orphaned, and later calls work: a screenshot of a frozen page fails with its own TimeoutError, page.close() and context.close() complete, and a new page opens in the same browser.
Doing this inside Playwright (for example, when a BaseException escapes run_until_complete in greenlet_main while a call is pending, throw it into the waiting greenlet and keep the loop running) would make sync Playwright safe under pytest-timeout's default signal method. If that is out of scope, a note in the docs that signal-based timeouts break the sync API would help too.
Environment
- Operating System: macOS 15.7.9 (repro above); Linux in Docker (CI, 1.52.0)
- CPU: arm64 (repro above)
- Browser: Chromium
- Python Version: 3.11.14
- Other info: greenlet 3.5.6, pytest 9.1.1, pytest-timeout 2.4.0
- Lingua principale
- Python
- Stelle
- 15k
- Fork
- 1.2k
- Merge medio
- 5g 12h
- PR unite (30g)
- 13
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Nessun modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/playwright-python
-
[Feature]: relax the pyee upper-bound pin (currently <14) in pyproject.tomlForse già presa @Ankitraj-sharma l’ha presa 6 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
microsoft/playwright-python#3208 · 3 commenti ·
I maintainer di solito rispondono entro 7 giorni
-
[Bug]: 0.0 is serialized as -0 when passed to evaluate()Forse già presa @wasim-builds l’ha presa 37 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
microsoft/playwright-python#3186 ·
I maintainer di solito rispondono entro 7 giorni
-
Difficoltà 4/5 3-5 giorni Idoneità per principianti 45/100
microsoft/playwright-python#3216 ·
I maintainer di solito rispondono entro 7 giorni
-
[Feature]: Official APIs to resolve/map ephemeral aria-refs to stable locators for LLM automationAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 35/100
microsoft/playwright-python#3207 ·
I maintainer di solito rispondono entro 7 giorni
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 1/100
microsoft/playwright-python#3195 ·
I maintainer di solito rispondono entro 7 giorni
Tutte le issue di microsoft/playwright-python
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 3 giorni
-
Negation with "not" and "no" is ignored during sentiment analysisForse già presa @vivek-3728 l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
techcsispit/mess-mood#11 · 1 commento ·
-
changelog investigate
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
ramnes/notion-sdk-py#408 ·
-
good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 83/100
btclib-org/btclib-wallet#267 ·
I maintainer di solito rispondono entro 1 giorno
-
good first issue tech-debt
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
knnmelprop/YAADO#111 ·