EPU Parser: potential race condition
@vredchenko ya está trabajando en esto.
Desde el 14/2/2025.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Status
The startup race window is confirmed to exist in current main. Actual data loss is
plausible but unproven — it has never been reproduced. The original description of this
issue overstated both the certainty and the size of the failure, which is likely why attempts
to reproduce it have come up empty.
This issue is now scoped as: build a test that can catch it, then fix what the test proves.
Corrected mechanics
smartem-agent watch parses the existing tree to completion, and only then starts the observer
(src/smartem_agent/__main__.py):
logging.info("Parsing existing directory contents...") # ~273
watcher.datastore = EpuParser.parse_epu_output_dir(watcher.datastore)
logging.info("..done! Now listening for new filesystem events")
observer = Observer()
observer.schedule(watcher, str(path), recursive=True) # ~278
...
observer.start() # ~296
Two corrections to the original write-up:
-
The window is larger than previously stated. It is not "parse-end to
schedule()".
EpuParser.parse_epu_output_diropens with
list(datastore.root_dir.glob("**/*EpuSession.dm"))— a full recursive walk materialised
at t0 — and the OS-level watch is not established untilobserver.start(). The exposure is
the entire parse duration. -
The window is far narrower in effect than "any data arriving is dropped".
SmartEMWatcherV2.watched_event_typesis["created", "modified"](fs_watcher.py:39),
not["created"]. Any file still being written, or ever rewritten, when the observer starts
is still picked up via itsmodifiedevent. The only genuinely lost file is one that is
created, finalised, and never touched again, entirely inside the window.
There is no mitigation elsewhere: on_any_event is the only event method in fs_watcher.py,
and there is no rescan, reconciliation, or periodic full-scan anywhere in the watcher. The
orphan manager handles out-of-order events, not missing ones.
Working hypothesis: only reachable on mid-session restart
On a fresh acquisition the initial glob finds almost nothing, so the window is near-zero and
the bug is probably not reachable at all. It should only open meaningfully when the agent
is restarted against a large, in-progress session, where the parse can run for a long time
while EPU keeps writing.
If correct, this explains the issue's history: anyone reproducing it the obvious way (start the
agent, drop files in) would see nothing and reasonably conclude it was not real.
Unknown, and not answerable from this repository
Whether EPU ever writes a matched file write-once-never-touch is the deciding fact, and it
is ThermoFisher's proprietary behaviour. If every EPU file gets a later modified event, this
bug does not exist in practice. That question is what the test needs to settle empirically.
Plan
- Catch it first. Add a regression test that reproduces the window deterministically
before changing any production code. Shape: pre-populate a large EPU tree, begin
watch, write additional files during the parse, then assert the resulting datastore is
equivalent to a clean full parse of the final tree. EPUPlayer
(smartem-devtools/packages/smartem-epuplayer/) is the natural driver for the concurrent
writes. Home:tests/smartem_agent/test_fs_watcher.py.
To make the window wide enough to hit reliably, the parse should be slowed via injection
(a hook or monkeypatchedparse_epu_output_dir) rather than by building a genuinely huge
fixture. - Fix it. Only once the test fails for the right reason. Likely direction: start the
observer before the parse and buffer events until the parse completes, reconciling
against the parsed datastore (writes are already idempotent by UUID). Alternatives
considered: post-start revalidation scan; atomic transition. - Retain the test. It stays in the suite as a permanent regression guard.
Run it on Windows as well as Linux
Production agents run on Windows EPU workstations, and the two platforms fail differently
here — in both directions. A green Linux suite is not evidence about production.
- Linux (inotify): inotify is not recursive. Watchdog registers a watch per subdirectory by
walking the tree at emitter start, and only adds watches for new subdirectories when it
processes their creation event. EPU creates deep subdirectories constantly
(GridSquare_*/FoilHoles/...), so there is a second, independent race: files landing in a
brand-new subdirectory before its watch is registered. Linux-only. - Windows (
ReadDirectoryChangesW): a single handle withbWatchSubtree=TRUEcovers the
whole subtree, so the per-subdirectory race does not arise. However the kernel notification
buffer can overflow under burst load and drop events wholesale — a documented characteristic
of that API that watchdog surfaces poorly. Windows-only, and closer to real acquisition load.
The test should assert the same invariant on both platforms.
This is a small CI change, not a new pipeline:
.github/workflows/ci.yml:44currently reads
runs-on: ["ubuntu-latest"] # can add windows-latest, macos-latest._test.ymlalready takesruns-onas an input, so the matrix is parameterised.build_win_smartem_agent.ymlalready runs onwindows-latest, so the Windows toolchain is
proven to install.
Note that adding windows-latest to the full test matrix will also surface any unrelated
Windows failures in the existing suite; consider a dedicated job for the filesystem tests if
that turns out to be noisy.
Code references
src/smartem_agent/__main__.py—watchcommand, ~273-296src/smartem_agent/fs_watcher.py:39—watched_event_typessrc/smartem_agent/fs_watcher.py:433—on_any_event(sole event method)src/smartem_agent/fs_parser.py:963—parse_epu_output_dir
- Lenguaje dominante
- TypeScript
- Estrellas
- 0
- Forks
- 0
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de DiamondLightSource/smartem-devtools
-
security
Dificultad 4/5 3-5 días Aptitud para principiantes 48/100
-
Dependency DashboardAbierto
Dificultad 5/5 Más de una semana Aptitud para principiantes 15/100
-
research security
Dificultad 5/5 Más de una semana Aptitud para principiantes 35/100
-
devops research smartem-agent
Dificultad 5/5 Más de una semana Aptitud para principiantes 25/100
-
Make Claude's project knowledge portable: private memory/transcripts, derived public AGENTS.mdAbiertoenhancement smartem-devtools:claude
Dificultad 5/5 Más de una semana Aptitud para principiantes 45/100
Todos los issues de DiamondLightSource/smartem-devtools
Issues similares
-
Dependencies view: `getParent` loops forever on untitled documents, extension host runs out of memoryPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 Menos de una hora Aptitud para principiantes 82/100
awslabs/visual-asset-management-system#413 ·
Los mantenedores suelen responder en 1 día
-
bug confirmed perf
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
videojs/video.js#9400 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug pending triage scope/agent
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
Los mantenedores suelen responder en 1 día
-
good first issue hacktoberfest
Dificultad 2/5 Medio día Aptitud para principiantes 70/100
HelpCode-ai/anythingmcp#996 ·
Los mantenedores suelen responder en 1 día