Bug: Segmentation Fault During Periodic Read-Only Ladybug Database Reconnection
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 25/100
Direzione di ricerca
Start with ladybug/database.py around init_pybind_database, then trace the follower_refresh_loop to _refresh_follower_main_database and _open_main_database. Review the static core dump and logs alongside the repeated read_pool_generation changes and checkpoint/WAL errors. Done means the native crash has a reproducible diagnosis and the reconnect behavior is validated under concurrent checkpoint/WAL activity.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Ladybug version
main
What operating system are you using?
ubuntu24.04
What happened?
Conclusion: This was not an OOM event. It was a SIGSEGV (segmentation fault) triggered by the LadybugDB native. The most likely cause is a concurrency or lifecycle race between repeatedly reopening the read-only database, closing old connection pools, and concurrent checkpoint/WAL operations.
Key evidence:
-
The log explicitly reports:
Fatal Python error: Segmentation fault -
The crashing thread was in:
ladybug/database.py -> init_pybind_databaseCall chain:
follower_refresh_loop -> _refresh_follower_main_database -> _open_main_database -> Ladybug Database initializationThis indicates that the crash occurred in Ladybug’s native/pybind layer, rather than in normal Python application code.
-
The database was repeatedly reconnected, approximately once every 15 seconds. The
read_pool_generationincreased from 1 to 19. Each refresh performed the following operations:- Opened the same
main_ontology.lbugdatabase. - Created 8 read connections.
- Switched the active connection pool.
- Closed the previous database instance.
- Opened the same
-
Before the crash, the log showed database-file concurrency/state errors:
Cannot open database in read-only mode while checkpoint is in progressCannot open file ... main_ontology.lbug.wal: No such file or directory
These errors indicate that the database was being opened for reading while its files may have been changing due to a checkpoint, WAL removal/rotation, or another process modifying the database.
-
Memory usage did not reach the container limit:
- cgroup memory limit: 16 GiB
- Last recorded cgroup usage: approximately 2.31 GiB
- Process RSS: approximately 1.97 GiB
- Memory monitor status:
normal
Therefore, this does not resemble a Kubernetes OOM kill. An OOM kill normally results in
OOMKilledand exit code 137, while a segmentation fault commonly results in exit code 139.
Most likely failure sequence:
Checkpoint/WAL state changes
↓
Read-only database opening fails and triggers retries
↓
Repeated open + connection-pool switch + close operations every 15 seconds
↓
A thread-safety or resource-lifecycle issue occurs in the Ladybug native layer
↓
SIGSEGV / core dump
Based on the current logs, it can be confirmed that the segmentation fault occurred during Ladybug native database initialization. The strongest suspected root cause is frequent read-database reconnection while checkpoint/WAL operations are changing the same database files, causing a thread-safety or resource-release issue in the native layer.
Are there known steps to reproduce?
I’m not sure how to go about fixing this issue. Every time I observe the problem, I only get a static core dump, and I cannot reproduce the failure reliably. As a result, it is unclear where exactly the root cause lies. Could you share some suggestions and troubleshooting ideas? @adsharma
- Lingua principale
- C++
- Stelle
- 1.8k
- Fork
- 148
- Merge medio
- 14h 57m
- PR unite (30g)
- 124
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di LadybugDB/ladybug
-
Bug: SET p.prop = NULL ... RETURN p.prop returns the old value for every row except the firstAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
I maintainer di solito rispondono entro 1 giorno
-
Bug: is_sorted assertion in scanCommittedInMem after a failed checkpoint with deleted in-memory relsAperta
Difficoltà 4/5 3-5 giorni Idoneità per principianti 48/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 4/5 3-5 giorni Idoneità per principianti 58/100
LadybugDB/ladybug#1117 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di LadybugDB/ladybug
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
I maintainer di solito rispondono entro 1 giorno
-
hipRTC lit tests compile against /opt/rocm's LLVM instead of the ROCm under test (ci/ hardcodes LLVM_PATH)Forse già presa @bernardogv l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
brndnmtthws/conky#2486 ·
I maintainer di solito rispondono entro 1 giorno
-
請增加教學:數字後的句號Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 70/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100