flushPendingStores runs refreshSearchStatistics under the process lock — multi-minute write stall on daemon restart
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 45/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- python, sqlite
- Domain
- backend, databases, performance
Research direction
Start at flushPendingStores and trace its call to BrainDatabase.refreshSearchStatistics, focusing on the process lock and write transaction around the call. Compare the three proposed approaches and inspect related issue #689's load-sensitive prepush test. Done means daemon restart flushing no longer holds the live database lock during the multi-minute statistics walk, with queue and chunk processing continuing safely.
Written by the indexing model from the issue text.
Description
Observed 2026-08-10 ~11:40-11:46: BrainBarDaemon restarted at 11:39:52 with pending stores queued; its flush path called BrainDatabase.refreshSearchStatistics() inside flushPendingStores while holding withPendingStoreProcessLock + a write txn. The stats refresh walks a full sqlite3BtreeCount over the 744k-chunk table (cold cache → ~6 minutes of disk btree walking). During that window: drain daemon stuck in SQLITE_BUSY open-retry (attempt 6/13 observed), watcher logging database-is-locked, maintenance seat's brain_store hard-failed at 120s transport abort. Self-resolved when the count finished; queue and chunks resumed immediately.
Stack sample (885/885 in one count): flushPendingStores → refreshSearchStatistics → sqlite3_exec → sqlite3BtreeCount → moveToChild → getAndInitPage.
Fix candidates: (1) move stats refresh OUTSIDE the process lock / write txn; (2) make it incremental or approximate (sqlite_stat1 or cached count with dirty flag); (3) skip stats refresh during flush entirely — flush is a replay path, not a query path.
Related: #692 (VACUUM freeze — same class: heavyweight ops holding write locks on the live DB), #689 (load-sensitive prepush test).
Filed by brainlayerClaude lead (Fable 5).
🤖 Generated with Claude Code
- Dominant language
- Python
- Stars
- 9
- Forks
- 7
- Avg merge
- 2h 8m
- Merged PRs (30d)
- 211
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from EtanHey/brainlayer
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
EtanHey/brainlayer#999 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
EtanHey/brainlayer#986 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
EtanHey/brainlayer#985 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
EtanHey/brainlayer#982 ·
Maintainers usually reply within 1 day
-
Difficulty 1/5 Under an hour Newbie friendliness 88/100
EtanHey/brainlayer#676 ·
Maintainers usually reply within 1 day
All issues in EtanHey/brainlayer
Similar issues
-
[Bug] @deck.gl/arcgis dist import resolves to unpublished @deck.gl/core source path (9.3.11, 9.4.0)Open
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
workflow: a tick's dispatch counts as 'only this step', and no review self-grants a round unattendedOpenworkflow
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
kristofdegrave/homeassistant-smart-charging#1505 ·
Maintainers usually reply within 1 day
-
metadata submission
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
canonical/content-cache-operator#163 · 1 comment ·
Maintainers usually reply within 1 day
-
[submission]Opensubmission
Difficulty 1/5 Under an hour Newbie friendliness 65/100
leanprover/lean-eval-submissions#1852 ·
Maintainers usually reply within 1 day