Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[fleet-standing harvest] Deferred-retention in the watcher is unbounded and invisible — #633's reshaped cap/backpressure + telemetry lane

Open
#672 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
38/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Quiet
Tech stack
python

Research direction

Start in watcher.py:864-871 and 937-961, then trace the watcher health and flush telemetry mentioned in the issue. Define the retention cap or backpressure behavior and the actionable threshold, ensuring retained count and age are visible. Done means normal deferrals cannot grow without bound and their accumulation is observable; Item 24 test isolation is intended to come first.

Written by the indexing model from the issue text.

Description

#633 round 4: ITERATE, then reshaped rather than iterated a fifth time. The blocker is real. At cdd0a742, deferred retention is unbounded and invisible:

callback deferring every input retained 1,000/1,000  while BRAINLAYER_WATCHER_FLUSH_RETAIN_LIMIT resolved to 3
_flush_failures == 0  →  retained_failed_input_count() == 0  →  health: alerting=false, failed=0/min
normalized_jsonl_entries_per_minute = 60   BUT   active_jsonl_entries_per_minute = 0

The retain limit only applies to raised exceptions (watcher.py:937-961), not to normal deferrals; and tick() retries the ENTIRE buffer (:864-871) under the shared indexer lock. So a sustained broken T3 DB grows a global queue that is rescanned every tick — turning F1's availability win into memory/CPU starvation for the very plain-Claude ingestion F1 exists to protect.

The reshape:

  1. #633 stays OPEN and stays UNMERGED. Merging an unbounded accumulator that can starve ingestion is not available, and deferring the bound to a follow-up is the "known issues is a permission slip" move.
  2. The retention bound + observability becomes its own clearly-scoped lane: cap or backpressure, retained count + age in watcher health and flush telemetry, actionable threshold. The generic T3 schema alarm does not observe accumulation.
  3. Item 24 (test isolation) goes FIRST — it is the multiplier.

Ruled out by execution: the restart-loss risk does not reproduce; the unadvanced offset plus the source JSONL is the durable replay log.

Archive reference: collab/archive/FLEET-STANDING-archive-2026-08-08.md line 13868

— EtanHey's maintenanceClaude (lead) · claude-code/claude-fable-5

Dominant language
Python
Stars
9
Forks
7
Avg merge
2h 8m
Merged PRs (30d)
211

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from EtanHey/brainlayer

All issues in EtanHey/brainlayer

Similar issues

More Python issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.