startHarper readiness deadline and teardown loopback recycle cause intermittent failures under sharded CI
@heskew đang làm issue này rồi.
Từ ngày 4/6/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Summary
Two races in the Harper lifecycle helpers (src/harperLifecycle.ts / dist/harperLifecycle.js) are a recurring source of intermittent failures for consumers that run sharded, concurrent integration suites — most visibly HarperFast/harper's Integration Tests workflow, which currently has to work around both. Filing here because the fix lives in this package.
Tracked downstream in HarperFast/harper#1139.
1. Fixed startup deadline (DEFAULT_STARTUP_TIMEOUT_MS)
startHarper() resolves only when Harper prints successfully started on stdout; otherwise it rejects after DEFAULT_STARTUP_TIMEOUT_MS, which defaults to 60s (harperLifecycle.js:27, read in runHarperCommand at :155-162). Under shard contention on shared CI runners, install + start regularly exceeds 60s, so the consumer papers over it with env bumps:
- HarperFast/harper sets
HARPER_INTEGRATION_TEST_STARTUP_TIMEOUT_MS=120000(Linux) and180000(Windows), with workflow comments explicitly chasing the deadline ("Slow runners ... may need more than the default 60s"; "Windows runners are noticeably slower than the Linux pool").
The fixed deadline is the flake lever: a slow-but-healthy boot is indistinguishable from a hang. Worth weighing a readiness poll / health-probe with backoff instead of a single wall-clock deadline, and/or a higher CI-aware default.
2. Teardown SIGKILL window + immediate loopback recycle
killHarper() (harperLifecycle.js:310-329) sends SIGTERM, then escalates to SIGKILL after only 200ms, then resolves. teardownHarper() (:351-358) then immediately calls releaseLoopbackAddress(), returning the loopback IP to the pool for the next suite to grab. Issues:
- 200ms is short for a graceful shutdown —
teardownHarper's own comment notes "rocksdb may be flushing." The process is frequently SIGKILL'd mid-flush. - Harper spawns worker threads / child processes;
proc.kill()signals only the direct child. Lingering workers can keep holding the fixed ports (9925/9926/9927/1883/8883) on that loopback IP after the parent is gone. - Because the ports are fixed and only the loopback address rotates between suites, the next suite that acquires the just-released IP can hit
EADDRINUSE/ connection races against sockets still inTIME_WAITor held by an orphaned worker.
Related observation: removeDeadProcessesFromPool (loopbackAddressPool.js:289-300) keys liveness on the test-runner PID stored in the pool, not the Harper child PID — so a runner killed while holding addresses (e.g. a timed-out shard) can leave its slot marked in-use until another process reaps it.
Impact
Downstream, these contribute the "ECONNREFUSED on restart" / "address already in use" class of intermittent failures and are the reason the startup-timeout env workarounds exist. (HarperFast/harper Integration Tests have been red on main 7 of the last 8 runs; the dominant single cause is a specific suite, tracked separately downstream, but the harness races are a real secondary contributor.)
Not proposing a specific fix here
Just capturing the mechanics. Options to weigh: readiness polling instead of a fixed deadline; a configurable, longer SIGTERM grace before SIGKILL; verifying the ports are actually free (not just the address bindable) before recycling; tree-kill of worker children on teardown.
🤖 Generated with Claude Code
- Ngôn ngữ chính
- TypeScript
- Star
- 1
- Fork
- 0
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của HarperFast/integration-testing
-
setupHarperWithFixture overwrites ctx.harper, dropping pre-set hostname (breaks multi-node add_node)Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
-
enhancement good first issue
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 74/100
-
A listener on all interfaces silently receives a test node's connections on macOS; the conflict canary cannot see itCó thể đã có người làm @dawsontoth đã nhận hôm nay. Đang mở
HarperFast/integration-testing#38 · 1 bình luận · 1 người được giao ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 58/100
Tất cả issue của HarperFast/integration-testing
Issue tương tự
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 88/100
supabase/agent-skills#611 ·
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 68/100
polka-codes/test#345 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 92/100
GoogleChromeLabs/project-sesame#217 ·
Maintainer thường phản hồi trong vòng 12 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
solana-foundation/solana-com#2202 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 85/100