Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

[bug]: silent mock drops at recording time — outChan cap of 100 collapses under modest parallel load

オープン
#4,175 コメント 0 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

まだ誰も着手していません。

評価

難易度
3/5
見積もり時間
1〜2日
初心者へのやさしさ
76/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
静か
技術スタック
go
領域
testing

調査の方向性

pkg/service/agent/agent.go の outgoingMockChanCap と pkg/agent/proxy/syncMock/syncMock.go の sendToOutChan から始めます。提供された、並行する producer と遅い consumer のシナリオを使って、pkg/agent/proxy/syncMock/ に対象を絞ったカバレッジを追加します。完了条件は、大きいバッファによって通常のワークロードでの損失が回避され、最初にドロップされた mock が警告を出力し、サンプリングされたエラーのカバレッジも維持されることです。

索引モデルが issue の本文から書いたものです。

説明

👀 Is there an existing issue for this?
  • I have searched and didn't find a matching issue. (PR #4171/#4172 raises the relay-layer buffers; this is a separate, downstream buffer with the same failure mode but a different drop reason.)
👍 Current behavior

The agent → CLI mock channel — make(chan *models.Mock, outgoingMockChanCap) in pkg/service/agent/agent.go:182 — is hard-capped at 100 slots (const outgoingMockChanCap = 100). When the producer (per-test mock buffer flush in syncMockManager) outpaces the consumer (single-goroutine gob encoder + HTTP/2 multipart stream to k8s-proxy + disk write), the 100-slot ceiling is reached fast and sendToOutChan drops mocks after a 200 ms sendBudget.

Two independently observable problems:

  1. The cap is too tight for any non-trivial parallel load. Under sustained back-to-back outbound activity (multiple connections, each emitting tens of mocks per second — the steady-state shape of any Cypress-driven test suite), the channel pegs at 100 within hundreds of milliseconds and stays full for the rest of the recording session. Direct measurement with a synthetic harness (4 producers × 200 mocks, slow consumer at 50 ms/mock) shows 9.8% silent mock loss with cap=100; same workload with cap=16384 gives 0% loss (numbers in the linked PR).
  2. Drops are nearly invisible. The drop log fires only at n == 1 || n%1024 == 0. By the time an operator sees the first sampled Error line, 2048 mocks have already been dropped (n=1 fires "somewhere", n=1024 next, n=2048 third — in any sufficiently long log window the first observed dropsSoFar value is 2048, not 1).

The pair compounds: the cap is too small AND the drops are too quiet, so a recording session can quietly degrade to losing 80%+ of mocks without any operator-visible signal until tests start failing at replay time with keploy-pg-v3: no recorded invocation matched (... reason=no recorded invocation shares this SQL hash).

👟 Steps to Replicate

Real-world repro is "run a Cypress suite against a recorded service for ~5 minutes" — but for a tight unit-test repro (no eBPF, no kubernetes, no real app), the following self-contained Go test in pkg/agent/proxy/syncMock/ reproduces the cliff deterministically:

package manager

import (
    "sync"
    "sync/atomic"
    "testing"
    "time"

    "go.keploy.io/server/v3/pkg/models"
)

func TestOutChanCliffRepro(t *testing.T) {
    for _, tc := range []struct{ name string; cap int }{
        {"old_cap_100", 100},
        {"new_cap_16384", 16384},
    } {
        t.Run(tc.name, func(t *testing.T) {
            ch := make(chan *models.Mock, tc.cap)
            mgr := &SyncMockManager{
                buffer: make([]*models.Mock, 0, defaultMockBufferCapacity),
            }
            mgr.SetOutputChannel(ch)

            // Stalled consumer matching the gob+HTTP+disk pipeline rate.
            var consumed atomic.Int64
            done := make(chan struct{})
            go func() {
                defer close(done)
                t := time.NewTimer(3 * time.Second); defer t.Stop()
                for {
                    select {
                    case _, ok := <-ch:
                        if !ok { return }
                        consumed.Add(1)
                        time.Sleep(50 * time.Millisecond)
                        if !t.Stop() { select { case <-t.C: default: } }
                        t.Reset(3 * time.Second)
                    case <-t.C:
                        return
                    }
                }
            }()

            const producers, perProducer = 4, 200
            var wg sync.WaitGroup; wg.Add(producers)
            for p := 0; p < producers; p++ {
                go func() {
                    defer wg.Done()
                    for i := 0; i < perProducer; i++ {
                        mgr.sendToOutChan(&models.Mock{Spec: models.MockSpec{ReqTimestampMock: time.Now()}})
                    }
                }()
            }
            wg.Wait()
            <-done

            sent := producers * perProducer
            t.Logf("cap=%d sent=%d consumed=%d DROPPED=%d loss=%.1f%%",
                tc.cap, sent, consumed.Load(), mgr.DropCount(),
                float64(mgr.DropCount())*100/float64(sent))
        })
    }
}

Output:

cap=100   sent=800 consumed=722 DROPPED=78 loss=9.8%
cap=16384 sent=800 consumed=800 DROPPED=0  loss=0.0%
📜 Logs (if any)

Real production agent logs from a recording session that hit the cliff (sanitized: pod identifiers and bundle IDs removed):

2026-05-04T20:22:35.666936741Z 🐰 Keploy: ... ERROR
syncMock outChan overflow; mock dropped — raise consumer throughput or increase outChan capacity
{"dropsSoFar": 2048, "outChanCap": 100, "budget": "200ms"}

That 2048 is a floor: it's the third milestone the rate sampler hits (after n==1 and n==1024, both of which were older than the captured log window). Realistic per-pod drop count over a recording session is 2k–3k; across a 4-pod deployment, 8k–12k mocks lost.

The downstream symptom — visible only after recording completes and replay begins:

keploy-pg-v3: no recorded invocation matched (hash=...);
reason=no recorded invocation shares this SQL hash

41 of 41 tests in the affected bundle failed with this shape, all of them caused by missing mocks for queries that did happen during recording but whose mock entries got dropped at the agent → CLI handoff.

💻 Operating system

Linux

🧾 System Info (uname -a)

Linux 6.12.x x86_64 (production: EKS node).

🎲 Version

Reproduces on main (current pkg/service/agent/agent.go:179 still has const outgoingMockChanCap = 100). PR #4171/#4172 makes the relay-layer buffers configurable but does NOT touch this constant — see "Relationship to PR #4172" below.

📦 Repository

keploy

🤔 What use case were you trying?

Recording a real service under realistic Cypress E2E load. The service emitted bursts of outbound HTTP calls (boto3 → AWS SNS/STS) plus dozens of Postgres queries per inbound test request. Across 4 mutated pods the cumulative mock production rate exceeded the consumer's drain rate within minutes, the channel pegged, and the rest of the session dropped mocks silently.

🔧 Relationship to PR #4172 / #4171

PR #4172 raises relay-layer buffers (DefaultPerConnCap 8 → 64 MiB, DefaultTeeChanBuf 64 → 1024) and exposes them as --max-memory-per-conn / --queue-size flags. Those defaults solve the case of a single connection producing a very large response (e.g. 10 MB Postgres blob).

This bug is at a different layer:

real socket → relay forwarder → tee staging chan → FakeConn → parser → EmitMock
                                       ↑
                              PR #4172's TeeChanBuf (raised to 1024)
                                                                ↓
parser → syncMockManager.AddMock → outChan ← THIS BUG (still 100)
                                       ↓
                                gob encoder → HTTP/2 → k8s-proxy → disk

The outChan cap is not configurable through any of #4172's flags. Verified by:

  1. Building PR #4172's binary locally.
  2. Running the burst-load Go client (16 workers × 80 outbound HTTP requests = 1280 mocks).
  3. Bumping --queue-size 100000 --max-memory-per-conn 512MiB (extreme overrides).
  4. Result: still ~70% mocks dropped — because the bottleneck is one layer downstream.

The two PRs are complementary, not redundant.

Suggested fix

Two parts, both no-flag (mirrors how #4172 handles its inline defaults):

  1. pkg/service/agent/agent.go: bump outgoingMockChanCap from 100 to 16384. Sized against production data: per-mock YAML sizes are median ~1 KiB / p95 ~10 KiB / max ~27 KiB, so 16384 slots ≈ 50 MiB at typical sizes, ~5% of the cgroup limits we saw. Same drop policy — drops still happen if the consumer truly stalls — but the cliff disappears for normal parallel test workloads.

  2. pkg/agent/proxy/syncMock/syncMock.go::sendToOutChan: promote a Warn line on the first drop alongside the existing per-1024 sampled Error. Operators stop having to wait through 2047 silent drops before getting an actionable signal that capture is now lossy. The sampled Error continues for sustained-drop telemetry so a stuck consumer doesn't flood the log.

Both changes are in the linked PR.

主要言語
Go
スター
18.5k
フォーク
2.4k
平均マージ
13時間 48分
マージ済み PR(30日)
101

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

keploy/keploy のほかの issue

keploy/keploy の issue をすべて見る

似ている issue

Go の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。