Failed flush retries re-insert already-committed rows (duplicate usage events)
还没有人认领这个 Issue。
评估
调研方向
Start by tracing ClickHouse::addBatch(), Accumulator::flush(), insert(), and generateId() to understand how chunk failures and retries affect buffered rows. Reproduce a mid-batch failure and a timeout-after-server-commit if the existing setup permits. Done means failed flush retries do not duplicate rows already committed to ClickHouse, including when a client timeout follows a successful insert.
由索引模型根据 Issue 内容生成。
描述
Summary
When a flush fails partway through, the retry re-inserts rows that already landed in ClickHouse, producing duplicate rows (and over-counted usage). Three behaviors combine to cause this, observed on 0.14.0:
-
ClickHouse::addBatch()is chunked but all-or-nothing. It splits the batch into 1000-rowINSERTs and only returnstrueafter every chunk succeeds. If chunk N fails, chunks 1..N-1 are already committed server-side, but the caller sees a thrown exception. -
Accumulator::flush()retains the whole buffer on failure. Buffer entries are only cleared whenaddBatch()returnstrue, so after a mid-batch failure the next flush re-sends all entries — including the ones whose chunks already succeeded. -
Inserts are not idempotent.
insert()'s own comment says so: MergeTree has no row-level dedup, and each retry callsgenerateId()again, so re-sent rows get fresh ids and can't be deduplicated by block hash either.
There's a second path to the same outcome with no chunking involved: a client-side timeout on an insert that the server actually completed (we observed a burst of Operation timed out after 30s in production while the server was demonstrably healthy and ingesting). The retry then duplicates the full batch.
Observed in production
Appwrite Cloud stats-usage workers logging, e.g.:
ClickHouse insert failed: Operation timed out
[Operation: addBatch(), Table: projects_usage_events, Query: INSERT INTO projects_usage_events (1000 rows)]
ClickHouse insert failed: Connection reset by peer
[Operation: addBatch(), Table: projects_usage_events, Query: INSERT INTO projects_usage_events (890 rows)]
The 890 rows failure is a tail chunk — the preceding 1000-row chunks of that same addBatch() call had already been inserted, and were re-inserted on the next flush.
Suggested directions
- Make
Accumulator::flush()/addBatch()clear buffer entries per successful chunk rather than per call, so a tail-chunk failure doesn't re-send committed chunks; and/or - Make retries idempotent: deterministic row ids derived from the buffered entry (not
generateId()at encode time) plus stable insert blocks, so ClickHouse'sinsert_deduplicateblock-hash dedup can absorb replays (or aReplacingMergeTree/dedup-on-read scheme).
The timeout-after-server-commit case can only be fully solved by idempotency, not by smarter chunk bookkeeping.
- 主要语言
- PHP
- 星标
- 0
- 派生
- 0
- 平均合并
- 1 小时 56 分钟
- 30 天内合并 PR
- 2
环境准备
- 提供 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
相似的 Issue
-
难度 2/5 1-3 小时 新手友好度 72/100
thephpleague/commonmark#1159 ·
-
bug
难度 2/5 1-3 小时 新手友好度 76/100
awslabs/aidlc-workflows#1879 ·
维护者通常 1 天内回复
-
[Bug] The PHP file that lists DNS records truncates records to 12 characters?可能已有人在做 @sahsanu 今天认领。 未关闭bug
难度 2/5 1-3 小时 新手友好度 78/100
hestiacp/hestiacp#5769 · 2 条评论 ·
维护者通常 1 天内回复
-
bug customer-reported
难度 2/5 1-3 小时 新手友好度 82/100
MagnaCapax/PMSS#1011 ·
维护者通常 5 天内回复
-
Talk Review
难度 2/5 1-3 小时 新手友好度 66/100
socallinuxexpo/scale-drupal#351 ·