Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Failed flush retries re-insert already-committed rows (duplicate usage events)

已关闭
#31 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
5/5
预计耗时
一周以上
新手友好度
35/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
clickhouse, php
领域
databases

调研方向

Start by tracing ClickHouse::addBatch(), Accumulator::flush(), insert(), and generateId() to understand how chunk failures and retries affect buffered rows. Reproduce a mid-batch failure and a timeout-after-server-commit if the existing setup permits. Done means failed flush retries do not duplicate rows already committed to ClickHouse, including when a client timeout follows a successful insert.

由索引模型根据 Issue 内容生成。

描述

Summary

When a flush fails partway through, the retry re-inserts rows that already landed in ClickHouse, producing duplicate rows (and over-counted usage). Three behaviors combine to cause this, observed on 0.14.0:

  1. ClickHouse::addBatch() is chunked but all-or-nothing. It splits the batch into 1000-row INSERTs and only returns true after every chunk succeeds. If chunk N fails, chunks 1..N-1 are already committed server-side, but the caller sees a thrown exception.

  2. Accumulator::flush() retains the whole buffer on failure. Buffer entries are only cleared when addBatch() returns true, so after a mid-batch failure the next flush re-sends all entries — including the ones whose chunks already succeeded.

  3. Inserts are not idempotent. insert()'s own comment says so: MergeTree has no row-level dedup, and each retry calls generateId() again, so re-sent rows get fresh ids and can't be deduplicated by block hash either.

There's a second path to the same outcome with no chunking involved: a client-side timeout on an insert that the server actually completed (we observed a burst of Operation timed out after 30s in production while the server was demonstrably healthy and ingesting). The retry then duplicates the full batch.

Observed in production

Appwrite Cloud stats-usage workers logging, e.g.:

ClickHouse insert failed: Operation timed out
  [Operation: addBatch(), Table: projects_usage_events, Query: INSERT INTO projects_usage_events (1000 rows)]
ClickHouse insert failed: Connection reset by peer
  [Operation: addBatch(), Table: projects_usage_events, Query: INSERT INTO projects_usage_events (890 rows)]

The 890 rows failure is a tail chunk — the preceding 1000-row chunks of that same addBatch() call had already been inserted, and were re-inserted on the next flush.

Suggested directions

  • Make Accumulator::flush()/addBatch() clear buffer entries per successful chunk rather than per call, so a tail-chunk failure doesn't re-send committed chunks; and/or
  • Make retries idempotent: deterministic row ids derived from the buffered entry (not generateId() at encode time) plus stable insert blocks, so ClickHouse's insert_deduplicate block-hash dedup can absorb replays (or a ReplacingMergeTree/dedup-on-read scheme).

The timeout-after-server-commit case can only be fully solved by idempotency, not by smarter chunk bookkeeping.

主要语言
PHP
星标
0
派生
0
平均合并
1 小时 56 分钟
30 天内合并 PR
2

环境准备

  • 提供 Dockerfile 或 Docker Compose 文件
  • 没有 Pull Request 模板
  • 阅读贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

相似的 Issue

更多 PHP Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。