Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

[Bug]: SpannerIO Change Streams may skip ChildPartitionsRecord when advancing tracker to artificial query end timestamp

未关闭
#40,232 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

维护者通常 1 天内回复

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
45/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
google-cloud, java

调研方向

The issue is in QueryChangeStreamAction.java around line 360. Start by understanding the restriction tracker and how it claims timestamps. Look at the logic for unbounded queries when stopAfterQuerySucceeds is false. The fix is to avoid advancing the tracker to the artificial changeStreamQueryEndTimestamp and instead leave it at the last claimed position. Run tests related to Spanner Change Streams to verify the fix doesn't break existing behavior.

由索引模型根据 Issue 内容生成。

描述

io java P1
What happened?

When reading from an unbounded Spanner Change Stream (or a Mutable Change Stream with a capped query end timestamp), QueryChangeStreamAction issues change stream queries with an artificial changeStreamQueryEndTimestamp (now + 2 minutes).

Previously, when the query completed and needed to resume (!stopAfterQuerySucceeds), QueryChangeStreamAction called tracker.tryClaim(changeStreamQueryEndTimestamp) before returning ProcessContinuation.resume():

https://github.com/apache/beam/blob/ac4acbb6282/sdks/java/io/google-cloud-platform/src/main/java/org/apache/beam/sdk/io/gcp/spanner/changestreams/action/QueryChangeStreamAction.java#L360-L365

In a very rare timing corner case where a change stream query finishes without returning the ChildPartitionsRecord at the end of a partition's range, claiming changeStreamQueryEndTimestamp can advance the restriction tracker past the partition's actual end timestamp. On the next continuation, the query resumes from changeStreamQueryEndTimestamp + 1ns, which falls outside the partition's valid timestamp range and returns an out-of-range start_timestamp error, causing QueryChangeStreamAction to mark the partition FINISHED before scheduling its child partitions.

Workaround & Scope

This is a client-side workaround in Apache Beam for this rare edge case while a fix is being implemented on the Spanner server side:

  • Unbounded queries: When !stopAfterQuerySucceeds, leaving the restriction tracker at the last claimed position (from the last processed data or heartbeat record) instead of advancing to changeStreamQueryEndTimestamp ensures the subsequent query resumes from lastClaimedTimestamp + 1ns and reads any remaining records (including ChildPartitionsRecord) before the partition ends. This workaround addresses the issue for unbounded queries.
  • Bounded queries: For bounded queries where changeStreamQueryEndTimestamp reaches endTimestamp (stopAfterQuerySucceeds == true), this client-side workaround does not apply and will be addressed by the Spanner server-side fix.
Issue Priority

Priority: 2 (default / most bugs should be filed as P2)

Issue Components
  • Component: Java SDK
  • Component: IO Connectors
主要语言
Java
星标
8.7k
派生
4.7k
平均合并
2 天 8 小时
30 天内合并 PR
246

环境准备

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

apache/beam 的其他 Issue

查看 apache/beam 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。