Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

ZooKeeper znode controller: release finalizer without connecting when the parent is deleting

未关闭
#1,049 1 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
3/5
预计耗时
1-2 天
新手友好度
68/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
kubernetes, rust

调研方向

从 finalizer::Event::Cleanup 分支开始,跟踪它如何处理带有 deletionTimestamp 的被引用 ZookeeperCluster。检查 ensure_znode_missing 清理路径,以及暴露出卡在 Terminating 状态的 namespace 的集成测试。完成的标准是删除操作跳过 ZooKeeper 连接,并在没有数分钟 backoff 的情况下释放 finalizer。

由索引模型根据 Issue 内容生成。

描述

Problem

In some scenarios, the ZookeeperZnode finalizer (zookeeper.stackable.tech/znode) can take around 130–150s to release. That's what sometimes leaves namespaces stuck in Terminating in our integration tests.

This can happen when the operator tries connection to Zookeeper to delete the node, while ZooKeeper is being torn down, the delete/cleanup path (ensure_znode_missing) runs its errors through controller-runtime's exponential backoff, and that's where the multi-minute stall comes from.

Mechanism

  1. znode cleanup (ensure_znode_missing) tries to connect to the ZK service and gets Connection refused — the ZK pods/endpoints are already gone — so the reconcile errors out.
  2. The ZookeeperCluster CR then drops out of the watch cache. At this point the finalizer's fast path (cluster doesn't exist → assume the znode is gone → drop the finalizer without connecting) would kick in, but the failed reconcile is already sitting in exponential backoff.
  3. ~139s gap: nothing re-runs, even though the fast-path condition is now true.
  4. Backoff finally expires, the reconcile re-runs, the fast path fires, the finalizer is removed, and the namespace deletes.

So it comes down to queue ordering under load. If cleanup runs after the CR leaves the store, it's instant if it runs before, it errors, hits backoff, and takes 130s+.

Fix

In the finalizer::Event::Cleanup arm: if the referenced ZookeeperCluster has a deletionTimestamp, drop the finalizer straight away without connecting to ZK. Retrying an unreachable server makes sense on the create path; on the delete path it shouldn't be allowed to block teardown.

主要语言
Rust
星标
37
派生
12
平均合并
21 小时
30 天内合并 PR
9

环境准备

  • 没有 Dockerfile 或 Docker Compose 文件
  • 有 Pull Request 模板
  • 没有贡献指南

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

stackabletech/zookeeper-operator 的其他 Issue

查看 stackabletech/zookeeper-operator 的全部 Issue

相似的 Issue

更多 Rust Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。