pull-npd-e2e-test failing ssh handshake
还没有人认领这个 Issue。
评估
- 难度
- 4/5
- 预计耗时
- 3-5 天
- 新手友好度
- 35/100
- Issue 类型
- 缺陷
- 描述清晰度
- 需要澄清
- 活跃度
- 停滞
- 技术栈
- go, kubernetes
- 领域
- ci-cd, infrastructure, testing
调研方向
从 test/e2e/metriconly/metrics_test.go 第 158 行开始,复现链接的 TestGrid 作业中的 pull-npd-e2e-test 失败。检查用于 prow 主机的 SSH 访问,以及收集指标和 journal 日志的命令。当 e2e 测试在没有 SSH 握手失败的情况下完成,并且可以存储调试数据时,即表示完成。
由索引模型根据 Issue 内容生成。
描述
https://testgrid.k8s.io/presubmits-node-problem-detector#pull-npd-e2e-test starts to fail recently.
[1] NPD should export Prometheus metrics. When OOM kills and docker hung happen
[1] NPD should update problem_counter and problem_gauge
[1] /home/prow/go/src/k8s.io/node-problem-detector/test/e2e/metriconly/metrics_test.go:158
[2] error dialing [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:54804->35.184.209.153:22: read: connection reset by peer', retrying
[2] error dialing [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:52980->35.184.209.153:22: read: connection reset by peer', retrying
[2] error dialing [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:53002->35.184.209.153:22: read: connection reset by peer', retrying
[2] error dialing [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:44696->35.184.209.153:22: read: connection reset by peer', retrying
[2] Error storing debugging data to test artifacts: [Error running command: {prow 35.184.209.153 curl http://localhost:20257/metrics 0 error getting SSH client to [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:52990->35.184.209.153:22: read: connection reset by peer'}
[2] Error running command: {prow 35.184.209.153 sudo journalctl -u node-problem-detector.service 0 error getting SSH client to [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:44688->35.184.209.153:22: read: connection reset by peer'}
[2] Error running command: {prow 35.184.209.153 sudo journalctl -k 0 error getting SSH client to [email protected]:22: 'ssh: handshake failed: read tcp 10.32.2.7:44708->35.184.209.153:22: read: connection reset by peer'}
[2] ]
This is affecting several different PRs: https://github.com/kubernetes/node-problem-detector/pull/955, https://github.com/kubernetes/node-problem-detector/pull/961, https://github.com/kubernetes/node-problem-detector/pull/969.
- 主要语言
- Go
- 星标
- 3.5k
- 派生
- 704
- 平均合并
- 12 小时 21 分钟
- 30 天内合并 PR
- 3
环境准备
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
kubernetes/node-problem-detector 的其他 Issue
-
Flag --enable-k8s-exporter=false causes loss of /healthz, /conditions, and /debug/pprof endpoints未关闭
难度 4/5 3-5 天 新手友好度 56/100
kubernetes/node-problem-detector#1332 ·
-
kind/cleanup sig/node
难度 5/5 一周以上 新手友好度 45/100
kubernetes/node-problem-detector#1331 ·
-
NPD needs significant unit test coverage可能重新可做 @DigitalVeer 于 64 天前认领,目前没有进行中的 PR。 未关闭kind/cleanup sig/node
kubernetes/node-problem-detector#1328 · 2 条评论 · 已指派 1 人 ·
-
kind/bug
难度 3/5 1-2 天 新手友好度 58/100
kubernetes/node-problem-detector#1176 · 7 条评论 ·
-
kind/feature lifecycle/frozen
难度 5/5 一周以上 新手友好度 35/100
kubernetes/node-problem-detector#1112 · 8 条评论 ·
查看 kubernetes/node-problem-detector 的全部 Issue
相似的 Issue
-
agent-butler-finding chore
难度 1/5 1 小时以内 新手友好度 88/100
jordansmall/spindrift#4146 ·
维护者通常 1 天内回复
-
security
难度 2/5 1-3 小时 新手友好度 68/100
IBM/ibmcloud-volume-file-vpc#119 ·
-
security
难度 2/5 1-3 小时 新手友好度 66/100
IBM/networking-go-sdk#339 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 88/100
kubernetes-sigs/mcp-lifecycle-operator#439 ·
维护者通常 1 天内回复
-
area: global bug dx priority: low
难度 2/5 1-3 小时 新手友好度 88/100
维护者通常 1 天内回复