Gauge for FilesystemIsReadOnly not downgraded to 0 after fixing the problem
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 45/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 静か
- 技術スタック
- go
調査の方向性
config/kernel-monitor.json と 160 行目付近の log_monitor.go から始め、永続的な FilesystemIsReadOnly 条件が problem_counter と problem_gauge をどのように更新するかを追跡します。指定されたメッセージを /dev/kmsg に注入して問題を再現し、ファイルシステムを修正する前後でメトリクスを比較します。永続的な条件に対して期待されるリセット動作が、pod を削除せずに文書化または修正されれば完了です。
索引モデルが issue の本文から書いたものです。
説明
The problem occurred when filesystem went to read only mode. That was fixed, but still in the metrics I was able to see the counter and gauge set up to 1.
I conducted a test and multiple times injected the FileSystemIsReadOnly to the /dev/kmsg (https://github.com/kubernetes/node-problem-detector/blob/master/config/kernel-monitor.json):
1 log_monitor.go:160] New status generated: &{Source:kernel-monitor Events:[{Severity:info Timestamp:2020-10-08 06:44:16.09315274 +0000 UTC m=+1331754.148888064 Reason:FilesystemIsReadOnly Message:Node condition ReadonlyFilesystem is now: True, reason: FilesystemIsReadOnly}] Conditions:[{Type:KernelDeadlock Status:False Transition:2020-09-22 20:48:21.98500453 +0000 UTC m=+0.040739839 Reason:KernelHasNoDeadlock Message:kernel has no deadlock} {Type:ReadonlyFilesystem Status:True Transition:2020-10-08 06:44:16.09315274 +0000 UTC m=+1331754.148888064 Reason:FilesystemIsReadOnly Message:Remounting filesystem read-only}]}
Still the metrics were shown as 1 and it did not downgraded to 0. Even the the issue with ro filesystem was fixed, still the metric was 1:
problem_counter{reason="FilesystemIsReadOnly"} 1
problem_gauge{reason="FilesystemIsReadOnly",type="ReadonlyFilesystem"} 1
As a workaround the pod was deleted and after that metrics were reset to 0.
What is the reason of that behaviour? The type "permanent"? Is deleting a pod the only solution?
kernel-monitor.json
{
"type": "permanent",
"condition": "ReadonlyFilesystem",
"reason": "FilesystemIsReadOnly",
"pattern": "Remounting filesystem read-only"
}
- 主要言語
- Go
- スター
- 3.5k
- フォーク
- 704
- 平均マージ
- 12時間 21分
- マージ済み PR(30日)
- 3
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
kubernetes/node-problem-detector のほかの issue
-
Flag --enable-k8s-exporter=false causes loss of /healthz, /conditions, and /debug/pprof endpointsオープン
難易度 4/5 3〜5日 初心者へのやさしさ 56/100
kubernetes/node-problem-detector#1332 ·
-
kind/cleanup sig/node
難易度 5/5 1週間以上 初心者へのやさしさ 45/100
kubernetes/node-problem-detector#1331 ·
-
NPD needs significant unit test coverage再び着手できるかも @DigitalVeer が 63 日前に担当しましたが、オープン中のプルリクエストはありません。 オープンkind/cleanup sig/node
kubernetes/node-problem-detector#1328 · コメント 2 件 · 担当者 1 名 ·
-
kind/bug
難易度 3/5 1〜2日 初心者へのやさしさ 58/100
kubernetes/node-problem-detector#1176 · コメント 7 件 ·
-
kind/feature lifecycle/frozen
難易度 5/5 1週間以上 初心者へのやさしさ 35/100
kubernetes/node-problem-detector#1112 · コメント 8 件 ·
kubernetes/node-problem-detector の issue をすべて見る
似ている issue
-
agentic-workflows
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 1 日以内に返信
-
priority/4/normal status/needs-triage type/bug/unconfirmed
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
authelia/authelia#13292 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
blinklabs-io/actions#138 ·
メンテナーはふだん 1 日以内に返信
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpointオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信