Unable to auto-scale Kubernetes cluster
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 35/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 停滞
- 技術スタック
- kubernetes
- 領域
- cloud, infrastructure
調査の方向性
Issue に記載されている CloudStack 4.18 のセットアップ、Kubernetes 1.24 ISO、cluster-autoscaler のデプロイで問題を再現します。まず cluster-autoscaler pod のログ、特に RBAC と unregistered-node の警告を確認します。原因を特定し、報告されたエラーなしで期待される autoscaling を復旧できれば完了です。
索引モデルが issue の本文から書いたものです。
説明
Hi!
I am unable to auto-scale Kubernetes clusters. As I understand, it create a "cluster-autoscaler" deployment that decides whether to scale or not. However, it does not seem to work, since it logs multiple errors and warnings in the pod, even though it is a completely clean cluster.
Normal scaling seems to work just fine.
Setup
A "default" CloudStack setup 4.18 running KVMs.
Settings (relevant)
- Cloud kubernetes service enabled true
- Cloud kubernetes cluster experimental features enabled true
- Cloud kubernetes cluster max size 50
The nodes uses the following service offering:
- 2 CPU x 2.05 Ghz
- 2048 MB memory
- 8 GB root disk
Replicate
-
Create a new cluster using Kubernets 1.24 ISO found here:
http://download.cloudstack.org/cks/ -
Enable forced auto-scaling
Since the cluster starts with only one worker node, auto-scaling with 3-5 nodes should trigger an upscale (I assume)
-
Check the logs for cluster-autoscaler in the Kubernetes cluster
Some notable entries:
E0807 14:41:30.317148 1 reflector.go:138] k8s.io/client-go/informers/factory.go:134: Failed to watch *v1.CSIDriver: failed to list *v1.CSIDriver: csidrivers.storage.k8s.io is forbidden: User "system:serviceaccount:kube-system:cluster-autoscaler" cannot list resource "csidrivers" in API group "storage.k8s.io" at the cluster scope
E0807 14:41:32.388828 1 reflector.go:138] k8s.io/client-go/informers/factory.go:134: Failed to watch *v1beta1.CSIStorageCapacity: failed to list *v1beta1.CSIStorageCapacity: csistoragecapacities.storage.k8s.io is forbidden: User "system:serviceaccount:kube-system:cluster-autoscaler" cannot list resource "csistoragecapacities" in API group "storage.k8s.io" at the cluster scope
Even though I have not edited anything myself (just a clean CKS cluster), I get these weird logs:
W0807 14:41:43.251280 1 clusterstate.go:590] Failed to get nodegroup for 6a4c91a3-9694-4596-9ddd-dc86e60136ff: Unable to find node 6a4c91a3-9694-4596-9ddd-dc86e60136ff in cluster
W0807 14:41:43.251361 1 clusterstate.go:590] Failed to get nodegroup for bd0b855f-6dc6-4678-9bea-b52329333024: Unable to find node bd0b855f-6dc6-4678-9bea-b52329333024 in cluster
I0807 14:57:06.667061 1 static_autoscaler.go:341] 2 unregistered nodes present
The IDs are correct in CloudStack
The entire log:
logs-from-cluster-autoscaler-in-cluster-autoscaler-5bf887ddd8-hxg2g.log
Please tell me if you need more logs to look at, or if I should try some other configuration.
Thanks!
- 主要言語
- Go
- スター
- 52
- フォーク
- 34
- 平均マージ
- 6日 1時間
- マージ済み PR(30日)
- 3
コントリビューションガイド
このリポジトリのコントリビューションガイドは索引されていません
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
apache/cloudstack-kubernetes-provider のほかの issue
-
Network ACL entries are shared per tier but deleted per Service, taking other Services' ports down オープン
難易度 4/5 3〜5日 初心者へのやさしさ 45/100
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
apache/cloudstack-kubernetes-provider#75 · コメント 2 件 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
apache/cloudstack-kubernetes-provider#73 · コメント 2 件 ·
-
難易度 5/5 1週間以上 初心者へのやさしさ 20/100
-
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
apache/cloudstack-kubernetes-provider の issue をすべて見る
似ている issue
-
bug github_actions
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
registrystack/registry-stack#1393 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
JakeChampion/lang#10213 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
oasisprotocol/oasis-sdk#2523 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100