kubernetes: a single 1-vCPU worker capped at one node cannot fit the cert-manager install
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- Mezza giornata
- Idoneità per principianti
- 65/100
- Tipo di issue
- Bug
- Chiarezza
- Specificata chiaramente
- Stato di attività
- Attiva
- Stack tecnologico
- helm, kubernetes
- Ambito
- devops
Direzione di ricerca
The issue is in the cert-manager addon chart deployment. Start by reviewing the CPU requests for all addons in the Helm values to confirm the overcommit on 1-vCPU nodes. The fix is either to add a validation warning in the chart when a 1-vCPU pool is used with addons enabled, or to document a minimum 2-vCPU requirement. Key files: packages/apps/kubernetes-nodes/values.yaml and the cert-manager chart values.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
A tenant cluster whose only worker pool is one 1-vCPU node that cannot scale out (maxReplicas: 1) can't finish installing the cert-manager addon. The other addons already take that node's whole allocatable CPU. cert-manager's startupapicheck hook Job then has nowhere to run, the install times out, and the reinstall after remediation hits the same full node, so the HelmRelease never turns Ready.
1 vCPU is what a pool gets by default. The default instanceType is u1.medium (kubernetes-nodes values), and every *.nano, *.micro, *.small and *.medium type in packages/system/kubevirt-instancetypes has one vCPU. The default maxReplicas: 10 hides the problem, because the autoscaler adds a node for the Pending pod. It shows up once a pool is capped at one node.
The numbers, at d36784603:
- The kubelet reservations for a 1-vCPU worker are 50m system plus 50m kube (_talosconfigtemplate.tpl), which leaves 900m allocatable.
helm templateof the addon charts with the values a ComputePlane cluster passes them gives these per-node CPU requests: cilium agent, envoy and operator 300m, coredns 2 x 100m, kubevirt-csi-node 20m, ingress-nginx controller plus protobuf-exporter 200m and defaultbackend 10m, metrics-server 100m, cert-manager controller, webhook and cainjector 70m. That is 900m.- The startupapicheck Job asks for 10m more.
The pod only fits if it gets scheduled before the node fills up, and it then completes and gives the 10m back. So the result depends on install order. I saw this in the computeplane e2e suite, which ran one such worker and failed once and passed once on the same charts. The failing run's HelmRelease reported timeout waiting for: [Job/cozy-cert-manager/cert-manager-startupapicheck status: 'InProgress']. No FailedScheduling event was observed, since nothing inside the tenant cluster was collected (#4773). Pending is inferred from the arithmetic and from the cert-manager webhook answering the tenant apiserver for the whole remaining window. The suite now runs two workers.
A ComputePlane cluster always enables cert-manager and ingress-nginx, so it hits this with any single 1-vCPU pool. A plain Kubernetes cluster hits it once addons.certManager.enabled is on.
Possible fixes: have the chart reject or warn about a pool with maxReplicas: 1 and less than about 2 vCPU when the addons are on, or document a minimum pool size.
- Lingua principale
- Go
- Stelle
- 2.2k
- Fork
- 209
- Merge medio
- 3g 10h
- PR unite (30g)
- 237
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di cozystack/cozystack
-
chore(cilium): drop externalIPs.enabled and nodePort.enabled, the Cilium chart does not read themApertaarea/cilium kind/cleanup triage/needs-triage
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
cozystack/cozystack#4816 · 1 reazione ·
I maintainer di solito rispondono entro 2 giorni
-
e2e: cozyreport lists a shared zpool twice and ships an empty zfs-pools.txt that does not say whyApertatriage/needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
cozystack/cozystack#4809 · 1 reazione ·
I maintainer di solito rispondono entro 2 giorni
-
area/ci kind/bug triage/needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
cozystack/cozystack#4806 · 1 commento · 2 reazioni ·
I maintainer di solito rispondono entro 2 giorni
-
triage/needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
cozystack/cozystack#4787 · 1 reazione ·
I maintainer di solito rispondono entro 2 giorni
-
kubernetes-nodes: raising minReplicas does not raise the MachineDeployment's replicas on upgradeApertatriage/needs-triage
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
cozystack/cozystack#4786 · 1 reazione ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di cozystack/cozystack
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
I maintainer di solito rispondono entro 1 giorno
-
Broken Claude manifestAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
-
[Chore] Remove dead AutogenV2 feature flagForse già presa @geeknishantkyeus l’ha presa oggi. Apertabug triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
kyverno/kyverno#17936 · 1 commento · 1 reazione ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100