[U-05][J-05] Document upgrades and multi-AZ behavior
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 68/100
- Tipo di issue
- Documentazione
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- helm, kubernetes
- Ambito
- devops, documentation, infrastructure
Direzione di ricerca
Inizia da charts/graylog/README.md e dal relativo indice, quindi esamina values.yaml, values.schema.json, templates/config/sc/aws-gp3.yaml e docs/RELEASING.md per verificare il comportamento documentato. Aggiungi le sezioni richieste Upgrading e Multi-AZ, gli esempi e la soluzione alternativa per lo schema, crea la voce upgrade-notes.md per ogni release e collegala alla documentazione delle release. Il lavoro è completo quando la README copre sia le procedure operative sia le relative limitazioni senza modificare il comportamento del chart.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The chart exposes updateStrategy.rollingUpdate.partition and type: OnDelete, but the README only lists them in the values tables. There is no procedural upgrade documentation, and nothing covers how availability zones interact with WaitForFirstConsumer storage, which pins each existing ordinal to a zone at first volume bind. Add two README sections.
Details
Current state: partition exists for both tiers (values.yaml:246-250 and 395-399) and renders in both StatefulSets. The README describes it in one line per values table, there is no Upgrading section (Maintenance covers only MongoDB backup and restore), and no AZ content exists in the README or docs/. topologySpreadConstraints is not exposed. The scheduling knobs are nodeSelector, tolerations, and affinity, and setting affinity replaces the default anti-affinity for that tier. One gotcha: values.schema.json (lines 352 and 540) types partition as string or null, so a plain --set with an integer fails validation. Use --set-string or a quoted values-file entry.
Section 1, Upgrading:
- Partition canary: set partition to N-1, upgrade, verify the canary via
GET /api/system/cluster/nodesplus actual throughput, then walk partition down to 0. Same for the Datanode, adding OpenSearch health checks. -
type: OnDeletefor full manual cadence control. - Note that StatefulSets update from the highest ordinal down, so pod-0 goes last automatically.
- Mixed-version UI churn during rolls: recommend ingress session affinity via the existing
ingress.web.annotations, with one controller-specific example. - Start a per-release
upgrade-notes.mdwith the next release and wire it intodocs/RELEASING.md.
Section 2, Multi-AZ:
- Zone pinning: PVCs bind where the pod first schedules and zonal volumes cannot cross zones, so existing ordinals need replacement capacity in their bound zone or sit
Pending. The chart's own AWS gp3 class usesWaitForFirstConsumer(templates/config/sc/aws-gp3.yaml:19). - Non-graceful node loss: roughly 6 minutes before the volume is force-detached, so replacement pods stall at least that long. State the order of magnitude so operators do not read it as a chart bug.
- A zone-spread example. Decide first: write it with the
affinityoverride (documenting that it replaces the default anti-affinity), or add atopologySpreadConstraintspassthrough to both StatefulSets and document that instead. - Scope note: applies to dynamically provisioned per-ordinal PVCs, not
existingClaimor disabled persistence.
Reference: U-05, U-06, J-05 (Production Readiness Review)
Impact
Operators upgrade without a canary because nothing tells them how, hit the --set schema error with no documented workaround, and first learn about zone pinning when a pod sticks in Pending after a node failure.
Notes for maintainers
Both sections go in charts/graylog/README.md plus its table of contents. The partition schema typing deserves a small follow-up fix to also accept integers. Related: #14 tracks HA behavior changes, while this issue only documents current behavior.
- Lingua principale
- Go Template
- Stelle
- 12
- Fork
- 4
- Merge medio
- 6g 2h
- PR unite (30g)
- 3
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di Graylog2/graylog-helm
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
Graylog2/graylog-helm#186 ·
-
Invalid nodeSelectorAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
Graylog2/graylog-helm#97 ·
-
feature
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
Graylog2/graylog-helm#61 ·
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 35/100
Graylog2/graylog-helm#190 · 1 commento ·
-
feature improvement infrastructure
Difficoltà 4/5 3-5 giorni Idoneità per principianti 52/100
Graylog2/graylog-helm#172 ·
Tutte le issue di Graylog2/graylog-helm
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 7 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
Flagsmith/flagsmith-charts#603 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
lightpanda-io/browser#3708 ·
I maintainer di solito rispondono entro 1 giorno
-
area:build enhancement good first issue P2
Difficoltà 2/5 1-3 ore Idoneità per principianti 92/100
uttrflow/uttrflow-swift#3212 ·
I maintainer di solito rispondono entro 1 giorno
-
area: ci type: chore
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
eknowledger/toolbench#140 ·
I maintainer di solito rispondono entro 1 giorno