Make Kubernetes job network readiness checks bounded and deployment-independent
I maintainer di solito rispondono entro 1 giorno
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- kubernetes, python
- Ambito
- devops, infrastructure
Direzione di ricerca
Prima ottieni la decisione del maintainer e dell’operatore Kubernetes sulla strategia di readiness e sui valori di timing, poi esamina kernelci/kbuild.py e config/runtime/base/python.jinja2. Aggiungi test unitari e di rendering dei template mirati, con requests e attese simulati, e verifica che i fallimenti delimitati siano classificati come errori di infrastruttura; per considerarlo completato è inoltre necessaria un’esecuzione di kbuild in staging su un nodo appena avviato, collegata qui.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Problem
KernelCI contains two startup workarounds for Kubernetes nodes whose networking or DNS may not be immediately ready.
KBuild._verify_network() performs an HTTP request to Google before generating the build script:
requests.get("https://google.com")
The current implementation has two failure modes:
- the request has no timeout and can hang indefinitely;
- a non-200 HTTP response does not decrement
retries, causing an infinite loop.
The generic Python job template contains a separate DNS workaround that resolves www.google.com up to 30 times before starting a job.
Both checks depend on an unrelated third-party service. Reaching Google does not prove that the KernelCI API, artifact storage, source repositories, or other services required by the job are reachable. Conversely, a Google outage or network policy blocking Google can prevent otherwise healthy KernelCI jobs from starting.
The HTTP check was originally introduced because networking could be slow to initialize on newly started Kubernetes nodes.
Relevant code
kernelci/kbuild.py:KBuild._verify_network()and its call fromwrite_script()config/runtime/base/python.jinja2: DNS readiness loop inmain()
ChromeOS/Tast-specific Google dependencies are being removed separately in #3196 and are outside this issue.
Desired outcome
Network initialization handling must:
- have a strict upper time limit;
- not depend on an unrelated public service;
- test connectivity relevant to the operation the job is about to perform;
- produce an actionable infrastructure error when readiness is not achieved;
- behave consistently between kbuild and other Python jobs.
Maintainer decision required
Before implementation, a KernelCI maintainer or Kubernetes operator should choose one of these strategies:
-
Retry the first required operation — recommended
Remove the generic preflight probes and apply bounded retry handling to the first real API, storage, source, or artifact request.
-
Use a configurable readiness target
Keep a preflight check, but provide the hostname or URL through runtime configuration. The default should be a KernelCI-controlled service required by the job.
The selected maximum startup delay, request timeout, and retry interval must also be recorded in this issue.
Implementation outline
After the strategy is confirmed:
- Remove both hard-coded Google readiness checks.
- Ensure every attempt has a connection and read timeout.
- Use a fixed attempt count or monotonic deadline so every failure path terminates.
- Handle expected network exceptions explicitly rather than catching every
Exception. - Avoid calling
sys.exit()from a low-level helper; return or raise an error that the job runner can classify. - Report exhausted network readiness as an infrastructure failure rather than a kernel build failure.
- Include the attempted service and elapsed time in the final error without exposing credentials or tokens.
- Add focused unit and template-rendering tests.
Required tests
Depending on the selected strategy, cover:
- immediate success;
- initial failures followed by success;
- DNS failure;
- connection timeout;
- read timeout;
- repeated non-success HTTP responses;
- retry exhaustion;
- enforcement of the maximum elapsed time;
- correct infrastructure-error classification.
Tests must mock network requests and waiting so they do not contact external services or introduce real delays.
Human participation required
This issue intentionally needs two concise operational checkpoints:
- A Kubernetes operator confirms whether delayed networking still occurs and approves the readiness strategy and timing values.
- After implementation, an operator runs a kbuild job on a newly scaled or cold Kubernetes node and links the result here.
Repository tests can prove that retry handling is bounded, but they cannot reproduce the real network initialization behavior of the production clusters.
Acceptance criteria
- Generic KernelCI startup code no longer uses
google.comas a connectivity probe. - No readiness or retry path can loop or block indefinitely.
- The total readiness wait is bounded by the maintainer-approved duration.
- Failure identifies the relevant unavailable service and is reported as an infrastructure error.
- Unit tests cover success, recovery, timeout, non-success response, and exhaustion.
- Generated runtime templates contain the selected bounded behavior.
- A staging kbuild job succeeds on a newly started Kubernetes node.
- The staging validation result is linked in this issue.
- Existing test and lint checks pass.
Out of scope
- Removing Google services that are required for an actual GKE or Google Cloud operation.
- Reworking unrelated download and upload retry policies.
- Changing Kubernetes cluster networking.
- Removing ChromeOS/Tast dependencies tracked by #3196.
- Lingua principale
- Python
- Stelle
- 120
- Fork
- 108
- Merge medio
- 1g 12h
- PR unite (30g)
- 21
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di kernelci/kernelci-core
-
good first issue
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
kernelci/kernelci-core#2591 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 25/100
kernelci/kernelci-core#3234 ·
I maintainer di solito rispondono entro 1 giorno
-
chromeos techdebt
Difficoltà 5/5 Più di una settimana Idoneità per principianti 35/100
kernelci/kernelci-core#3196 ·
I maintainer di solito rispondono entro 1 giorno
-
kubernetes runners missing logs and test naming wrongForse di nuovo libera @nuclearcat l’ha presa 67 giorni fa e non c’è nessuna pull request aperta. Aperta
kernelci/kernelci-core#3170 · 2 commenti · 1 reazione · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 20/100
kernelci/kernelci-core#3131 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di kernelci/kernelci-core
Issue simili
-
needs-human needs-triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
gke-labs/kube-agents#2400 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno
-
Device Details tables: FS/SF columns contradict each other (nfet_01v8 Vt row, pfet_01v8 Idsat row)Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
google/skywater-pdk#450 ·
-
Drained trajectory arrays are overwritten when the sequence buffer is reusedForse già presa @sylvesterkaczmarek l’ha presa oggi. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
google-deepmind/bsuite#56 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
LearningCircuit/local-deep-research#7206 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
chingu-voyages/V62-tier3-team-33#285 ·
I maintainer di solito rispondono entro 1 giorno