Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Make Kubernetes job network readiness checks bounded and deployment-independent

Abierto
#3,197 3 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
35/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
kubernetes, python

Línea de trabajo

Primero obtén la decisión del maintainer y del operador de Kubernetes sobre la estrategia de readiness y los valores de timing; después inspecciona kernelci/kbuild.py y config/runtime/base/python.jinja2. Añade pruebas unitarias y de renderizado de plantillas específicas, con requests y esperas simulados, y verifica que los fallos acotados se clasifiquen como errores de infraestructura; para darlo por terminado también se requiere una ejecución de kbuild en staging sobre un nodo recién iniciado, enlazada aquí.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

good first issue techdebt

Problem

KernelCI contains two startup workarounds for Kubernetes nodes whose networking or DNS may not be immediately ready.

KBuild._verify_network() performs an HTTP request to Google before generating the build script:

requests.get("https://google.com")

The current implementation has two failure modes:

  • the request has no timeout and can hang indefinitely;
  • a non-200 HTTP response does not decrement retries, causing an infinite loop.

The generic Python job template contains a separate DNS workaround that resolves www.google.com up to 30 times before starting a job.

Both checks depend on an unrelated third-party service. Reaching Google does not prove that the KernelCI API, artifact storage, source repositories, or other services required by the job are reachable. Conversely, a Google outage or network policy blocking Google can prevent otherwise healthy KernelCI jobs from starting.

The HTTP check was originally introduced because networking could be slow to initialize on newly started Kubernetes nodes.

Relevant code

  • kernelci/kbuild.py: KBuild._verify_network() and its call from write_script()
  • config/runtime/base/python.jinja2: DNS readiness loop in main()

ChromeOS/Tast-specific Google dependencies are being removed separately in #3196 and are outside this issue.

Desired outcome

Network initialization handling must:

  • have a strict upper time limit;
  • not depend on an unrelated public service;
  • test connectivity relevant to the operation the job is about to perform;
  • produce an actionable infrastructure error when readiness is not achieved;
  • behave consistently between kbuild and other Python jobs.

Maintainer decision required

Before implementation, a KernelCI maintainer or Kubernetes operator should choose one of these strategies:

  1. Retry the first required operation — recommended

    Remove the generic preflight probes and apply bounded retry handling to the first real API, storage, source, or artifact request.

  2. Use a configurable readiness target

    Keep a preflight check, but provide the hostname or URL through runtime configuration. The default should be a KernelCI-controlled service required by the job.

The selected maximum startup delay, request timeout, and retry interval must also be recorded in this issue.

Implementation outline

After the strategy is confirmed:

  1. Remove both hard-coded Google readiness checks.
  2. Ensure every attempt has a connection and read timeout.
  3. Use a fixed attempt count or monotonic deadline so every failure path terminates.
  4. Handle expected network exceptions explicitly rather than catching every Exception.
  5. Avoid calling sys.exit() from a low-level helper; return or raise an error that the job runner can classify.
  6. Report exhausted network readiness as an infrastructure failure rather than a kernel build failure.
  7. Include the attempted service and elapsed time in the final error without exposing credentials or tokens.
  8. Add focused unit and template-rendering tests.

Required tests

Depending on the selected strategy, cover:

  • immediate success;
  • initial failures followed by success;
  • DNS failure;
  • connection timeout;
  • read timeout;
  • repeated non-success HTTP responses;
  • retry exhaustion;
  • enforcement of the maximum elapsed time;
  • correct infrastructure-error classification.

Tests must mock network requests and waiting so they do not contact external services or introduce real delays.

Human participation required

This issue intentionally needs two concise operational checkpoints:

  1. A Kubernetes operator confirms whether delayed networking still occurs and approves the readiness strategy and timing values.
  2. After implementation, an operator runs a kbuild job on a newly scaled or cold Kubernetes node and links the result here.

Repository tests can prove that retry handling is bounded, but they cannot reproduce the real network initialization behavior of the production clusters.

Acceptance criteria

  • Generic KernelCI startup code no longer uses google.com as a connectivity probe.
  • No readiness or retry path can loop or block indefinitely.
  • The total readiness wait is bounded by the maintainer-approved duration.
  • Failure identifies the relevant unavailable service and is reported as an infrastructure error.
  • Unit tests cover success, recovery, timeout, non-success response, and exhaustion.
  • Generated runtime templates contain the selected bounded behavior.
  • A staging kbuild job succeeds on a newly started Kubernetes node.
  • The staging validation result is linked in this issue.
  • Existing test and lint checks pass.

Out of scope

  • Removing Google services that are required for an actual GKE or Google Cloud operation.
  • Reworking unrelated download and upload retry policies.
  • Changing Kubernetes cluster networking.
  • Removing ChromeOS/Tast dependencies tracked by #3196.
Lenguaje dominante
Python
Estrellas
120
Forks
108
Merge medio
1 d 12 h
PR fusionados (30 d)
21

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de kernelci/kernelci-core

Todos los issues de kernelci/kernelci-core

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.