isRunning()/Signal()/Kill()/joinSandboxNetNs() trust a raw PID with no liveness/identity check, so PID reuse after the VMM exits lets urunc signal (and even network-namespace-join) an unrelated host process
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 4/5
- Tiempo estimado
- 3-5 días
- Aptitud para principiantes
- 52/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- go, linux
- Área
- networking, operating-systems, security
Línea de trabajo
Empieza siguiendo Create y los métodos del ciclo de vida en pkg/unikontainers/unikontainers.go, especialmente isRunning, Signal, Kill y joinSandboxNetNs. Después, revisa checkValidNsPath en pkg/unikontainers/utils.go y killProcess en pkg/unikontainers/hypervisors/utils.go. Revisa cómo se almacena u.State.Pid e identifica las pruebas existentes o los puntos de entrada de prueba para estas rutas. La tarea estará terminada cuando la reutilización de PID se rechace de forma consistente antes de enviar señales, detener, eliminar o unirse a un espacio de nombres de red.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Description
urunc records the monitor (VMM) process's PID once, in u.State.Pid, when the container is created, and then uses that raw integer for the rest of the container's lifecycle: Signal(), Kill(), Delete() via isRunning(), and joinSandboxNetNs(). None of these call sites verify that the PID still refers to the same process that was originally launched, they only check whether some process currently holds that PID number.
On Linux, PIDs are recycled. Once the VMM exits and gets reaped, that PID becomes eligible for reuse by any new process on the host. Issue #716 shows urunc containers can sit around for a long time after their VMM has already died, which is exactly the kind of window that makes PID recycling realistic on a busy node.
Relevant code:
pkg/unikontainers/unikontainers.go:1407-1415 (isRunning):
func (u *Unikontainer) isRunning() bool {
vmmType := hypervisors.VmmType(u.State.Annotations[annotHypervisor])
if vmmType != hypervisors.HedgeVmm {
return syscall.Kill(u.State.Pid, syscall.Signal(0)) == nil
}
...
}
pkg/unikontainers/unikontainers.go:794-838 (Signal, Kill):
func (u *Unikontainer) Signal(signal unix.Signal) error {
...
return vmm.Signal(u.State.Pid, signal)
}
func (u *Unikontainer) Kill() error {
err := u.joinSandboxNetNs()
...
err = vmm.Stop(u.State.Pid)
...
err = network.CleanupAllUruncTaps()
...
}
pkg/unikontainers/unikontainers.go:921-948 (joinSandboxNetNs):
if netNsPath == "" {
netNsPath = fmt.Sprintf("/proc/%d/ns/net", u.State.Pid)
err := checkValidNsPath(netNsPath)
...
}
...
fd, err := unix.Open(netNsPath, unix.O_RDONLY|unix.O_CLOEXEC, 0)
...
err = unix.Setns(int(fd), unix.CLONE_NEWNET)
checkValidNsPath (pkg/unikontainers/utils.go:209-220) only does an os.Lstat(path) existence check, never an identity check.
pkg/unikontainers/hypervisors/utils.go:69-92 (killProcess, used by every hypervisor's Stop):
func killProcess(pid int) error {
const timeout = 2 * time.Second
err := unix.Kill(pid, unix.SIGKILL)
...
}
Sends SIGKILL to whatever process currently owns that PID, with no identity check.
None of these five call sites cross check the PID against anything that proves it is still the original VMM, for example /proc/<pid>/stat starttime or cgroup membership, which is a technique other runtimes such as runc rely on for this exact reason.
If the VMM's PID gets reused by an unrelated host process before urunc kill or urunc delete runs:
isRunning()reports true for the unrelated process, soDelete()permanently refuses to clean up the already dead container.Kill()callsjoinSandboxNetNs(), which opens the unrelated process's network namespace and joins it, then callsvmm.Stop(), which SIGKILLs the unrelated process, then runsnetwork.CleanupAllUruncTaps()inside the wrong namespace, removing tap devices that do not belong to this container.
This causes two real failure modes that need no attacker or malicious image author, just ordinary OS PID recycling:
- Host level collateral damage: an unrelated process gets killed, and unrelated network resources are torn down in the wrong netns.
- Stuck containers:
Delete()can refuse indefinitely to remove an already dead container whose PID was reused, leaving orphaned state and blocking CRI or Kubernetes garbage collection, which compounds known issue #716.
Suggested fix: record a lightweight identity token for the VMM PID at Create() time, for example the /proc/<pid>/stat starttime field (which the kernel guarantees changes across PID reuse), and validate it in isRunning(), Signal(), Kill(), and joinSandboxNetNs() before treating the PID as belonging to this container. A mismatch should be treated the same as the process no longer existing.
System info
- Urunc version: main branch
- Arch: any
- VMM: any (Qemu, Firecracker, Cloud Hypervisor, HVT, SPT), any backend going through
killProcess - Unikernel: any
Steps to reproduce
- Start a urunc container with any hypervisor.
u.State.Pidis set to the VMM's host PID, for example 12345. - Let the VMM exit without going through
urunc killorurunc delete(guest panic, VMM OOM kill, or normal completion), so PID 12345 gets reaped and freed. - On a host with enough process churn, an unrelated process ends up receiving PID 12345. This is normal PID reuse behavior on Linux.
- Call
urunc kill <id> 9orurunc delete <id>and observe:isRunning()returns true for the unrelated process, so delete is refused.Kill()joins the unrelated process's network namespace, sends it SIGKILL, and cleans up tap devices in the wrong namespace.
Note on LLM usage
This issue was drafted with the assistance of an LLM (Claude Sonnet 5, Anthropic), which was used to trace the code paths and structure this report. The underlying code, call sites, and failure scenario described above were read and verified directly against the repository by me before filing; the analysis and conclusion are my own responsibility.
- Lenguaje dominante
- Go
- Estrellas
- 298
- Forks
- 205
- Merge medio
- 2 d 18 h
- PR fusionados (30 d)
- 27
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de urunc-dev/urunc
-
The containerd version detection in urunc-deploy does not parse correctly the version stringAbiertobug K8s/Tools
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día
-
Core enhancement OCI
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
urunc-dev/urunc#1066 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
dev
Dificultad 2/5 1-3 horas Aptitud para principiantes 75/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
urunc-dev/urunc#1021 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
Los mantenedores suelen responder en 1 día
Todos los issues de urunc-dev/urunc
Issues similares
-
agent-butler-finding chore
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
jordansmall/spindrift#4146 ·
Los mantenedores suelen responder en 1 día
-
security
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
IBM/ibmcloud-volume-file-vpc#119 ·
-
security
Dificultad 2/5 1-3 horas Aptitud para principiantes 66/100
IBM/networking-go-sdk#339 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
kubernetes-sigs/mcp-lifecycle-operator#439 ·
Los mantenedores suelen responder en 1 día
-
area: global bug dx priority: low
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 1 día