AKS Windows Server 2022: RSTs to a deleted pod's IP loop between the synthetic NIC and the accelerated networking VF, pinning cores at high DPC
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Stack tecnologico
- azure, kubernetes
- Ambito
- infrastructure, networking
Direzione di ricerca
Start by reviewing the packet counters, DPC measurements, and pktmon capture in the issue, then compare the reported behavior with related issue #631. The report does not identify a code entry point or a reproducible test, so first establish whether maintainers can reproduce or diagnose the loop. Done means identifying a supported cause, fix, or mitigation for the reported Windows Server 2022 networking behavior.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Please fill out all the sections below for bug issues, otherwise it'll be closed as it won't be actionable for us to address.
Describe the bug
On AKS Windows Server 2022 nodes, after a pod is torn down while it holds outbound TCP connections to an in-cluster LoadBalancer VIP, the host gets stuck forwarding TCP RST packets addressed to the dead pod's IP. The packets enter on the synthetic Hyper-V adapter and are sent straight back out on the accelerated networking VF, indefinitely: 18,000–183,000 packets/sec of ~60-byte packets. They never reach any container.
The kernel DPC work this generates pins specific cores at 25–96%, and every workload on the node slows 1.5x–13x. In our fleet 18 of 138 Windows nodes were in this state at once after a weekend of redeployments. We have seen this class of Windows-node degradation for several years and have worked around it by rotating to fresh nodes on every release.
The dead pod's IP is still present as an ipConfiguration on the node's Azure NIC (Azure CNI node-subnet preallocation), but HNS has no endpoint for it. Only replacing the node has cleared it so far.
To Reproduce
We cannot reproduce it on demand: it is not every teardown, and the same node has had pods torn down with no loop afterwards. The sequence we observed each time:
- An AKS cluster with Windows Server 2022 nodes, Azure CNI (node subnet, IPs preallocated on the node NIC), accelerated networking enabled.
- An ingress controller (ingress-nginx, on Linux nodes) exposed as an internal
type: LoadBalancerService,externalTrafficPolicy: Cluster,ipMode: VIP. - Windows pods (IIS / .NET) calling other in-cluster services through hostnames that CoreDNS resolves to that LB VIP, so each Windows pod holds keep-alive HTTPS client connections to
<LB_VIP>:443. - One of those pods goes away while its connections are open. We saw it after either of:
- a
rollout restart/ redeploy that moves the pod to another node; - a node reboot with the pod still scheduled, where the pod came back with a new IP and the old IP began looping.
- a
- On the original node, RST,ACK packets from
<LB_VIP>:443to<DEAD_POD_IP>:<ephemeral ports>start circulating, with the receive rate on the synthetic adapter equal to the send rate on the VF. They continue until the node is replaced.
Evidence from an affected node.
Packet counters, 3 × 2 s samples:
\network interface(microsoft hyper-v network adapter _2)\packets received/sec = 177,859
\network interface(mellanox connectx-4 lx virtual ethernet adapter)\packets sent/sec = 177,990
Hyper-V Virtual Switch ext pkts/s = ~336,000
An unaffected node in the same pool reads 27–28 packets/sec.
Per-core DPC on an 8-core node:
cpu 0 dpc% 91.4 cpu 2 dpc% 72.0 cpu 4 dpc% 85.5 cpu 6 dpc% 67.0
cpu 1,3,5,7 dpc% 0.0 _total dpc% 39.5 dpcq/s 97,282
No process accounts for this CPU.
pktmon capture, 8 s, headers only:
<LB_VIP>:443 > <DEAD_POD_IP>:51198 Flags [R.] n=25,739
<LB_VIP>:443 > <DEAD_POD_IP>:51222 Flags [R.] n=25,388
<LB_VIP>:443 > <DEAD_POD_IP>:51197 Flags [R.] n=25,388
... 7 flows, 145,294 packets in 8 s
We confirmed the same pattern on a second cluster: 61,224 packets in 5 s to one dead IP.
<DEAD_POD_IP>has no pod and no HNS endpoint (Get-HnsEndpoint).- It is still listed in the node NIC's ipConfigurations.
- Pod inventory shows its last owner was a pod on that same node, torn down at the times above.
Expected behavior
A packet arriving for a NIC-owned secondary IP that has no HNS endpoint behind it should be dropped, or answered once, by the host. Instead it is forwarded back out through the VF, apparently without decrementing TTL, delivered back to the same node by Azure, and looped indefinitely.
Configuration:
- Edition: Windows Server 2022 (AKS node image
AKSWindows-2022-containerd-20348.5139.260513, OS build 10.0.20348.5139) - Base Image being used: Windows Server Core,
mcr.microsoft.com/dotnet/framework/aspnet:4.8.1-windowsservercore-ltsc2022(container OS 10.0.20348.5622) - Container engine: containerd (AKS), kubelet v1.35.4
- Container Engine version: containerd 1.7.20+azure
Also relevant:
- Network: Azure CNI node subnet; maxPods 250, so 251 ipConfigurations per NIC
- Accelerated networking: enabled; Mellanox ConnectX-4 Lx Virtual Ethernet Adapter, driver 23.4.26054.1
- VM sizes affected: Standard_D8d_v4, Standard_E8as_v4, Standard_F16s_v2
Additional context
- Impact: IIS request latency for pods on affected nodes went from p50 ~250 ms to ~2,700 ms. Two replicas of one deployment, on an affected and an unaffected node, measured 2,769 ms and 253 ms p50 over the same 6 hours. Slowdown correlated with node DPC (r ≈ 0.78 across 12 workloads).
- Normal on affected nodes: disk latency (~0.2 ms), Defender and Windows Update activity, the System event log, HNS network and endpoint counts, kubelet, containerd and kube-proxy state. The node image is identical to healthy nodes.
- Not reproduced elsewhere: a smaller, low-traffic cluster with the same pod churn and node reuse (node image 20348.5020) shows no loop.
- Questions:
- Is this a known HNS / VFP / accelerated networking defect, and is there a fix in a newer Windows Server 2022 or 2025 node image?
- Is there a supported way to clear an affected node without replacing it, e.g. an
hnsdiag/vfpctrlaction or releasing the stale IP? - Is there a supported configuration that prevents it, such as a CNI setting, VFP policy, or kube-proxy flag (DSR)?
- Does Azure CNI Overlay on Windows avoid it, given pod IPs would no longer be NIC ipConfigurations?
- Possibly related: #631 (stale HNS endpoints after pod deletion); Azure/AKS#5483 (intermittent Windows pod networking failures, closed without root cause).
- Our current workaround: replace affected nodes, and alert on node
% DPC Time> 2% (any core > 15%) or synthetic-NIC receive > 5,000 pps with VF send within 10% of it.
Fullpktmon --comp allcaptures,collectlogs.ps1output and the NIC ipConfiguration listing are available on request.
- Lingua principale
- PowerShell
- Stelle
- 552
- Fork
- 75
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/Windows-Containers
-
enhancement triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
microsoft/Windows-Containers#630 · 6 commenti ·
-
enhancement triage
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
microsoft/Windows-Containers#653 · 1 commento ·
-
enhancement triage
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
microsoft/Windows-Containers#652 · 1 commento ·
-
CimFS-backed overlay mount performance parity with Linux overlayfs for container image layersApertaenhancement triage
Difficoltà 5/5 Più di una settimana Idoneità per principianti 30/100
microsoft/Windows-Containers#651 · 1 commento ·
-
enhancement triage
Difficoltà 5/5 Più di una settimana Idoneità per principianti 25/100
microsoft/Windows-Containers#650 · 2 commenti ·
Tutte le issue di microsoft/Windows-Containers
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
voxpupuli/puppet-quadlets#122 · 5 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 85/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
ci needs-ac
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
Ikalus1988/MisakaNet#2930 ·
I maintainer di solito rispondono entro 1 giorno