AKS Windows Server 2022: RSTs to a deleted pod's IP loop between the synthetic NIC and the accelerated networking VF, pinning cores at high DPC
まだ誰も着手していません。
評価
- 難易度
- 5/5
- 見積もり時間
- 1週間以上
- 初心者へのやさしさ
- 30/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
- 技術スタック
- azure, kubernetes
調査の方向性
Start by reviewing the packet counters, DPC measurements, and pktmon capture in the issue, then compare the reported behavior with related issue #631. The report does not identify a code entry point or a reproducible test, so first establish whether maintainers can reproduce or diagnose the loop. Done means identifying a supported cause, fix, or mitigation for the reported Windows Server 2022 networking behavior.
索引モデルが issue の本文から書いたものです。
説明
Please fill out all the sections below for bug issues, otherwise it'll be closed as it won't be actionable for us to address.
Describe the bug
On AKS Windows Server 2022 nodes, after a pod is torn down while it holds outbound TCP connections to an in-cluster LoadBalancer VIP, the host gets stuck forwarding TCP RST packets addressed to the dead pod's IP. The packets enter on the synthetic Hyper-V adapter and are sent straight back out on the accelerated networking VF, indefinitely: 18,000–183,000 packets/sec of ~60-byte packets. They never reach any container.
The kernel DPC work this generates pins specific cores at 25–96%, and every workload on the node slows 1.5x–13x. In our fleet 18 of 138 Windows nodes were in this state at once after a weekend of redeployments. We have seen this class of Windows-node degradation for several years and have worked around it by rotating to fresh nodes on every release.
The dead pod's IP is still present as an ipConfiguration on the node's Azure NIC (Azure CNI node-subnet preallocation), but HNS has no endpoint for it. Only replacing the node has cleared it so far.
To Reproduce
We cannot reproduce it on demand: it is not every teardown, and the same node has had pods torn down with no loop afterwards. The sequence we observed each time:
- An AKS cluster with Windows Server 2022 nodes, Azure CNI (node subnet, IPs preallocated on the node NIC), accelerated networking enabled.
- An ingress controller (ingress-nginx, on Linux nodes) exposed as an internal
type: LoadBalancerService,externalTrafficPolicy: Cluster,ipMode: VIP. - Windows pods (IIS / .NET) calling other in-cluster services through hostnames that CoreDNS resolves to that LB VIP, so each Windows pod holds keep-alive HTTPS client connections to
<LB_VIP>:443. - One of those pods goes away while its connections are open. We saw it after either of:
- a
rollout restart/ redeploy that moves the pod to another node; - a node reboot with the pod still scheduled, where the pod came back with a new IP and the old IP began looping.
- a
- On the original node, RST,ACK packets from
<LB_VIP>:443to<DEAD_POD_IP>:<ephemeral ports>start circulating, with the receive rate on the synthetic adapter equal to the send rate on the VF. They continue until the node is replaced.
Evidence from an affected node.
Packet counters, 3 × 2 s samples:
\network interface(microsoft hyper-v network adapter _2)\packets received/sec = 177,859
\network interface(mellanox connectx-4 lx virtual ethernet adapter)\packets sent/sec = 177,990
Hyper-V Virtual Switch ext pkts/s = ~336,000
An unaffected node in the same pool reads 27–28 packets/sec.
Per-core DPC on an 8-core node:
cpu 0 dpc% 91.4 cpu 2 dpc% 72.0 cpu 4 dpc% 85.5 cpu 6 dpc% 67.0
cpu 1,3,5,7 dpc% 0.0 _total dpc% 39.5 dpcq/s 97,282
No process accounts for this CPU.
pktmon capture, 8 s, headers only:
<LB_VIP>:443 > <DEAD_POD_IP>:51198 Flags [R.] n=25,739
<LB_VIP>:443 > <DEAD_POD_IP>:51222 Flags [R.] n=25,388
<LB_VIP>:443 > <DEAD_POD_IP>:51197 Flags [R.] n=25,388
... 7 flows, 145,294 packets in 8 s
We confirmed the same pattern on a second cluster: 61,224 packets in 5 s to one dead IP.
<DEAD_POD_IP>has no pod and no HNS endpoint (Get-HnsEndpoint).- It is still listed in the node NIC's ipConfigurations.
- Pod inventory shows its last owner was a pod on that same node, torn down at the times above.
Expected behavior
A packet arriving for a NIC-owned secondary IP that has no HNS endpoint behind it should be dropped, or answered once, by the host. Instead it is forwarded back out through the VF, apparently without decrementing TTL, delivered back to the same node by Azure, and looped indefinitely.
Configuration:
- Edition: Windows Server 2022 (AKS node image
AKSWindows-2022-containerd-20348.5139.260513, OS build 10.0.20348.5139) - Base Image being used: Windows Server Core,
mcr.microsoft.com/dotnet/framework/aspnet:4.8.1-windowsservercore-ltsc2022(container OS 10.0.20348.5622) - Container engine: containerd (AKS), kubelet v1.35.4
- Container Engine version: containerd 1.7.20+azure
Also relevant:
- Network: Azure CNI node subnet; maxPods 250, so 251 ipConfigurations per NIC
- Accelerated networking: enabled; Mellanox ConnectX-4 Lx Virtual Ethernet Adapter, driver 23.4.26054.1
- VM sizes affected: Standard_D8d_v4, Standard_E8as_v4, Standard_F16s_v2
Additional context
- Impact: IIS request latency for pods on affected nodes went from p50 ~250 ms to ~2,700 ms. Two replicas of one deployment, on an affected and an unaffected node, measured 2,769 ms and 253 ms p50 over the same 6 hours. Slowdown correlated with node DPC (r ≈ 0.78 across 12 workloads).
- Normal on affected nodes: disk latency (~0.2 ms), Defender and Windows Update activity, the System event log, HNS network and endpoint counts, kubelet, containerd and kube-proxy state. The node image is identical to healthy nodes.
- Not reproduced elsewhere: a smaller, low-traffic cluster with the same pod churn and node reuse (node image 20348.5020) shows no loop.
- Questions:
- Is this a known HNS / VFP / accelerated networking defect, and is there a fix in a newer Windows Server 2022 or 2025 node image?
- Is there a supported way to clear an affected node without replacing it, e.g. an
hnsdiag/vfpctrlaction or releasing the stale IP? - Is there a supported configuration that prevents it, such as a CNI setting, VFP policy, or kube-proxy flag (DSR)?
- Does Azure CNI Overlay on Windows avoid it, given pod IPs would no longer be NIC ipConfigurations?
- Possibly related: #631 (stale HNS endpoints after pod deletion); Azure/AKS#5483 (intermittent Windows pod networking failures, closed without root cause).
- Our current workaround: replace affected nodes, and alert on node
% DPC Time> 2% (any core > 15%) or synthetic-NIC receive > 5,000 pps with VF send within 10% of it.
Fullpktmon --comp allcaptures,collectlogs.ps1output and the NIC ipConfiguration listing are available on request.
- 主要言語
- PowerShell
- スター
- 556
- フォーク
- 75
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
microsoft/Windows-Containers のほかの issue
-
enhancement triage
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
microsoft/Windows-Containers#630 · コメント 6 件 ·
-
Whatsappオープンbug triage
難易度 5/5 1週間以上 初心者へのやさしさ 5/100
microsoft/Windows-Containers#656 · コメント 1 件 ·
-
enhancement triage
難易度 5/5 1週間以上 初心者へのやさしさ 25/100
microsoft/Windows-Containers#653 · コメント 1 件 ·
-
enhancement triage
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
microsoft/Windows-Containers#652 · コメント 1 件 ·
-
enhancement triage
難易度 5/5 1週間以上 初心者へのやさしさ 30/100
microsoft/Windows-Containers#651 · コメント 1 件 ·
microsoft/Windows-Containers の issue をすべて見る
似ている issue
-
[BUG] Container scenario crashes without expected_recovery_time, kube DNS example uses retry_waitオープンneeds-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 77/100
krkn-chaos/krkn#1627 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
int128/typescript-action#1575 ·
メンテナーはふだん 1 日以内に返信
-
agent/sec-check hive/hosted-available-lke648397-260827-5n31 security
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
femiwiki/docker-mediawiki#1497 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
EPFL-ENAC/co2-calculator#3055 ·
メンテナーはふだん 1 日以内に返信