[Feature]: Ignore already cordoned nodes during driver upgrades
Maintainer thường phản hồi trong vòng 1 ngày
@JunAr7112 đang làm issue này rồi.
Từ ngày 10/9/2026.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
Requestor: @kvalliyurnatt
Summary
Ignore nodes that are already cordoned (not by any of the GPU operator components) during upgrades from upgrade controller/k8s-driver-manager
Motivation
If there is a node that has already been cordoned (not by any of the GPU operator components), then that means the node is not accepting any workloads at the moment, and there is no guarantee of the state that node is in, it could be down or in an unresponsive state. So it feels like the upgrade controller/k8s-driver-manager should ignore nodes that are already cordoned and not take any action on such nodes until they are uncordoned.
Proposal
We will need a way differentiate nodes that are cordoned by the GPU operator components, for which we could add an annotation on the nodes when the upgrade controller/k8s-driver-manager cordons a node(I believe the upgrade controller already does this) and then uncordon nodes based on that annotation. Then the upgrade controller/k8s-driver-manager can ignore nodes that are in a cordoned state but don't have the annotation on them.
- Ngôn ngữ chính
- Go
- Star
- 2.9k
- Fork
- 552
- Merge trung bình
- 1 ngày 21 giờ
- Pull request đã merge (30 ngày)
- 78
Chuẩn bị môi trường
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/gpu-operator
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
NVIDIA/gpu-operator#2968 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 88/100
NVIDIA/gpu-operator#2955 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
lifecycle/stale question
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 64/100
NVIDIA/gpu-operator#2280 · 2 bình luận · 1 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[Bug]: GPU Operator MPS config-manager cannot signal MPS daemon due to process-target mismatchĐang mởbug needs-triage
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 68/100
NVIDIA/gpu-operator#2970 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 68/100
NVIDIA/gpu-operator#2957 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của NVIDIA/gpu-operator
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
siderolabs/terraform-provider-talos#414 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
JuliaComputing/jh#63 · 1 bình luận ·
-
area/proxy kind/bug priority/backlog triage/accepted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
lexfrei/cloudflare-tunnel-gateway-controller#840 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Priority: Normal Type: Bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
cloudflare/cloudflared#1747 ·