NetworkGarbageCollector releases a network's VLAN while it still has live NICs

Đang mở
#14,177 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
35/100
Loại issue
Lỗi
Độ rõ ràng
Khá rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
java

Hướng nghiên cứu

Trace NetworkGarbageCollector and the PowerReportMissing recovery path, then inspect nics and op_networks state during the reproduction steps. Verify live-NIC detection, power-state reconciliation, and VirtualRouter accounting against the described scenario. Done means garbage collection cannot release a VLAN while an unremoved NIC remains live.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

bug
problem

When a VM's domain is detected missing (PowerReportMissing) past the graceful period (vm.op.wait.interval), CloudStack releases its NIC (broadcast_uri/isolation_uri -> NULL, nics_count) and marks the VM Stopped. If the VM's power state later resyncs back to Running (domain restored, HA restart, etc.), the NIC is never re-reserved and nics_count is never restored.
Separately, op_networks.nics_count never counts the network's own VirtualRouter NIC, so it under-counts from network creation. Combined, a single NIC release event can drive nics_count to 0 while the network still has live NICs (including on a Running VM).
NetworkGarbageCollector trusts nics_count==0 (plus a check of CloudStack's own DB-tracked "no non-Stopped instances") without verifying against the actual nics table or hypervisor state, and proceeds to stop the VR and release the VLAN back to the dynamic allocation pool - while it may still be bridged to a live VM. In our environment, GC's cleanup step also removes the host-level bridge, making the VM unrecoverable via a normal restart.

versions

tested on 4.22 (Maybe be observed on older versions too)

The steps to reproduce the bug
  1. Deploy a VM on an isolated network - note the nics_count = 1 (instead of 2 for VM and VR)
  2. Disable HA on the VM (as HA masks the issue)
  3. Perform virsh destroy <vm-name> to simulate domain loss - backup the dumpxml virsh dumpxml <vm-name> > backup.xml
  4. After the graceful period, CloudStack releases the NIC and decrements nics_count to 0.
  5. Restore the domain via virsh create <dumped-xml> . VM resyncs to Running, but NIC stays unreserved, counter stays 0
  6. Lower network.gc.interval/network.gc.wait to accelerate the scavenger; restart management server.
  7. virsh destroy again to bring the VM's tracked state back to Stopped
  8. GC fires: stops the VR, releases the VLAN, network transitions to Allocated broadcast_uri=NULL, despite 2 live NICs still present in nics.
    ...
What to do about it?
  • On VM power-state resync to Running after a missing-VM event, reconcile the NIC (broadcast_uri/isolation_uri/state=Reserved) and restore nics_count.
  • Harden NetworkGarbageCollector to verify live NIC count from the nics table (WHERE network_id=? AND removed IS NULL) before tearing down a network, instead of trusting the cached counter alone
  • Include the VirtualRouter's own NIC in nics_count from network creation
Ngôn ngữ chính
Java
Star
3.1k
Fork
1.4k
Merge trung bình
7 ngày 5 giờ
Pull request đã merge (30 ngày)
28

Hướng dẫn đóng góp

Mở hướng dẫn đóng góp

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của apache/cloudstack

Tất cả issue của apache/cloudstack

Issue tương tự

Thêm issue về Java

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.