Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

NetworkGarbageCollector releases a network's VLAN while it still has live NICs

Open
#14,177 1 comment 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
35/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
java

Research direction

Trace NetworkGarbageCollector and the PowerReportMissing recovery path, then inspect nics and op_networks state during the reproduction steps. Verify live-NIC detection, power-state reconciliation, and VirtualRouter accounting against the described scenario. Done means garbage collection cannot release a VLAN while an unremoved NIC remains live.

Written by the indexing model from the issue text.

Description

bug
problem

When a VM's domain is detected missing (PowerReportMissing) past the graceful period (vm.op.wait.interval), CloudStack releases its NIC (broadcast_uri/isolation_uri -> NULL, nics_count) and marks the VM Stopped. If the VM's power state later resyncs back to Running (domain restored, HA restart, etc.), the NIC is never re-reserved and nics_count is never restored.
Separately, op_networks.nics_count never counts the network's own VirtualRouter NIC, so it under-counts from network creation. Combined, a single NIC release event can drive nics_count to 0 while the network still has live NICs (including on a Running VM).
NetworkGarbageCollector trusts nics_count==0 (plus a check of CloudStack's own DB-tracked "no non-Stopped instances") without verifying against the actual nics table or hypervisor state, and proceeds to stop the VR and release the VLAN back to the dynamic allocation pool - while it may still be bridged to a live VM. In our environment, GC's cleanup step also removes the host-level bridge, making the VM unrecoverable via a normal restart.

versions

tested on 4.22 (Maybe be observed on older versions too)

The steps to reproduce the bug
  1. Deploy a VM on an isolated network - note the nics_count = 1 (instead of 2 for VM and VR)
  2. Disable HA on the VM (as HA masks the issue)
  3. Perform virsh destroy <vm-name> to simulate domain loss - backup the dumpxml virsh dumpxml <vm-name> > backup.xml
  4. After the graceful period, CloudStack releases the NIC and decrements nics_count to 0.
  5. Restore the domain via virsh create <dumped-xml> . VM resyncs to Running, but NIC stays unreserved, counter stays 0
  6. Lower network.gc.interval/network.gc.wait to accelerate the scavenger; restart management server.
  7. virsh destroy again to bring the VM's tracked state back to Stopped
  8. GC fires: stops the VR, releases the VLAN, network transitions to Allocated broadcast_uri=NULL, despite 2 live NICs still present in nics.
    ...
What to do about it?
  • On VM power-state resync to Running after a missing-VM event, reconcile the NIC (broadcast_uri/isolation_uri/state=Reserved) and restore nics_count.
  • Harden NetworkGarbageCollector to verify live NIC count from the nics table (WHERE network_id=? AND removed IS NULL) before tearing down a network, instead of trusting the cached counter alone
  • Include the VirtualRouter's own NIC in nics_count from network creation
Dominant language
Java
Stars
3.1k
Forks
1.4k
Avg merge
6d 20h
Merged PRs (30d)
27

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/cloudstack

All issues in apache/cloudstack

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.