Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

KVM: destroy with vm.destroy.forcestop leaves the instance running when its host is briefly disconnected

未關閉
#14,232 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
45/100
Issue 類型
缺陷
描述清晰度
描述清楚
活躍度
活躍
技術堆疊
java
領域
backend, cloud

研究方向

The issue is in the destroy path when vm.destroy.forcestop=true. Start by examining the UserVmManagerImpl.destroyVm, VirtualMachineManagerImpl.destroy(), and VirtualMachineManagerImpl.advanceExpunge() methods, focusing on advanceStop() and releaseVmResources(). Look at how host states (Disconnected, Connecting, Alert, Rebalancing, Up) are handled versus terminal states (Down, Removed). The fix likely involves checking the host state before deciding to force-stop and release resources. Run tests related to VM destruction and host state transitions to verify the behavior.

由索引模型根據 Issue 內容生成。

描述

ISSUE TYPE
  • Bug Report
COMPONENT NAME
engine-orchestration, destroyVirtualMachine, KVM
CLOUDSTACK VERSION
4.22
CONFIGURATION

vm.destroy.forcestop=true. KVM hosts. Any rolling restart of the agents or the management servers while instances are being destroyed.

OS / ENVIRONMENT

KVM / libvirt.

SUMMARY

With vm.destroy.forcestop=true, destroying an instance whose host is briefly disconnected releases the instance's NICs, IP addresses and volumes without stopping it. The domain keeps running on the host with no record in CloudStack. Its IP is handed to the next instance, so two live instances answer for the same address, and its root volume stays in Destroy because the delete fails while the domain holds the image.

Same end state as #14206, different trigger. #14206 is a stale power report during start; this one is the destroy path itself.

STEPS TO REPRODUCE
  1. Set vm.destroy.forcestop=true.
  2. Restart cloudstack-agent on a KVM host (or a management server, which disconnects the agents connected to it while they rebalance).
  3. While the host is Disconnected, destroy a running instance on it with expunge=true.

Any client that destroys instances routinely hits this on every rolling upgrade.

EXPECTED RESULTS

The destroy fails, the instance stays Running, and the caller retries once the host is back. Or the destroy waits for the host.

ACTUAL RESULTS

From a real occurrence on 4.22, during a rolling package upgrade:

13:39:02 WARN  Unable to stop VM instance {"id":3645123,...,"state":"Stopping"} due to [AgentUnavailableException:
               Resource [Host:110] is unreachable: Host 110: Host with specified id is not in the right state: Disconnected]
13:39:02 WARN  Unable to actually stop VM instance {"id":3645123,...} but continue with release because it's a force stop
13:39:02 DEBUG VM instance {"id":3645123,...} is stopped on the host.  Proceeding to release resource held.
13:39:02 DEBUG Successfully released network resources for the VM ...
13:39:15 DEBUG Expunged VM instance {"id":3645123,...}
13:41:58 WARN  Host reports 9 instance(s) that do not exist in CloudStack DB, they are running unmanaged. host: hv104, instances: [i-625-3645123-VM, ...]

Across two rolling upgrades, 31 instances on 6 hosts were left running this way, and new instances were given their addresses within hours.

CAUSE

UserVmManagerImpl.destroyVm(DestroyVMCmd), VirtualMachineManagerImpl.destroy() and VirtualMachineManagerImpl.advanceExpunge() all stop the instance with cleanUpEvenIfUnableToStop = vm.destroy.forcestop. In advanceStop(), a forced stop that gets AgentUnavailableException or OperationTimedoutException releases the resources and marks the instance stopped:

} catch (AgentUnavailableException | OperationTimedoutException e) {
    logger.warn("Unable to stop {} due to [{}].", ...);
} finally {
    if (!stopped) {
        if (!cleanUpEvenIfUnableToStop) { ... throw ... }
        else { logger.warn("Unable to actually stop {} but continue with release because it's a force stop", vm); ... }
    }
}
releaseVmResources(profile, cleanUpEvenIfUnableToStop);

A forced stop means "the host cannot tell us, treat the instance as stopped". That is right when the host is gone (Down, Removed). It is wrong when the host is only unreachable for a while, whatever its status says: Disconnected, Connecting, Alert or Rebalancing, and also Up while a crashed management server's hosts have not yet been taken over or a disconnect investigation is inconclusive. Those hosts are expected back with their domains still running.

vm.destroy.forcestop is a global setting applied to every destroy, not a statement by the caller that it knows the host is gone. It should not bypass that distinction.

IMPACT
  • an instance runs unmanaged and invisible
  • its IP is reassigned, giving an address conflict between two live instances
  • its root volume stays in Destroy and the storage cleanup fails on it every run
主要語言
Java
星號
3.1k
分支
1.4k
平均合併
6 天 20 小時
30 天內合併 PR
27

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

apache/cloudstack 的其他 Issue

查看 apache/cloudstack 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。