Hacktoberfest 2026:維護者為十月標記出來的 issue,仍然開放、適合新手。 瀏覽 Hacktoberfest issue

Multi-arch zone: system VM arch-fallback leaves an orphaned ROOT volume, making the system VM unstartable on every host

未關閉
#14,235 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
4/5
預估耗時
3-5 天
新手友好度
45/100
Issue 類型
缺陷
描述清晰度
描述清楚
活躍度
活躍
技術堆疊
java, sql
領域
backend, cloud

研究方向

Look at the system VM allocation logic, likely in ClusteredVirtualMachineManagerImpl or VolumeOrchestrator, where the ROOT volume is created before the architecture fallback. The bug is that the first volume is not cleaned up when the template changes. Search for 'Retrying with another template' logs and the volume creation path. The fix must ensure the orphaned ROOT volume is removed when allocation retries with a different architecture template. Test by simulating a multi-arch zone with an unschedulable host of the preferred architecture.

由索引模型根據 Issue 內容生成。

描述

bug
problem

In a multi-arch zone (aarch64 + x86_64 clusters), if a system VM (SSVM/CPVM) cannot be
deployed on the architecture chosen by system.vm.preferred.architecture, CloudStack retries
allocation using the other architecture's SystemVM template — logged as "Retrying with another
template"
. A ROOT volume has already been created for the first template at that point, and
after the retry the VM has two ROOT volumes, both at device_id = 0, one per architecture.
Nothing in the logs shows the first volume being released.

Every subsequent start attempt produces libvirt domain XML with two disks on target vda,
which libvirt rejects:

XML error: target 'vda' duplicated for disk sources '<VOL_X86_UUID>' and '<VOL_AARCH64_UUID>'

The system VM therefore cannot start on any host of either architecture, and enters a
permanent create → fail → destroy → recreate loop.

Impact

Losing all hosts of the preferred architecture in a mixed-arch zone leaves the zone unable to
run system VMs at all — no SSVM (no template/ISO/snapshot operations) and no CPVM (no console
access) — even though healthy hosts of the other architecture and a ready SystemVM template
for that architecture are available. The zone does not self-heal.

Determinism

Reproduced on every recreation attempt (VM ids 235, 236, 241, 242, 243). Control cases confirm
the inverse: VMs allocated without hitting the fallback path (239, 240, 244) each received
exactly one ROOT volume and started normally.

Setup

Multi-arch zone, both architectures in the same zone:

ZONE1 (id 1)
├── POD_AARCH64 (id 1)
│   └── CLUSTER_AARCH64 (id 1)  arch=aarch64  → AARCH64_HOST01 (host id 1)
└── POD_X86 (id 3)
    ├── CLUSTER_X86_A (id 5)    arch=x86_64   → X86_HOST01 (host id 9)
    └── CLUSTER_X86_B (id 6)    arch=x86_64   → X86_HOST02 (host id 10), X86_HOST03 (host id 11)

Host state during reproduction — the only aarch64 host is not schedulable, all x86_64 hosts
are healthy:

AARCH64_HOST01  aarch64  Up / ErrorInMaintenance
X86_HOST01      x86_64   Up / Enabled
X86_HOST02      x86_64   Up / Enabled
X86_HOST03      x86_64   Up / Enabled

Both SystemVM templates are registered and Ready, so a valid x86_64 template is available:

SELECT id, unique_name, name, type, hypervisor_type, arch, state FROM vm_template WHERE id IN (3,202);
+-----+-------------+-------------------------+--------+-----------------+---------+--------+
| id  | unique_name | name                    | type   | hypervisor_type | arch    | state  |
+-----+-------------+-------------------------+--------+-----------------+---------+--------+
|   3 | routing-3   | SystemVM Template (KVM) | SYSTEM | KVM             | x86_64  | Active |
| 202 | routing-202 | SystemVM Template (KVM) | SYSTEM | KVM             | aarch64 | Active |
+-----+-------------+-------------------------+--------+-----------------+---------+--------+

Both isready=true, both QCOW2, both Debian GNU/Linux 12 (64-bit), both on the shared pool.

Relevant configuration at reproduction time (system.vm.preferred.architecture is
zone-scoped; the value below is the effective zone value):

system.vm.preferred.architecture  = aarch64     <- the trigger (zone-scoped)
migrate.vm.across.clusters        = false
recreate.systemvm.enabled         = false
systemvm.root.disk.size           = -1
system.vm.default.hypervisor      = NULL
Evidence — two ROOT volumes on one VM, both device_id = 0
SELECT v.id vol, v.name, v.instance_id, v.volume_type, v.device_id, v.template_id,
       t.unique_name, t.arch, v.pool_id, v.state, v.created
FROM volumes v LEFT JOIN vm_template t ON t.id = v.template_id
WHERE v.instance_id IN (240,241) AND v.removed IS NULL ORDER BY v.instance_id, v.id;
+-----+----------+-------------+-------------+-----------+-------------+-------------+---------+---------+-------+---------------------+
| vol | name     | instance_id | volume_type | device_id | template_id | unique_name | arch    | pool_id | state | created             |
+-----+----------+-------------+-------------+-----------+-------------+-------------+---------+---------+-------+---------------------+
| 262 | ROOT-240 |         240 | ROOT        |         0 |         202 | routing-202 | aarch64 |       1 | Ready | 2026-09-23 18:16:20 |
| 263 | ROOT-241 |         241 | ROOT        |         0 |         202 | routing-202 | aarch64 |       1 | Ready | 2026-09-23 18:18:44 |
| 264 | ROOT-241 |         241 | ROOT        |         0 |           3 | routing-3   | x86_64  |       1 | Ready | 2026-09-23 18:18:45 |
+-----+----------+-------------+-------------+-----------+-------------+-------------+---------+---------+-------+---------------------+
  • VM 240 (CPVM, allocated while the aarch64 host was still available) — 1 ROOT volume.
  • VM 241 (SSVM, allocated after the aarch64 host became unschedulable) — 2 ROOT volumes,
    created one second apart, both at device_id = 0.
SELECT vm.id, vm.instance_name, vm.type, vm.state, vm.vm_template_id, vm.host_id, vm.created
FROM vm_instance vm WHERE vm.id IN (240,241);
+-----+---------------+--------------------+----------+----------------+---------+---------------------+
| id  | instance_name | type               | state    | vm_template_id | host_id | created             |
+-----+---------------+--------------------+----------+----------------+---------+---------------------+
| 240 | v-240-VM      | ConsoleProxy       | Running  |            202 |       1 | 2026-09-23 18:16:19 |
| 241 | s-241-VM      | SecondaryStorageVm | Starting |              3 |      11 | 2026-09-23 18:18:44 |
+-----+---------------+--------------------+----------+----------------+---------+---------------------+

vm_instance.vm_template_id = 3 — the VM record was correctly re-templated to x86_64; only the
aarch64 ROOT volume (263) from the abandoned first attempt was left behind.

The two volumes are backed by different-architecture templates:

$ qemu-img info <POOL_MOUNT>/<VOL_AARCH64_UUID>
  virtual size: 4.88 GiB   backing file: <POOL_MOUNT>/<TMPL_AARCH64_UUID>   # aarch64 template (id 202)
$ qemu-img info <POOL_MOUNT>/<VOL_X86_UUID>
  virtual size: 4.88 GiB   backing file: <POOL_MOUNT>/<TMPL_X86_UUID>       # x86_64 template (id 3)
Evidence — the start failure

Management server, identical on every x86_64 host:

Unable to start VM on Host {"id":10,"name":"X86_HOST02","type":"Routing","uuid":"<HOST02_UUID>"} due to XML error: target 'vda' duplicated for disk sources '<POOL_MOUNT>/<VOL_X86_UUID>' and '<POOL_MOUNT>/<VOL_AARCH64_UUID>'

Unable to start VM on Host {"id":11,"name":"X86_HOST03","type":"Routing","uuid":"<HOST03_UUID>"} due to XML error: target 'vda' duplicated for disk sources '<POOL_MOUNT>/<VOL_X86_UUID>' and '<POOL_MOUNT>/<VOL_AARCH64_UUID>'

Agent side (/var/log/cloudstack/agent/agent.log):

2026-09-23 18:25:52,529 WARN  [resource.wrapper.LibvirtStartCommandWrapper] (AgentRequest-Handler-5:[]) (logid:62824aef) LibvirtException org.libvirt.LibvirtException: XML error: target 'vda' duplicated for disk sources '<POOL_MOUNT>/<VOL_X86_UUID>' and '<POOL_MOUNT>/<VOL_AARCH64_UUID>'
2026-09-23 18:26:22,295 WARN  [resource.wrapper.LibvirtStartCommandWrapper] (AgentRequest-Handler-3:[]) (logid:62824aef) LibvirtException ... (same)
2026-09-23 18:26:38,965 WARN  [resource.wrapper.LibvirtStartCommandWrapper] (AgentRequest-Handler-4:[]) (logid:62824aef) LibvirtException ... (same)
Full management-server trace of the failing allocation

Single allocation pass for s-241-VMlogid:343a509f, thread secstorage-1. Note the two
Allocating Instance from Template lines and the two Adding disk object [ROOT-241] lines
within ~450 ms.

2026-09-23 18:18:44,716 INFO  [c.c.v.ClusteredVirtualMachineManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Allocating Instance from Template: Template {"format":"QCOW2","id":202,"name":"SystemVM Template (KVM)","uniqueName":"routing-202","uuid":"<TMPL_AARCH64_UUID>"} with hostname: s-241-VM and 3 networks

2026-09-23 18:18:44,976 INFO  [o.a.c.e.o.VolumeOrchestrator] (secstorage-1:[ctx-804dea2b, ctx-c112c21a]) (logid:343a509f) Adding disk object [ROOT-241] to VM [VM instance {"id":241,"instanceName":"s-241-VM","state":"Stopped","type":"SecondaryStorageVm","uuid":"<SSVM_UUID>"}]
        ^-- creates volume 263 (template 202, aarch64)

2026-09-23 18:18:45,037 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Excluding 2 clusters as they are x86_64 arch, conflicting with the requested arch aarch64

2026-09-23 18:18:45,045 DEBUG [o.a.c.s.PremiumSecondaryStorageManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Unable to allocate secondary storage VM instance {"id":241,...} with Template {"format":"QCOW2","id":202,...,"uniqueName":"routing-202",...} due to [Unable to create a deployment for VM instance {"id":241,...}]. Retrying with another template com.cloud.exception.InsufficientServerCapacityException: Unable to create a deployment for VM instance {"id":241,...}Scope=interface com.cloud.dc.DataCenter; id=1
        ^-- "Retrying with another template": no line shows volume 263 being released

2026-09-23 18:18:45,074 INFO  [c.c.v.ClusteredVirtualMachineManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Allocating Instance from Template: Template {"format":"QCOW2","id":3,"name":"SystemVM Template (KVM)","uniqueName":"routing-3","uuid":"<TMPL_X86_UUID>"} with hostname: s-241-VM and 3 networks

2026-09-23 18:18:45,164 INFO  [o.a.c.e.o.VolumeOrchestrator] (secstorage-1:[ctx-804dea2b, ctx-70e344c6]) (logid:343a509f) Adding disk object [ROOT-241] to VM [VM instance {"id":241,...}]
        ^-- creates volume 264 (template 3, x86_64) — VM now has TWO ROOT volumes

2026-09-23 18:18:45,239 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Excluding 1 clusters as they are aarch64 arch, conflicting with the requested arch x86_64

2026-09-23 18:18:45,292 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Checking suitable pools for volume [Volume {"id":263,"instanceId":241,"name":"ROOT-241","uuid":"<VOL_AARCH64_UUID>","volumeType":"ROOT"}, ROOT] of VM [VM instance {"id":241,...}].
2026-09-23 18:18:45,349 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (secstorage-1:[ctx-804dea2b]) (logid:343a509f) Checking suitable pools for volume [Volume {"id":264,"instanceId":241,"name":"ROOT-241","uuid":"<VOL_X86_UUID>","volumeType":"ROOT"}, ROOT] of VM [VM instance {"id":241,...}].
        ^-- the planner now carries BOTH ROOT volumes for the same VM

Every subsequent start attempt plans both volumes and is rejected by the agent:

2026-09-23 18:19:26,965 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[ctx-8acced15, job-1069/job-2490, ctx-db3c7a5c]) (logid:62824aef) Excluding 1 clusters as they are aarch64 arch, conflicting with the requested arch x86_64
2026-09-23 18:19:27,025 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[...]) (logid:62824aef) Checking suitable pools for volume [Volume {"id":263,...,"volumeType":"ROOT"}, ROOT] of VM [VM instance {"id":241,...,"state":"Starting"}].
2026-09-23 18:19:27,030 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[...]) (logid:62824aef) Checking suitable pools for volume [Volume {"id":264,...,"volumeType":"ROOT"}, ROOT] of VM [VM instance {"id":241,...,"state":"Starting"}].
2026-09-23 18:19:33,557 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[...]) (logid:62824aef) Excluding 1 clusters as they are aarch64 arch, conflicting with the requested arch x86_64
2026-09-23 18:19:33,596 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[...]) (logid:62824aef) Checking suitable pools for volume [Volume {"id":263,...}] ...
2026-09-23 18:19:33,601 DEBUG [c.c.d.DeploymentPlanningManagerImpl] (Work-Job-Executor-91:[...]) (logid:62824aef) Checking suitable pools for volume [Volume {"id":264,...}] ...
Observations

No source was inspected; the following is only what the logs and database show.

  1. PremiumSecondaryStorageManagerImpl reports the preferred-arch attempt failing with
    InsufficientServerCapacityException and logs "Retrying with another template". This
    occurs after VolumeOrchestrator has already logged Adding disk object [ROOT-241]
    for the first template.
  2. Immediately after, ClusteredVirtualMachineManagerImpl logs a second Allocating Instance from Template — same VM instance id, different template id — and VolumeOrchestrator logs
    Adding disk object [ROOT-241] a second time.
  3. Two ROOT volume rows are persisted for that VM, one per template/arch, both with
    device_id = 0. No log line indicates the first volume being released, detached or
    expunged between the two attempts.
  4. From then on DeploymentPlanningManagerImpl evaluates both volumes on every start
    attempt, and the agent rejects the resulting domain XML with target 'vda' duplicated.
  5. The condition does not clear by itself. Destroying the VM expunges both volumes, but the
    next recreation reproduces it identically as long as the fallback path is reached.

The observable defect: the arch-fallback retry results in a second ROOT volume at the same
device_id on the same VM instance, and the volume from the abandoned attempt is never
removed.
Whether the right fix is to clean up before retrying, reuse the existing volume, or
prevent a duplicate device_id is for maintainers to judge.

Possibly related: #9981 (multi-arch zone — wrong-arch hosts offered for live migration),
another case of incomplete arch handling. Multi-arch zones were introduced in #9619 (4.20).

Workaround (verified)

Set the preferred architecture to one that actually has capacity, so the first allocation
attempt succeeds and the fallback retry is never reached.

⚠️ system.vm.preferred.architecture is zone-scoped. Changing only the global value has no
effect while a zone-level override exists. Verified: the global was set to x86_64 at
18:36:02, and a system VM created at 18:37:56 — 114 s later — still allocated the aarch64
template first, because data_center_details still held aarch64.

-- check for an overriding zone-scoped value first
SELECT dc_id, name, value FROM data_center_details
WHERE name = 'system.vm.preferred.architecture';
update configuration name=system.vm.preferred.architecture value=x86_64 zoneid=<ZONE_UUID>

Then expunge the stuck system VM so it is recreated. After doing so, the next recreation
(s-244-VM) allocated exactly one ROOT volume from template 3 and reached Running on an
x86_64 host, with the aarch64 host still unavailable:

v-240-VM  Running  tmpl=202  roots=1   (CPVM, on AARCH64_HOST01)
s-244-VM  Running  tmpl=3    roots=1   (SSVM, on X86_HOST03)

(Summarised from vm_instance / volumes; roots = count of non-removed ROOT volumes.)

This isolates the defect to the fallback retry path — when the first allocation attempt
succeeds, a system VM runs correctly on the non-preferred architecture with a single ROOT
volume. It is not a general problem with x86_64 system VMs.

The reverse direction is untested: losing all x86_64 hosts while preferring x86_64 would be
expected to trigger the identical failure toward aarch64.

versions

Apache CloudStack: 4.22.1.0 (3 management servers, all Up; agent 4.22.1.0)

Hypervisor: KVM

  • libvirt 8.0.0
  • QEMU 6.2.0
  • Host OS Ubuntu 22.04, kernel 5.15.0-191-generic

Storage: NFS primary storage (NetworkFilesystem), single shared pool, reachable from all
hosts of both architectures. Secondary storage also NFS.

Network: Advanced zone networking, single shared bridge per host (no per-traffic-type
separation).

The steps to reproduce the bug
1. Build a multi-arch zone containing at least one aarch64 cluster and one or more x86_64
   clusters, sharing primary storage.

2. Register both the aarch64 and the x86_64 KVM SystemVM templates, and confirm both are
   Ready.

3. Set system.vm.preferred.architecture = aarch64 at ZONE scope.

4. Allow the SSVM and CPVM to start normally. They land on the aarch64 host, each with a
   single ROOT volume from the aarch64 template. This is correct behaviour.

5. Make every aarch64 host unschedulable — putting the only aarch64 host into maintenance is
   sufficient.

6. Destroy the SSVM and/or CPVM so that CloudStack must recreate them.

7. Observe the recreated system VM:
     - the management log shows "Retrying with another template", with a second
       "Allocating Instance from Template" and a second "Adding disk object [ROOT-<id>]"
     - the VM now has TWO ROOT volumes, both device_id = 0, one per architecture:
         SELECT id, instance_id, volume_type, device_id, template_id
         FROM volumes WHERE instance_id = <vm id> AND removed IS NULL;
     - every start attempt fails on every host with:
         XML error: target 'vda' duplicated for disk sources '...' and '...'
     - the VM loops Starting -> Stopped indefinitely and never runs

Expected: the system VM is recreated from the x86_64 template with exactly one ROOT volume and
starts on an available x86_64 host.
What to do about it?

make sure the change in template type is taken into account when the recreate happens due to a change in architecture

主要語言
Java
星號
3.1k
分支
1.4k
平均合併
6 天 20 小時
30 天內合併 PR
27

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

apache/cloudstack 的其他 Issue

查看 apache/cloudstack 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。