CloudStack 4.22: `prepareHostForMaintenance` throws NPE when stale destroyed volume references removed storage pool
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 55/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Ít trao đổi
- Công nghệ
- java, mariadb, mysql
- Lĩnh vực
- backend, cloud, infrastructure
Hướng nghiên cứu
Bắt đầu tại UserVmManagerImpl.isAnyVmVolumeUsingLocalStorage ở dòng 7558 và lần theo các nơi gọi nó trong isVMUsingLocalStorage và ResourceManagerImpl.doMaintain. Tái hiện trạng thái của một volume đã bị hủy nhưng vẫn còn tồn đọng và một storage pool đã bị xóa, sau đó xác minh rằng quá trình bảo trì không còn phát sinh NullPointerException không được xử lý và hành vi được chọn cho volume tồn đọng đó được một test phù hợp bao phủ.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
problem
Description
While attempting to place a KVM host into Maintenance mode in CloudStack 4.22, the maintenance operation failed with a NullPointerException.
The issue appears to occur when a VM has a stale/destroyed volume entry in the volumes table that still references a removed storage pool.
Instead of gracefully ignoring the stale volume metadata or returning a user-facing validation error, CloudStack crashes during maintenance preparation.
Environment
- CloudStack version: 4.22
- Hypervisor: KVM
- Primary storage: NFS
- Database: MySQL/MariaDB
Steps to reproduce
-
Have a VM with:
- one valid active ROOT volume
- one stale/destroyed ROOT volume entry still present in the
volumestable
-
The stale volume references a storage pool which has already been removed.
-
Attempt to place the host running the VM into Maintenance mode.
Example problematic DB state:
VM
SELECT id, uuid, name, instance_id, pool_id, state, removed
FROM volumes
WHERE instance_id = 446
ORDER BY id;
Result:
+------+--------------------------------------+----------+-------------+---------+---------+---------+
| id | uuid | name | instance_id | pool_id | state | removed |
+------+--------------------------------------+----------+-------------+---------+---------+---------+
| 928 | 6dea3d6f-bd6d-4e8b-9524-6e99c029694c | ROOT-446 | 446 | 4 | Destroy | NULL |
| 1554 | 80acf9ae-b047-41b8-bded-cdceb6de7051 | ROOT-446 | 446 | 2 | Ready | NULL |
+------+--------------------------------------+----------+-------------+---------+---------+---------+
Storage pool state
The stale volume references storage pool ID 4, which is already removed:
storage_pool_name: Export-Domain
storage_pool_removed: 2026-04-23 13:44:08
Actual result
Host maintenance fails with:
java.lang.NullPointerException: Cannot invoke
"org.apache.cloudstack.storage.datastore.db.StoragePoolVO.isLocal()"
because "storagePool" is null
Relevant stack trace:
at com.cloud.vm.UserVmManagerImpl.isAnyVmVolumeUsingLocalStorage(UserVmManagerImpl.java:7558)
at com.cloud.vm.UserVmManagerImpl.isVMUsingLocalStorage(UserVmManagerImpl.java:7121)
at com.cloud.resource.ResourceManagerImpl.doMaintain(ResourceManagerImpl.java:1553)
at com.cloud.resource.ResourceManagerImpl.maintain(ResourceManagerImpl.java:1653)
at org.apache.cloudstack.api.command.admin.host.PrepareForHostMaintenanceCmd.execute(PrepareForHostMaintenanceCmd.java:99)
Expected result
CloudStack should not throw an unhandled NullPointerException.
Possible expected behavior:
- ignore destroyed/removed stale volumes during maintenance evaluation
- skip volumes attached to removed pools
- or return a proper validation error identifying the problematic VM/volume
Workaround
Marking the stale destroyed volume row as removed allowed maintenance to proceed:
UPDATE volumes
SET removed = NOW()
WHERE id = 928
AND instance_id = 446
AND state = 'Destroy'
AND removed IS NULL
AND pool_id = 4;
Additional notes
The issue appears to be triggered specifically by:
- stale destroyed volume rows
- still linked to an active/running VM
- referencing removed storage pools
- while evaluating VM local-storage usage during host maintenance
versions
The versions ACS 4.22, KVM (should not be relevant)
The steps to reproduce the bug
...
What to do about it?
No response
- Ngôn ngữ chính
- Java
- Star
- 3.1k
- Fork
- 1.4k
- Merge trung bình
- 6 ngày 20 giờ
- Pull request đã merge (30 ngày)
- 27
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của apache/cloudstack
-
bug
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
apache/cloudstack#14222 ·
-
create-kubernetes-binaries-iso.sh builds the ISO without setting a volume ID on EL8 based os's Đang mởbug component:kubernetes
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
apache/cloudstack#14180 ·
-
bug component:projects component:UI
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 88/100
apache/cloudstack#14070 · 5 bình luận ·
-
component:backup
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
apache/cloudstack#14013 ·
-
KVM agent fails to connect to Ceph RBD storage pool after upgrading Ceph client to Tentacle 20.2.4 Đang mởbug component:ceph
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
apache/cloudstack#13989 · 3 bình luận ·
Tất cả issue của apache/cloudstack
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
elastic/gradle-plugins#157 ·
-
enhancement Tools
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 75/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
apache/rocketmq-dashboard#5008 ·
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
DETECT_PARAMETER_NAMES=false silently disables @ConstructorProperties-based Creator detection too Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
FasterXML/jackson-databind#6229 ·