VMware VM import fails when primary storage is a datastore cluster with multiple LUNs
Mantenedores costumam responder em até 1 dia
Ninguém assumiu esta issue ainda.
- #13886 de @DaanHoogland — fechado sem integrar
Avaliação
- Dificuldade
- 4/5
- Tempo estimado
- 3-5 dias
- Facilidade para iniciantes
- 48/100
- Tipo de issue
- Bug
- Clareza
- Razoavelmente clara
- Status de atividade
- Ativa
- Stack de tecnologia
- java
- Domínio
- cloud, infrastructure
Direção de pesquisa
Comece rastreando syncStoragePool e syncDatastoreClusterStoragePool em server/src/main/java/com/cloud/storage/StorageManagerImpl.java e, em seguida, inspecione o tratamento de ModifyStoragePoolCommand em VmwareResource.java e getDatastoresInDatastoreCluster() em StoragepodMO.java. Compare o caminho de conexão do host em DefaultHostListener.java com a busca de VMs não gerenciadas em UnmanagedVMsManagerImpl.java. Considera-se concluído quando um resultado incompleto do cluster de datastores não for mais aceito silenciosamente e o caminho afetado de sincronização/importação tiver verificação.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
The required feature described as a wish
claude generated description for problem asked about in #12561
problem
VMware "Datastore Cluster" (Storage DRS pod) primary storage pools can end up with zero child-datastore entries in the storage_pool table, even though syncStoragePool reports success. When this happens, importing an existing/unmanaged VM whose disk lives on one of the cluster's underlying datastores fails with:
Storage pool for disk Hard disk 1 (460-2000) with datastore: <datastore-name> not found in zone ID: <zone-uuid>
Originally reported at: https://github.com/apache/cloudstack/discussions/12561
Root cause (from code inspection)
Child storage_pool rows for a datastore-cluster pool are only created when a live ModifyStoragePoolCommand round-trip to a connected ESXi host returns a non-empty datastoreClusterChildren list:
plugins/hypervisors/vmware/src/main/java/com/cloud/hypervisor/vmware/resource/VmwareResource.java(execute(ModifyStoragePoolCommand), around line 5291) callsStoragepodMO.getDatastoresInDatastoreCluster()to enumerate child datastores.vmware-base/src/main/java/com/cloud/hypervisor/vmware/mo/StoragepodMO.java(getDatastoresInDatastoreCluster(), lines ~41-44) reads thechildEntityproperty with no validation or retry — if vCenter's property collector returns an empty/partial list (e.g. permissions/RBAC restricting which datastores the CloudStack service account can see on that specific pod, or a transient inventory glitch), no exception is raised.server/src/main/java/com/cloud/storage/StorageManagerImpl.java(syncDatastoreClusterStoragePool(...), lines ~2760-2815) persists whatever child list comes back. An empty list simply results in "0 added, 0 removed" — there is no check that the returned child count is non-zero or matches expectations.- This sync path is triggered by host-connect events (
DefaultHostListener.hostConnect,engine/storage/volume/src/main/java/org/apache/cloudstack/storage/datastore/provider/DefaultHostListener.java, lines ~136-186) and by the explicitsyncStoragePoolAPI (StorageManagerImpl.syncStoragePool, lines ~2643-2699), which also only samples one arbitrary connected host to answer.
Net effect: if the first (and every subsequent) sync attempt for this specific pod happens to get an empty answer from whichever host was asked, the pool is silently left with no child rows forever, while the API/UI report the sync as successful. Other datastore clusters in the same environment are unaffected because their sync happened to succeed.
Downstream, getStoragePool(...) in server/src/main/java/org/apache/cloudstack/vm/UnmanagedVMsManagerImpl.java (lines ~516-551) does a generic path/name substring match over all pools in the cluster/zone — it works fine for datastore-cluster children when their row exists, so this is not a separate bug in the import code; it's a downstream symptom of the missing storage_pool row.
What versions of cloudstack and any infra components are you using
- CloudStack: 4.20.0.0
- Hypervisor: VMware (vCenter, ESXi)
- Primary storage: Datastore Cluster (Storage DRS pod) containing 4 LUNs
- Storage: HPE Alletra 5050 (FC/iSCSI LUNs presented as VMFS datastores)
The steps to reproduce the bug
- Create a VMware Datastore Cluster (Storage DRS pod) in vCenter containing multiple LUNs/datastores.
- Add the datastore cluster as CloudStack primary storage.
- Confirm in the
storage_pooltable that no child rows were created for the individual datastores (parent= the cluster pool's ID) — only the single parent/cluster row exists. - Run
syncStoragePoolagainst the pool — it completes successfully but still does not populate child rows. - Attempt to import an existing/unmanaged VM whose disk resides on one of the underlying datastores in that cluster.
- Import fails:
Storage pool for disk <disk> with datastore: <name> not found in zone ID: <zone>.
Note: other datastore clusters in the same environment, added the same way, do have correctly populated child rows — the failure appears to depend on whether the very first sync happened to get a complete answer from vCenter.
What to do about it?
Suggested fix directions:
- In
StoragepodMO.getDatastoresInDatastoreCluster()/VmwareResource.execute(ModifyStoragePoolCommand), treat an empty (or unexpectedly short)childEntity/datastoreClusterChildrenresult as suspicious rather than a normal success — e.g. log a warning, or retry against a different host before persisting. - In
StorageManagerImpl.syncDatastoreClusterStoragePool(...)/syncStoragePool(...), surface a warning (API response and/or logs) when a datastore-cluster sync returns zero child datastores for a pool that previously had none, instead of silently reporting success. - Consider having
syncStoragePooltry more than one connected host when the first host's answer looks incomplete, since host-level permissions/visibility to the storage pod can vary.
As a workaround, users can try re-triggering sync from a different host in the cluster, verify the vCenter service account's permissions on the affected storage pod/datastores, or remove and re-add the primary storage pool to force a fresh full discovery.
- Linguagem predominante
- Java
- Estrelas
- 3.1k
- Forks
- 1.4k
- Merge médio
- 6d 11h
- PRs com merge (30d)
- 13
Preparar o ambiente
- Sem Dockerfile nem arquivo Docker Compose
- Tem um modelo de pull request
- Ler o guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de apache/cloudstack
-
listPublicIpAddresses NullPointerException on shared networks with a VR (4.22)Talvez já em andamento Um pull request vinculado a esta issue está aberto ou já foi mesclado. Abertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 80/100
apache/cloudstack#14248 ·
Mantenedores costumam responder em até 1 dia
-
CKS: upgradeKubernetesCluster fails on control node when binaries ISO ships headlamp.yaml instead of dashboard.yamlTalvez já em andamento @mw-0 assumiu há 11 dias. Abertabug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
apache/cloudstack#14244 · 1 comentário ·
Mantenedores costumam responder em até 1 dia
-
Resize volume API validation errors are not displayed in the UITalvez já em andamento @sathvikaragi assumiu há 14 dias. Abertabug
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 90/100
apache/cloudstack#14222 ·
Mantenedores costumam responder em até 1 dia
-
create-kubernetes-binaries-iso.sh builds the ISO without setting a volume ID on EL8 based os'sAbertabug component:kubernetes
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
apache/cloudstack#14180 ·
Mantenedores costumam responder em até 1 dia
-
bug component:projects component:UI
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
apache/cloudstack#14070 · 5 comentários ·
Mantenedores costumam responder em até 1 dia
Todas as issues de apache/cloudstack
Issues semelhantes
-
[destination-snowflake] Custom domains rejected unlike source connectionsTalvez já em andamento @kuza55 assumiu hoje. Abertaautoteam community connectors/destination/snowflake team/use
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
area-dashboard
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 68/100
Mantenedores costumam responder em até 1 dia
-
component/operate kind/feature-request
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
Mantenedores costumam responder em até 1 dia
-
Forge coverage prompts carry text the agent cannot act onTalvez já em andamento @graalvmbot assumiu hoje. Aberta
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 85/100
oracle/graalvm-reachability-metadata#10572 ·
Mantenedores costumam responder em até 1 dia
-
[CI] Core CI doesn't run for changes to amoro-format-lance (and amoro-web)Talvez já em andamento @MarkAlex1234 assumiu hoje. Aberta
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 88/100
Mantenedores costumam responder em até 2 dias