SystemVM agent fails TLS to management server
Maintainer antworten meist innerhalb von 1 Tag
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 65/100
- Issue-Typ
- Bug
- Klarheit
- Klar beschrieben
- Aktivitätsstatus
- Aktiv
- Tech-Stack
- bash, java
- Bereich
- backend, cloud, infrastructure
Rechercherichtung
Beginnen Sie mit SetupCertificateCommand.java und CAManagerImpl.java, um den Zertifikat-Payload nachzuverfolgen, und prüfen Sie anschließend den SetupCertificateCommand-Zweig in VmwareResource.java sowie scripts/util/keystore-cert-import. Reproduzieren oder testen Sie die VMware-SSH-Weiterleitung und verifizieren Sie, dass die CA-Argumente das Skript erreichen, dass fehlende CA-Eingaben mit einem Exit-Code ungleich null fehlschlagen oder den dokumentierten Disk-Fallback verwenden und dass cloud.jks einen trustedCertEntry erhält.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
SystemVM agent fails TLS to management server — keystore-cert-import silently exits 0, leaving cloud.jks without trustedCertEntry (VMware, 4.22)
problem
After a fresh system VM (ConsoleProxy / SecondaryStorageVm) boots, its agent cannot complete the TLS handshake with the management server on port 8250. The console proxy never binds its public HTTP/HTTPS listener (browser gets ERR_CONNECTION_REFUSED on the public IP), the SSVM never registers as a host, and ACS keeps destroying and re-creating the system VMs — each replacement exhibits the same failure.
Root cause: on the VMware SSH dispatch path, SetupCertificateCommand invokes scripts/util/keystore-cert-import without positional arguments $7 (CACERT_FILE path) and $8 (CACERT content). The script then falls through to the elif [ ! -f "$CACERT_FILE" ] branch and calls a bare exit — which returns exit code 0. The management server receives SetupCertificateAnswer: result=true and believes the push succeeded. In reality the awk ... "$CACERT_FILE" and the subsequent keytool -import -trustcacerts -alias cloudca calls never run, so the CA certificate is never added to cloud.jks as a trustedCertEntry.
Evidence observed on the affected deployment.
Management-server log, every ~10s while the system VM is Running:
ERROR [c.c.u.n.Link] (AgentManager-SSLHandshakeHandler-1:[])
SSL error caught during unwrap data: (certificate_unknown)
Received fatal alert: certificate_unknown,
for local address=/<MS>:8250, remote address=/<SYSVM>:xxxxx.
The client may have invalid ca-certificates.
System VM /var/log/cloud.log, reciprocally:
ERROR [utils.nio.Link] SSL error caught during wrap data:
No trusted certificate found, for local address=/<SYSVM>:xxxxx,
remote address=/<MS>:8250.
Keystore inspection inside the VM (using the passphrase stored in agent.properties):
$ keytool -list -keystore /usr/local/cloud/systemvm/conf/cloud.jks \
-storepass "$(sed -n 's/^keystore.passphrase=//p' /usr/local/cloud/systemvm/conf/agent.properties)"
Your keystore contains 1 entry
cloud, <date>, PrivateKeyEntry, ...
Only the leaf PrivateKeyEntry is present — no trustedCertEntry for the root CA. The file /usr/local/cloud/systemvm/conf/cloud.ca.crt is correctly written on disk (byte-for-byte identical to the live MS root CA), but Java does not read loose PEM files on disk; it only trusts entries inside cloud.jks.
The relevant piece of scripts/util/keystore-cert-import in 4.22.0.0:
CACERT_FILE="$7"
CACERT=$(echo "$8" | tr '^' '\n' | tr '~' ' ')
...
# Import ca certs
if [ ! -z "${CACERT// }" ]; then
echo "$CACERT" > "$CACERT_FILE"
elif [ ! -f "$CACERT_FILE" ]; then
echo "Cannot find ca certificate file: $CACERT_FILE, exiting!"
exit # <-- bare exit -> exit 0 -> MS sees success
fi
awk '/-----BEGIN CERTIFICATE-----?/{n++}{print > "cloudca." n }' "$CACERT_FILE"
for caChain in $(ls cloudca.*); do
keytool -delete -noprompt -alias "$caChain" -keystore "$KS_FILE" -storepass "$KS_PASS" > /dev/null 2>&1 || true
keytool -import -noprompt -storepass "$KS_PASS" -trustcacerts -alias "$caChain" -file "$caChain" -keystore "$KS_FILE" > /dev/null 2>&1
done
Example arguments actually logged by MS for a v-26-VM SetupCertificateCommand dispatch (6 positional args, not 10):
Run command on VR: <SYSVM_IP>, script: keystore-cert-import with args:
/usr/local/cloud/systemvm/conf/agent.properties <KS_PASS>
/usr/local/cloud/systemvm/conf/cloud.jks
ssh
/usr/local/cloud/systemvm/conf/cloud.crt
"-----BEGIN~CERTIFICATE-----^...leaf cert content..."
$7 (CACERT_FILE) and $8 (CACERT content) are absent.
versions
- Apache CloudStack: 4.22.0.0 (package
cloudstack-management-4.22.0.0-shapeblue0, ShapeBlue RPM build) - Management server host OS: Rocky Linux 10.0
- Management server JDK: OpenJDK 21.0.10 (for
cloudstack-management.service) - SystemVM template:
systemvmtemplate-4.22.0-x86_64-vmware.ova(official 4.22.0 vSphere OVA, built Tue Oct 14 11:01:03 UTC 2025) - SystemVM guest OS: Debian 12 (bookworm), OpenJDK 17.0.16 (
keytoolas shipped in the image) - Hypervisor: VMware vSphere (ESXi cluster, 6 hosts)
- Primary storage: VMFS
- Secondary storage: NFS
- Network: VMware DVS, shared Public VLAN for system VMs, Basic zone networking
- CA framework:
ca.framework.provider.plugin=root,ca.plugin.root.auth.strictness=false
The steps to reproduce the bug
- Install Apache CloudStack 4.22.0.0 management server on a clean host (e.g. Rocky 10).
- Register a VMware zone and upload the official
systemvmtemplate-4.22.0-x86_64-vmware.ovaas the SystemVM template (typeSYSTEM). - Let ACS auto-provision the ConsoleProxy and SecondaryStorageVm for that zone.
- Wait ~1 minute after both system VMs reach
state=Running,power_state=PowerOn. - Observe in
management-server.logthatSetupCertificateCommandis dispatched via the VMware SSH path and returnsSetupCertificateAnswer: result=true. - Despite that "success",
management-server.logkeeps emittingSSL error caught during unwrap data: (certificate_unknown) Received fatal alert: certificate_unknown, ... remote address=/<SYSVM>:xxxxxevery ~10 s. - SSH into the system VM on port 3922 using the MS key (
/var/cloudstack/management/.ssh/id_rsa) and run:
Expected (buggy) output: the keystore contains exactly 1 entry (PASS=$(sed -n 's/^keystore.passphrase=//p' /usr/local/cloud/systemvm/conf/agent.properties) keytool -list -keystore /usr/local/cloud/systemvm/conf/cloud.jks -storepass "$PASS"cloud,PrivateKeyEntry) — notrustedCertEntry. - Verify
/usr/local/cloud/systemvm/conf/cloud.ca.crtexists on disk and matches the MS root CA (same serial asca.plugin.root.ca.certificatein theconfigurationtable), but it was never imported into the keystore. - Result: the agent never completes TLS, the host never reaches
Up, and the console proxy never opens its public HTTP listener — browser hitsERR_CONNECTION_REFUSEDon the system VM public IP.
What to do about it?
Two fixes, ideally both.
A. Primary — ensure the VMware SSH dispatch of SetupCertificateCommand passes the CA args.
The dispatch path that invokes scripts/util/keystore-cert-import needs to also pass positional arguments $7 (/usr/local/cloud/systemvm/conf/cloud.ca.crt) and $8 (CA certificate content, same ^/~-encoded form used for the leaf cert in $6). The CA material is already available in the Certificate object used by CAManagerImpl (see server/src/main/java/org/apache/cloudstack/ca/CAManagerImpl.java around the SetupCertificateCommand cmd = new SetupCertificateCommand(certificate) call) and in ca.plugin.root.ca.certificate in the configuration table — it just isn't being carried into the script invocation on the VMware SSH path.
Relevant files:
core/src/main/java/org/apache/cloudstack/ca/SetupCertificateCommand.java— command payload; may need to explicitly carry the CA certificate to the VR path.server/src/main/java/org/apache/cloudstack/ca/CAManagerImpl.java— where the command is built; has theCertificateobject with CA material.plugins/hypervisors/vmware/src/main/java/com/cloud/hypervisor/vmware/resource/VmwareResource.java— VMwareexecuteInVRdispatch; check the branch that handlesSetupCertificateCommandand routes it tokeystore-cert-import.
B. Secondary — make scripts/util/keystore-cert-import fail loud, with a safe disk fallback.
Current (buggy) behaviour:
elif [ ! -f "$CACERT_FILE" ]; then
echo "Cannot find ca certificate file: $CACERT_FILE, exiting!"
exit
fi
Suggested patch:
- elif [ ! -f "$CACERT_FILE" ]; then
- echo "Cannot find ca certificate file: $CACERT_FILE, exiting!"
- exit
- fi
+ elif [ ! -f "$CACERT_FILE" ]; then
+ # Fall back to the CA cert that cloud-early-config placed on disk.
+ if [ -f /usr/local/cloud/systemvm/conf/cloud.ca.crt ]; then
+ CACERT_FILE=/usr/local/cloud/systemvm/conf/cloud.ca.crt
+ else
+ echo "Cannot find ca certificate file; neither arg \$7 nor /usr/local/cloud/systemvm/conf/cloud.ca.crt available, aborting!" >&2
+ exit 1
+ fi
+ fi
A non-zero exit lets the management server see the failure (today, SetupCertificateAnswer.result=true completely masks it). The disk fallback is defense-in-depth against future dispatch-side regressions of this same shape.
Workaround currently in use on the affected deployment (for anyone hitting this before a fix is merged):
For each system VM, extract the existing leaf key+cert with openssl pkcs12, rebuild the PKCS12 with openssl pkcs12 -export ... -certfile cloud.ca.crt, then keytool -import -noprompt -trustcacerts -alias cloudca -file cloud.ca.crt -keystore cloud.jks -storepass <agent.properties passphrase>, then systemctl restart cloud. Note: the keytool shipped in the 4.22 systemvm template (OpenJDK 17.0.16, Debian 12) rejected -import of both PEM and DER forms of the CA with CertificateParsingException: signed overrun, bytes = 115, so the rebuilt keystore had to be generated on the management server (OpenJDK 21) and copied back. That broken keytool -import behaviour on the systemvm image may deserve a separate issue.
- Vorherrschende Sprache
- Java
- Sterne
- 3.1k
- Forks
- 1.4k
- Ø Merge
- 5 T. 19 Std.
- Gemergte PRs (30 T.)
- 17
Entwicklungsumgebung
- Kein Dockerfile und keine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus apache/cloudstack
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 80/100
apache/cloudstack#14248 ·
Maintainer antworten meist innerhalb von 1 Tag
-
bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
apache/cloudstack#14244 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
bug
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 90/100
apache/cloudstack#14222 ·
Maintainer antworten meist innerhalb von 1 Tag
-
create-kubernetes-binaries-iso.sh builds the ISO without setting a volume ID on EL8 based os'sOffenbug component:kubernetes
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 88/100
apache/cloudstack#14180 ·
Maintainer antworten meist innerhalb von 1 Tag
-
bug component:projects component:UI
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 88/100
apache/cloudstack#14070 · 5 Kommentare ·
Maintainer antworten meist innerhalb von 1 Tag
Alle Issues in apache/cloudstack
Ähnliche Issues
-
type: possible bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 72/100
grimmory-tools/grimmory#2850 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 90/100
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 65/100
aoqia194/leaf-loader#19 ·