UI extremely slow when db.ha.enabled=true after MySQL source failure (4.22.1.0)
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 48/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Active
- Tech stack
- java, mysql
- Domain
- backend, database, performance
Research direction
Start by reproducing the failure with db.ha.enabled=true, db.cloud.replicas and db.usage.replicas set in db.properties, then inspect the MySQL connector failover behavior and its retry settings. Compare requests after the source is isolated with the documented manual-promotion procedure. Done means the Management Server fails over promptly to the replica and the UI remains usable without repeated source retries.
Written by the indexing model from the issue text.
Description
problem
With Database High Availability enabled (db.ha.enabled=true), after the MySQL source becomes unreachable the Management Server UI becomes extremely slow / barely usable.
Failover to the configured replica seems to happen via the MySQL connector, but each (or almost each) request appears to retry the dead source first, causing severe UI latency.
Workaround that restores normal performance:
- disable db.ha.enabled
- manually promote the replica (STOP REPLICA / RESET REPLICA ALL / read_only=OFF)
- point db.cloud.host / db.usage.host to the promoted DB in db.properties
- restart cloudstack-management
This matches the install guide "Failover" procedure better than the client-side db.ha mechanism.
Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way replication.
Ref: https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability
versions
CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
OS: Ubuntu 24.04
MySQL: 8.0.46 (async replication source→replica, GTID)
Hypervisor: VMware
Topology: 2 Management Servers (multi-site), reverse proxy in front of UI
The steps to reproduce the bug
- Deploy 2 MS pointing to MySQL source; configure two-way replica (GTID)
- Set in db.properties:
db.ha.enabled=true
db.cloud.replicas=replica-ip
db.usage.replicas=replica-ip - Restart cloudstack-management on both nodes
- Stop / isolate the MySQL source (simulate site/DB failure)
- Use UI/API through the remaining Management Server
What to do about it?
Expected: MS fail over to replica promptly; UI remains usable.
Actual: UI becomes very slow after source outage (connector retries against unreachable source; defaults secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 3600/5000/3600).
Suggestions:
- fix connector failover so replica is used without per-request source retries hanging the UI
- and/or document clearly that db.ha.enabled is not recommended for production on MySQL 8.x and that manual promotion (install guide Failover) is the supported path
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.4k
- Avg merge
- 6d 20h
- Merged PRs (30d)
- 27
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/cloudstack
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 90/100
apache/cloudstack#14222 ·
-
bug component:kubernetes
Difficulty 1/5 Under an hour Newbie friendliness 88/100
apache/cloudstack#14180 ·
-
bug component:projects component:UI
Difficulty 1/5 Under an hour Newbie friendliness 88/100
apache/cloudstack#14070 · 5 comments ·
-
component:backup
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
apache/cloudstack#14013 ·
-
KVM agent fails to connect to Ceph RBD storage pool after upgrading Ceph client to Tentacle 20.2.4 Openbug component:ceph
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/cloudstack#13989 · 3 comments ·
All issues in apache/cloudstack
Similar issues
-
area/plugin
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
kestra-io/plugin-kestra#190 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
google-ai-edge/LiteRT-LM#3739 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
integra-team-red/meet-map#249 ·
-
[Studio][Bug] Cancelled create-user dialog keeps the password and admin switch for the next attempt Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
apache/rocketmq-dashboard#5064 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
wso2/dpdp-accelerator#287 ·