Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

UI extremely slow when db.ha.enabled=true after MySQL source failure (4.22.1.0)

Open
#14,172 0 comments 0 reactions 0 assignees View on GitHub

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
java, mysql

Research direction

Start by reproducing the failure with db.ha.enabled=true, db.cloud.replicas and db.usage.replicas set in db.properties, then inspect the MySQL connector failover behavior and its retry settings. Compare requests after the source is isolated with the documented manual-promotion procedure. Done means the Management Server fails over promptly to the replica and the UI remains usable without repeated source retries.

Written by the indexing model from the issue text.

Description

bug
problem

With Database High Availability enabled (db.ha.enabled=true), after the MySQL source becomes unreachable the Management Server UI becomes extremely slow / barely usable.

Failover to the configured replica seems to happen via the MySQL connector, but each (or almost each) request appears to retry the dead source first, causing severe UI latency.

Workaround that restores normal performance:

  • disable db.ha.enabled
  • manually promote the replica (STOP REPLICA / RESET REPLICA ALL / read_only=OFF)
  • point db.cloud.host / db.usage.host to the promoted DB in db.properties
  • restart cloudstack-management

This matches the install guide "Failover" procedure better than the client-side db.ha mechanism.

Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way replication.
Ref: https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability

versions

CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
OS: Ubuntu 24.04
MySQL: 8.0.46 (async replication source→replica, GTID)
Hypervisor: VMware
Topology: 2 Management Servers (multi-site), reverse proxy in front of UI

The steps to reproduce the bug
  1. Deploy 2 MS pointing to MySQL source; configure two-way replica (GTID)
  2. Set in db.properties:
    db.ha.enabled=true
    db.cloud.replicas=replica-ip
    db.usage.replicas=replica-ip
  3. Restart cloudstack-management on both nodes
  4. Stop / isolate the MySQL source (simulate site/DB failure)
  5. Use UI/API through the remaining Management Server
What to do about it?

Expected: MS fail over to replica promptly; UI remains usable.

Actual: UI becomes very slow after source outage (connector retries against unreachable source; defaults secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 3600/5000/3600).

Suggestions:

  • fix connector failover so replica is used without per-request source retries hanging the UI
  • and/or document clearly that db.ha.enabled is not recommended for production on MySQL 8.x and that manual promotion (install guide Failover) is the supported path
Dominant language
Java
Stars
3.1k
Forks
1.4k
Avg merge
6d 20h
Merged PRs (30d)
27

Contributor guide

Open the contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/cloudstack

All issues in apache/cloudstack

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.