UI extremely slow when db.ha.enabled=true after MySQL source failure (4.22.1.0)

Offen
#14,172 0 Kommentare 0 Reaktionen 0 zugewiesene Personen Auf GitHub ansehen

Dieses Issue hat noch niemand übernommen.

Bewertung

Schwierigkeit
4/5
Geschätzter Aufwand
3-5 Tage
Anfängerfreundlichkeit
48/100
Issue-Typ
Bug
Klarheit
Größtenteils klar
Aktivitätsstatus
Aktiv
Tech-Stack
java, mysql

Rechercherichtung

Start by reproducing the failure with db.ha.enabled=true, db.cloud.replicas and db.usage.replicas set in db.properties, then inspect the MySQL connector failover behavior and its retry settings. Compare requests after the source is isolated with the documented manual-promotion procedure. Done means the Management Server fails over promptly to the replica and the UI remains usable without repeated source retries.

Vom Indexierungsmodell aus dem Issue-Text verfasst.

Beschreibung

bug
problem

With Database High Availability enabled (db.ha.enabled=true), after the MySQL source becomes unreachable the Management Server UI becomes extremely slow / barely usable.

Failover to the configured replica seems to happen via the MySQL connector, but each (or almost each) request appears to retry the dead source first, causing severe UI latency.

Workaround that restores normal performance:

  • disable db.ha.enabled
  • manually promote the replica (STOP REPLICA / RESET REPLICA ALL / read_only=OFF)
  • point db.cloud.host / db.usage.host to the promoted DB in db.properties
  • restart cloudstack-management

This matches the install guide "Failover" procedure better than the client-side db.ha mechanism.

Note: docs still say "Tested with MySQL 5.1 and 5.5" and suggest two-way replication.
Ref: https://docs.cloudstack.apache.org/en/4.22.1.0/adminguide/reliability.html#configuring-database-high-availability

versions

CloudStack: 4.22.1.0 (Ubuntu packages, download.cloudstack.org noble)
OS: Ubuntu 24.04
MySQL: 8.0.46 (async replication source→replica, GTID)
Hypervisor: VMware
Topology: 2 Management Servers (multi-site), reverse proxy in front of UI

The steps to reproduce the bug
  1. Deploy 2 MS pointing to MySQL source; configure two-way replica (GTID)
  2. Set in db.properties:
    db.ha.enabled=true
    db.cloud.replicas=replica-ip
    db.usage.replicas=replica-ip
  3. Restart cloudstack-management on both nodes
  4. Stop / isolate the MySQL source (simulate site/DB failure)
  5. Use UI/API through the remaining Management Server
What to do about it?

Expected: MS fail over to replica promptly; UI remains usable.

Actual: UI becomes very slow after source outage (connector retries against unreachable source; defaults secondsBeforeRetrySource/queriesBeforeRetrySource/initialTimeout = 3600/5000/3600).

Suggestions:

  • fix connector failover so replica is used without per-request source retries hanging the UI
  • and/or document clearly that db.ha.enabled is not recommended for production on MySQL 8.x and that manual promotion (install guide Failover) is the supported path
Vorherrschende Sprache
Java
Sterne
3.1k
Forks
1.4k
Ø Merge
7 T. 5 Std.
Gemergte PRs (30 T.)
28

Beitragsleitfaden

Beitragsleitfaden öffnen

Erste Schritte

  1. Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
  2. Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
  3. Forken Sie das Repository und arbeiten Sie in einem Branch.
  4. Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.

Mehr aus apache/cloudstack

Alle Issues in apache/cloudstack

Ähnliche Issues

Weitere Issues zu Java

Neue Issues direkt in Ihr Postfach

Eine kurze Übersicht über anfängerfreundliche GitHub-Issues.