UsageManagerImpl.parse() rewind to the oldest unprocessed usage event is unbounded
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 35/100
Direzione di ricerca
Inizia con UsageManagerImpl.parse() e UsageEventDaoImpl.listLatestEvents(), quindi esamina le query usage_job e usage_event descritte nell’issue. Riproduci lo start_date fissato e l’intervallo di aggregazione crescente, quindi definisci un comportamento di riavvolgimento limitato o di quarantena degli eventi i cui criteri di completamento includano impedire che un evento permanentemente non elaborato causi una riaggregazione senza limiti.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
problem
Split out of #13399 at @DaanHoogland's request.
UsageManagerImpl.parse() correctly derives its start date from the last successful job:
long lastSuccess = _usageJobDao.getLastJobSuccessDateMillis();
if (lastSuccess != 0) {
startDateMillis = lastSuccess + 1;
}
and then overwrites it with the create date of the oldest unprocessed event, moving it only ever earlier:
// make sure start date is before all of our un-processed events
Date oldestEventDate = events.get(0).getCreateDate();
if (oldestEventDate.getTime() < startDateMillis) {
startDateMillis = oldestEventDate.getTime();
startDate = new Date(startDateMillis);
}
events comes from UsageEventDaoImpl.listLatestEvents() — WHERE processed = 0 AND created <= ? ORDER BY createDate ASC, with no limit and no floor.
Consequence: any event that can never be successfully processed pins the aggregation start date permanently. Each subsequent run re-aggregates from that date to the present, growing by one aggregation range per range elapsed. The job continues to report success = 1 throughout, so nothing surfaces as an error.
Measured on an affected 4.22.1.0 deployment:
- 1,678 consecutive successful jobs whose
start_millisnever advanced past a single event 71 days earlier exec_timeof 2,555,055 ms (42.6 min) per hourly run, growing by 24 aggregation periods per daycloud_usageat 54.4M rows / 14 GB, of which roughly 150k rows were genuine; duplicate counts were an exact multiple of the number of replays
Not specific to one event type: #13112 was reported on 4.21.0.0 against usage_type = 13, predating the volume-specific trigger in #13399. The rewind is the common mechanism; the triggering event type varies.
versions
Observed on CloudStack 4.22.1.0 (EL9 packages), MySQL 8.x / InnoDB, with usage.stats.job.aggregation.range = 60 and usage.stats.job.exec.time = 00:15.
#13112 reports the same mechanism on 4.21.0.0, so this is not new in 4.22.
The steps to reproduce the bug
- Have any row in
cloud_usage.usage_eventthat cannot be successfully processed, #13399 gives one reliable route, but the mechanism is independent of cause. - Let the usage job run on its normal schedule.
SELECT id, start_date, end_date, exec_time FROM cloud_usage.usage_job WHERE success = 1 ORDER BY id DESC LIMIT 5;—start_datestays pinned at the oldest unprocessed event'screatedwhileend_dateadvances.TIMESTAMPDIFF(HOUR, start_date, end_date)grows without bound, as doesexec_time.- Row counts in
cloud_usage.cloud_usageper period equal the number of times that period has been re-aggregated.
What to do about it?
Bound the rewind, and/or provide a way to quarantine or age out events that repeatedly fail to process, so one unprocessable row cannot halt aggregation indefinitely.
This is a design decision rather than an obvious patch, which is why it's raised separately from #13399.
- Lingua principale
- Java
- Stelle
- 3.1k
- Fork
- 1.4k
- Merge medio
- 6g 20h
- PR unite (30g)
- 27
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di apache/cloudstack
-
bug
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
apache/cloudstack#14222 ·
-
create-kubernetes-binaries-iso.sh builds the ISO without setting a volume ID on EL8 based os's Apertabug component:kubernetes
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
apache/cloudstack#14180 ·
-
bug component:projects component:UI
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
apache/cloudstack#14070 · 5 commenti ·
-
component:backup
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
apache/cloudstack#14013 ·
-
KVM agent fails to connect to Ceph RBD storage pool after upgrading Ceph client to Tentacle 20.2.4 Apertabug component:ceph
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
apache/cloudstack#13989 · 3 commenti ·
Tutte le issue di apache/cloudstack
Issue simili
-
awaiting triage bug Causes friction Hop Gui P1 P2 Transforms
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
apache/flink-agents#1152 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 70/100
jenkinsci/blueocean-plugin#5417 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
objectionary/eo-graphs#75 ·