UsageManagerImpl.parse() rewind to the oldest unprocessed usage event is unbounded
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 35/100
Research direction
Start with UsageManagerImpl.parse() and UsageEventDaoImpl.listLatestEvents(), then inspect the usage_job and usage_event queries described in the issue. Reproduce the pinned start_date and growing aggregation range, and define a bounded rewind or event-quarantine behavior whose completion criteria include preventing one permanently unprocessed event from causing unbounded re-aggregation.
Written by the indexing model from the issue text.
Description
problem
Split out of #13399 at @DaanHoogland's request.
UsageManagerImpl.parse() correctly derives its start date from the last successful job:
long lastSuccess = _usageJobDao.getLastJobSuccessDateMillis();
if (lastSuccess != 0) {
startDateMillis = lastSuccess + 1;
}
and then overwrites it with the create date of the oldest unprocessed event, moving it only ever earlier:
// make sure start date is before all of our un-processed events
Date oldestEventDate = events.get(0).getCreateDate();
if (oldestEventDate.getTime() < startDateMillis) {
startDateMillis = oldestEventDate.getTime();
startDate = new Date(startDateMillis);
}
events comes from UsageEventDaoImpl.listLatestEvents() — WHERE processed = 0 AND created <= ? ORDER BY createDate ASC, with no limit and no floor.
Consequence: any event that can never be successfully processed pins the aggregation start date permanently. Each subsequent run re-aggregates from that date to the present, growing by one aggregation range per range elapsed. The job continues to report success = 1 throughout, so nothing surfaces as an error.
Measured on an affected 4.22.1.0 deployment:
- 1,678 consecutive successful jobs whose
start_millisnever advanced past a single event 71 days earlier exec_timeof 2,555,055 ms (42.6 min) per hourly run, growing by 24 aggregation periods per daycloud_usageat 54.4M rows / 14 GB, of which roughly 150k rows were genuine; duplicate counts were an exact multiple of the number of replays
Not specific to one event type: #13112 was reported on 4.21.0.0 against usage_type = 13, predating the volume-specific trigger in #13399. The rewind is the common mechanism; the triggering event type varies.
versions
Observed on CloudStack 4.22.1.0 (EL9 packages), MySQL 8.x / InnoDB, with usage.stats.job.aggregation.range = 60 and usage.stats.job.exec.time = 00:15.
#13112 reports the same mechanism on 4.21.0.0, so this is not new in 4.22.
The steps to reproduce the bug
- Have any row in
cloud_usage.usage_eventthat cannot be successfully processed, #13399 gives one reliable route, but the mechanism is independent of cause. - Let the usage job run on its normal schedule.
SELECT id, start_date, end_date, exec_time FROM cloud_usage.usage_job WHERE success = 1 ORDER BY id DESC LIMIT 5;—start_datestays pinned at the oldest unprocessed event'screatedwhileend_dateadvances.TIMESTAMPDIFF(HOUR, start_date, end_date)grows without bound, as doesexec_time.- Row counts in
cloud_usage.cloud_usageper period equal the number of times that period has been re-aggregated.
What to do about it?
Bound the rewind, and/or provide a way to quarantine or age out events that repeatedly fail to process, so one unprocessable row cannot halt aggregation indefinitely.
This is a design decision rather than an obvious patch, which is why it's raised separately from #13399.
- Dominant language
- Java
- Stars
- 3.1k
- Forks
- 1.4k
- Avg merge
- 6d 20h
- Merged PRs (30d)
- 27
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from apache/cloudstack
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 90/100
apache/cloudstack#14222 ·
-
bug component:kubernetes
Difficulty 1/5 Under an hour Newbie friendliness 88/100
apache/cloudstack#14180 ·
-
bug component:projects component:UI
Difficulty 1/5 Under an hour Newbie friendliness 88/100
apache/cloudstack#14070 · 5 comments ·
-
component:backup
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
apache/cloudstack#14013 ·
-
KVM agent fails to connect to Ceph RBD storage pool after upgrading Ceph client to Tentacle 20.2.4 Openbug component:ceph
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
apache/cloudstack#13989 · 3 comments ·
All issues in apache/cloudstack
Similar issues
-
area/plugin
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
kestra-io/plugin-kestra#190 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
google-ai-edge/LiteRT-LM#3739 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
integra-team-red/meet-map#249 ·
-
[Studio][Bug] Cancelled create-user dialog keeps the password and admin switch for the next attempt Open
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
apache/rocketmq-dashboard#5064 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
wso2/dpdp-accelerator#287 ·