VM snapshot usage records are attributed to the wrong snapshot; older snapshots stop accruing usage entirely

未關閉 適合新手
#13,921 0 則留言 0 個 reaction 已指派 0 人 在 GitHub 檢視

還沒有人認領這個 Issue。

評估

難度
2/5
預估耗時
1-3 小時
新手友好度
84/100
Issue 類型
缺陷
描述清晰度
描述清楚
活躍度
活躍
技術堆疊
java
領域
backend

研究方向

先定位 VMSnapshotUsageParser,並檢查第 68 行附近的鍵建構,以及第 88 行附近的 createUsageRecord。重現兩個 snapshot 的情況,並查詢 issue 中描述的使用量記錄。當每個 snapshot 的時間間隔都帶有自己的 ID,且較舊的 snapshot 持續累積使用量而不會被逐出時,即表示完成。

由索引模型根據 Issue 內容生成。

描述

bug
problem

When a VM has more than one VM snapshot, VMSnapshotUsageParser mis-attributes usage and silently drops usage for all but the newest snapshot. Two distinct defects combine:

  1. The usage interval is labelled with the wrong snapshot ID. Duration, size and disk offering are taken from the previous event, but the snapshot ID is taken from the current one (line 88):

createUsageRecord(UsageTypes.VM_SNAPSHOT, duration, previousCreated, createDate, account, volId, zoneId,
previousEvent.getDiskOfferingId(), vmId,
previousEvent.getSize(), // previous snapshot
usageRec.getVmSnapshotId()); // current snapshot ← wrong label

  1. Concurrent snapshots of the same volume overwrite each other. unprocessedUsage is keyed on VM + volume only (line 68), so a newer event evicts the older one and the older snapshot never reaches the closing loop:

String key = vmId + ":" + volId;
...
unprocessedUsage.put(key, usageRec); // line 96 — evicts the previous snapshot

Net effect: the first snapshot's usage is credited to the second snapshot, and the first snapshot stops accruing usage even though it still exists and is Ready.

Observed in a lab with two snapshots on one VM, neither deleted (both Ready, removed IS NULL):

DB entries from repro environment confirming the above

mysql> select * from cloud.vm_snapshots where removed is null;
+----+--------------------------------------+------------------------------+--------------+-------------+-------+------------+-----------+---------------------+------------------+-------+--------+---------+--------------+---------------------+---------------------+---------+
| id | uuid | name | display_name | description | vm_id | account_id | domain_id | service_offering_id | vm_snapshot_type | state | parent | current | update_count | updated | created | removed |
+----+--------------------------------------+------------------------------+--------------+-------------+-------+------------+-----------+---------------------+------------------+-------+--------+---------+--------------+---------------------+---------------------+---------+
| 4 | 09bd5ac1-6a90-487f-afed-0798d5440107 | i-2-193-VM_VS_20260818150857 | Testsnap1 | Testsnap1 | 193 | 2 | 1 | 1 | Disk | Ready | NULL | 0 | 2 | 2026-08-18 15:08:59 | 2026-08-18 15:08:57 | NULL |
| 5 | 71bcbec9-0760-4b34-8551-9fcc5a891027 | i-2-193-VM_VS_20260818154342 | Testsnap2 | Testsnap2 | 193 | 2 | 1 | 1 | Disk | Ready | 4 | 1 | 2 | 2026-08-18 15:43:45 | 2026-08-18 15:43:42 | NULL |
+----+--------------------------------------+------------------------------+--------------+-------------+-------+------------+-----------+---------------------+------------------+-------+--------+---------+--------------+---------------------+---------------------+---------+
2 rows in set (0.00 sec)

mysql> select * from cloud_usage where usage_type=25 and usage_id=5;
+---------+---------+------------+-----------+-------------------------------------------------------------------------------+---------------+------------+--------------------+----------------+---------+-------------+-------------+----------+------+------------+------------+---------------------+---------------------+--------------+-----------+-----------+--------+------------------+-----------+-------+
| id | zone_id | account_id | domain_id | description | usage_display | usage_type | raw_usage | vm_instance_id | vm_name | offering_id | template_id | usage_id | type | size | network_id | start_date | end_date | virtual_size | cpu_speed | cpu_cores | memory | quota_calculated | is_hidden | state |
+---------+---------+------------+-----------+-------------------------------------------------------------------------------+---------------+------------+--------------------+----------------+---------+-------------+-------------+----------+------+------------+------------+---------------------+---------------------+--------------+-----------+-----------+--------+------------------+-----------+-------+
| 1803640 | 1 | 2 | 1 | VMSnapshot Id: 5 Usage: VM Id: 193 Volume Id: 223 Size: (8.00 GB) 8589934592 | 0.579445 Hrs | 25 | 0.5794447064399719 | 193 | NULL | NULL | NULL | 5 | NULL | 8589934592 | NULL | 2026-08-18 15:08:59 | 2026-08-18 15:43:45 | NULL | NULL | NULL | NULL | 0 | 0 | NULL |
| 1803641 | 1 | 2 | 1 | VMSnapshot Id: 5 Usage: VM Id: 193 Volume Id: 223 Size: (8.00 GB) 8589934592 | 8.270833 Hrs | 25 | 8.270833015441895 | 193 | NULL | NULL | NULL | 5 | NULL | 8589934592 | NULL | 2026-08-18 15:43:45 | 2026-08-18 23:59:59 | NULL | NULL | NULL | NULL | 0 | 0 | NULL |
+---------+---------+------------+-----------+-------------------------------------------------------------------------------+---------------+------------+--------------------+----------------+---------+-------------+-------------+----------+------+------------+------------+---------------------+---------------------+--------------+-----------+-----------+--------+------------------+-----------+-------+
2 rows in set (2.45 sec)

Record 1803640 spans Testsnap1's creation to Testsnap2's creation, Testsnap1's usage interval , but carries usage_id = 5. Testsnap1 (id 4) has no usage records at all.

Below is the list of Usage expected vs Actual

Testsnap1 (15:08:59 (created time) --> 23:59:59) - Expected Usage: 8.85 hrs, Actual Usage: 0
Testsnap2 (15:43:45 (created time) --> 23:49:59) - Expected Usage: 8.27 Hrs, Actual Usage: 8.85 Hrs

Note: The visible symptom depends on aggregation timing. If both snapshots are created within one aggregation window, the first snapshot gets no records at all. If the second is created in a later window, the first gets records only up to the last completed bucket, and the interval between that boundary and the second snapshot's creation is credited to the second snapshot.

Snapshot chaining makes this the normal case rather than an edge case, note parent = 4 on Testsnap2.

versions

Cloudstack Version: 4.22/4.22.1
Hypervisor: KVM
Primary Storage: NFS
Database: MySQL 8.4

Usage Settings:

+-----------------------------------+-------+
| name | value |
+-----------------------------------+-------+
| usage.aggregation.timezone | GMT |
| usage.execution.timezone | GMT |
| usage.sanity.check.interval | NULL |
| usage.snapshot.virtualsize.select | false |
| usage.stats.job.aggregation.range | 1440 |
| usage.stats.job.exec.time | 00:15 |
+-----------------------------------+-------+

The steps to reproduce the bug
  1. Deploy an instance on any hypervisor supporting VM snapshots.
  2. Take a VM snapshot (Testsnap1).
  3. Take a second VM snapshot on the same instance (Testsnap2). Do not delete either.
  4. Wait for the usage aggregation job to run.
  5. Query the usage records:

SELECT usage_id, description, usage_display, start_date, end_date
FROM cloud_usage.cloud_usage
WHERE usage_type = 25 ORDER BY start_date;

SELECT id, vm_id, volume_id, vm_snapshot_id, size, created, processed
FROM cloud_usage.usage_vmsnapshot WHERE vm_id = ;

What to do about it?

Two changes in VMSnapshotUsageParser:

  1. Line 88 — label the record with the snapshot the interval belongs to:

previousEvent.getSize(), previousEvent.getVmSnapshotId());

  1. Line 68 — include the snapshot ID in the key so concurrent snapshots of one volume don't evict each other:

String key = vmId + ":" + volId + ":" + usageRec.getVmSnapshotId();

Note the snapshot ID is already populated correctly on insert. createUsageVMSnapshot() reads it from usage_event_details and calls vsVO.setVmSnapshotId(...), so the data needed for both fixes is present; it's only ignored by the parser.

主要語言
Java
星號
3.1k
分支
1.4k
平均合併
7 天 5 小時
30 天內合併 PR
28

貢獻指南

開啟貢獻指南

從這裡開始

  1. 先讀完整個 Issue,再讀專案的貢獻指南。
  2. 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
  3. Fork 儲存庫,在一個分支上完成修改。
  4. 送出 Pull Request,並在描述裡引用這個 Issue 編號。

apache/cloudstack 的其他 Issue

查看 apache/cloudstack 的全部 Issue

相似的 Issue

更多 Java Issue

把新 issue 寄到你的電子郵件信箱

精選適合新手參與的 GitHub issue 摘要。