MultiProcessCollector._accumulate_metrics always drops timestamps
還沒有人認領這個 Issue。
評估
- 難度
- 3/5
- 預估耗時
- 1-2 天
- 新手友好度
- 45/100
- Issue 類型
- 缺陷
- 描述清晰度
- 描述清楚
- 活躍度
- 停滯
- 技術堆疊
- python
研究方向
從 prometheus_client/multiprocess.py 開始,沿著 _read_metrics 和 _accumulate_metrics 追蹤 MultiProcessCollector.merge。確認在每種多程序模式下如何追蹤所選 sample 的時間戳,以及如何寫入重新建立的 sample。完成的標準是壓縮保留後續選擇和 exposition 所需的時間戳,同時不改變不使用時間戳的模式。
由索引模型根據 Issue 內容生成。
描述
Summary
MultiProcessCollector._accumulate_metrics drops timestamps for all exported metrics, regardless of whether the selected sample or multiproc_mode used or needs them. This causes metrics file compaction routines (to cleanup old DB files from long-running web worker processes, for example) to drop samples that they otherwise wouldn't.
Details
MultiProcessCollector.merge calls _read_metrics to read all metrics (including dupes) from the DB files, followed by _accumulate_metrics to dedupe these based on multiproc_mode.
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L35-L44
In the docstring, it explicitly calls out the use case this affects (compaction by writing back to the mmap files):
But if writing the merged data back to mmap files, use
accumulate=False to avoid compound accumulation.
_accumulate_metrics considers the timestamp when deciding which sample to keep:
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L97-L116
However, it does not include this timestamp in the recreated sample it returns:
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L153
Impact
We have an internal compaction tool that effectively just periodically calls the MultiProcessCollector.merge method periodically, wrapped with an flock, and given the prevalence of open issues such as #568 I suspect others may too. Without this, long-running gunicorn processes with worker rotation settings will accumulate large numbers of stale files that slow scrape times significantly. This compaction ignores live pids, only compacting DBs from dead ones.
This issue results in this compaction potentially dropping samples that would otherwise have been the newest timestamp, simply because they've been compacted from a dead pid. Consider the following:
t1. process 1 and 2 spawn and define gauge my_gauge with multiproc_mode="mostrecent"
t2. process 1 samples my_gauge with timestamp=time.time()
t3. process 2 samples my_gauge with timestamp=time.time()
t4. process 2 dies
t5. compaction calls MultiProcessCollector.merge as part of stale DB compaction.
t6. MultiProcessCollector._accumulate_metrics returns the process 2 sample without a timestamp, which is then written to the compacted DB
t7. both future compaction and regular metrics exposition (via a scrape or otherwise) now drop the process 2 sample despite it being newest
Proposed Solution
Since the collector already tracks the sample timestamp in a defaultdict(float) and collection considers 0.0 and None equivalent, I think it should be as simple as changing the exported sample to this (or the equivalent):
timestamped_samples = []
for (name, label), value in samples.items():
without_pid_key = (name, tuple(l for l in labels if l[0] != 'pid'))
timestamped_samples.append(
prometheus_client.samples.Sample(
name_, dict(labels), value, sample_timestamps[without_pid_key]
)
)
metric.samples = timestamped_samples
- 主要語言
- Python
- 星號
- 4.4k
- 分支
- 879
- 平均合併
- 8 天 4 小時
- 30 天內合併 PR
- 1
環境準備
- 沒有 Dockerfile 或 Docker Compose 檔案
- 沒有 Pull Request 範本
- 閱讀貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
prometheus/client_python 的其他 Issue
-
bug
難度 2/5 1-3 小時 新手友好度 70/100
prometheus/client_python#1177 · 1 則留言 ·
-
難度 4/5 3-5 天 新手友好度 45/100
prometheus/client_python#1210 ·
-
難度 2/5 1-3 小時 新手友好度 58/100
prometheus/client_python#1199 · 1 個 reaction ·
-
難度 5/5 一週以上 新手友好度 35/100
prometheus/client_python#1176 ·
-
難度 1/5 1-3 小時 新手友好度 52/100
prometheus/client_python#1126 · 2 則留言 ·
查看 prometheus/client_python 的全部 Issue
相似的 Issue
-
難度 2/5 1-3 小時 新手友好度 85/100
mozilla/bedrock#17413 · 1 個 reaction ·
維護者通常 2 天內回覆
-
instance instance add
難度 2/5 1-3 小時 新手友好度 68/100
searxng/searx-instances#943 · 1 則留言 ·
-
難度 2/5 1-3 小時 新手友好度 68/100
維護者通常 1 天內回覆
-
bug tools
難度 2/5 1-3 小時 新手友好度 88/100
維護者通常 1 天內回覆
-
bug
難度 2/5 1-3 小時 新手友好度 86/100
lance-format/lance#9655 ·
維護者通常 2 天內回覆