MultiProcessCollector._accumulate_metrics always drops timestamps
Ninguém assumiu esta issue ainda.
Avaliação
- Dificuldade
- 3/5
- Tempo estimado
- 1-2 dias
- Facilidade para iniciantes
- 45/100
- Tipo de issue
- Bug
- Clareza
- Claramente especificada
- Status de atividade
- Estagnada
- Stack de tecnologia
- python
- Domínio
- observability-sre
Direção de pesquisa
Comece em prometheus_client/multiprocess.py, acompanhando MultiProcessCollector.merge por _read_metrics e _accumulate_metrics. Verifique como o timestamp do sample selecionado é rastreado para cada modo multiprocesso e como o sample recriado é escrito. Está concluído quando a compactação preserva o timestamp necessário para a seleção e a exposição posteriores, sem alterar os modos que não usam timestamps.
Escrita pelo modelo de indexação a partir do texto da issue.
Descrição
Summary
MultiProcessCollector._accumulate_metrics drops timestamps for all exported metrics, regardless of whether the selected sample or multiproc_mode used or needs them. This causes metrics file compaction routines (to cleanup old DB files from long-running web worker processes, for example) to drop samples that they otherwise wouldn't.
Details
MultiProcessCollector.merge calls _read_metrics to read all metrics (including dupes) from the DB files, followed by _accumulate_metrics to dedupe these based on multiproc_mode.
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L35-L44
In the docstring, it explicitly calls out the use case this affects (compaction by writing back to the mmap files):
But if writing the merged data back to mmap files, use
accumulate=False to avoid compound accumulation.
_accumulate_metrics considers the timestamp when deciding which sample to keep:
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L97-L116
However, it does not include this timestamp in the recreated sample it returns:
https://github.com/prometheus/client_python/blob/master/prometheus_client/multiprocess.py#L153
Impact
We have an internal compaction tool that effectively just periodically calls the MultiProcessCollector.merge method periodically, wrapped with an flock, and given the prevalence of open issues such as #568 I suspect others may too. Without this, long-running gunicorn processes with worker rotation settings will accumulate large numbers of stale files that slow scrape times significantly. This compaction ignores live pids, only compacting DBs from dead ones.
This issue results in this compaction potentially dropping samples that would otherwise have been the newest timestamp, simply because they've been compacted from a dead pid. Consider the following:
t1. process 1 and 2 spawn and define gauge my_gauge with multiproc_mode="mostrecent"
t2. process 1 samples my_gauge with timestamp=time.time()
t3. process 2 samples my_gauge with timestamp=time.time()
t4. process 2 dies
t5. compaction calls MultiProcessCollector.merge as part of stale DB compaction.
t6. MultiProcessCollector._accumulate_metrics returns the process 2 sample without a timestamp, which is then written to the compacted DB
t7. both future compaction and regular metrics exposition (via a scrape or otherwise) now drop the process 2 sample despite it being newest
Proposed Solution
Since the collector already tracks the sample timestamp in a defaultdict(float) and collection considers 0.0 and None equivalent, I think it should be as simple as changing the exported sample to this (or the equivalent):
timestamped_samples = []
for (name, label), value in samples.items():
without_pid_key = (name, tuple(l for l in labels if l[0] != 'pid'))
timestamped_samples.append(
prometheus_client.samples.Sample(
name_, dict(labels), value, sample_timestamps[without_pid_key]
)
)
metric.samples = timestamped_samples
- Linguagem predominante
- Python
- Estrelas
- 4.4k
- Forks
- 876
- Merge médio
- 8d 4h
- PRs com merge (30d)
- 1
Guia de contribuição
Primeiros passos
- Leia a issue inteira e depois o guia de contribuição do projeto.
- Comente na issue dizendo que vai assumir — evita que duas pessoas façam o mesmo trabalho.
- Faça um fork do repositório e trabalhe em uma branch.
- Abra um pull request que referencie o número da issue.
Mais de prometheus/client_python
-
bug
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
prometheus/client_python#1177 · 1 comentário ·
-
Dificuldade 4/5 3-5 dias Facilidade para iniciantes 45/100
prometheus/client_python#1210 ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 58/100
prometheus/client_python#1199 · 1 reação ·
-
Dificuldade 5/5 Mais de uma semana Facilidade para iniciantes 35/100
prometheus/client_python#1176 ·
-
Dificuldade 1/5 1-3 horas Facilidade para iniciantes 52/100
prometheus/client_python#1126 · 2 comentários ·
Todas as issues de prometheus/client_python
Issues semelhantes
-
bug confirmed issue
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
open-webui/open-webui#30750 · 1 comentário ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
-
enhancement
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 75/100
OpenwaterHealth/openmotion-bloodflow-app#604 · 1 comentário ·
-
Dificuldade 2/5 1-3 horas Facilidade para iniciantes 70/100
-
good first issue
Dificuldade 1/5 Menos de uma hora Facilidade para iniciantes 90/100