Duplicated timeseries in CollectorRegistry with Multiprocess Gunicorn
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 25/100
- Issue 類型
- 缺陷
- 描述清晰度
- 需要釐清
- 活躍度
- 停滯
- 技術堆疊
- docker, prometheus, python
研究方向
先從 README 中的 multiprocess metrics 範例和所參照的 gunicorn.conf.py 設定開始,然後使用所示的 Dockerfile 設定和 Gunicorn 指令重現該故障。完成的標準是找出 duplicated timeseries 錯誤的原因,並記錄或修正設定,讓 /metrics 繼續回傳 metrics。
由索引模型根據 Issue 內容生成。
描述
I know this is a subject that comes up somewhat frequently but for the love of me I can't figure out what I'm doing wrong.
-
I have a service in Amazon ECS thats running a single task with multiple workers (actually the problem happens in my other service that just has one worker also).
-
I've created the directory and set the
PROMETHEUS_MULTIPROC_DIRin the Dockerfile:
RUN mkdir -p /tmp/prom-metrics
ENV PROMETHEUS_MULTIPROC_DIR /tmp/prom-metrics
- I'm using the sample code in the README to create the registry in the
/metricsrequest and return it:
registry = CollectorRegistry()
if getenv('PROMETHEUS_MULTIPROC_DIR'):
multiprocess.MultiProcessCollector(registry)
data = generate_latest(registry)
status = '200 OK'
response_headers = [
('Content-type', CONTENT_TYPE_LATEST),
('Content-Length', str(len(data))),
]
return Response(data, status, response_headers)
- I've created the
gunicorn.conf.pyfile with the sample from the README and passed it into my gunicorn startup script via-c:
from prometheus_client import multiprocess
def child_exit(server, worker):
multiprocess.mark_process_dead(worker.pid)
In my two services, gunicorn starts them as follows:
# app 1 with workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 --workers 3 app:app
# app 2 without workers
gunicorn -c /app/utils/gunicorn.conf.py -b :5000 -t 3600 --keep-alive 60 --threads 8 app:app
The service boots successfully and accepts some metrics which are definitely collected in multiprocess mode, seeing as the HELP line simply displays Multiprocess metric.
This works for a few calls but eventually I get the dreaded Duplicated timeseries in CollectorRegistry error and no additional metrics are populated.
What might I be doing wrong?
- 主要語言
- Python
- 星號
- 4.4k
- 分支
- 876
- 平均合併
- 8 天 4 小時
- 30 天內合併 PR
- 1
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
prometheus/client_python 的其他 Issue
-
bug
難度 2/5 1-3 小時 新手友好度 70/100
prometheus/client_python#1177 · 1 則留言 ·
-
難度 4/5 3-5 天 新手友好度 45/100
prometheus/client_python#1210 ·
-
難度 2/5 1-3 小時 新手友好度 58/100
prometheus/client_python#1199 · 1 個 reaction ·
-
難度 5/5 一週以上 新手友好度 35/100
prometheus/client_python#1176 ·
-
難度 1/5 1-3 小時 新手友好度 52/100
prometheus/client_python#1126 · 2 則留言 ·
查看 prometheus/client_python 的全部 Issue
相似的 Issue
-
agent-ready documentation needs-triage
難度 1/5 1-3 小時 新手友好度 88/100
-
documentation
難度 1/5 1 小時以內 新手友好度 91/100
-
workflow-status page template still says reusable workflows are "triggered only by workflow_call:" 未關閉
難度 1/5 1 小時以內 新手友好度 92/100
-
instance instance add
難度 1/5 1 小時以內 新手友好度 72/100
searxng/searx-instances#939 · 1 則留言 ·
-
area-deployment area-integrations triage:bot-seen
難度 2/5 半天 新手友好度 86/100