Evaluate and implement comprehensive API monitoring
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
Hướng nghiên cứu
Bắt đầu bằng việc lập danh mục các thao tác API của PolicyEngine được nêu trong issue và xem xét các kiểm tra cấp máy chủ hiện có, bao gồm độ bao phủ nhóm route từ #2155. So sánh các tùy chọn giám sát tùy chỉnh, self-hosted và hosted với các yêu cầu được liệt kê, sau đó ghi lại quyết định, độ bao phủ endpoint, việc triển khai, thông báo, hướng dẫn vận hành và các kiểm tra được giữ lại hoặc hoãn lại.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Problem
PolicyEngine needs functional monitoring across substantially more API behavior than the Better Stack free tier can cover. Existing host-level checks do not detect failures limited to particular route groups, and realistic calculation monitoring may require authenticated requests, response validation, and polling asynchronous operations.
Adding paid Better Stack coverage immediately would commit us to a service and recurring cost before we have compared it with an internally operated solution or other hosted products.
This issue supersedes #2155 and should preserve its required coverage while broadening the work to include option evaluation and platform design.
Objective
Evaluate and implement a cost-effective monitoring approach for PolicyEngine APIs. The selected approach may be:
- a small custom service operated by PolicyEngine;
- an existing self-hosted monitoring product; or
- a paid hosted monitoring service, including Better Stack, if its cost and capabilities are the best fit.
Do not select an implementation until the alternatives have been compared against explicit requirements.
Investigation
- Inventory the API endpoints and complete user operations that need monitoring, including metadata, household, policy, user-profile, household-calculation, economy, and budget-window behavior.
- Define request frequency, acceptable response time, failure conditions, environments, and required geographic coverage.
- Determine requirements for:
- public and authenticated requests;
- safe test data and credential storage;
- multi-request and asynchronous operations;
- response status, content, and schema validation;
- Slack failure and recovery notifications, including duplicate suppression;
- execution history, latency reporting, and diagnostic evidence;
- configuration stored in version control;
- maintenance ownership, security updates, reliability, and data retention.
- Compare custom, self-hosted, and hosted approaches. Document expected recurring cost, implementation effort, operational burden, service limits, and vendor dependence.
- Prototype the leading candidates where documentation alone does not establish whether they satisfy the requirements.
Implementation
After recording the decision, implement the selected approach and migrate or supplement the existing checks. Monitoring should exercise representative successful operations with safe inputs rather than only checking API root responses.
Completion criteria
- The requirements and endpoint coverage inventory are documented.
- The alternatives and their expected costs and maintenance requirements are compared in writing.
- The selected approach and reasons for selecting it are documented.
- Required public, authenticated, and asynchronous API operations can be checked where applicable.
- Failures and recoveries produce actionable Slack notifications without excessive duplicate messages.
- Monitor definitions, credentials guidance, operating instructions, and ownership are documented.
- Existing checks are retained, migrated, or deliberately removed with the reason recorded.
- The route-group coverage previously tracked in #2155 is implemented or explicitly deferred in follow-up issues.
- Ngôn ngữ chính
- Python
- Star
- 18
- Fork
- 33
- Merge trung bình
- 1 ngày 4 giờ
- Pull request đã merge (30 ngày)
- 21
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của PolicyEngine/policyengine-api
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
PolicyEngine/policyengine-api#3853 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 65/100
PolicyEngine/policyengine-api#3851 ·
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 45/100
PolicyEngine/policyengine-api#3847 ·
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 38/100
PolicyEngine/policyengine-api#3823 ·
-
PolicyEngine/policyengine-api#3818 · 1 người được giao ·
Tất cả issue của PolicyEngine/policyengine-api
Issue tương tự
-
[Bug] reef-hermes tells me to resume with hermes --resume, which does not work from my shell Đang mởarea: harness bug status: needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Human-Agent-Society/reef#625 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
learningequality/kolibri#15351 · 2 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Name consistency Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
eellak/triplestore#65 · 1 bình luận ·