Implement trace validation to distinguish environmental artifacts from induced faults
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 5/5
- Thời gian dự kiến
- Hơn một tuần
- Mức phù hợp với người mới
- 25/100
- Loại issue
- Tính năng
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- kubernetes, python
- Lĩnh vực
- ai, testing-qa
Hướng nghiên cứu
Issue không cung cấp tệp mã nguồn, bài kiểm thử hoặc điểm vào để bắt đầu; hãy bắt đầu bằng cách xem lại kubectl trace được cung cấp và quy trình đánh giá độ chính xác của việc phát hiện. Công việc được xem là hoàn tất khi phân biệt được một docker namespace trống do môi trường kiểm thử gây ra với một lỗi Kubernetes có bằng chứng, tránh tạo các issue report dương tính giả.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Description
In the test environment, running kubectl get pods -n docker or kubectl get all -n docker sometimes returns:
No resources found in docker namespace.
The detection agent interprets this as a K8s issue and marks it as such.
However, in the production env this namespace always contains resources, so this case should not be treated as an issue. This leads to false positives and inflates detection accuracy metrics.
Trace
"agent": "openrouter",
"session_id": "292a425f-0b28-45f5-96bc-e4f02fa42623",
"problem_id": "flower_model_misconfig-detection",
"start_time": 1756135764.7457962,
"end_time": 1756135765.7294734,
"trace": [
{
"role": "assistant",
"content": "```\nexec_shell(\"kubectl get pods -n docker\")\n```"
},
{
"role": "env",
"content": "No resources found in docker namespace.\n"
},
{
"role": "assistant",
"content": "```\nexec_shell(\"kubectl get all -n docker\")\n```"
},
{
"role": "env",
"content": "No resources found in docker namespace.\n"
},
{
"role": "assistant",
"content": "```\nsubmit(\"Yes\")\n```"
},
{
"role": "env",
"content": "1"
}
],
"results": {
"Detection Accuracy": "Correct",
"TTD": 0.9836771488189697,
"steps": 3,
"in_tokens": 15,
"out_tokens": 33
}
Open Question
Would it make sense to introduce a supervisor agent to validate conversations before marking them as issues?
The idea is that if no actual problem is evidenced, the supervisor would return NoIssue or Inconclusive instead of allowing a false positive.
One option could be adding a lightweight GPT-5–based supervisor to evaluate detection tasks.
Any suggestions or alternative approaches are welcome. 🤔
- Ngôn ngữ chính
- Python
- Star
- 989
- Fork
- 176
- Merge trung bình
- 7 ngày 15 giờ
- Pull request đã merge (30 ngày)
- 3
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/AIOpsLab
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
-
Regarding to the case #162 Đang mở
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 35/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 25/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 35/100
Tất cả issue của microsoft/AIOpsLab
Issue tương tự
-
[Bug] reef-hermes tells me to resume with hermes --resume, which does not work from my shell Đang mởarea: harness bug status: needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
Human-Agent-Society/reef#625 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 70/100
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 80/100
learningequality/kolibri#15351 · 2 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
-
Name consistency Đang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
eellak/triplestore#65 · 1 bình luận ·