Ansible Operator: owner references of created jobs not matching the actual owner CR frequently
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 35/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Cần làm rõ
- Mức độ hoạt động
- Đình trệ
- Công nghệ
- ansible
- Lĩnh vực
- infrastructure
Hướng nghiên cứu
Bắt đầu với cấu hình watches.yml và /opt/ansible/playbook.yml được tham chiếu, sau đó tái hiện việc xử lý nhiều Optimization CRs trong vòng vài giây với watchDependentResources được bật. Theo dõi cách mỗi Job được tạo nhận ownerReferences của nó và xác minh rằng mọi Job đều nhất quán tham chiếu đến CR đã kích hoạt lần chạy playbook của nó.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Bug Report
What did you do?
Created Ansible Operator to create jobs based on custom CRs. When a CR appears, the playbook triggers the creation of a job (via ansible k8s module) and deletes the CR after job completion is detected.
What did you expect to see?
When a batch of CRs is detected and processed, it is (naturally) expected that a job, which is created during the playbook run for a particular CR, has an owner reference to exactly the CR for which the playbook was started.
What did you see instead? Under which circumstances?
The actual owner references are wrongly assigned about 50% of the time when multiple CRs are created in a short time frame. They seem to be randomly pointing to one of the CRs created in bulk. This appears to be a severe issue (unless I am doing something totally wrong?). I did not find any issue related to this when searching.
Example of two jobs that were created from two watched CRs. The CRs were created within 3 seconds of each other. Note that the job name is set equal to the CR name for which it was created. The name is a UID. While the actual job data is correctly derived from the CR, it is weirdly apparent that owner references are actually switched here, the first job has the second CR as owner, while the second job has the first CR as owner:
Job 1:
apiVersion: batch/v1
kind: Job
metadata:
name: 439964c7-6941-43b2-b2ff-3a8676eca868-20250502173039
namespace: tenant-d4af8bbf-dfa2-41d2-a91a-1f4092f0222a
uid: 6242278d-9c18-4a1a-8655-4e3a360c9904
resourceVersion: '24780744'
generation: 1
creationTimestamp: '2025-02-06T16:57:50Z'
labels:
optimization_id: 439964c7-6941-43b2-b2ff-3a8676eca868
optimization_instance_id: 439964c7-6941-43b2-b2ff-3a8676eca868-20250502173039
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: 'false'
ownerReferences:
- apiVersion: abc.xyz.com/v1alpha1
kind: Optimization
name: bf1178cf-5788-43a4-98fe-c422705a037c-20250502173042
uid: 3cfd7c28-098d-4553-83e4-140b37f73977
...
Job 2:
apiVersion: batch/v1
kind: Job
metadata:
name: bf1178cf-5788-43a4-98fe-c422705a037c-20250502173042
namespace: tenant-d4af8bbf-dfa2-41d2-a91a-1f4092f0222a
uid: a1a93678-9f55-4032-8dbd-f5fa5bdc0be0
resourceVersion: '24780937'
generation: 1
creationTimestamp: '2025-02-06T16:57:44Z'
labels:
optimization_id: bf1178cf-5788-43a4-98fe-c422705a037c
optimization_instance_id: bf1178cf-5788-43a4-98fe-c422705a037c-20250502173042
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: 'false'
ownerReferences:
- apiVersion: abc.xyz.com/v1alpha1
kind: Optimization
name: 439964c7-6941-43b2-b2ff-3a8676eca868-20250502173039
uid: 86e8ca4f-c286-4f81-a0e3-5be500ea9deb
...
I observed the assignment of job ownership to CRs to be anything of the following:
- the assignments may be switched around like above
- all three jobs may be marked as owned by one CR
- the ownership may be correctly assigned
From several tests, job ownership assignment to CRs seems to be rather undeterministic behavior for CRs created in a short time frame (< 5 seconds).
Environment
Kubernetes cluster type:
DigitalOcean DOKS with k8s 1.31
$ operator-sdk version
quay.io/operator-framework/ansible-operator:v1.37.1
$ kubectl version
1.31
Possible Solution
It seems as if there is no clear back reference from playbook being executed to the CR that triggered it? Seems like when the job is being created it gets assigned one owner which may be currently "active" in another thread or similar?
A possible workaround may be to manually assign the ownerships in the playbook, assuming that when watchDependentResources: false there are no ownerships automatically injected?
Additional context
watches.yml:
---
- version: v1alpha1
group: abc.xyz.com
kind: Optimization
playbook: /opt/ansible/playbook.yml
reconcilePeriod: "10s"
watchDependentResources: true
manageStatus: true
Thanks for looking into this, I feel this is a quite critical bug?
- Ngôn ngữ chính
- Go
- Star
- 12
- Fork
- 33
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
- Không có Dockerfile hay tệp Docker Compose
- Có mẫu pull request
- Không có hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của operator-framework/ansible-operator-plugins
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
-
Upgrade to go 1.26.6 to fix CVEsĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 65/100
operator-framework/ansible-operator-plugins#237 · 3 bình luận ·
-
Use runtime.GOMAXPROCS(0) instead of runtime.NumCPU() to set --max-concurrent-reconciles defaultĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 74/100
-
language/ansible
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 78/100
-
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
operator-framework/ansible-operator-plugins#240 · 1 bình luận ·
Tất cả issue của operator-framework/ansible-operator-plugins
Issue tương tự
-
agentic-workflows
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
-
priority/4/normal status/needs-triage type/bug/unconfirmed
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
authelia/authelia#13292 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
blinklabs-io/actions#138 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
[UI] AlbumDetails collapses multi-genre list to single primary genre on viewports < lg breakpointĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày