bug: dangling schema in an APIExport is neither removed nor reported
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 55/100
調査の方向性
Start in internal/controller/apiexport/reconciler.go at mergeResourceSchemas(), then read internal/controller/apiexport/controller.go where readySchemaNames is built. Reproduce the dangling reference with the APIExport and APIResourceSchema commands in the issue, and inspect existing event and condition handling. Done means a missing schema is reported on the owning APIExport with enough detail to identify the entry, without breaking managed-resource reconciliation.
索引モデルが issue の本文から書いたものです。
説明
What happens
If an APIExport contains a schema name (spec.latestResourceSchemas, spec.resources[].schema
in v1alpha2) that does not resolve to an existing APIResourceSchema, nothing in the export
owner's view tells them about it:
- the Sync Agent leaves the entry in place forever (
mergeResourceSchemas()keeps every
existing entry that belongs to a group/resource noPublishedResourcemanages — deliberately,
see below), and it does not validate that the referenced ARS exists; - the
APIExport's own status does not mention it either — it only carriesIdentityValid; - the only signal is on the consumer side: every
APIBindingto that export goes
APIExportValid=Falsewithreason: InternalError/
message: Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "<name>" not found,
the API group stops being served, and the export owner is being paged with a message that
does not name the entry they have to remove.
So the broken entry can only be found by reading the export by hand, and the export owner —
who is usually also the Sync Agent — gets no event and no condition about it.
Why it happens
internal/controller/apiexport/reconciler.go (mergeResourceSchemas()):
// Now we include all other existing ARS that use unknown resources;
// this both allows an APIExport to contain "unmanaged" ARS, and also
// will purposefully leave behind ARS for deleted PublishedResources,
// allowing cleanup to take place outside of the agent's control.
"Outside of the agent's control" means an admin has to notice — but the agent owns the export,
so that admin has no place to look. Two realistic ways to get there:
- A
PublishedResourceis deleted, and its ARS is removed later (by hand, or by re-creating
the installation on a new cluster) — the name stays in the export. - An admin-managed export, which the docs explicitly allow
("The APIExport above might look like it is defining all resources for the API test.example.com API
group, but in reality it might contain resource schemas like v1.crontabs.initech.com"),
references a schema that was never created (typo, ordering, a schema that failed to create).
Related: readySchemaNames in internal/controller/apiexport/controller.go is built from
PublishedResource.status.resourceSchemaName, so while the ARS controller cannot project a
resource (CRD missing on the service cluster, ARS creation failing), the agent leaves the whole
export untouched, broken entries included.
Reproduction
api-syncagent v0.7.0 and main (same code path). kcp v0.32.x, agent pointed at a workspace with
an APIExport/repro-export and an APIExportEndpointSlice of the same name, one
PublishedResource for widgets.repro.pax.dev on the service cluster.
-
Publish the resource, the agent creates its schema and adds it to the export:
$ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
Delete the
PublishedResource(the resource is no longer managed by the agent) and then
delete the ARS itself — the export now references something that does not exist:$ kubectl get apiresourceschemas No resources found $ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
The agent keeps running and reconciling; it never removes the entry and never reports it —
no event on the export, no log line, and the export's status stays:$ kubectl get apiexport repro-export -o jsonpath='{.status.conditions}' [{"status":"True","type":"IdentityValid"}] -
The consumer only sees the failure:
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason} {.message}{"\n"}{end}' Ready=True APIExportValid=False InternalError Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "v4ed05578.widgets.repro.pax.dev" not found
Recovery (verified)
Removing the entry from the export restores it — APIExportValid=True, Ready=True, and the
agent does not put the name back as long as no PublishedResource refers to that group/resource:
$ kubectl patch apiexport repro-export --type=merge -p '{"spec":{"resources":[]}}'
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'
Ready=True
APIExportValid=True
Note that this does not stick for a resource the agent still manages: the next reconcile of
that PublishedResource adds its own schema name back (and removes the entry an admin had put
there — Warning RemovingResourceSchemas), so there the underlying reason for the missing ARS
has to be fixed instead.
Suggestion
- Have the
apiexportcontroller (it already lists the ARS/PublishedResources of the export)
emit aWarningevent and/or a status condition likeSchemaMissingon theAPIExportwhen
an entry inspec.latestResourceSchemas/spec.resources[]has no matching ARS. That would
put the signal where the owner of the export actually looks, instead of only on the consumers. - Alternatively/additionally: document the cleanup in a troubleshooting page — I will send a
docs PR for that in a moment, since the procedure above is not written down anywhere today.
- 主要言語
- Go
- スター
- 25
- フォーク
- 29
- 平均マージ
- 5日 18時間
- マージ済み PR(30日)
- 1
環境構築
- Dockerfile または Docker Compose ファイルあり
- プルリクエストのテンプレートあり
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
kcp-dev/api-syncagent のほかの issue
-
kind/bug
難易度 2/5 1〜3時間 初心者へのやさしさ 74/100
kcp-dev/api-syncagent#163 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
kcp-dev/api-syncagent#161 ·
-
kind/feature
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
kcp-dev/api-syncagent#186 · リアクション 1 件 ·
-
bug: PublishedResource updates leave previous sync controllers active再び着手できるかも @adoi が 71 日前に担当しましたが、オープン中のプルリクエストはありません。 オープンkind/bug
kcp-dev/api-syncagent#182 · コメント 1 件 · 担当者 1 名 ·
-
bug: finalizer "syncagent.kcp.io/cleanup" on related resources block deletion of primary objectオープンkind/bug
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
kcp-dev/api-syncagent#172 · コメント 1 件 · リアクション 1 件 ·
kcp-dev/api-syncagent の issue をすべて見る
似ている issue
-
automation models
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
txn2/mcp-data-platform#1984 ·
メンテナーはふだん 1 日以内に返信
-
agentic-workflows
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
メンテナーはふだん 1 日以内に返信
-
kind/docs prio/P2
難易度 1/5 1時間未満 初心者へのやさしさ 95/100
agent-substrate/substrate#1986 ·
メンテナーはふだん 1 日以内に返信