bug: dangling schema in an APIExport is neither removed nor reported
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 4/5
- Geschätzter Aufwand
- 3-5 Tage
- Anfängerfreundlichkeit
- 55/100
Rechercherichtung
Start in internal/controller/apiexport/reconciler.go at mergeResourceSchemas(), then read internal/controller/apiexport/controller.go where readySchemaNames is built. Reproduce the dangling reference with the APIExport and APIResourceSchema commands in the issue, and inspect existing event and condition handling. Done means a missing schema is reported on the owning APIExport with enough detail to identify the entry, without breaking managed-resource reconciliation.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
What happens
If an APIExport contains a schema name (spec.latestResourceSchemas, spec.resources[].schema
in v1alpha2) that does not resolve to an existing APIResourceSchema, nothing in the export
owner's view tells them about it:
- the Sync Agent leaves the entry in place forever (
mergeResourceSchemas()keeps every
existing entry that belongs to a group/resource noPublishedResourcemanages — deliberately,
see below), and it does not validate that the referenced ARS exists; - the
APIExport's own status does not mention it either — it only carriesIdentityValid; - the only signal is on the consumer side: every
APIBindingto that export goes
APIExportValid=Falsewithreason: InternalError/
message: Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "<name>" not found,
the API group stops being served, and the export owner is being paged with a message that
does not name the entry they have to remove.
So the broken entry can only be found by reading the export by hand, and the export owner —
who is usually also the Sync Agent — gets no event and no condition about it.
Why it happens
internal/controller/apiexport/reconciler.go (mergeResourceSchemas()):
// Now we include all other existing ARS that use unknown resources;
// this both allows an APIExport to contain "unmanaged" ARS, and also
// will purposefully leave behind ARS for deleted PublishedResources,
// allowing cleanup to take place outside of the agent's control.
"Outside of the agent's control" means an admin has to notice — but the agent owns the export,
so that admin has no place to look. Two realistic ways to get there:
- A
PublishedResourceis deleted, and its ARS is removed later (by hand, or by re-creating
the installation on a new cluster) — the name stays in the export. - An admin-managed export, which the docs explicitly allow
("The APIExport above might look like it is defining all resources for the API test.example.com API
group, but in reality it might contain resource schemas like v1.crontabs.initech.com"),
references a schema that was never created (typo, ordering, a schema that failed to create).
Related: readySchemaNames in internal/controller/apiexport/controller.go is built from
PublishedResource.status.resourceSchemaName, so while the ARS controller cannot project a
resource (CRD missing on the service cluster, ARS creation failing), the agent leaves the whole
export untouched, broken entries included.
Reproduction
api-syncagent v0.7.0 and main (same code path). kcp v0.32.x, agent pointed at a workspace with
an APIExport/repro-export and an APIExportEndpointSlice of the same name, one
PublishedResource for widgets.repro.pax.dev on the service cluster.
-
Publish the resource, the agent creates its schema and adds it to the export:
$ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
Delete the
PublishedResource(the resource is no longer managed by the agent) and then
delete the ARS itself — the export now references something that does not exist:$ kubectl get apiresourceschemas No resources found $ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
The agent keeps running and reconciling; it never removes the entry and never reports it —
no event on the export, no log line, and the export's status stays:$ kubectl get apiexport repro-export -o jsonpath='{.status.conditions}' [{"status":"True","type":"IdentityValid"}] -
The consumer only sees the failure:
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason} {.message}{"\n"}{end}' Ready=True APIExportValid=False InternalError Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "v4ed05578.widgets.repro.pax.dev" not found
Recovery (verified)
Removing the entry from the export restores it — APIExportValid=True, Ready=True, and the
agent does not put the name back as long as no PublishedResource refers to that group/resource:
$ kubectl patch apiexport repro-export --type=merge -p '{"spec":{"resources":[]}}'
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'
Ready=True
APIExportValid=True
Note that this does not stick for a resource the agent still manages: the next reconcile of
that PublishedResource adds its own schema name back (and removes the entry an admin had put
there — Warning RemovingResourceSchemas), so there the underlying reason for the missing ARS
has to be fixed instead.
Suggestion
- Have the
apiexportcontroller (it already lists the ARS/PublishedResources of the export)
emit aWarningevent and/or a status condition likeSchemaMissingon theAPIExportwhen
an entry inspec.latestResourceSchemas/spec.resources[]has no matching ARS. That would
put the signal where the owner of the export actually looks, instead of only on the consumers. - Alternatively/additionally: document the cleanup in a troubleshooting page — I will send a
docs PR for that in a moment, since the procedure above is not written down anywhere today.
- Vorherrschende Sprache
- Go
- Sterne
- 25
- Forks
- 29
- Ø Merge
- 5 T. 18 Std.
- Gemergte PRs (30 T.)
- 1
Entwicklungsumgebung
- Enthält ein Dockerfile oder eine Docker-Compose-Datei
- Hat eine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus kcp-dev/api-syncagent
-
kind/bug
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 74/100
kcp-dev/api-syncagent#163 ·
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 78/100
kcp-dev/api-syncagent#161 ·
-
kind/feature
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
kcp-dev/api-syncagent#186 · 1 Reaktion ·
-
bug: PublishedResource updates leave previous sync controllers activeEvtl. wieder frei @adoi hat das vor 72 Tagen übernommen, und es ist kein Pull Request offen. Offenkind/bug
kcp-dev/api-syncagent#182 · 1 Kommentar · 1 zugewiesene Person ·
-
bug: finalizer "syncagent.kcp.io/cleanup" on related resources block deletion of primary objectOffenkind/bug
Schwierigkeit 4/5 3-5 Tage Anfängerfreundlichkeit 48/100
kcp-dev/api-syncagent#172 · 1 Kommentar · 1 Reaktion ·
Alle Issues in kcp-dev/api-syncagent
Ähnliche Issues
-
agent-butler-finding chore
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 88/100
jordansmall/spindrift#4146 ·
Maintainer antworten meist innerhalb von 1 Tag
-
security
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 68/100
IBM/ibmcloud-volume-file-vpc#119 ·
-
security
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 66/100
IBM/networking-go-sdk#339 ·
Maintainer antworten meist innerhalb von 1 Tag
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
kubernetes-sigs/mcp-lifecycle-operator#439 ·
Maintainer antworten meist innerhalb von 1 Tag
-
area: global bug dx priority: low
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Maintainer antworten meist innerhalb von 1 Tag