bug: dangling schema in an APIExport is neither removed nor reported
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 55/100
Research direction
Start in internal/controller/apiexport/reconciler.go at mergeResourceSchemas(), then read internal/controller/apiexport/controller.go where readySchemaNames is built. Reproduce the dangling reference with the APIExport and APIResourceSchema commands in the issue, and inspect existing event and condition handling. Done means a missing schema is reported on the owning APIExport with enough detail to identify the entry, without breaking managed-resource reconciliation.
Written by the indexing model from the issue text.
Description
What happens
If an APIExport contains a schema name (spec.latestResourceSchemas, spec.resources[].schema
in v1alpha2) that does not resolve to an existing APIResourceSchema, nothing in the export
owner's view tells them about it:
- the Sync Agent leaves the entry in place forever (
mergeResourceSchemas()keeps every
existing entry that belongs to a group/resource noPublishedResourcemanages — deliberately,
see below), and it does not validate that the referenced ARS exists; - the
APIExport's own status does not mention it either — it only carriesIdentityValid; - the only signal is on the consumer side: every
APIBindingto that export goes
APIExportValid=Falsewithreason: InternalError/
message: Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "<name>" not found,
the API group stops being served, and the export owner is being paged with a message that
does not name the entry they have to remove.
So the broken entry can only be found by reading the export by hand, and the export owner —
who is usually also the Sync Agent — gets no event and no condition about it.
Why it happens
internal/controller/apiexport/reconciler.go (mergeResourceSchemas()):
// Now we include all other existing ARS that use unknown resources;
// this both allows an APIExport to contain "unmanaged" ARS, and also
// will purposefully leave behind ARS for deleted PublishedResources,
// allowing cleanup to take place outside of the agent's control.
"Outside of the agent's control" means an admin has to notice — but the agent owns the export,
so that admin has no place to look. Two realistic ways to get there:
- A
PublishedResourceis deleted, and its ARS is removed later (by hand, or by re-creating
the installation on a new cluster) — the name stays in the export. - An admin-managed export, which the docs explicitly allow
("The APIExport above might look like it is defining all resources for the API test.example.com API
group, but in reality it might contain resource schemas like v1.crontabs.initech.com"),
references a schema that was never created (typo, ordering, a schema that failed to create).
Related: readySchemaNames in internal/controller/apiexport/controller.go is built from
PublishedResource.status.resourceSchemaName, so while the ARS controller cannot project a
resource (CRD missing on the service cluster, ARS creation failing), the agent leaves the whole
export untouched, broken entries included.
Reproduction
api-syncagent v0.7.0 and main (same code path). kcp v0.32.x, agent pointed at a workspace with
an APIExport/repro-export and an APIExportEndpointSlice of the same name, one
PublishedResource for widgets.repro.pax.dev on the service cluster.
-
Publish the resource, the agent creates its schema and adds it to the export:
$ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
Delete the
PublishedResource(the resource is no longer managed by the agent) and then
delete the ARS itself — the export now references something that does not exist:$ kubectl get apiresourceschemas No resources found $ kubectl get apiexport repro-export -o jsonpath='{.spec.resources}' [{"group":"repro.pax.dev","name":"widgets","schema":"v4ed05578.widgets.repro.pax.dev","storage":{"crd":{}}}] -
The agent keeps running and reconciling; it never removes the entry and never reports it —
no event on the export, no log line, and the export's status stays:$ kubectl get apiexport repro-export -o jsonpath='{.status.conditions}' [{"status":"True","type":"IdentityValid"}] -
The consumer only sees the failure:
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason} {.message}{"\n"}{end}' Ready=True APIExportValid=False InternalError Invalid APIExport. Please contact the APIExport owner to resolve: APIResourceSchema "v4ed05578.widgets.repro.pax.dev" not found
Recovery (verified)
Removing the entry from the export restores it — APIExportValid=True, Ready=True, and the
agent does not put the name back as long as no PublishedResource refers to that group/resource:
$ kubectl patch apiexport repro-export --type=merge -p '{"spec":{"resources":[]}}'
$ kubectl get apibinding repro-export -o jsonpath='{range .status.conditions[*]}{.type}={.status}{"\n"}{end}'
Ready=True
APIExportValid=True
Note that this does not stick for a resource the agent still manages: the next reconcile of
that PublishedResource adds its own schema name back (and removes the entry an admin had put
there — Warning RemovingResourceSchemas), so there the underlying reason for the missing ARS
has to be fixed instead.
Suggestion
- Have the
apiexportcontroller (it already lists the ARS/PublishedResources of the export)
emit aWarningevent and/or a status condition likeSchemaMissingon theAPIExportwhen
an entry inspec.latestResourceSchemas/spec.resources[]has no matching ARS. That would
put the signal where the owner of the export actually looks, instead of only on the consumers. - Alternatively/additionally: document the cleanup in a troubleshooting page — I will send a
docs PR for that in a moment, since the procedure above is not written down anywhere today.
- Dominant language
- Go
- Stars
- 25
- Forks
- 29
- Avg merge
- 5d 18h
- Merged PRs (30d)
- 1
Getting set up
- Ships a Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from kcp-dev/api-syncagent
-
kind/bug
Difficulty 2/5 1-3 hours Newbie friendliness 74/100
kcp-dev/api-syncagent#163 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
kcp-dev/api-syncagent#161 ·
-
kind/feature
Difficulty 4/5 3-5 days Newbie friendliness 48/100
kcp-dev/api-syncagent#186 · 1 reaction ·
-
bug: PublishedResource updates leave previous sync controllers activeMay be free again @adoi claimed this 71 days ago, and no pull request is open. Openkind/bug
kcp-dev/api-syncagent#182 · 1 comment · 1 assignee ·
-
bug: finalizer "syncagent.kcp.io/cleanup" on related resources block deletion of primary objectOpenkind/bug
Difficulty 4/5 3-5 days Newbie friendliness 48/100
kcp-dev/api-syncagent#172 · 1 comment · 1 reaction ·
All issues in kcp-dev/api-syncagent
Similar issues
-
bug needs-triage
Difficulty 2/5 1-3 hours Newbie friendliness 86/100
DataDog/dd-trace-go#5469 ·
Maintainers usually reply within 1 day
-
bug tests
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
l3montree-dev/devguard#3101 ·
Maintainers usually reply within 1 day
-
area:*of bug
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
oapi-codegen/oapi-codegen#2593 ·
Maintainers usually reply within 1 day
-
bug
Difficulty 1/5 Under an hour Newbie friendliness 85/100
DaoCloud/DaoCloud-docs#7432 ·
Maintainers usually reply within 1 day