PostgreSql mesh query provider stall blocks ProvisionPlan's deployment-record existence check — provider-stall family (#5315) persists on memex
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 25/100
Research direction
Start with the parent investigation in #5315 and read the QueryFanInStallTerminal documentation for the budget ladder, pool behavior, and fleet history. Then inspect the Initial-emission path of PostgreSqlPartitionedMeshQuery and the fan-in merge that raises QueryProviderStalledException, using PostgreSQL, ThreadPool, and pg-read telemetry for the named pod around 2026-10-05T03:52Z. Done means identifying and resolving the provider-side stall; do not widen the ProvisionPlan timeout. The issue notes this may be folded into #5315.
Written by the indexing model from the issue text.
Description
What is failing
The deployment-provisioning plan for memex-cloud aborted in its early "read deployment record" phase (phase 2/8): the existence check for the Hosting/Deployment record at Deployments/memex-cloud could not be established because the Postgres mesh query provider (MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery) did not emit an Initial within the query fan-in's 15 s bound, on two queries — the exact-path probe and the Deployments namespace listing. The plan correctly failed closed and reported the read as unavailable (retryable) rather than treating the record as absent — the consumer behaved as designed; the defect is the stalled provider.
Probable cause (medium confidence)
The recurring provider-stall family tracked under #5315 — process-level starvation (ThreadPool / the process-wide pg-read:Postgres pool) that the fan-in sees at the provider, not a defect in the plan or the query shape. Supporting evidence: identical exception and provider in fleet occurrences across deployments for the past three weeks (#5315, #5345, #5390, #5393, #6071, #6104), with single occurrences at unrelated consumer sites — this is another site of the same underlying stall, not a new defect at this log site. The timing (2026-10-05 03:52Z) is well after the live-requery coalescing (#5615) and exact-probe batching (2026-09-25) fixes shipped, matching #6071/#6104's reading that the memex deployment remains affected. No provider-side stack exists in the sample, so the precise queue this time is not establishable from this incident alone.
Impact
Low on its own numbers: one occurrence, one pod, one provisioning run, retryable and non-corrupting. The systemic weight belongs on #5315. If these stalls start bursting on memex and begin blocking provisioning runs repeatedly, a human should re-rule severity upward.
Where to look
- The
Initial-emission path ofMeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQueryand the fan-in merge that raisesQueryProviderStalledException— parent investigation in #5315. - QueryFanInStallTerminal — the budget ladder, the
pg-read:pool reading, and the fleet history behind this diagnosis. - The
ProvisionPlandeployment-record phase is the consumer surface only; per policy the stalled provider must be fixed, never the consumer's timeout widened. - PostgreSQL / ThreadPool /
pg-read:pool telemetry for podmemex-portal-deployment-67c67b456d-255rqaround 2026-10-05T03:52Z.
Likely duplicate / related. Same exception, same provider, same single-occurrence-on-memex shape as #6104 (filed 2026-10-04 from the PlatformBuildInboxWatcher site) and #6071 — this is the provider-stall family (#5315) recurring at one more consumer site after the coalescing and batching fixes. A human can probably fold this into #5315; the one line worth keeping is that the family still fires on memex and now blocks deployment provisioning.
Evidence
| Fingerprint | 58fe8ab83526d1da |
| Category | ProvisionPlan |
| Severity | Error |
| Exception | QueryProviderStalledException |
| Namespace | memex |
| Pods | memex-portal-deployment-67c67b456d-255rq |
| Occurrences | 1 |
| First seen | 2026-10-05 03:52:51Z |
| Last seen | 2026-10-05 03:52:51Z |
| Routing | not determined — no configured route matches the category ProvisionPlan. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject. |
Recent log lines
2026-10-05 03:52:51Z memex-portal-deployment-67c67b456d-255rq fail: ProvisionPlan[0]
[ProvisionPlan] 'memex-cloud' FAILED in phase 2/8 'Read deployment record': could NOT ESTABLISH whether a Hosting/Deployment record exists for 'Deployments/memex-cloud' — the index was asked (path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1; namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500) and listed none, but that is a FLOOR, not an answer: 'path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1': the read failed — QueryProviderStalledException: Query provider(s) [MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery] did not emit an Initial within the query fan-in's 15s bound for query 'path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1' (user 'system-security'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.; 'namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500': the read failed — QueryProviderStalledException: Query provider(s) [MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery] did not emit an Initial within the query fan-in's 15s bound for query 'namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500' (user 'system-security'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.. Evidence: path:Deployments/meme…[truncated]
Opened automatically from Admin/_LogIncident/58fe8ab83526d1da. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site 5779601b0f38e51d: other fingerprints of this site fold in here as comments rather than opening tickets of their own.
- Dominant language
- C#
- Stars
- 12
- Forks
- 5
- Avg merge
- 3h 59m
- Merged PRs (30d)
- 975
Getting set up
- No Dockerfile or Docker Compose file
- Has a pull request template
- No contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from Systemorph/MeshWeaver
-
sev:L
Difficulty 1/5 Under an hour Newbie friendliness 72/100
Systemorph/MeshWeaver#6233 ·
Maintainers usually reply within 1 day
-
documentation feedback sev:L
Difficulty 1/5 Under an hour Newbie friendliness 82/100
Systemorph/MeshWeaver#6033 ·
Maintainers usually reply within 1 day
-
area:search documentation
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Systemorph/MeshWeaver#6030 ·
Maintainers usually reply within 1 day
-
ApiTokenService.RevokeToken posts its revocation SaveMeshNodeRequest from the mesh (router) hub instead of a node-operation hubPossibly taken A pull request linked to this issue is open or already merged. Openbug sev:M
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Systemorph/MeshWeaver#6026 · 1 comment ·
Maintainers usually reply within 1 day
-
area:hosting bug sev:L
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
Systemorph/MeshWeaver#6025 ·
Maintainers usually reply within 1 day
All issues in Systemorph/MeshWeaver
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
stryker-mutator/stryker-net#3892 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
MobiFlight/MobiFlight-Connector#3419 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
Kryptos-FR/MarkView.Avalonia#105 ·
Maintainers usually reply within 1 day
-
[辞書]Open提案 辞書
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
microsoft/fluentui-blazor#5410 ·
Maintainers usually reply within 1 day