Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

PostgreSql mesh query provider stall blocks ProvisionPlan's deployment-record existence check — provider-stall family (#5315) persists on memex

Closed
#6,117 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
5/5
Estimated time
Over a week
Newbie friendliness
25/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
csharp, postgresql
Domain
databases

Research direction

Start with the parent investigation in #5315 and read the QueryFanInStallTerminal documentation for the budget ladder, pool behavior, and fleet history. Then inspect the Initial-emission path of PostgreSqlPartitionedMeshQuery and the fan-in merge that raises QueryProviderStalledException, using PostgreSQL, ThreadPool, and pg-read telemetry for the named pod around 2026-10-05T03:52Z. Done means identifying and resolving the provider-side stall; do not widen the ProvisionPlan timeout. The issue notes this may be folded into #5315.

Written by the indexing model from the issue text.

Description

bug sev:M

What is failing

The deployment-provisioning plan for memex-cloud aborted in its early "read deployment record" phase (phase 2/8): the existence check for the Hosting/Deployment record at Deployments/memex-cloud could not be established because the Postgres mesh query provider (MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery) did not emit an Initial within the query fan-in's 15 s bound, on two queries — the exact-path probe and the Deployments namespace listing. The plan correctly failed closed and reported the read as unavailable (retryable) rather than treating the record as absent — the consumer behaved as designed; the defect is the stalled provider.

Probable cause (medium confidence)

The recurring provider-stall family tracked under #5315 — process-level starvation (ThreadPool / the process-wide pg-read:Postgres pool) that the fan-in sees at the provider, not a defect in the plan or the query shape. Supporting evidence: identical exception and provider in fleet occurrences across deployments for the past three weeks (#5315, #5345, #5390, #5393, #6071, #6104), with single occurrences at unrelated consumer sites — this is another site of the same underlying stall, not a new defect at this log site. The timing (2026-10-05 03:52Z) is well after the live-requery coalescing (#5615) and exact-probe batching (2026-09-25) fixes shipped, matching #6071/#6104's reading that the memex deployment remains affected. No provider-side stack exists in the sample, so the precise queue this time is not establishable from this incident alone.

Impact

Low on its own numbers: one occurrence, one pod, one provisioning run, retryable and non-corrupting. The systemic weight belongs on #5315. If these stalls start bursting on memex and begin blocking provisioning runs repeatedly, a human should re-rule severity upward.

Where to look

  • The Initial-emission path of MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery and the fan-in merge that raises QueryProviderStalledException — parent investigation in #5315.
  • QueryFanInStallTerminal — the budget ladder, the pg-read: pool reading, and the fleet history behind this diagnosis.
  • The ProvisionPlan deployment-record phase is the consumer surface only; per policy the stalled provider must be fixed, never the consumer's timeout widened.
  • PostgreSQL / ThreadPool / pg-read: pool telemetry for pod memex-portal-deployment-67c67b456d-255rq around 2026-10-05T03:52Z.

Likely duplicate / related. Same exception, same provider, same single-occurrence-on-memex shape as #6104 (filed 2026-10-04 from the PlatformBuildInboxWatcher site) and #6071 — this is the provider-stall family (#5315) recurring at one more consumer site after the coalescing and batching fixes. A human can probably fold this into #5315; the one line worth keeping is that the family still fires on memex and now blocks deployment provisioning.


Evidence
Fingerprint 58fe8ab83526d1da
Category ProvisionPlan
Severity Error
Exception QueryProviderStalledException
Namespace memex
Pods memex-portal-deployment-67c67b456d-255rq
Occurrences 1
First seen 2026-10-05 03:52:51Z
Last seen 2026-10-05 03:52:51Z
Routing not determined — no configured route matches the category ProvisionPlan. This repository is the configured fallback, not a finding about who owns the fault; the category names the LOGGER, which may not be the subject.
Recent log lines
2026-10-05 03:52:51Z memex-portal-deployment-67c67b456d-255rq fail: ProvisionPlan[0]
      [ProvisionPlan] 'memex-cloud' FAILED in phase 2/8 'Read deployment record': could NOT ESTABLISH whether a Hosting/Deployment record exists for 'Deployments/memex-cloud' — the index was asked (path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1; namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500) and listed none, but that is a FLOOR, not an answer: 'path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1': the read failed — QueryProviderStalledException: Query provider(s) [MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery] did not emit an Initial within the query fan-in's 15s bound for query 'path:Deployments/memex-cloud nodeType:Hosting/Deployment select:path,id limit:1' (user 'system-security'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.; 'namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500': the read failed — QueryProviderStalledException: Query provider(s) [MeshWeaver.Hosting.PostgreSql.PostgreSqlPartitionedMeshQuery] did not emit an Initial within the query fan-in's 15s bound for query 'namespace:Deployments nodeType:Hosting/Deployment select:path,id limit:500' (user 'system-security'). The merged Initial gates on EVERY provider, so this query has NO snapshot to answer with — it is reported as unavailable (retryable) rather than left hanging with no error. This is an availability failure, never a permission verdict: a consumer deciding access must fail CLOSED and say it could not establish the answer. Fix the stalled provider; never bump the consumer's timeout.. Evidence: path:Deployments/meme…[truncated]

Opened automatically from Admin/_LogIncident/58fe8ab83526d1da. Recurrences are folded into this issue rather than opening new ones.
It also stands for the whole log site 5779601b0f38e51d: other fingerprints of this site fold in here as comments rather than opening tickets of their own.

Dominant language
C#
Stars
12
Forks
5
Avg merge
3h 59m
Merged PRs (30d)
975

Getting set up

  • No Dockerfile or Docker Compose file
  • Has a pull request template
  • No contributing guide

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Systemorph/MeshWeaver

All issues in Systemorph/MeshWeaver

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.