[BUG] Configuration subscription stops permanently after the daprd sidecar restarts
Maintainers usually reply within 3 days
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
Research direction
Start in dapr/clients/grpc/_response.py with ConfigurationWatcher, then trace subscribe_configuration() in both dapr.clients.DaprClient and dapr.aio.clients.DaprClient. Reproduce the stream failure by restarting daprd, and verify that both subscription paths recover and deliver later updates without restarting the application.
Written by the indexing model from the issue text.
Description
Expected Behavior
DaprClient.subscribe_configuration() keeps delivering configuration updates for the lifetime of the client.
If the daprd sidecar process exits and comes back (for example an OOM kill of the sidecar container while the application container keeps running), the SDK should re-establish SubscribeConfigurationAlpha1 against the new sidecar and continue calling the handler. Updates that happened while the sidecar was down should be visible on the new subscription, or the application should at least be able to observe that the previous stream ended.
Unary calls such as get_configuration() already work again once the new sidecar is listening. The long-lived configuration subscription should recover the same way.
Actual Behavior
The subscription works after a fresh application start. After daprd restarts, the handler is never called again. The application process stays healthy, so a Kubernetes pod is not recreated, and dynamic config updates are silently lost until the application process itself restarts.
ConfigurationWatcher in dapr/clients/grpc/_response.py reads SubscribeConfigurationAlpha1 on a daemon thread. When the stream raises, the thread prints to stdout and returns:
except Exception:
print(f'{self.store_name} configuration watcher for keys {self.keys} stopped.')
pass
There is no retry. subscribe_configuration() does not retain the ConfigurationWatcher, so nothing is left to start a new stream. The failure is not reported through the Python logger, and the returned subscription id belongs to the dead sidecar.
Seen with dapr Python SDK 1.18.3. Both dapr.clients.DaprClient.subscribe_configuration and dapr.aio.clients.DaprClient.subscribe_configuration use this watcher.
Steps to Reproduce the Problem
- Run an app with a
configuration.rediscomponent and the Python SDK 1.18.3. - Subscribe once at startup:
from dapr.clients import DaprClient
def handler(subscription_id, response):
print("update", subscription_id, {k: v.value for k, v in response.items.items()})
with DaprClient() as client:
subscription_id = client.subscribe_configuration(
store_name="configstore",
keys=["my-key"],
handler=handler,
)
print("subscribed", subscription_id)
# keep the process alive
- Confirm a config change invokes
handler. - Restart only daprd. In Kubernetes this is a sidecar OOM (
OOMKilled) while the app container keeps the same process. Locally, stop and start thedaprdprocess without restarting the app. - Change
my-keyagain after daprd is healthy.
handleris not called. Stdout may contain{store} configuration watcher for keys [...] stopped.Recreating the app process makes updates work again until the next sidecar restart.
Release Note
RELEASE NOTE: FIX Reconnect configuration subscriptions after the Dapr sidecar stream drops.
- Dominant language
- Python
- Stars
- 273
- Forks
- 155
- Avg merge
- 3d 13h
- Merged PRs (30d)
- 11
Getting set up
Starts the project's dev container in your browser, under your own GitHub account.
- No Dockerfile or Docker Compose file
- Has a pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from dapr/python-sdk
-
dapr-ext-workflow good first issue kind/enhancement P2
Difficulty 2/5 1-2 days Newbie friendliness 72/100
dapr/python-sdk#853 · 4 comments ·
Maintainers usually reply within 3 days
-
[WORKFLOW SDK FEATURE REQUEST] Retry WaitForInstanceCompletion/Start on a server-sent CANCELLEDOpendapr-ext-workflow kind/enhancement
Difficulty 4/5 3-5 days Newbie friendliness 55/100
dapr/python-sdk#1246 ·
Maintainers usually reply within 3 days
-
kind/bug
Difficulty 4/5 3-5 days Newbie friendliness 65/100
dapr/python-sdk#1233 ·
Maintainers usually reply within 3 days
-
kind/bug
Difficulty 4/5 3-5 days Newbie friendliness 45/100
dapr/python-sdk#1230 ·
Maintainers usually reply within 3 days
-
dapr-ext-workflow kind/enhancement
Difficulty 3/5 1-2 days Newbie friendliness 78/100
dapr/python-sdk#1213 ·
Maintainers usually reply within 3 days
Similar issues
-
customer-reported
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
Azure/azure-cli#34150 · 1 comment ·
Maintainers usually reply within 1 day
-
community-request
Difficulty 1/5 Under an hour Newbie friendliness 95/100
NVIDIA-NeMo/Curator#2464 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
WeblateOrg/translation-finder#1099 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
trezor/trezor-firmware#7997 ·
Maintainers usually reply within 2 days
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
Maintainers usually reply within 1 day