No Management Proxy Node: Coordinator randomly goes down
Nobody has claimed this yet.
Assessment
- Difficulty
- 5/5
- Estimated time
- Over a week
- Newbie friendliness
- 20/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- azure, hadoop, kubernetes, postgresql, rust
- Domain
- backend, cloud, distributed-systems
Research direction
No source files, tests, or entry points are named. Begin by reproducing the failure on Stackable 24.3 with Apache Druid 28.0.1 in AKS, using the reported HDFS and PostgreSQL configuration; done means identifying the cause of the missing management proxy connection and confirming recovery without restarting all services.
Written by the indexing model from the issue text.
Description
Affected Stackable version
24.3
Affected Apache Druid version
28.0.1
Current and expected behavior
After roughly 3-4 days, the router will display "No Management Proxy Node." It seems, from testing, that the error is that the router cannot connect to the coordinator. However, all services display healthy logs and there are no clear errors, nor error codes from the panel.
The difficulty to debug comes from the fact that there are no errors.
Possible solution
The only way we have to recover from this state is to restart all services.
Additional context
- Extensions:
'["druid-kafka-indexing-service", "druid-datasketches", "prometheus-emitter", "druid-basic-security", "druid-opa-authorizer", "postgresql-metadata-storage", "druid-hdfs-storage", "druid-stats"]' - Deep Storage: HDFS
- Metadata Store: Postgres
Environment
AKS
Would you like to work on fixing this bug?
None
- Dominant language
- Rust
- Stars
- 12
- Forks
- 1
- Avg merge
- 1d 17h
- Merged PRs (30d)
- 10
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from stackabletech/druid-operator
-
Difficulty 4/5 3-5 days Newbie friendliness 35/100
stackabletech/druid-operator#692 ·
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
stackabletech/druid-operator#647 ·
-
Difficulty 3/5 1-2 days Newbie friendliness 30/100
stackabletech/druid-operator#646 ·
-
Server failing to create PoolableConnectionFactory. Failing with SCRAM-based authentication error. Opentype/bug
Difficulty 4/5 3-5 days Newbie friendliness 25/100
stackabletech/druid-operator#605 ·
-
Dependency Dashboard Open
Difficulty 2/5 1-3 hours Newbie friendliness 30/100
stackabletech/druid-operator#569 ·
All issues in stackabletech/druid-operator
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Eynzof/Hermes-CN-Desktop#610 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 68/100
gitbutlerapp/gitbutler#15998 · 1 comment ·
-
bug triage:deciding
Difficulty 1/5 Under an hour Newbie friendliness 88/100
open-telemetry/otel-arrow#4132 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100