Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

[Bug] PD task REST: balanceLeaders throws a bare HTTP 500 during the balance-shard window, and task responses cannot be told apart from a no-op or a follower answer

Open Beginner friendly
#3,231 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
70/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
java
Domain
api, backend

Research direction

The bug is in TaskAPI.java lines 92-95 where balanceLeaders throws PDException without a handler. Compare with sibling endpoints like patrolPartitions that catch and return toJSON(e). Also examine TaskScheduleService.java:478 for the error message. After fixing, test the REST endpoints /v1/task/balanceLeaders and /v1/task/patrolPartitions to ensure they return proper JSON instead of a bare 500.

Written by the indexing model from the issue text.

Description

Bug Type (问题类型)

rest-api (结果不合预期)

Before submit
  • I have confirmed and searched that there are no similar problems in the historical issue and documents
Environment (环境信息)
  • Server Version: master 83ef9f3 (pd image sha256:69dc1d4e8625)
  • Backend: HStore, 3 PD + 3 Store, 12 shard groups
  • OS: kind v0.33 (Kubernetes 1.37.0)
Expected & Actual behavior (期望与实际表现)

Measured 2026-09-22 while executing the documented recovery sequence (patrolPartitions, balanceLeaders, balancePartitions) on the PD leader with a follower as control:

  1. GET /v1/task/balanceLeaders within 180 s of a balancePartitions call (which sets the balance-shard key even when it moves nothing) answers {"timestamp":...,"status":500,"error":"Internal Server Error","path":"/v1/task/balanceLeaders"} with no reason. The reason (PDException: balance shard is processing, please try later!, TaskScheduleService.java:478) appears only in the PD log. Cause: TaskAPI.balanceLeaders (TaskAPI.java:92-95) declares throws PDException with no handler, while its siblings catch and return toJSON(e). One-line fix.
  2. GET /v1/task/patrolPartitions returns the byte-identical body {"status": 0,"partitions": [ ]} on the leader and on a follower, whether or not it repaired anything; a follower does no work and reports success. balancePartitions differs by two bytes ({} from the leader, empty body from a follower). Operators driving recovery over REST cannot tell "ran and repaired", "ran, nothing to do" and "hit a follower, did nothing" apart without reading PD logs.

Ask: catch the PDException in balanceLeaders like the sibling endpoints; and have the task endpoints either redirect to the leader (as the read APIs do via getMembers) or answer an explicit error naming the leader, plus include what was done (groups reallocated, leaders moved, partitions moved) in the body.

Dominant language
Java
Stars
3.2k
Forks
637
Avg merge
3d 17h
Merged PRs (30d)
22

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from apache/hugegraph

All issues in apache/hugegraph

Similar issues

More Java issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.