Zombie simulated datanode
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 25/100
- Issue type
- Bug
- Clarity
- Needs clarification
- Activity status
- Stale
- Tech stack
- java
- Domain
- distributed-systems, testing
Research direction
Start by reproducing the issue with start-dynamometer-cluster.sh and the production audit-log replay described in the report. Inspect the simulated datanode and ResourceManager behavior after the Yarn application is killed, then check the WebHDFS datanode list and NameNode errors. Done means disconnected simulated datanodes no longer remain running or send stale blocks after the application ends.
Written by the indexing model from the issue text.
Description
After running start-dynamometer-cluster.sh and replay the prod audit log for some time, some simulated datanodes (containers) lost connection to the RM and when the Yarn application is killed, these containers are still running, which will sending their blocks to the Namenode.
In this case, since datanode has gone through some changes with the replay where Namenode started from a fresh fsimage. Below errors will show up in the webhdfs page after the Namenode starts up.
Safe mode is ON. The reported blocks 1526116 needs additional 395902425 blocks to reach the threshold 0.9990 of total blocks 397826363. The number of live datanodes 3 has reached the minimum number 0. Name node detected blocks with generation stamps in future. This means that Name node metadata is inconsistent.This can happen if Name node metadata files have been manually replaced. Exiting safe mode will cause loss of 7141 byte(s). Please restart name node with right metadata or use "hdfs dfsadmin -safemode forceExitif you are certain that the NameNode was started with thecorrect FsImage and edit logs. If you encountered this duringa rollback, it is safe to exit with -safemode forceExit.
and checking datanode tab in the webhdfs page, a list of a couple datanodes will show up.
- Dominant language
- Java
- Stars
- 134
- Forks
- 34
- PR merge metrics
- No merged PRs in 30d
Contributor guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from linkedin/dynamometer
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
linkedin/dynamometer#106 ·
-
Difficulty 4/5 3-5 days Newbie friendliness 25/100
linkedin/dynamometer#104 · 4 comments ·
-
linkedin/dynamometer#101 · 2 comments · 1 assignee ·
-
Difficulty 3/5 1-2 days Newbie friendliness 30/100
linkedin/dynamometer#100 · 3 comments ·
-
Difficulty 3/5 1-2 days Newbie friendliness 45/100
linkedin/dynamometer#70 · 1 comment ·
All issues in linkedin/dynamometer
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
infinispan/infinispan#18150 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
-
untriaged
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
opensearch-project/k-NN#3597 ·
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 82/100