Revisit acceptable snapshot threshold for joiners
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 4/5
- Estimated time
- 3-5 days
- Newbie friendliness
- 68/100
- Issue type
- Bug
- Clarity
- Mostly clear
- Activity status
- Quiet
- Tech stack
- cpp
- Domain
- distributed-systems, security
Research direction
Start in node_state.h at find_local_startup_snapshot() and trace how the snapshot-age bound and verify_snapshot() are applied when previous_service_identity is absent. Compare this with the fetch path using join_config.service_cert and the existing retry logic. Done means a joiner accepts only a locally held snapshot from the current service, while falling back to the existing fetch and join retries when that condition is not met.
Written by the indexing model from the issue text.
Description
We currently allow joiners to use a snapshot from a predecessor service, without checking the signature on that snapshot, and then relying on later checks (and potentially waiting until consensus, for the merkle checks in signature verification) to recognise it was a legitimate place to start.
Sketching a timeline:
- Service A
- Node A.p
- Writes transactions 1 through 500
- Creates and writes `snapshot_495_500.committed`
- Dies
- Service B
- Node B.r
- Trying to recover from service A
- Finds `snapshot_495_500.committed` locally
- Confirms `snapshot_495_500.committed` was signed by A
- Recovers from `snapshot_495_500.committed`
- Writes `service_identity = B` in transaction 501
- Completes recovery and opens
- Writes transactions 501 through 600
- Node B.j
- Tries to join Service B (with no knowledge of A)
- Finds `snapshot_495_500.committed`
- Doesn't verify the signature on `snapshot_495_500.committed`
- Starts from `snapshot_495_500.committed`
- Submits a join request, is accepted
- Receives 496 through 600 via consensus
There's a risk that if Node B.j for some reason found a snapshot from a different service X (in practice, this is most likely to be some failed recovery B', but for these purposes its equivalent to a totally unrelated service), it doesn't recognise that until the last step here fails.
Specifically, in the code:
node_state.h : find_local_startup_snapshot()
for (const auto& [snapshot_seqno, snapshot_path] : committed_snapshots)
{
...
try
{
verify_snapshot(segments, config.recover.previous_service_identity);
}
NB: for a joiner, config.recover.previous_service_identity is std::nullopt.
This is specifically when looking for a local/already-held snapshot. The fetch path is different and always calls verify_snapshot(segments, join_config.service_cert); (ie - "the snapshot you[service] served me better be signed by you[service]").
We think we can improve this, by raising the lower-bound that the join-target uses to decide whether a joiner's startup seqno is "recent enough". Specifically, where we currently have a bound on "how many snapshots in the past" they may be, we should also require that this is a snapshot from the current service. This should be cheap to apply, falling back to the existing retry logic for fetches and joins if it fails. This may introduce a slight delay after recovery, where joiners must wait for a snapshot to be created before they can fetch it and join, but ensures they can locally verify that snapshot, and we never proceed with unverified contents.
- Dominant language
- C++
- Stars
- 876
- Forks
- 260
- Avg merge
- 1d 13h
- Merged PRs (30d)
- 160
Getting set up
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from microsoft/CCF
-
Difficulty 3/5 1-2 days Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 45/100
microsoft/CCF#8439 · 1 comment ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 35/100
Maintainers usually reply within 1 day
-
Dead code removalOpen
Difficulty 5/5 Over a week Newbie friendliness 25/100
microsoft/CCF#8326 · 1 reaction ·
Maintainers usually reply within 1 day
-
Difficulty 5/5 Over a week Newbie friendliness 25/100
Maintainers usually reply within 1 day
Similar issues
-
bug
Difficulty 1/5 1-3 hours Newbie friendliness 88/100
isl-org/Open3D#7585 · 1 comment ·
Maintainers usually reply within 2 days
-
Unconfirmed bug
Difficulty 1/5 Under an hour Newbie friendliness 88/100
luanti-org/luanti#17605 · 1 comment ·
Maintainers usually reply within 2 days
-
area: config area: firmware priority: P2 - medium size: S type: bug
Difficulty 2/5 1-3 hours Newbie friendliness 76/100
Mizithra/ActiveTerrain#16 ·
-
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
grumpycoders/pcsx-redux#2171 ·
Maintainers usually reply within 1 day
-
Difficulty 2/5 1-3 hours Newbie friendliness 70/100
Maintainers usually reply within 2 days