Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

32-bit truncation/overflow in WAL length arithmetic wipes or panics on WALs near 4 GiB

Open
#560 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
3/5
Estimated time
1-2 days
Newbie friendliness
58/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
go

Research direction

Read wal/wal.go and wal/record.go first, then check the WAL wiring at instance.go:429 and rotation behavior in wal/gc.go. Update framing arithmetic and bounds validation without truncation or panics, reject unrepresentable writes, and add replay and oversized-payload tests covering the stated 4 GiB boundaries and length fields.

Written by the indexing model from the issue text.

Description

medium security wal

Details

The record framing code performs all length arithmetic in uint32 while file sizes are int64 and payload lengths are Go ints. Four wraparounds result, all in the crash-recovery path that runs at every node startup:

  1. wal.go ReadAll: readRecord(w.file, uint32(bytesToRead)) truncates the remaining file size to 32 bits. For ANY single WAL file >= 2^32 bytes, uint32(remaining) understates the true remaining size as reads approach a 4 GiB multiple: with near-certainty (probability ~1 - 12/recordSize) the next genuine record's payloadLen exceeds the truncated maxSize, readRecord errors, and ReadAll treats the perfectly valid record as a torn tail, calling truncateAt(fileSize - remaining). Everything past offset ~(fileSize - 2^32) — about 4 GiB of valid, fsynced consensus records — is silently destroyed and ReadAll returns nil error. When fileSize mod 2^32 is smaller than the first record's payload length, the failure occurs on the very first read and the ENTIRE log is truncated to zero. The consensus layer (simplex/epoch.go restoreFromWal) then starts as if the lost records never existed, so the node can re-vote in rounds it already voted in (equivocation risk).

  2. record.go readRecord: make([]byte, payloadLen+recordChecksumLen) wraps for payloadLen in [2^32-8, 2^32-1]; the buffer becomes 0..7 bytes, io.ReadFull succeeds, and payloadAndChecksum[:payloadLen] panics with slice bounds out of range. The payloadLen > maxSize guard does not prevent this because maxSize is itself the truncated uint32 cast: when the remaining size's low 32 bits are >= 2^32-8 (file within 8 bytes below a 4 GiB multiple), a corrupted or maliciously written length field converts the intended graceful truncate-and-recover behavior into a deterministic panic on every startup.

  3. record.go readRecord return value: recordSizeLen + payloadLen + recordChecksumLen wraps in uint32 for payloadLen >= 2^32-12, so ReadAll's bytesToRead accounting drifts and a later truncateAt offset falls inside valid records (requires a crafted ~4 GiB record with valid CRC, i.e. file tampering).

  4. record.go writeRecord: uint32(len(payload)) silently truncates the length prefix for payloads >= 4 GiB, writing a self-inconsistent record that fails CRC on replay and poisons the log from that point.

Reachability: in this repository's wiring (instance.go:429) the file WAL is always wrapped by GarbageCollectedWAL, which rotates once accumulated payload bytes exceed maxWalSize — default 100 MB (wal/gc.go) when ParameterConfig.WALMaxEntryCount is 0 — so a single file reaches 4 GiB only if the embedder passes a cap above 2^32 bytes or uses the exported wal.WriteAheadLog directly (it enforces no size bound). The wipe variant (1) needs no file tampering — only a >= 4 GiB file and a restart; variants 2 and 3 additionally require a corrupted/crafted length field on disk. Dominant impacts: silent wholesale destruction of durable consensus state (divergent replay/equivocation risk) and a persistent startup crash loop.

Evidence

  1. wal/wal.go:82–94
    bytesToRead is the int64 file size; it is truncated to uint32 when passed to readRecord as maxSize. For ANY file >= 2^32 bytes the truncated maxSize understates the remaining size: just before the remaining size crosses a 4 GiB multiple, the next valid record's payloadLen near-certainly exceeds uint32(remaining), the record is rejected as corrupt, and truncateAt(fileInfo.Size() - bytesToRead) silently destroys everything from that offset onward (~4 GiB of fsynced records; the ENTIRE log when fileSize mod 2^32 is below the first record's payload length) while returning a nil error.
  2. wal/record.go:59–63
    payloadLen is a uint32 read from disk. payloadLen + recordChecksumLen (8) is computed in uint32 arithmetic: for payloadLen >= 2^32-8 it wraps, so make() allocates a tiny buffer (e.g. 7 bytes for payloadLen 0xFFFFFFFF), ReadFull succeeds, and payloadAndChecksum[:payloadLen] panics with slice bounds out of range. The line 55 payloadLen > maxSize check does not prevent this because maxSize is itself the truncated cast: reachable when the remaining size's low 32 bits are >= 2^32-8 (file within 8 bytes below a 4 GiB multiple) and the on-disk length field is corrupted/crafted to a huge value. The node then crash-loops on every startup instead of performing its intended graceful truncation recovery.
  3. wal/record.go:73
    The returned bytesRead = recordSizeLen + payloadLen + recordChecksumLen also wraps in uint32 for payloadLen >= 2^32-12, causing ReadAll's bytesToRead accounting to drift and the eventual truncateAt offset to fall inside valid records (mistruncation of durable data). Requires a crafted ~4 GiB record whose CRC verifies, i.e. local file tampering.
  4. wal/record.go:28
    Write side: uint32(len(payload)) silently truncates the length prefix for payloads >= 4 GiB, producing a record whose declared length mismatches its data; on replay it fails CRC and everything from that record onward is truncated away. Unreachable with realistic consensus record sizes but the same root-cause family in the exported API.

Impact

Integrity: an entirely valid durable WAL is silently truncated to zero (or mistruncated mid-log) with a success return, rolling back consensus safety state and enabling divergent replay/equivocation. Availability: the payloadLen+8 wrap turns recovery into a deterministic slice-bounds panic on every startup, keeping the node down until manual file surgery.

Reproduction steps

  1. Triggered when the node restarts and replays its WAL. Requires a single WAL file near/above 4 GiB, which the wired GarbageCollectedWAL prevents by default (100 MB payload rotation) — needs an embedder-supplied cap above 2^32 bytes or direct use of the exported WriteAheadLog, hence AT PRESENT. The whole-log wipe variant needs no corruption: only file size >= 2^32 and a restart. The panic and mistruncation variants additionally need a corrupted/crafted length field on disk. No privileges or user interaction; the trigger is local on-disk state, not a network-deliverable input (AV LOCAL).

Recommended fix

  1. Remaining-file-size and record-length arithmetic is performed in uint32 while actual sizes are 64-bit, so casts and additions wrap: uint32(bytesToRead) in ReadAll, payloadLen+recordChecksumLen and the bytesRead sum in readRecord, and uint32(len(payload)) in writeRecord. Fix criteria: All framing arithmetic must be carried out in a width that cannot wrap for any representable file/payload size (e.g. int64 with explicit bounds checks), and declared payload lengths must be validated against the true remaining size minus header/checksum overhead before allocation or slicing. Verified by replay tests with file sizes straddling 4 GiB and with length fields of 0xFFFFFFF0-0xFFFFFFFF: no panic, no truncation of valid records, oversize declarations rejected as corruption at the correct offset.
  2. writeRecord silently truncates the length prefix for payloads >= 4 GiB, producing an unreadable record. Fix criteria: Append must reject payloads whose length cannot be represented in the record length field, returning an error instead of writing a malformed record. Verified by unit test that an oversized payload is refused and the log remains readable.

Severity: MEDIUM
Status: Open
Category: Integer overflow
CWE: CWE-197
Repository: ava-labs/Simplex
Branch: main
Date created: 2026-08-21


Dominant language
Go
Stars
22
Forks
4
Avg merge
3d 1h
Merged PRs (30d)
19

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from ava-labs/Simplex

All issues in ava-labs/Simplex

Similar issues

More Go issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.