flightcontrol: nil pointer dereference in (*progressState).run when a waiter is cancelled as the call finishes
Maintainers usually reply within 1 day
Nobody has claimed this yet.
Assessment
- Difficulty
- 2/5
- Estimated time
- 1-3 hours
- Newbie friendliness
- 90/100
- Issue type
- Bug
- Clarity
- Clearly specified
- Activity status
- Active
- Tech stack
- go
- Domain
- build-system
Research direction
The data race is in util/flightcontrol/flightcontrol.go. Look at (*progressState).run around lines 303-308 where it closes writers without holding the lock, and at progressState.close around line 356 where slices.Delete modifies ps.writers. Apply the suggested fix: snapshot ps.writers under the lock, then close the snapshot outside the critical section. Run the provided reproducer test to confirm the panic no longer occurs.
Written by the indexing model from the issue text.
Description
Summary
dockerd crashes with a nil pointer dereference in util/flightcontrol.(*progressState).run. The cause is a data race between run and progressState.close.
When run reads io.EOF, it sets ps.done under ps.mu, releases the lock, and then ranges over ps.writers and calls Close on each one without the lock (flightcontrol.go:303-308 at v0.33.0). Meanwhile, a waiter whose context is cancelled while other waiters still hold the call reaches progressState.close. That function removes its writer under the lock with ps.writers = slices.Delete(ps.writers, i, i+1) (flightcontrol.go:356).
Since Go 1.22, slices.Delete zeroes the vacated tail element. The unlocked range in run still uses the old length, so it reads a nil rawProgressWriter and calls Close on it. addr=0x18 is the offset of the first method in the itab, and Close is that method.
The unlocked read itself was reported as a data race in #3923, which was closed without a code change. It became a crash with b5286f8dcb ("apply x/tools/modernize fixes", 2025-03-07). That commit replaced the append-based removal with slices.Delete. It first shipped in v0.21.0 (Docker 28.1.0), and the code is byte-identical in v0.33.0, v0.33.1 and master as of 2026-10-06.
Observed
Docker 29.8.0, vendoring BuildKit v0.33.0, as a rootless daemon with the containerd snapshotter, running several concurrent docker buildx bake sessions (docker driver) whose targets share identical stages:
panic: runtime error: invalid memory address or nil pointer dereference
[signal SIGSEGV: segmentation violation code=0x1 addr=0x18 pc=0x...]
goroutine ... [running]:
github.com/moby/moby/v2/vendor/github.com/moby/buildkit/util/flightcontrol.(*progressState).run(...)
.../vendor/github.com/moby/buildkit/util/flightcontrol/flightcontrol.go:307 +0x201
created by github.com/moby/moby/v2/vendor/github.com/moby/buildkit/util/flightcontrol.newCall[...] in goroutine ...
.../vendor/github.com/moby/buildkit/util/flightcontrol/flightcontrol.go:113 +0x245
The daemon exits, and every running container and build on it dies with it. On this host it happened once in about 5.5 days of continuous build load.
Reproducer
This is a standalone test against github.com/moby/buildkit at v0.33.0. It uses one shared call, one waiter that stays, and 32 waiters whose contexts are cancelled at about the moment the shared function returns:
package repro
import (
"context"
"math/rand"
"sync"
"testing"
"time"
"github.com/moby/buildkit/util/flightcontrol"
"github.com/moby/buildkit/util/progress"
)
func TestProgressStateCloseRace(t *testing.T) {
const iterations = 20000
const cancelled = 32
for i := 0; i < iterations; i++ {
var g flightcontrol.Group[int]
release := make(chan struct{})
fn := func(ctx context.Context) (int, error) {
<-release
return 1, nil
}
var wg sync.WaitGroup
_, keepCtx, keepClose := progress.NewContext(context.Background())
wg.Add(1)
go func() { defer wg.Done(); _, _ = g.Do(keepCtx, "k", fn) }()
cancels := make([]context.CancelFunc, 0, cancelled)
for j := 0; j < cancelled; j++ {
_, pctx, _ := progress.NewContext(context.Background())
ctx, cancel := context.WithCancel(pctx)
cancels = append(cancels, cancel)
wg.Add(1)
go func() { defer wg.Done(); _, _ = g.Do(ctx, "k", fn) }()
}
time.Sleep(200 * time.Microsecond)
go func() {
time.Sleep(time.Duration(rand.Intn(40)) * time.Microsecond)
close(release)
}()
for _, c := range cancels {
c()
}
wg.Wait()
keepClose(nil)
}
}
Results with Go 1.26:
- v0.33.0 panics within a fraction of a second with the same trace (flightcontrol.go:307 and :113,
addr=0x18). Under-raceit reports a DATA RACE betweenslices.DeleteincloseandCloseinrun. - v0.21.0 panicked in 3 of 3 runs.
- v0.20.2, which uses append-based removal, passed in 3 of 3 runs.
Suggested fix
Snapshot the writers while holding the lock, then close the snapshot:
if errors.Is(err, io.EOF) {
ps.mu.Lock()
ps.done = true
writers := slices.Clone(ps.writers)
ps.mu.Unlock()
for _, w := range writers {
w.Close()
}
}
With this change, v0.33.0 passed the reproducer in 3 of 3 runs of 20,000 iterations. Closing the writers while holding the lock would also work, but cloning keeps Close outside the critical section.
- Dominant language
- Go
- Stars
- 10.3k
- Forks
- 1.5k
- Avg merge
- 2d 5h
- Merged PRs (30d)
- 65
Getting set up
- Ships a Dockerfile or Docker Compose file
- No pull request template
- Read the contributing guide
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from moby/buildkit
-
docs: rootless.md should recommend a per-binary AppArmor profile over `apparmor_restrict_unprivileged_userns=0Possibly taken @shashankvarma499 claimed this 23 days ago. Open
Difficulty 2/5 1-3 hours Newbie friendliness 84/100
Maintainers usually reply within 1 day
-
area/dockerfile
Difficulty 2/5 1-3 hours Newbie friendliness 65/100
Maintainers usually reply within 1 day
-
area/source needs/maintainer-decision
Difficulty 5/5 Over a week Newbie friendliness 30/100
Maintainers usually reply within 1 day
-
Avoid pulling whole Nydus layers to merge cached imagesPossibly taken @shayonj claimed this 8 days ago. Open
Difficulty 4/5 3-5 days Newbie friendliness 48/100
Maintainers usually reply within 1 day
-
Difficulty 4/5 3-5 days Newbie friendliness 52/100
Maintainers usually reply within 1 day
Similar issues
-
status:approved type:bug
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Gentleman-Programming/gentle-ai#5326 ·
Maintainers usually reply within 1 day
-
area/testing kind/cleanup priority/backlog triage/accepted
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
lexfrei/cloudflare-tunnel-gateway-controller#999 ·
Maintainers usually reply within 1 day
-
automation models
Difficulty 2/5 1-3 hours Newbie friendliness 75/100
Maintainers usually reply within 1 day
-
bug
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
Maintainers usually reply within 1 day
-
Can the search results with Year Released be in enclosed in the ( ) like Movie/TV(Year Released)Open
Difficulty 2/5 1-3 hours Newbie friendliness 85/100
Dhairya3391/kari#32 ·