[Bug]: flock() on a mounted volume hangs forever (mount is missing `nolock`)
Maintainer antworten meist innerhalb von 3 Tagen
Dieses Issue hat noch niemand übernommen.
Bewertung
- Schwierigkeit
- 2/5
- Geschätzter Aufwand
- 1-3 Stunden
- Anfängerfreundlichkeit
- 88/100
- Issue-Typ
- Bug
- Klarheit
- Klar beschrieben
- Aktivitätsstatus
- Aktiv
- Tech-Stack
- go
- Bereich
- infrastructure
Rechercherichtung
Beginne in packages/envd/internal/api/init.go:607 und untersuche, wie nfsOptions für Sandbox-Volume-Mounts verwendet wird. Vergleiche dies mit packages/orchestrator/pkg/nfsproxy/e2e_start.sh:29 und den Filestore-Mount-Konfigurationen und führe anschließend die flock-Reproduktion aus, um zu bestätigen, dass die Sperre des gemounteten Volumes umgehend zurückgegeben wird, statt zu hängen.
Vom Indexierungsmodell aus dem Issue-Text verfasst.
Beschreibung
Sandbox ID or Build ID
iu5wt6ki121fvvkv6cklc (stuck process), iqop9y9ik26oaud4l4tky (remount test)
Environment
e2b JS SDK 2.46.0, sandbox guest Ubuntu 24.04, host macOS 15.6.
Volumes mounted through volumeMounts on Sandbox.create.
Timestamp of the issue
2026-09-04 13:47 UTC
Frequency
Happens every time
Expected behavior
flock() on a file in a mounted volume either succeeds as a node-local lock, or fails promptly with an error such as ENOLCK.
Actual behavior
It blocks forever in uninterruptible sleep (state D). The process cannot be killed, not even with SIGKILL, and it holds the sandbox until the sandbox itself is destroyed.
Reads and writes on the volume are fine. Only locking hangs.
Kernel stack of the stuck process:
__do_sys_flock → nfs_flock → nfs3_proc_lock → nlmclnt_lock → nlmclnt_call → rpc_wait_bit_killable
Issue reproduction
- Create a volume and mount it on a sandbox:
const volume = await Volume.create('flock-repro')
const sbx = await Sandbox.create('<template>', {
volumeMounts: { '/mnt/vol': volume },
})
- In the sandbox, take a lock on the volume and on local disk for comparison:
touch /mnt/vol/f /tmp/f
timeout 10 flock /mnt/vol/f -c true; echo $? # 124, blocked
timeout 10 flock /tmp/f -c true; echo $? # 0
- The first command never returns.
cat /proc/<pid>/stackshows the trace above, andcat /proc/<pid>/wchanreadsrpc_wait_bit_killable.
Additional context
Cause. nfsOptions in packages/envd/internal/api/init.go:607 sets neither nolock
nor local_lock, so local_lock defaults to none and the kernel sends every lock to
the proxy as an NLM request. go-nfs has no lock manager, and the portmapper answers the
lookup for one with port 0 (pkg/portmap/main.go:59). Because the mount is hard, the
client retries the bind forever instead of failing.
Fix. Add "nolock" to nfsOptions. Locks then resolve locally, which is the correct
semantics anyway given the proxy has no lock manager.
Why this looks like an omission. The same server is mounted with nolock everywhere
else in this repo:
packages/orchestrator/pkg/nfsproxy/e2e_start.sh:29mounts the nfsproxy withnolock,
so the e2e suite never runs the configuration production uses.- The Filestore chunk cache mount sets
"nolock", // do not use locking. - The host Filestore mount sets
"lock", "local_lock=none"deliberately, where a real
NLM-capable server is on the other end.
Only the sandbox volume mount leaves the choice unmade. Upstream has the same report
against go-nfs, closed by the reporter with "Using nolock solved this":
https://github.com/willscott/go-nfs/issues/98
Use case. We hit it with the Codex CLI, which takes an flock in $CODEX_HOME. The failure is silent: no error, no timeout, no log.
- Vorherrschende Sprache
- Go
- Sterne
- 1.6k
- Forks
- 448
- PR-Merge-Kennzahlen
- Keine gemergten PRs in 30 T.
Entwicklungsumgebung
Startet den Dev-Container des Projekts im Browser, mit Ihrem eigenen GitHub-Konto.
- Kein Dockerfile und keine Docker-Compose-Datei
- Keine Pull-Request-Vorlage
- Beitragsleitfaden lesen
Erste Schritte
- Lesen Sie das ganze Issue und danach den Beitragsleitfaden des Projekts.
- Schreiben Sie ins Issue, dass Sie es übernehmen — das erspart doppelte Arbeit.
- Forken Sie das Repository und arbeiten Sie in einem Branch.
- Öffnen Sie einen Pull Request, der die Issue-Nummer nennt.
Mehr aus e2b-dev/runtime
-
sandbox cache: StartRemoving state transition not broadcast, all allocations see stale Running stateOffen
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 86/100
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 86/100
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 88/100
Maintainer antworten meist innerhalb von 3 Tagen
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 82/100
Maintainer antworten meist innerhalb von 3 Tagen
Alle Issues in e2b-dev/runtime
Ähnliche Issues
-
bug
Schwierigkeit 1/5 Unter einer Stunde Anfängerfreundlichkeit 92/100
open-telemetry/opentelemetry-go-compile-instrumentation#1417 ·
Maintainer antworten meist innerhalb von 2 Tagen
-
agent-research-finding agent-research-recommend chore ready-for-agent
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 85/100
jordansmall/spindrift#4068 · 1 Kommentar ·
Maintainer antworten meist innerhalb von 1 Tag
-
Type/Task
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 74/100
OpenNSW/nsw-srilanka#537 ·
Maintainer antworten meist innerhalb von 1 Tag
-
security
Schwierigkeit 2/5 1-2 Tage Anfängerfreundlichkeit 62/100
-
Schwierigkeit 2/5 1-3 Stunden Anfängerfreundlichkeit 90/100
Maintainer antworten meist innerhalb von 1 Tag