Self-hosted Embed: failed pause or checkpoint loses the sandbox
@svalleru ci sta già lavorando.
Dal 23/9/2026.
Valutazione
Questa issue non è ancora stata valutata.
Descrizione
On self-hosted Embed, a failed pause or checkpoint can stop the original sandbox and leave no snapshot the SDK can restore. Its files and process state become inaccessible.
Reproduce
Pause: disk sync failure
- Create a sandbox and write a file inside it.
- Call
sandbox.pause(keep_memory=True). - Use
straceto makefsyncreturnEIOfor the rootfs diff file during the pause.
Pause returns 500: Error pausing sandbox; the orchestrator reports failed to sync file: input/output error. Reproduced 2/2 times. Without the injected error, pause/resume preserves files and process memory.
Fault injection command
Attach after the diff file appears and before it is synced (about 1–2 seconds). Replace the path with the actual diff file path; a watcher was used to catch this window.
strace -f -p "$(pgrep -x orchestrator)" -qq -yy \
-P "/orchestrator/build/<build-id>-rootfs.ext4-<random-suffix>" \
-e trace=fsync,fdatasync \
-e inject=fsync:error=EIO:when=1 \
-o /tmp/pause-fsync.strace
The original incident followed repeated host No space left on device errors. Those logs do not establish the cause of the later fsync error.
Checkpoint: memory allocation failure
Exhaust the host's reserved hugepage memory pool, then call create_snapshot(). Starting the replacement sandbox fails with mmap memfd: cannot allocate memory. Reproduced 1/1 time with real memory exhaustion on the same release binary. Normal checkpoint/restore works with memory available.
Actual result
In both cases:
- The original sandbox's Firecracker process stops;
get_inforeturns “not found.” - The snapshot build is marked
failed; restoring it returns404: template not found.
The pause failure deletes the local snapshot artifacts. The checkpoint failure leaves local files, but they cannot be restored through the SDK.
Expected result
A failed pause or checkpoint should leave the sandbox recoverable: either the original sandbox remains available or a usable snapshot can be restored.
Environment
- Single-node Embed, local storage (
file:///var/lib/e2b/storage/templates) - Ubuntu 26.04 ARM64 VM with nested virtualization
- Orchestrator
v0.16.202609130627-59497eb9134(unmodified release binary) - API
v0.14.202609170000-908833e4c12; Python SDK2.51.0 - Compose revision:
a065a4ddb3f2c6a4149634d9acb14b62f65839ac
Related: #3659 covers a snapshot marked successful that crashes on restore.
- Lingua principale
- Go
- Stelle
- 1.6k
- Fork
- 438
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Guida per i contributori
Apri la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di e2b-dev/runtime
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
-
sandbox cache: StartRemoving state transition not broadcast, all allocations see stale Running state Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 86/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
Tutte le issue di e2b-dev/runtime
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
prometheus/procfs#872 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
bazel-contrib/rules_go#4726 · 1 commento ·
-
area/auto-scaling area/monitoring area/ops-productivity kind/enhancement
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100