Guest heap corruption in a program exec'd by a large parent (dpkg maintainer script)
メンテナーはふだん 2 日以内に返信
@jserv がすでに取り組んでいます。
2026年9月1日 から。
評価
この issue はまだ評価されていません。
説明
Guest heap corruption in a program exec'd by a large parent (dpkg maintainer script)
Summary
When dpkg execs a maintainer script whose child allocates a non-trivial amount
of memory, that child's heap is corrupted. Depending on the workload it dies
either with SIGSEGV inside the CPython eval loop or with glibc's
free(): invalid pointer abort.
The exact same command, run from /bin/sh in the same sysroot and the same
filesystem state, completes normally every time. Inserting one extra exec
between dpkg and the program also makes the failure disappear.
Reproducible 100% of the time on the affected code path.
Environment
- elfuse
b1aac9a - macOS 26 (Darwin 25.6.0), Apple silicon
- Sysroot: Debian 12 (bookworm) aarch64, glibc 2.36, dpkg 1.21.23,
python3.11 3.11.2-6+deb12u8 - Invocation form:
elfuse --fakeroot --sysroot ROOT /bin/sh -c '…'
Prerequisite: the worktree carried one local patch in src/syscall/fs-stat.c,
unrelated to this report but needed to get this far on a Debian 12 sysroot —
glibc 2.36 spells fstat(fd) as fstatat(fd, "", &st, AT_EMPTY_PATH), and
stat_at_path() runs path_translate_at() before its AT_EMPTY_PATH branch,
so the empty path is measured against dirfd and returns ENOTDIR for any
descriptor with no host path (pipe, socket). That makes cat into a pipe fail
with cat: standard output: Not a directory, which breaks apt/dpkg early. The
fix is to answer AT_EMPTY_PATH + empty path from the descriptor before any
path translation, the way sys_fchmodat/sys_fchownat already do. Filed
separately; mentioned here only so the reproduction is possible.
Reproduction
In a Debian 12 aarch64 sysroot with python3.11 installed:
elfuse --fakeroot --sysroot ROOT /bin/sh -c '
cd /tmp && apt-get download python3.11 &&
DEBIAN_FRONTEND=noninteractive dpkg --unpack /tmp/python3.11_*.deb &&
DEBIAN_FRONTEND=noninteractive dpkg --configure python3.11'
Output:
Setting up python3.11 (3.11.2-6+deb12u8) ...
Segmentation fault (core dumped)
dpkg: error processing package python3.11 (--configure):
installed python3.11 package post-installation script subprocess returned error exit status 139
The package's postinst is just:
files=$(dpkg -L libpython3.11-stdlib:arm64 | sed -n '/^\/usr\/lib\/python3.11\/.*\.py$/p')
/usr/bin/python3.11 -E -S /usr/lib/python3.11/py_compile.py $files
i.e. one python process byte-compiling 270 stdlib .py files. Takes about two
minutes under elfuse.
An immediate retry (dpkg --configure -a) usually succeeds, so inside a long
apt-get install run this presents as an intermittent failure that leaves
packages half-configured.
Two symptoms, same trigger
Same postinst, only the file list differs:
workload (from dpkg -L libpython3.11-stdlib) |
result |
|---|---|
| all 270 files | Segmentation fault (core dumped), exit 139 |
| last 160 files | free(): invalid pointer → Aborted (core dumped), exit 134 |
| first 112 files | Segmentation fault (core dumped) |
| first 64 files | completes normally |
The glibc free(): invalid pointer abort is the clearer signal: the child's
heap metadata is being corrupted. The SIGSEGV looks like the same corruption
landing on a pointer instead.
The threshold behavior (64 files fine, 112 files crash, and the last 160 files
crash as well) says this scales with how much the child allocates, not with any
particular input file.
Fault detail
Taken from a build with the SIGSEGV branch in src/syscall/proc.c logging at
warn level instead of if (verbose) log_debug — identical on every run:
elfuse: worker: EL0 data fault at 0x2008e8010 PC=0x4a09c4 (ESR=0x92000007 FSC=0x7) -> SIGSEGV/MAPERR
- ESR
0x92000007: EC=0x24 (data abort from a lower EL), IL=1, WnR=0 (read),
DFSC=0b000111 → translation fault, level 3. So the page is simply not mapped. - FAR
0x2008e8010is in the guest mmap area (mmap allocations in this process
run upward from0x200000000). - PC
0x4a09c4is in the non-PIE/usr/bin/python3.11image (load base
0x400000,.text0x420380–0x6c3b04). The binary is stripped; the
nearest exported symbol isPyEval_EvalCode+0x5c4, so the containing function
is probably a static one in the eval loop. The faulting instruction is a
table load out of a pointer that was just read from the frame:
4a09bc: b9406be3 ldr w3, [sp, #0x68]
4a09c0: f9000415 str x21, [x0, #0x8]
4a09c4: f863d821 ldr x1, [x1, w3, sxtw #3] <-- faults
guest_materialize_lazy() (src/core/guest.c:3661) had already declined the
address and the stale-TLB retry did not apply, so at fault time the VA is either
outside every noreserve region or its page-table extension failed.
What it is not
Each of these was tested, so they need not be re-checked:
- Not apt. Plain
dpkg -i/dpkg --unpack+dpkg --configurereproduce
it with apt out of the picture. - Not the apt pty.
apt-get -o DPkg::Use-Pty=false install --reinstall
still crashes. - Not the workload. The identical command line, run from
/bin/shin the
same sysroot and the same freshly-unpacked state, completes — repeatedly, for
all 270 files, with a cold__pycache__. - Not argv/env block size. argv is ~13 KB in both the working and the
failing case; padding the environment from a shell by 0–1024 bytes never
crashes. - Not a python-under-dpkg problem per se. Replacing the postinst's command
withpython3.11 -E -S -c 'print("ok")'under the same dpkg configure step
works fine — the child has to allocate first. - One extra exec makes it go away. Replacing
/usr/bin/python3.11with a
#!/bin/shwrapper thatexecs the real binary with identical argv and
environment: no crash, for the full 270-file workload.
That last point is the interesting one: what the exec'd image ends up with
appears to depend on the address space of the process that exec'd it. dpkg is a
large process (it reads its status database and file lists into memory);
/bin/sh (dash) and the shell wrapper are tiny. When the tiny shell is the one
that execs python, python is fine.
Notes for whoever picks this up
sys_execve()resetsbrk_base/brk_currentand rebuilds the fixed
regions atsrc/syscall/exec.c:1725-1820; that is the first place I would
look for state carried over from the previous image.- Diagnostics gap worth closing regardless of this bug: the fault detail above
is only logged under--verbose, which is not usable during a long package
install (gigabytes of syscall trace). Logging it at warn level when
signal_deliver_fault()reports SIG_DFL/core — i.e. only when the guest
process is about to die of the fault — would make a crash like this
self-reporting with no tracing.
- 主要言語
- C
- スター
- 271
- フォーク
- 28
- 平均マージ
- 2日 18時間
- マージ済み PR(30日)
- 19
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
sysprog21/elfuse のほかの issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 52/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 35/100
メンテナーはふだん 2 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 78/100
メンテナーはふだん 2 日以内に返信
sysprog21/elfuse の issue をすべて見る
似ている issue
-
[P2] Workspace updates silently ignore forbidden assignments while staging the row対応中かも このイシューにリンクされたプルリクエストがオープン中、またはマージ済みです。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
メンテナーはふだん 1 日以内に返信
-
難易度 1/5 1時間未満 初心者へのやさしさ 70/100
メンテナーはふだん 4 日以内に返信
-
category:port-update
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
microsoft/vcpkg#54338 · コメント 1 件 ·
メンテナーはふだん 2 日以内に返信
-
area:http-gateway good first issue priority:low type:docs
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
crazy-goat/php-fpm-ng#828 ·
メンテナーはふだん 1 日以内に返信
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 63/100
SunDevilRocketry/Flight-Computer-Firmware#347 ·
メンテナーはふだん 3 日以内に返信