sys_mmap: eager zero-fill on MAP_ANONYMOUS makes sparse mappings ~5x slower than Linux
メンテナーはふだん 2 日以内に返信
@Max042004 がすでに取り組んでいます。
2026年7月7日 から。
評価
この issue はまだ評価されていません。
説明
Summary
sys_mmap zeroes the entire region synchronously for every MAP_ANONYMOUS mapping (src/syscall/mem.c:2439):
/* Zero the mapped region */
memset((uint8_t *) g->host_base + result_off, 0, length);
This is functionally correct, but pays the full zero-fill cost up front. Linux backs new anonymous pages with the shared zero page and zeroes lazily on first touch, so workloads that map a region and touch only a fraction of it (e.g. CPython pymalloc 256 KiB arenas) pay for pages they never use.
Measurements (main @ e6b3455, Apple M-series)
Static aarch64 binary looping mmap(256 KiB, MAP_ANON|MAP_PRIVATE) → touch → munmap, 5000 iterations:
| Workload | elfuse | OrbStack (Linux) |
|---|---|---|
| 1-byte touch | 7,041 ns/cycle | 1,392 ns/cycle |
| touch every 4 KiB page (64 pages) | 6,934 ns/cycle | 8,904 ns/cycle |
On elfuse the 1-byte-touch cost equals the full-touch cost — the whole region is paid for at mmap time. On Linux the cost scales with pages actually touched.
Downstream effect: CPython for _ in range(5M): pass runs ~11.5 ns/iter on elfuse vs ~7.3 on OrbStack with the same python3.12 binary; the gap mostly disappears with PYTHONMALLOC=malloc (bypassing pymalloc's mmap-backed arenas), confirming it is mmap-bound rather than interpreter-bound.
Notes toward a fix
- The lazy machinery already exists:
guest_materialize_lazy()(src/core/guest.c:3265) plus the translation-fault path insrc/syscall/proc.calready perform fault-time PTE creation, per-2MiB-block zeroing, and TLBI — but only for regions carrying thenoreserveflag. Extending this to all anonymous mappings is the natural fix. Open questions: stale-PTE invalidation cost on address reuse (currentlyguest_invalidate_ptesper mmap on the NORESERVE path), and concurrent faults on the same range from multiple vCPUs. - Host
madvise(MADV_DONTNEED)is not a shortcut: on macOS it behaves likeMADV_FREE, and replacing the memset with it regressed the benchmark to 14–30 µs/cycle in a local experiment. - The memset itself is only ~25% of the gap (measured by unsoundly disabling it: ~5.6 µs/cycle remains). The rest is the per-syscall region-tracking /
guest_extend_page_tables/guest_update_perms/ TLBI path — lazy materialization also skips most of that work for pages never touched.
Reproduction
#include <stdio.h>
#include <stdlib.h>
#include <sys/mman.h>
#include <time.h>
int main(int argc, char **argv) {
int N = (argc > 1) ? atoi(argv[1]) : 5000;
size_t SZ = 256 * 1024;
struct timespec t0, t1;
clock_gettime(CLOCK_MONOTONIC, &t0);
for (int i = 0; i < N; i++) {
char *p = mmap(NULL, SZ, PROT_READ | PROT_WRITE,
MAP_ANON | MAP_PRIVATE, -1, 0);
p[0] = 1; /* or: for (size_t j = 0; j < SZ; j += 4096) p[j] = 1; */
munmap(p, SZ);
}
clock_gettime(CLOCK_MONOTONIC, &t1);
double ns = (t1.tv_sec - t0.tv_sec) * 1e9 + (t1.tv_nsec - t0.tv_nsec);
printf("%.1f ns/cycle\n", ns / N);
return 0;
}
aarch64-linux-musl-gcc -O2 -static -o bench_mmap bench_mmap.c
build/elfuse bench_mmap 5000
- 主要言語
- C
- スター
- 271
- フォーク
- 28
- 平均マージ
- 2日 16時間
- マージ済み PR(30日)
- 15
環境構築
- Dockerfile・Docker Compose ファイルなし
- プルリクエストのテンプレートなし
- コントリビューションガイドを読む
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
sysprog21/elfuse のほかの issue
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
メンテナーはふだん 2 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 30/100
メンテナーはふだん 2 日以内に返信
-
難易度 4/5 3〜5日 初心者へのやさしさ 25/100
メンテナーはふだん 2 日以内に返信
-
難易度 3/5 1〜2日 初心者へのやさしさ 65/100
メンテナーはふだん 2 日以内に返信
-
bug
難易度 4/5 3〜5日 初心者へのやさしさ 48/100
メンテナーはふだん 2 日以内に返信
sysprog21/elfuse の issue をすべて見る
似ている issue
-
IO.get_env on Node truncates names at embedded NUL対応中かも @Yi-111-a が今日担当しました。 オープン
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
HigherOrderCO/Bend#1449 · コメント 1 件 ·
-
bug C/C++ code
難易度 2/5 1〜3時間 初心者へのやさしさ 70/100
webarkit/WebARKitLib#84 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
OpenPrinting/cups#1751 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
es-ude/OnDeviceTraining#488 ·
メンテナーはふだん 1 日以内に返信
-
enhancement good first issue
難易度 2/5 1〜3時間 初心者へのやさしさ 66/100
メンテナーはふだん 1 日以内に返信