Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Linux expert arena (0.1.39): MADV_HUGEPAGE + defrag=madvise makes the arena's faults 20x slower (25 s start -> 432 s), and 'loaded ... at X GiB/s' does not show it

Open
#771 1 comment 1 reaction 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
4/5
Estimated time
3-5 days
Newbie friendliness
48/100
Issue type
Bug
Clarity
Mostly clear
Activity status
Active
Tech stack
cpp, linux
Domain
backend

Research direction

Start with the arena huge-page policy and note in src/core/pinned_thp.cpp, then read src/core/pinned_thp_test.cpp and the comment's proposed behavior. Check how the existing large-page environment controls and defrag setting are handled; use the 64 MiB selftest rather than reproducing the 40 GiB report. Done means the note and test cover the requested and skipped THP cases, with the default behavior clarified by maintainers.

Written by the indexing model from the issue text.

Description

Update (measured after this report). The 3.06-3.12 GiB/s load line above is not what this bug costs.
The same profile that crawled here reached ready at 432 s while still printing loaded 39.97 GiB at 2.86 GiB/s,
and the same profile with the MADV_HUGEPAGE request skipped is ready in ~25 s. The isolated 20x madvise
measurement, the /proc numbers, the switch and the note change are in the comment below.

Symptom

Engine 0.1.39 (source build), arena profile (strata-iq3_xxs.json: IQ3_XXS native pack, experts in RAM, no
--mmap-experts, --max-context 131072 --kv-resident 32768, MTP and vision on).
Machine: Manjaro Linux, kernel 6.12.108-1-MANJARO, i7-11700B, 62 GiB RAM, RTX 5070 Ti (16.3 GiB).

With the RAM free the fill is fast, and the log names the request that is in play:

strata generate: expert arena: cudaHostRegister PORTABLE ok; MAP_HUGETLB unavailable (needed 20464 2 MiB pages, vm.nr_hugepages=0); transparent huge pages requested (MADV_HUGEPAGE)
strata generate: loaded 39.97 GiB at 3.12 GiB/s

(a later start of the same profile, same engine: strata generate: loaded 39.97 GiB at 3.06 GiB/s)

With the RAM already in use, the same 39.97 GiB fill crawls at 60-77 MB/s, one core at 100 %, no disk I/O,
and the PC becomes unusable while it fills. Two starts of the arena profile did that in one afternoon; both were
stopped, and the engine's output is buffered, so those logs end at the line before the arena note:

strata generate: GPU 0: NVIDIA GeForce RTX 5070 Ti, compute capability 12.0     <- last line in those logs

The numbers below are live /proc readings taken while the fill was in progress, not log lines:

  • RSS deltas: +484 MB in 8 s, +616 MB in 8 s (60-77 MB/s); RSS 31.7 GB of ~43 GB when sampled
  • /proc/vmstat, 8 s window: compact_stall +858, compact_fail +858, thp_fault_alloc +0,
    thp_fault_fallback +290, allocstall_normal +286, nr_free_pages -498 MB
  • the process: 18,304 minor faults/s; at 4 KB per page that is 75 MB/s, the same number the RSS deltas gave.
    CPU 95-100 % of one core. In the same window the device read 0-1 MB/s and wrote nothing.
  • what the user sees: the server repeats [strata] still starting (N s) - please wait ... (serve/server.py:307)
Why the MADV_HUGEPAGE request is my prime suspect

This machine: /sys/kernel/mm/transparent_hugepage/enabled = [always], defrag = [madvise],
vm.nr_hugepages = 0.

The kernel's own description of the defrag policies (Documentation/admin-guide/mm/transhuge.rst):

  • madvise: "will enter direct reclaim like always but only for regions that are have used madvise(MADV_HUGEPAGE). This is the default behaviour."
  • always: "... an application requesting THP will stall on allocation failure and directly reclaim pages and compact memory in an effort to allocate a THP immediately."

With defrag=madvise, the arena's own madvise(MADV_HUGEPAGE) is what can put this 40 GiB mapping into
synchronous direct reclaim and compaction in the faulting thread. The counters point the same way: in that 8 s
window a compaction was tried ~107 times per second and never succeeded (compact_fail == compact_stall), no
new THP was allocated (thp_fault_alloc +0), the faults fell back to 4 KB pages (thp_fault_fallback +290), and
at 18,304 faults/s the faulting thread paid ~50 µs of kernel work per page - 75 MB/s with one core burnt.
(The arena does get huge pages earlier in a run: 13.8 GB of a 31 GB fill were backed by them when sampled.)

With the RAM still free there is nothing to compact, and the same request costs nothing (3.06-3.12 GiB/s above).
For reference, the 0.1.38 build (which has no THP request: its note ends in ; using 4 KB pages) filled the same
profile on this machine at 3.32 GiB/s, so the request buys nothing in this phase here; the TLB argument in the
comment above it is about the CPU expert pool during decode, not this fill.

Not the cause (checked)
  • disk: the device did not move during the window (0-1 MB/s reads, no writes); the expert bytes come from the
    page cache and the pack
  • loader/pack: the same code path, same pack and same machine filled the same arena at 3.06-3.12 GiB/s minutes
    earlier
  • VRAM: not used in this phase (the fill is host memory; the cache slots are chosen afterwards)
  • an OOM: no swap activity and no OOM kill, the fill kept progressing - the pages it got were simply 4 KB ones
  • STRATA_NO_LARGEPAGES=1 is not a workaround: it also skips the hugetlb attempt, i.e. it changes two things
Same mechanism reported elsewhere
  • golang/go#61718 "runtime: MADV_HUGEPAGE causes stalls when allocating memory"; the Go runtime dropped its own
    huge-page policy in 1.21.4 ("the Go runtime would no longer try to impose a huge page policy itself")
  • cockroachdb/cockroach#130241 "server: provide guidance and/or software control over transparent huge pages"
  • coreos/bugs#2635 transparent huge pages set to [always] are sub-optimal for many applications
What I would propose (implemented locally, with a selftest - happy to open a PR)
  1. put the policy in the note, so a crawling start explains itself in the log:
    ...; transparent huge pages requested (MADV_HUGEPAGE; kernel defrag=madvise)
    (one read of /sys/kernel/mm/transparent_hugepage/defrag; the note is what users paste into reports)
  2. a same-run A/B switch for the request only: STRATA_NO_ARENA_THP=1 ->
    ; transparent huge pages skipped (STRATA_NO_ARENA_THP); using 4 KB pages. Default unchanged.
    Alternative if you prefer no new switch: skip the request, or warn, when
    /sys/kernel/mm/transparent_hugepage/defrag is always, madvise or defer+madvise - i.e. whenever a
    2 MiB block may be impossible to get synchronously.
  3. src/core/pinned_thp_test.cpp: a selftest for the note in both cases (64 MiB arena, no model, no 40 GiB).

If you want a different default than 2, say so and I will follow that instead. I can also report a before/after
arena run with the switch on this machine, but that is the run that makes the PC unusable here, so it is a
one-shot: I would rather do it once you tell me what direction you want.

Dominant language
C++
Stars
11.6k
Forks
1k
Avg merge
7h 46m
Merged PRs (30d)
30

Getting set up

This project ships no dev container, Dockerfile or contributing guide, so setting up is up to you: start from its README, and see our first-contribution guide for the general steps.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from Niko1221/Strata

All issues in Niko1221/Strata

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.