Repeated opcache_reset() under concurrency permanently grows ZCSG(map_ptr_last), inflating every newly spawned php-fpm worker
還沒有人認領這個 Issue。
評估
- 難度
- 4/5
- 預估耗時
- 3-5 天
- 新手友好度
- 45/100
- Issue 類型
- 缺陷
- 描述清晰度
- 描述清楚
- 活躍度
- 活躍
- 領域
- performance
研究方向
The issue is in ext/opcache/ZendAccelerator.c and zend_persist.c, focusing on ZCSG(map_ptr_last) and CG(map_ptr_last) handling during opcache_reset(). Start by examining accel_activate() and zend_reset_cache_vars() to see how map_ptr_last is reset. The reproducer script shows how to trigger the bug with concurrent requests. Look for where cache_script_in_shared_memory() writes back to SHM. Done means verifying that repeated opcache_reset() no longer causes permanent memory growth in new workers.
由索引模型根據 Issue 內容生成。
描述
Description
Description
On a busy php-fpm server, an application that calls opcache_reset() on many requests makes every newly spawned worker start bigger and bigger, without limit, until FPM is restarted.
The memory is a single large zero-filled anonymous mapping allocated by zend_map_ptr_extend() when a worker loads its first script from SHM. It is not visible to userland: on the affected production server PHP reported peak request memory of 2-38 MB while each worker's RSS was ~1.1 GB from birth.
Each opcache_reset() permanently increases ZCSG(map_ptr_last); it never shrinks. Workers born later allocate (and zero) the whole table.
Concurrency is required to reproduce. With resets happening while the pool is idle, growth is negligible (500 resets => +2 MB). With concurrent traffic and workers being spawned/reaped, the table grows on every reset.
I have not pinned down the exact code path for the concurrent case. My reading is that accel_activate() resets ZCSG(map_ptr_last) via zend_reset_cache_vars() only in the process that performs the restart, while workers that were busy keep their old (large) CG(map_ptr_last) and write it back to SHM on the next cache_script_in_shared_memory() (zend_persist.c: ZCSG(map_ptr_last) = CG(map_ptr_last);). That part is a hypothesis; the measurements below are not.
Real-world impact
A PHP application called opcache_reset() on ~4,700 requests/day (~1 every 16 s) on a shared server with ~850 req/min of traffic:
- growth measured in production: ~0.3 MB per reset (app of ~6,000 cached scripts)
- 21 hours after the last FPM restart, every newly spawned worker was ~960 MB at birth
- php-fpm total RSS: 52 GB across ~50 workers; a graceful reload brought it back to 2.4 GB
- the previous day the box (128 GB RAM) hit OOM, the kernel started killing processes and sites went down
- OPcache SHM is shared by all pools of that PHP version, so unrelated accounts on the same server were inflated too
Reproducer
Attached script: starts its own php-fpm (temp dir, own socket), serves a small app (2,500 autoloaded classes with inheritance/interfaces/traits) and sends concurrent FastCGI requests, optionally issuing opcache_reset() concurrently. Python 3.6+, stdlib only.
python3 php_opcache_reset_repro.py <php-fpm> <opcache.so|-> 600 noreset
python3 php_opcache_reset_repro.py <php-fpm> <opcache.so|-> 600 reset
It prints, before and after, the max worker RSS and the max anonymous mapping in a worker (where the map_ptr table lives).
Results (600 cycles, 6 concurrent requests each)
| PHP | Build | mode | max worker RSS | max anon mapping |
|---|---|---|---|---|
| 8.3.33 | Debian/Sury | noreset | 20 -> 22 MB | 4 -> 2 MB |
| 8.4.24 | Debian 13 | noreset | 21 -> 22 MB | 6 -> 2 MB |
| 8.5.8 | static-php-cli | noreset | 42 -> 44 MB | 19 -> 19 MB |
| 8.1.34 | cPanel ea-php | reset (598) | 19 -> 87 MB | 4 -> 57 MB |
| 8.1.34 | CloudLinux alt-php | reset (598) | 20 -> 85 MB | 4 -> 54 MB |
| 8.3.33 | Debian/Sury | reset (592) | 20 -> 94 MB | 4 -> 65 MB |
| 8.3.33 | cPanel ea-php | reset (598) | 19 -> 84 MB | 6 -> 55 MB |
| 8.4.24 | Debian 13 | reset (600) | 20 -> 97 MB | 6 -> 66 MB |
| 8.4.25 | cPanel ea-php | reset (600) | 20 -> 89 MB | 6 -> 58 MB |
| 8.4.25 | CloudLinux alt-php | reset (595) | 21 -> 87 MB | 6 -> 55 MB |
| 8.5.8 | static-php-cli | reset (597) | 43 -> 111 MB | 19 -> 61 MB |
Growth is linear and does not level off: 1,800 resets give ~200 MB (8.3.33) and ~206 MB (8.5.8) of anonymous mapping per worker.
Note this is distinct from GH-8646, whose fix (zend_map_ptr_reset() at request end when CG(interned_strings) is non-empty) only applies when OPcache is disabled. Here OPcache is enabled and the growth happens through opcache_reset().
Expected result
Repeated opcache_reset() should not permanently increase per-worker memory. After a reset the map_ptr table should return to a size proportional to the code actually cached.
Actual result
ZCSG(map_ptr_last) grows on every reset and never shrinks, so every worker spawned afterwards allocates and zero-fills an ever larger table, until php-fpm is restarted/reloaded.
PHP Version
Reproduced on 8.1.34, 8.3.33, 8.4.24, 8.4.25 and 8.5.8 (Debian, Sury, static-php-cli, cPanel ea-php, CloudLinux alt-php). The relevant code in ext/opcache/ZendAccelerator.c looks unchanged in master.
Operating System
AlmaLinux 10 (production), Debian 13 (local repro)
- 主要語言
- C
- 星號
- 40.4k
- 分支
- 8.2k
- 平均合併
- 2 天 17 小時
- 30 天內合併 PR
- 115
貢獻指南
從這裡開始
- 先讀完整個 Issue,再讀專案的貢獻指南。
- 在 Issue 下留言說明你要接手 —— 這能避免兩個人做同樣的事。
- Fork 儲存庫,在一個分支上完成修改。
- 送出 Pull Request,並在描述裡引用這個 Issue 編號。
php/php-src 的其他 Issue
-
Bug Status: Needs Triage
難度 2/5 1-3 小時 新手友好度 76/100
-
Bug Status: Needs Triage
難度 1/5 1 小時以內 新手友好度 90/100
-
Bug Status: Needs Triage
難度 2/5 1-3 小時 新手友好度 78/100
-
Bug Category: Tests Status: Verified
難度 2/5 1-3 小時 新手友好度 68/100
-
Bug SAPI: fpm Status: Needs Triage
難度 2/5 1-3 小時 新手友好度 65/100
相似的 Issue
-
bug
難度 2/5 1-3 小時 新手友好度 75/100
bradcypert/plum#53 ·
-
Component: GLib
難度 2/5 1-3 小時 新手友好度 70/100
-
難度 2/5 1-3 小時 新手友好度 75/100
-
Status: Opened
難度 2/5 1-3 小時 新手友好度 70/100
-
難度 2/5 1-3 小時 新手友好度 75/100
nextbsd/nextbsd-userland#285 ·