Segfault of the whole process when restarting an unhealthy thread: SG(server_context) is read after ts_free_thread()
Maintainer thường phản hồi trong vòng 1 ngày
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 76/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- backend, operating-systems
Hướng nghiên cứu
Bắt đầu trong frankenphp.c, tại php_thread(), đặc biệt là đường dẫn khởi động lại sau ts_free_thread(), và xem xét frankenphp_thread_index() cùng frankenphp_log_message(). Chạy Docker reproducer được cung cấp để quan sát sự cố crash, sau đó xác minh rằng các thread không lành mạnh được khởi động lại và container vẫn tiếp tục chạy qua các request lặp lại.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
What happened?
Disclosure: this issue was analyzed and written with Claude Opus 5.5 (Anthropic). The reproducer, the gdb session and the patch test below were run as described.
In classic mode, when a bailout escapes php_request_shutdown() (or php_request_startup() fails), php_thread() marks the thread unhealthy, tears it down and starts a new one. The teardown calls ts_free_thread() and then logs "Restarting unhealthy thread" through frankenphp_log_message(), which reads SG(server_context). At that point the thread no longer has TSRM storage, so the read faults and the whole process segfaults: every PHP thread dies and in-flight requests are dropped, instead of one thread being restarted.
frankenphp.c (v1.12.7; identical on main at e8752fa):
#ifdef ZTS
ts_free_thread();
#endif
if (thread_is_healthy) {
go_frankenphp_on_thread_shutdown(thread_index);
return NULL;
}
frankenphp_log_message("Restarting unhealthy thread", LOG_WARNING);
if (!frankenphp_new_php_thread(thread_index)) {
/* probably unreachable */
frankenphp_log_message("Failed to restart an unhealthy thread", LOG_ERR);
}
static inline uintptr_t frankenphp_thread_index(void) {
frankenphp_server_ctx *ctx = (frankenphp_server_ctx *)SG(server_context);
return ctx == NULL ? thread_index : ctx->thread_index;
}
static void frankenphp_log_message(const char *message, int syslog_type_int) {
go_log(frankenphp_thread_index(), (char *)message, syslog_type_int);
}
On Linux frankenphp.c is built without a static TSRMLS cache (ZEND_TSRMLS_CACHE_DEFINE() is only under PHP_WIN32), so SG(x) evaluates tsrm_get_ls_cache() + sapi_globals_offset on each access. ts_free_thread() ends with tsrm_tls_set(0), so tsrm_get_ls_cache() returns NULL and the read lands at sapi_globals_offset.
gdb on the stock dunglas/frankenphp:1.12.7-php8.5 image with the reproducer below:
Thread "php-0" received signal SIGSEGV, Segmentation fault.
call tsrm_get_ls_cache@plt
mov sapi_globals_offset(%rip),%rdx
=> mov (%rax,%rdx,1),%rax
test %rax,%rax
je ...
mov (%rax),%rdi
rax 0x0
rdx 0x20
si_addr 0x20
The same instruction sequence faults at the same address in a separate build of FrankenPHP 1.12.7 against PHP 8.5.10 ZTS, where the backtrace also shows ts_free_thread() returning on that thread immediately before the fault.
This started in v1.12.4. #2293 (v1.12.2, fixing #2268) added the restart path, at a time when frankenphp_log_message() called go_log(thread_index, ...) directly. #2438 (v1.12.4, fixing #2339) routed it through frankenphp_thread_index(), which turned the log call that follows ts_free_thread() into this access.
Reproducer
A convenient trigger is an opentelemetry post hook that runs out of memory. A function's observer frame is popped only after its end handlers return, so a hook that bails out is run a second time by zend_observer_fcall_end_all() at the start of php_request_shutdown(), which has no zend_try around it. The second fatal error escapes request shutdown and the thread is marked unhealthy. Any other bailout that marks a thread unhealthy reaches the same code.
Dockerfile:
FROM dunglas/frankenphp:1.12.7-php8.5
RUN install-php-extensions opentelemetry
ENV SERVER_NAME=:80
ENV FRANKENPHP_CONFIG="num_threads 2"
COPY index.php /app/public/index.php
index.php:
<?php
OpenTelemetry\Instrumentation\hook(null, 'f', post: static function (): void {
$x = str_repeat('x', (int) (memory_get_usage() + 256 * 1024 * 1024));
echo strlen($x);
});
function f(): void { echo ''; } // an empty body gets optimized away and is never observed
f();
$ docker build -t fp-crash . && docker run -d --name fp-crash -p 127.0.0.1:8099:80 fp-crash
$ for i in 1 2 3 4 5; do curl -s -o /dev/null -w '%{http_code}\n' http://127.0.0.1:8099/; done
$ docker inspect -f '{{.State.Status}} {{.State.ExitCode}}' fp-crash
exited 139
With num_threads 2, the container exited within the first three requests in each of four runs. With num_threads 1, one run kept the container up for 10 requests, but those requests appeared to hang rather than succeed (see #2558 below).
Expected: the unhealthy thread is replaced and the server keeps serving. Actual: the process exits on SIGSEGV (status 139).
Suggested fix
Log with the OS-thread-local thread_index on the path that runs after ts_free_thread(). The neighbouring go_frankenphp_on_thread_shutdown() and frankenphp_new_php_thread() calls already use it there:
go_log(thread_index, (char *)"Restarting unhealthy thread", LOG_WARNING);
if (!frankenphp_new_php_thread(thread_index)) {
go_log(thread_index, (char *)"Failed to restart an unhealthy thread", LOG_ERR);
}
With this change applied to the 1.12.7 build described above, a reproducer that made the unpatched build exit on the second request was survived for 8 consecutive requests, with the threads being restarted.
Related
- #2558 / #2595: when FrankenPHP is PID 1, a SIGSEGV in a PHP thread may leave the process up with a stuck thread instead of exiting, so this crash may show up as a hang.
Build Type
Docker (Debian Trixie)
Worker Mode
No
Operating System
GNU/Linux
CPU Architecture
x86_64
PHP configuration
phpinfo() output
FrankenPHP v1.12.7 PHP 8.5.11 Caddy v2.11.4
PHP 8.5.11 (cli) (built: Sep 24 2026 19:07:36) (ZTS)
Zend Engine v4.5.11, with Zend OPcache v8.5.11
opentelemetry extension version 1.4.2 (installed with install-php-extensions)
Image: dunglas/frankenphp:1.12.7-php8.5 (Debian GNU/Linux 13 trixie), otherwise defaults
memory_limit=128M, opcache.enable=1
[PHP Modules] Core ctype curl date dom fileinfo filter hash iconv json lexbor libxml mbstring mysqlnd openssl opentelemetry pcre PDO pdo_sqlite Phar posix random readline Reflection session SimpleXML sodium SPL sqlite3 standard tokenizer uri xml xmlreader xmlwriter Zend OPcache zlib
Relevant log output
Relevant log output
No log line precedes the exit: the process terminates with status 139 inside the "Restarting unhealthy thread" log call. The image does not log PHP errors by default (log_errors=0); with display_errors=1 the response body shows the "Allowed memory size ... exhausted" fatal error twice, the second time during shutdown.
- Ngôn ngữ chính
- Go
- Star
- 11.4k
- Fork
- 490
- Merge trung bình
- 2 ngày 22 giờ
- Pull request đã merge (30 ngày)
- 19
Chuẩn bị môi trường
- Có Dockerfile hoặc tệp Docker Compose
- Không có mẫu pull request
- Đọc hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của php/frankenphp
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 90/100
php/frankenphp#2671 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Include frankenphp version in logsCó thể đã có người làm @ousamabenyounes đã nhận 80 ngày trước. Đang mởenhancement
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
php/frankenphp#2483 · 3 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
suggestion about compile.mdCó thể làm lại được Pull request cho issue này đã bị đóng mà không được merge. Đang mởenhancement
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 68/100
php/frankenphp#601 · 5 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
v1.13.0: config reload returns HTTP 500 "server is not registered, you must first call frankenphp.Init() with the WithServer() option" while workers boot (regression from v1.12.7)Có thể làm lại được Pull request cho issue này đã bị đóng mà không được merge. Đang mở
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 54/100
php/frankenphp#2698 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Worker Mode regression in 1.13.0: Nextcloud worker fails to reach frankenphp_handle_request()Có thể đã có người làm @dunglas đã nhận hôm nay. Đang mởbug
Độ khó 4/5 3-5 ngày Mức phù hợp với người mới 48/100
php/frankenphp#2697 ·
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của php/frankenphp
Issue tương tự
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
Maintainer thường phản hồi trong vòng 1 ngày
-
duplication
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
openvibely/openvibely#1443 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 80/100
keyxmakerx/Chronicle#1179 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
raised-by:worker
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
medici-finance/assay#2486 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area/testing kind/bug triage/needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 75/100
cozystack/cozystack#4841 · 1 reaction ·
Maintainer thường phản hồi trong vòng 2 ngày