Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Misc. bug: llama-server second SIGTERM during shutdown calls exit() from the signal handler and can deadlock in glibc free()

Open Beginner friendly
#29,581 1 comment 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 1 day

Nobody has claimed this yet.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
82/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
cpp, linux

Research direction

Start in tools/server/server.cpp at signal_handler and inspect the second-interrupt path alongside shutdown_handler. Reproduce by sending SIGTERM followed by another SIGTERM during teardown, then verify that the immediate-termination path exits without waiting on cleanup or deadlocking in glibc; the issue provides the proposed POSIX behavior to check.

Written by the indexing model from the issue text.

Description

Name and Version

llama-server b11206 (commit 2b129ccfa), Vulkan (RADV) backend, Linux x86_64, glibc (Fedora-based container image), run under a container init that forwards signals. Also seen on b11200. The handler is unchanged on current master.

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

llama-server

Command line
llama-server -m <any model>
# then: SIGTERM, followed by a second SIGTERM a few ms later
Problem description & steps to reproduce

What happens. When llama-server receives SIGTERM, then a second
SIGTERM (or SIGINT) while it is already tearing down, the process
sometimes never exits. It logs

srv    operator(): operator(): cleaning up before exit...
Received second interrupt, terminating immediately.

and then sits in futex_wait indefinitely with the model already
released. It has to be SIGKILLed. In our eviction loop this happens on a
few percent of stops (7 of 72 stops over two builds, b11200 and b11206,
counting only stops that took more than 30 s).

Stack of the hung process (eu-stack, two samples 20 s apart, identical):

TID main:
#0  __lll_lock_wait_private
#1  _int_free_chunk
#2  __run_exit_handlers
#3  exit
#4  signal_handler(int)
#5  __restore_rt
#6  unlink_chunk.isra.0
#7  _int_free_merge_chunk
#8  _int_free_chunk
#9  llama_vocab::~llama_vocab()
#10 llama_model::~llama_model()
#11 llama_model_qwen3::~llama_model_qwen3()
#12 common_init_result::~common_init_result()
#13 server_context::~server_context()
#14 llama_server(common_params&, int, char**)
#15 llama_server(int, char**)
#16 __libc_start_call_main
TID log worker:
#0..#3 pthread_cond_wait
#4  std::condition_variable::wait(...)
#5  common_log::resume()::{lambda()#1}

/proc/<pid>/status at the same moment: SigBlk has SIGTERM set (the
thread is inside the SIGTERM handler), SigPnd/ShdPnd are empty.

Cause. tools/server/server.cpp:

static inline void signal_handler(int signal) {
    if (is_terminating.test_and_set()) {
        // in case it hangs, we can force terminate the server by hitting Ctrl+C twice
        // this is for better developer experience, we can remove when the server is stable enough
        fprintf(stderr, "Received second interrupt, terminating immediately.\n");
        exit(1);
    }
    shutdown_handler(signal);
}

The first signal starts the orderly shutdown, and the main thread reaches
server_context::~server_context(), which frees the vocab. The second
signal interrupts the main thread inside glibc _int_free while it holds
the malloc arena lock. The handler then calls exit(1), which is not
async-signal-safe: __run_exit_handlers runs atexit and static
destructors, one of them calls free(), and that free() waits on the
arena lock that the interrupted frame beneath it already holds. The thread
deadlocks on itself. fprintf in the same branch has the same class of
problem (stdio locks).

So the "hit Ctrl+C twice to force terminate" escape hangs exactly when
it is most likely to be used: while teardown is freeing memory.

Suggested fix. In the second-interrupt branch, use only
async-signal-safe calls:

static inline void signal_handler(int signal) {
    if (is_terminating.test_and_set()) {
        static const char msg[] = "Received second interrupt, terminating immediately.\n";
        (void) !write(STDERR_FILENO, msg, sizeof(msg) - 1);
        _exit(1);
    }
    shutdown_handler(signal);
}

_exit skips the atexit and destructor work, which is what "terminate
immediately" means anyway. (Windows goes through SetConsoleCtrlHandler
on its own thread, so the _exit change would be POSIX-only, or use
_exit there too.)

Repro sketch. Start llama-server with any small model. Send SIGTERM
and then a second SIGTERM a few milliseconds later, repeatedly (a process
supervisor that forwards and also signals directly does this). Most runs
exit 1 at once. The hang needs the second signal to land inside a free()
during teardown, so it takes a loop of stops to hit; ours hit about one
stop in ten.

First Bad Commit

The second-interrupt exit(1) branch came in with #5734 ("Server: Hit Ctrl+C twice to exit"). The deadlock needs a second signal to land during teardown, so it has probably been latent since then.

Relevant log output
srv    operator(): operator(): cleaning up before exit...
Received second interrupt, terminating immediately.
Dominant language
C++
Stars
129k
Forks
23.6k
Avg merge
2d 10h
Merged PRs (30d)
435

Getting set up

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from ggml-org/llama.cpp

All issues in ggml-org/llama.cpp

Similar issues

More C++ issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.