Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

flm serve: streaming /api/chat segfaults the server on every request (insert() twice, generate() never called)

Abierto Apto para principiantes
#757 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
2/5
Tiempo estimado
1-3 horas
Aptitud para principiantes
78/100
Tipo de issue
Error
Claridad
Bien especificado
Estado de actividad
Activo
Stack tecnológico
cpp
Área
api, backend

Línea de trabajo

Empieza en src/server/rest_handler.cpp, en la rama de streaming /api/chat alrededor de las líneas 792-823, y compárala con el bloque de streaming /api/generate alrededor de las líneas 690-697. Usa la solicitud curl del issue con llama3.2:1b y gemma4-it:e4b, comprobando que el streaming siga activo, devuelva chunks de tokens y un chunk done con stats, coincida con el texto no streaming y realice un prefill.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

What happens

Every POST /api/chat with "stream": true crashes flm serve with SIGSEGV. The client gets no bytes back, and every other client loses the server too.

flm serve llama3.2:1b
curl -sN http://127.0.0.1:52625/api/chat -H 'Content-Type: application/json' \
  -d '{"model":"llama3.2:1b","messages":[{"role":"user","content":"Count from 1 to 5."}],"stream":true}'
# connection closes with no response; the flm process has exited (SIGSEGV)

It reproduces on llama3.2:1b and gemma4-it:e4b. The server log shows one Prefill chunk 1/1 line, then nothing. Backtraces from the core dumps:

llama3.2:1b                                  gemma4-it:e4b
#0 Sampler::sample_greedy(buffer<bf16>&)     #0 Sampler::sample_greedy(buffer<bf16>&)
#1 AutoModel::_shared_insert(...)            #1 AutoModel::_shared_insert(...)
#2 Llama3::insert(...)                       #2 Gemma4e::insert(...)
#3 RestHandler::handle_chat(...)             #3 RestHandler::handle_chat(...)
Cause

The streaming branch of handle_chat runs the insert() try block twice (rest_handler.cpp:792-819). It never calls generate() before ostream.finalize_chat() (:823).

The first insert() writes the whole prompt to token_history. The second one then prefix-matches all of it in AutoModel::_shared_insert (automodel.cpp:184-201). It erases every token, prefills an empty vector, and samples an empty logits buffer (:234).

This looks like the same code as #436, which reported empty chunks on v0.9.36 and was closed because the Ollama API is no longer maintained. On v1.0.6 the result is no longer an empty stream: the whole process crashes, including for clients that only use /v1/*. If the Ollama routes are staying unmaintained, removing /api/chat would also close this.

Fix we are carrying

Replace the second insert() block with the generate() block that streaming /api/generate already uses (:690-697):
https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/pkgs/fastflowlm/patches/stream-chat-generate.patch

The diff is small. Our copy of the file carries other patches, so the context lines differ from upstream's. With the fix, both models stream token chunks and then one done: true chunk with stats. The streamed text is byte-identical to the non-streaming reply, and there is one prefill per request. I can send this as a PR against main if that's useful.

Probes and logs: https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/bench-logs/oflm-api-conformance-2026-09-28-after-187/README.md

Environment

FLM v1.0.6 built from the tag, with our downstream request-handling patches applied. In this block those patches change only the bodies of the catch clauses. The double insert() and the missing generate() are upstream code.
Ryzen AI MAX+ 395 (Strix Halo), XDNA2 NPU, 8 columns. NixOS, Linux 7.2.8, amdxdna 0.10, NPU FW 1.1.2.65.

Lenguaje dominante
C++
Estrellas
1.9k
Forks
156
Merge medio
2 d 2 h
PR fusionados (30 d)
11

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de ROCm/FastFlowLM

Todos los issues de ROCm/FastFlowLM

Issues similares

Más issues de C++

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.