flm serve: streaming /api/chat segfaults the server on every request (insert() twice, generate() never called)
Los mantenedores suelen responder en 1 día
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 2/5
- Tiempo estimado
- 1-3 horas
- Aptitud para principiantes
- 78/100
Línea de trabajo
Empieza en src/server/rest_handler.cpp, en la rama de streaming /api/chat alrededor de las líneas 792-823, y compárala con el bloque de streaming /api/generate alrededor de las líneas 690-697. Usa la solicitud curl del issue con llama3.2:1b y gemma4-it:e4b, comprobando que el streaming siga activo, devuelva chunks de tokens y un chunk done con stats, coincida con el texto no streaming y realice un prefill.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
What happens
Every POST /api/chat with "stream": true crashes flm serve with SIGSEGV. The client gets no bytes back, and every other client loses the server too.
flm serve llama3.2:1b
curl -sN http://127.0.0.1:52625/api/chat -H 'Content-Type: application/json' \
-d '{"model":"llama3.2:1b","messages":[{"role":"user","content":"Count from 1 to 5."}],"stream":true}'
# connection closes with no response; the flm process has exited (SIGSEGV)
It reproduces on llama3.2:1b and gemma4-it:e4b. The server log shows one Prefill chunk 1/1 line, then nothing. Backtraces from the core dumps:
llama3.2:1b gemma4-it:e4b
#0 Sampler::sample_greedy(buffer<bf16>&) #0 Sampler::sample_greedy(buffer<bf16>&)
#1 AutoModel::_shared_insert(...) #1 AutoModel::_shared_insert(...)
#2 Llama3::insert(...) #2 Gemma4e::insert(...)
#3 RestHandler::handle_chat(...) #3 RestHandler::handle_chat(...)
Cause
The streaming branch of handle_chat runs the insert() try block twice (rest_handler.cpp:792-819). It never calls generate() before ostream.finalize_chat() (:823).
The first insert() writes the whole prompt to token_history. The second one then prefix-matches all of it in AutoModel::_shared_insert (automodel.cpp:184-201). It erases every token, prefills an empty vector, and samples an empty logits buffer (:234).
This looks like the same code as #436, which reported empty chunks on v0.9.36 and was closed because the Ollama API is no longer maintained. On v1.0.6 the result is no longer an empty stream: the whole process crashes, including for clients that only use /v1/*. If the Ollama routes are staying unmaintained, removing /api/chat would also close this.
Fix we are carrying
Replace the second insert() block with the generate() block that streaming /api/generate already uses (:690-697):
https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/pkgs/fastflowlm/patches/stream-chat-generate.patch
The diff is small. Our copy of the file carries other patches, so the context lines differ from upstream's. With the fix, both models stream token chunks and then one done: true chunk with stats. The streamed text is byte-identical to the non-streaming reply, and there is one prefill per request. I can send this as a PR against main if that's useful.
Probes and logs: https://github.com/noamsto/nix-amd-ai/blob/46d7a73c1044435cec0ff08410e1a3bc87cbc544/bench-logs/oflm-api-conformance-2026-09-28-after-187/README.md
Environment
FLM v1.0.6 built from the tag, with our downstream request-handling patches applied. In this block those patches change only the bodies of the catch clauses. The double insert() and the missing generate() are upstream code.
Ryzen AI MAX+ 395 (Strix Halo), XDNA2 NPU, 8 columns. NixOS, Linux 7.2.8, amdxdna 0.10, NPU FW 1.1.2.65.
- Lenguaje dominante
- C++
- Estrellas
- 1.9k
- Forks
- 156
- Merge medio
- 2 d 2 h
- PR fusionados (30 d)
- 11
Preparar el entorno
- Incluye un Dockerfile o un archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de ROCm/FastFlowLM
-
No valid checkpoint to restore + Max length reached on multi-turn tool calls (Qwen3.6-MoE, NPU)Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
ROCm/FastFlowLM#744 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
ROCm/FastFlowLM#741 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Download timeoutAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 74/100
ROCm/FastFlowLM#520 · 7 comentarios ·
Los mantenedores suelen responder en 1 día
-
[Nice-to-have] /v1/pingAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
ROCm/FastFlowLM#434 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
No NPU device found on ubuntuAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
ROCm/FastFlowLM#400 · 8 comentarios · 10 reacciones ·
Los mantenedores suelen responder en 1 día
Todos los issues de ROCm/FastFlowLM
Issues similares
-
Dificultad 2/5 Medio día Aptitud para principiantes 84/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 1-3 horas Aptitud para principiantes 88/100
ROCm/rocm-libraries#12703 ·
Los mantenedores suelen responder en 2 días
-
bug
Dificultad 1/5 1-3 horas Aptitud para principiantes 88/100
isl-org/Open3D#7585 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Feature request
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
qbittorrent/qBittorrent#24975 ·
Los mantenedores suelen responder en 3 días