[Bug]: ollama branch overwrites extra_body, so GRAPHIFY_DISABLE_THINKING and providers.json extra_body are silently inert; plus no JSON mode and no seed, so extraction is not reproducible (73 vs 34 nodes on identical input)
Los mantenedores suelen responder en 1 día
@oleksii-tumanov ya está trabajando en esto.
Desde el 6/10/2026.
- #4168 de @oleksii-tumanov — abierto
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 38/100
- Tipo de issue
- Error
- Claridad
- Bastante claro
- Estado de actividad
- Activo
- Stack tecnológico
- python
- Área
- ai, backend-api-design
Línea de trabajo
Start in llm.py, especially _call_openai_compat around lines 1419-1473, and inspect the Ollama request construction. Reproduce the extra_body behavior with GRAPHIFY_DISABLE_THINKING, then compare the response_format and seed requests described in the issue. Done means the requested Ollama controls and supported prompt customization have clear, documented behavior without breaking existing provider options.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Pre-flight checks
- I have checked the Troubleshooting section in the README
- I searched existing issues. #1621 (closed) added
GRAPHIFY_DISABLE_THINKING; nobody noticed the ollama branch overwrites it. #3680 asks forextra_headersin providers.json, which is adjacent but different. No issue covers reproducibility or a seed.
What happened?
We run graphify nightly over a ~21k-document corpus with a local model (ollama + qwen2.5:14b, 3× Tesla P40). To get usable extraction we have ended up wrapping the openai.OpenAI client in our own code to inject two fields into every request. We would much rather not maintain a monkeypatch of a third-party library's private call path, so here is why we have to, with the measurements.
Three separate things on the ollama branch, one of which is a plain bug.
1. The ollama branch replaces extra_body, so GRAPHIFY_DISABLE_THINKING is silently inert there
_call_openai_compat builds extra_body through a chain of if/elif (llm.py:1419-1427):
if extra_body is not None:
kwargs["extra_body"] = extra_body # custom provider's own extra_body
elif "moonshot" in base_url:
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
elif _thinking_disabled_via_env(): # GRAPHIFY_DISABLE_THINKING
kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
and then, further down, unconditionally for ollama (llm.py:1473):
kwargs["extra_body"] = {"options": {"num_ctx": num_ctx}, "keep_alive": keep_alive}
That is an assignment, not a merge. So on the ollama backend:
GRAPHIFY_DISABLE_THINKING=1has no effect at all — it is documented as a user choice in_thinking_disabled_via_env's own docstring, and on this backend it is discarded a few lines later;- a custom provider's
extra_bodyfromproviders.jsonis discarded too, although the comment atllm.py:1416-1418explicitly promises the opposite ("When supplied, it wins over the moonshot default — the user has explicitly chosen the request shape for their endpoint"). The example in that comment ischat_template_kwargs.enable_thinking=falsefor a self-hosted Qwen, i.e. exactly the local-model case this branch breaks.
Fix is a merge instead of a replacement:
_eb = dict(kwargs.get("extra_body") or {})
_eb.setdefault("options", {})["num_ctx"] = num_ctx
_eb["keep_alive"] = keep_alive
kwargs["extra_body"] = _eb
2. No way to force JSON mode, and the difference is large
Ollama's OpenAI-compatible endpoint supports response_format={"type": "json_object"}, which constrains generation to valid JSON. graphify does not send it on any path (grep -rn response_format in 0.9.73 finds only _json_object_candidates, the tolerant parser), and there is no flag to ask for it.
Measured on the same two documents, same prompt, same model, only that field added:
| nodes | edges | output tokens | |
|---|---|---|---|
without response_format |
39 | 30 | 3,238 |
with response_format |
73 | 73 | 7,856 |
Tolerant parsing recovers a JSON object from prose, but it cannot recover the content the model never emitted because it drifted out of JSON.
3. No seed, so extraction is not reproducible — which makes any A/B comparison meaningless
There is no seed anywhere in the package (grep -c '"seed"' llm.py → 0). Ollama accepts options.seed, and with temperature: 0 already in the ollama BACKENDS entry a fixed seed would make extraction reproducible.
Without it, we measured this — identical input, identical prompt, identical configuration, same server, consecutive runs:
run 1 : 73 nodes / 73 edges / 7,856 output tokens / 1,125 s
run 2 : 34 nodes / 18 edges / 2,404 output tokens / 321 s
Input was byte-identical (6,715 prompt tokens, 2,687-character system prompt, verified equal), so this is not truncation — the generation itself varies by 3.3× in length and 4× in edge count. Environment checked to rule out the obvious causes: temperature: 0 is configured and sent, OLLAMA_NUM_PARALLEL=1, OLLAMA_FLASH_ATTENTION=0, OLLAMA_KV_CACHE_TYPE=f16.
The consequence is not only ours: graphify's own quality comparisons on a local backend cannot be trusted at n=1. A seed (and surfacing it in the docs) would fix that for everyone.
Why we are reporting this rather than just patching it
We also have to replace _EXTRACTION_SYSTEM, a private module global, because our corpus is Italian prose rather than code and the shipped prompt extracts ~1 node per document on it. There is no supported way to extend or replace it: no environment variable, no parameter, and --prompt-file only selects the cache namespace, not the prompt. So every integration like ours ends up reaching into module internals — and silently loses whatever the next release adds to that prompt. (We just caught ourselves dropping 0.9.73's own <untrusted_source> SECURITY paragraph that way, which would have turned our corpus into an instruction channel. That one is on us, but a supported hook would have prevented it.)
A documented extension point — a parameter, an env var, or a documented "append to the system prompt" hook — would let integrators stop monkeypatching private internals.
Steps to reproduce
Point 1, no model needed — inspect the request graphify builds:
GRAPHIFY_DISABLE_THINKING=1 python - <<'PY'
# observe that kwargs["extra_body"] ends up as
# {"options": {...}, "keep_alive": "..."}
# with no "thinking" key, on the ollama branch
PY
or simply read llm.py:1419-1427 against llm.py:1473.
Point 3, two identical runs against a local ollama:
export OLLAMA_BASE_URL=http://127.0.0.1:11436/v1
export GRAPHIFY_TOKEN_BUDGET=7000
for i in 1 2; do
python -c '
from pathlib import Path
from graphify.llm import extract_corpus_parallel
r = extract_corpus_parallel([Path("a.md"), Path("b.md")], backend="ollama",
model="qwen2.5:14b", root=Path("."),
max_concurrency=1, token_budget=7000)
print(len(r["nodes"]), len(r["edges"]), r["input_tokens"], r["output_tokens"])'
done
Error output or graph output
# point 3, two consecutive runs, identical input (6,715 prompt tokens both times):
73 73 6715 7856
34 18 6715 2404
# point 2, same documents and prompt, only response_format added:
without: 39 nodes, 30 edges, 3238 output tokens
with : 73 nodes, 73 edges, 7856 output tokens
# point 1 produces no error at all — the env var is simply discarded. That is the report.
Graphify version
0.9.73
Operating System
Linux
Additional context
Backend ollama, model qwen2.5:14b (Q4_K_M), 3× Tesla P40, Ubuntu, OLLAMA_CONTEXT_LENGTH=11000. Corpus ~21k documents, mostly Italian prose. Related to our #3982 (the cache/source_file side of the same local-model path) and to the closed #1621, which introduced the env var that point 1 shows is inert here.
- Lenguaje dominante
- Python
- Estrellas
- 124k
- Forks
- 11.9k
- Métricas de merge de PR
- Sin PR fusionados en 30 d
Preparar el entorno
Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de Graphify-Labs/graphify
-
[Bug]: `graphify export svg` writes a graph.svg that is not well-formed XML when a label contains a control characterPosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
Graphify-Labs/graphify#4241 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Graphify-Labs/graphify#3763 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Graphify-Labs/graphify#3611 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Nix supportAbierto
Dificultad 2/5 1-3 horas Aptitud para principiantes 62/100
Graphify-Labs/graphify#3193 · 2 reacciones ·
Los mantenedores suelen responder en 1 día
-
Dificultad 1/5 Menos de una hora Aptitud para principiantes 92/100
Graphify-Labs/graphify#2871 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
Todos los issues de Graphify-Labs/graphify
Issues similares
-
HTML: <template> content is extracted as document textPosiblemente ocupada @ryanmeowy la tomó hoy. Abiertobug html
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
docling-project/docling#4714 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
[BUG] Qdrant RAG client applies score_threshold to raw cosine similarity, not the 0-1 score it returnsPosiblemente ocupada @roydonsequeira la tomó hoy. Abiertobug
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 72/100
ashhart/TensorFold#536 ·
Los mantenedores suelen responder en 1 día
-
area/install-update comp/cli duplicate P2 python:uv sweeper:risk-compatibility type/bug
Dificultad 1/5 Menos de una hora Aptitud para principiantes 62/100
NousResearch/hermes-agent#135440 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
bug
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
Deepak3699/Ai_Mentor#244 ·
Los mantenedores suelen responder en 1 día