Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

[Bug]: ollama branch overwrites extra_body, so GRAPHIFY_DISABLE_THINKING and providers.json extra_body are silently inert; plus no JSON mode and no seed, so extraction is not reproducible (73 vs 34 nodes on identical input)

Abierto
#3,988 1 comentario 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 1 día

@oleksii-tumanov ya está trabajando en esto.

Desde el 6/10/2026.

  • #4168 de @oleksii-tumanov — abierto

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
38/100
Tipo de issue
Error
Claridad
Bastante claro
Estado de actividad
Activo
Stack tecnológico
python

Línea de trabajo

Start in llm.py, especially _call_openai_compat around lines 1419-1473, and inspect the Ollama request construction. Reproduce the extra_body behavior with GRAPHIFY_DISABLE_THINKING, then compare the response_format and seed requests described in the issue. Done means the requested Ollama controls and supported prompt customization have clear, documented behavior without breaking existing provider options.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

Pre-flight checks
  • I have checked the Troubleshooting section in the README
  • I searched existing issues. #1621 (closed) added GRAPHIFY_DISABLE_THINKING; nobody noticed the ollama branch overwrites it. #3680 asks for extra_headers in providers.json, which is adjacent but different. No issue covers reproducibility or a seed.
What happened?

We run graphify nightly over a ~21k-document corpus with a local model (ollama + qwen2.5:14b, 3× Tesla P40). To get usable extraction we have ended up wrapping the openai.OpenAI client in our own code to inject two fields into every request. We would much rather not maintain a monkeypatch of a third-party library's private call path, so here is why we have to, with the measurements.

Three separate things on the ollama branch, one of which is a plain bug.


1. The ollama branch replaces extra_body, so GRAPHIFY_DISABLE_THINKING is silently inert there

_call_openai_compat builds extra_body through a chain of if/elif (llm.py:1419-1427):

if extra_body is not None:
    kwargs["extra_body"] = extra_body            # custom provider's own extra_body
elif "moonshot" in base_url:
    kwargs["extra_body"] = {"thinking": {"type": "disabled"}}
elif _thinking_disabled_via_env():               # GRAPHIFY_DISABLE_THINKING
    kwargs["extra_body"] = {"thinking": {"type": "disabled"}}

and then, further down, unconditionally for ollama (llm.py:1473):

    kwargs["extra_body"] = {"options": {"num_ctx": num_ctx}, "keep_alive": keep_alive}

That is an assignment, not a merge. So on the ollama backend:

  • GRAPHIFY_DISABLE_THINKING=1 has no effect at all — it is documented as a user choice in _thinking_disabled_via_env's own docstring, and on this backend it is discarded a few lines later;
  • a custom provider's extra_body from providers.json is discarded too, although the comment at llm.py:1416-1418 explicitly promises the opposite ("When supplied, it wins over the moonshot default — the user has explicitly chosen the request shape for their endpoint"). The example in that comment is chat_template_kwargs.enable_thinking=false for a self-hosted Qwen, i.e. exactly the local-model case this branch breaks.

Fix is a merge instead of a replacement:

    _eb = dict(kwargs.get("extra_body") or {})
    _eb.setdefault("options", {})["num_ctx"] = num_ctx
    _eb["keep_alive"] = keep_alive
    kwargs["extra_body"] = _eb
2. No way to force JSON mode, and the difference is large

Ollama's OpenAI-compatible endpoint supports response_format={"type": "json_object"}, which constrains generation to valid JSON. graphify does not send it on any path (grep -rn response_format in 0.9.73 finds only _json_object_candidates, the tolerant parser), and there is no flag to ask for it.

Measured on the same two documents, same prompt, same model, only that field added:

nodes edges output tokens
without response_format 39 30 3,238
with response_format 73 73 7,856

Tolerant parsing recovers a JSON object from prose, but it cannot recover the content the model never emitted because it drifted out of JSON.

3. No seed, so extraction is not reproducible — which makes any A/B comparison meaningless

There is no seed anywhere in the package (grep -c '"seed"' llm.py → 0). Ollama accepts options.seed, and with temperature: 0 already in the ollama BACKENDS entry a fixed seed would make extraction reproducible.

Without it, we measured this — identical input, identical prompt, identical configuration, same server, consecutive runs:

run 1 :  73 nodes /  73 edges /  7,856 output tokens / 1,125 s
run 2 :  34 nodes /  18 edges /  2,404 output tokens /   321 s

Input was byte-identical (6,715 prompt tokens, 2,687-character system prompt, verified equal), so this is not truncation — the generation itself varies by 3.3× in length and 4× in edge count. Environment checked to rule out the obvious causes: temperature: 0 is configured and sent, OLLAMA_NUM_PARALLEL=1, OLLAMA_FLASH_ATTENTION=0, OLLAMA_KV_CACHE_TYPE=f16.

The consequence is not only ours: graphify's own quality comparisons on a local backend cannot be trusted at n=1. A seed (and surfacing it in the docs) would fix that for everyone.


Why we are reporting this rather than just patching it

We also have to replace _EXTRACTION_SYSTEM, a private module global, because our corpus is Italian prose rather than code and the shipped prompt extracts ~1 node per document on it. There is no supported way to extend or replace it: no environment variable, no parameter, and --prompt-file only selects the cache namespace, not the prompt. So every integration like ours ends up reaching into module internals — and silently loses whatever the next release adds to that prompt. (We just caught ourselves dropping 0.9.73's own <untrusted_source> SECURITY paragraph that way, which would have turned our corpus into an instruction channel. That one is on us, but a supported hook would have prevented it.)

A documented extension point — a parameter, an env var, or a documented "append to the system prompt" hook — would let integrators stop monkeypatching private internals.

Steps to reproduce

Point 1, no model needed — inspect the request graphify builds:

GRAPHIFY_DISABLE_THINKING=1 python - <<'PY'
# observe that kwargs["extra_body"] ends up as
#   {"options": {...}, "keep_alive": "..."}
# with no "thinking" key, on the ollama branch
PY

or simply read llm.py:1419-1427 against llm.py:1473.

Point 3, two identical runs against a local ollama:

export OLLAMA_BASE_URL=http://127.0.0.1:11436/v1
export GRAPHIFY_TOKEN_BUDGET=7000
for i in 1 2; do
  python -c '
from pathlib import Path
from graphify.llm import extract_corpus_parallel
r = extract_corpus_parallel([Path("a.md"), Path("b.md")], backend="ollama",
                            model="qwen2.5:14b", root=Path("."),
                            max_concurrency=1, token_budget=7000)
print(len(r["nodes"]), len(r["edges"]), r["input_tokens"], r["output_tokens"])'
done
Error output or graph output
# point 3, two consecutive runs, identical input (6,715 prompt tokens both times):
73 73 6715 7856
34 18 6715 2404

# point 2, same documents and prompt, only response_format added:
without: 39 nodes, 30 edges, 3238 output tokens
with   : 73 nodes, 73 edges, 7856 output tokens

# point 1 produces no error at all — the env var is simply discarded. That is the report.
Graphify version

0.9.73

Operating System

Linux

Additional context

Backend ollama, model qwen2.5:14b (Q4_K_M), 3× Tesla P40, Ubuntu, OLLAMA_CONTEXT_LENGTH=11000. Corpus ~21k documents, mostly Italian prose. Related to our #3982 (the cache/source_file side of the same local-model path) and to the closed #1621, which introduced the env var that point 1 shows is inert here.

Lenguaje dominante
Python
Estrellas
124k
Forks
11.9k
Métricas de merge de PR
Sin PR fusionados en 30 d

Preparar el entorno

Este proyecto no incluye contenedor de desarrollo, Dockerfile ni guía de contribución, así que la configuración corre por tu cuenta: empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de Graphify-Labs/graphify

Todos los issues de Graphify-Labs/graphify

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.