Log values inside lists and dicts skip the size bound from #12849
I maintainer di solito rispondono entro 1 giorno
@julian-risch ci sta già lavorando.
Dal 28/9/2026.
Valutazione
- Difficoltà
- 3/5
- Tempo stimato
- 1-2 giorni
- Idoneità per principianti
- 65/100
Direzione di ricerca
Look at the bound_event_dict_values function in the logging configuration (likely in haystack/logging/ or haystack/utils/). The issue is that it only bounds top-level values. Start by understanding the current implementation and the test case provided in the issue. Then, implement a recursive walker for lists, tuples, and dicts that applies the same bounding logic, converting nested exceptions with str and truncating containers that exceed MAX_LOG_VALUE_LENGTH. Ensure a depth limit to avoid infinite recursion. Run the reproduction script to verify the fix.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Describe the bug
bound_event_dict_values (added in #12849) only bounds top-level values. When a value is a list or dict, it goes to the renderer unchanged. So the renderer still falls back to repr() for anything inside it, and neither the container nor its items get truncated.
This leaves two gaps:
- A
UnicodeDecodeErrorinside a list still puts its whole buffer into the log line. That is the leak #12849 fixed for a bare exception. - A long list or dict is written out in full. This one happens in Haystack's own code: when an LLM reply is missing expected keys,
_parse_dict_from_jsoninhaystack/utils/misc.pylogskeys=list(parsed_json.keys()). Theeventstring gets truncated at 4096 characters, but thekeysfield does not.
Expected behavior
The limit applies to the rendered size of every field, the same way it does for top-level strings and exceptions, so exceptions nested in a container also go through str.
To Reproduce
On main (3c0dc45b), with structlog installed:
import contextlib
import io
import json
import logging
from haystack import logging as haystack_logging
from haystack.utils.misc import _parse_dict_from_json
def last_json_line(emit):
buffer = io.StringIO()
with contextlib.redirect_stderr(buffer):
haystack_logging.configure_logging(use_json=True)
emit()
return buffer.getvalue().strip().splitlines()[-1]
# 1. In-tree call: the `keys` field is a list, so it is not bounded.
reply = json.dumps({f"key_{i:05d}": i for i in range(3000)})
line = last_json_line(lambda: _parse_dict_from_json(reply, expected_keys=["score"], raise_on_failure=False))
print("missing-keys warning:", len(line), "chars")
# 2. The UnicodeDecodeError from #12849, wrapped in a list.
try:
(b"%PDF-1.7\r%\xe2\xe3\xcf\xd3" + b"A" * 100_000).decode("utf-8")
except UnicodeDecodeError as error:
decode_error = error
line = last_json_line(lambda: logging.getLogger("haystack.repro").warning("Conversion failed", extra={"error": decode_error}))
print("bare exception: ", len(line), "chars")
line = last_json_line(lambda: logging.getLogger("haystack.repro").warning("Conversion failed", extra={"errors": [decode_error]}))
print("exception in a list: ", len(line), "chars")
Output (same on two runs):
missing-keys warning: 43286 chars
bare exception: 247 chars
exception in a list: 100253 chars
Additional context
One way to fix it: walk list, tuple and dict values in bound_event_dict_values, convert nested exceptions with str, and if the rendered container is still over MAX_LOG_VALUE_LENGTH, replace it with a truncated string. The walk needs a depth limit so self-referencing containers can't recurse forever. I can open a PR for this if that approach works for you.
FAQ Check
- Have you had a look at our new FAQ page?
System:
- OS: Linux
- Haystack version: main (3c0dc45b)
- Lingua principale
- Python
- Stelle
- 26.6k
- Fork
- 3.2k
- Merge medio
- 1g 13h
- PR unite (30g)
- 263
Preparare l'ambiente
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di deepset-ai/haystack
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
deepset-ai/haystack#13029 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
deepset-ai/haystack#13022 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
deepset-ai/haystack#12994 · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
MarkdownHeaderSplitter treats headings inside longer closing fences as headersForse già presa @julian-risch l’ha presa 2 giorni fa. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
deepset-ai/haystack#12954 · 2 commenti · 1 assegnatario ·
I maintainer di solito rispondono entro 1 giorno
-
P3
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
deepset-ai/haystack#12945 ·
I maintainer di solito rispondono entro 1 giorno
Tutte le issue di deepset-ai/haystack
Issue simili
-
Claiming namespace `apoint`Apertanamespace operations
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
EclipseFdn/open-vsx.org#13573 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
collective/icalendar#1854 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100
rancher/rancher-ai-agent#412 ·
I maintainer di solito rispondono entro 6 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
TUDelftGeodesy/DePSI#134 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
HenriquesLab/rxiv-maker#335 ·