Python: text_chunker's was_split flag can never be True, so every separator tier always runs (dead early-exit)
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 82/100
Direzione di ricerca
Inizia in python/semantic_kernel/text/text_chunker.py leggendo _split_str_lines(), _split_str() e _split_list(), quindi esegui la riproduzione fornita di split_plaintext_lines per osservare il numero di chiamate per livello. È completato quando una suddivisione effettiva imposta il flag, il ciclo esce dopo il livello di separatore riuscito e l’output rimane invariato evitando le chiamate ridondanti.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
What happens
text_chunker.py's was_split / input_was_split flag, used to short-circuit the separator-tier loop once a good split has been found, can never actually become True. It is initialized to False in every function that returns it, and is only ever combined with other values that are (by the same argument) always False. As a result, split_plaintext_lines(), split_markdown_lines(), split_plaintext_paragraph() and split_markdown_paragraph() always run through every separator tier (newline, ., ?!, ;, :, ,, brackets, space, -, and finally the unconditional half-length hard split) on every already-split segment, instead of stopping once an earlier, more meaningful tier has produced a valid split.
Where
python/semantic_kernel/text/text_chunker.py
_split_str_lines()(used by bothsplit_plaintext_linesandsplit_markdown_lines):
lines: list[str] = []
was_split = False
for split_option in separators:
if not lines:
lines, was_split = _split_str(...)
else:
lines, was_split = _split_list(...)
if was_split:
break # pragma: no cover
Note the maintainers' own # pragma: no cover on the break — their coverage tooling has already flagged that this branch never executes as True, which matches the root cause below.
_split_str():
def _split_str(text, max_tokens, separators, trim, token_counter=_token_counter):
input_was_split = False
if not text:
return [], input_was_split
...
if token_counter(text) <= max_tokens:
return text_as_is, input_was_split # False
...
elif set(separators) & set(text) and len(text) > 2:
...
else:
return text_as_is, input_was_split # False
if 0 < cutpoint < len(text):
lines = []
for text_part in [text[:cutpoint], text[cutpoint:]]:
split, has_split = _split_str(text_part, max_tokens, separators, trim, token_counter)
lines.extend(split)
input_was_split = input_was_split or has_split # OR of two values that are False by induction
else:
return text_as_is, input_was_split # False
return lines, input_was_split
No line in this function (or in _split_list(), which OR's _split_str's always-False result over a list) ever assigns True to this flag. By induction over the recursion, every base case returns False, so every recursive case (which only ORs base-case results together) also returns False. The flag is dead: it is computed, threaded through three functions, and checked, but can never be True.
Why this matters
The whole point of TEXT_SPLIT_OPTIONS / MD_SPLIT_OPTIONS being an ordered list (prefer splitting on newlines, then sentence punctuation, then softer punctuation, then whitespace, then a hard character cut as last resort) is defeated: because was_split is always False, _split_str_lines never breaks early and always falls through every remaining tier for every segment already produced, even when a stronger separator already produced a fully within-budget split. This is a real perf/efficiency defect for a function whose entire purpose is chunking large documents for embeddings/RAG (SKPY-CHUNK), since it does up to ~9x more recursive splitting work than the algorithm is designed to do.
Repro
import semantic_kernel.text.text_chunker as tc
call_count = {"split_str": 0, "split_list": 0}
orig_split_str, orig_split_list = tc._split_str, tc._split_list
def counting_split_str(*a, **k):
call_count["split_str"] += 1
return orig_split_str(*a, **k)
def counting_split_list(*a, **k):
call_count["split_list"] += 1
return orig_split_list(*a, **k)
tc._split_str, tc._split_list = counting_split_str, counting_split_list
text = "\n".join(f"Line number {i} is short." for i in range(30))
lines = tc.split_plaintext_lines(text, max_token_per_line=50)
print("output lines:", len(lines))
print("call counts:", call_count)
Output on semantic-kernel==1.44.1 (byte-identical to current main):
output lines: 4
call counts: {'split_str': 43, 'split_list': 9}
The input already fully fits (30 short lines, each well under the 50-token budget) after the very first tier (["\n", "\r"]) splits it. If the early-exit worked as designed, only tier 1 would run. Instead, _split_list is invoked once for every one of the remaining 9 tiers (., ?!, ;, :, ,, brackets, space, -, None), each re-scanning the already-finished lines.
You can also confirm the flag is unconditionally False directly:
from semantic_kernel.text.text_chunker import _split_str
lines, was_split = _split_str("A sentence that is long enough to need splitting here.", max_tokens=5, separators=["."], trim=True)
print(was_split) # False, even though a real split happened
Expected vs actual
- Expected: once a separator tier fully resolves a segment (or the recursive split materially succeeds),
_split_str_linesshould stop trying weaker tiers for that segment, per the tiered design and theif was_split: breakintent. - Actual:
was_splitcan never beTrue, so every tier always runs against every already-split segment, regardless of how the input was chunked.
Suggested direction
In _split_str, the branch that actually performs a split (if 0 < cutpoint < len(text):) needs to record that a split happened at this level, e.g. input_was_split = True before/alongside the recursive OR, not only OR the recursive children's flags. The current code never marks "a split occurred here," only "a split occurred deeper," which can never be true from any base case.
Environment
semantic-kernel1.44.1 (PyPI), verified byte-for-byte identical againstpython/semantic_kernel/text/text_chunker.pyon the currentmainbranch.- Python 3.13, Windows.
- Lingua principale
- C#
- Stelle
- 28.6k
- Fork
- 4.8k
- Merge medio
- 12h 20m
- PR unite (30g)
- 12
Preparare l'ambiente
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/semantic-kernel
-
python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
microsoft/semantic-kernel#14491 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Python: [Python] structured_outputs_transform reuses ChatHistory across calls (prompt pollution)Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
microsoft/semantic-kernel#14483 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
.NET python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 74/100
microsoft/semantic-kernel#14482 · 2 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
Python: [Python] as_agent_framework_tool drops parameter defaults (optionals become required)Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
microsoft/semantic-kernel#14481 ·
I maintainer di solito rispondono entro 2 giorni
-
python triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
microsoft/semantic-kernel#14460 ·
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di microsoft/semantic-kernel
Issue simili
-
area-ai untriaged
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
dotnet/extensions#7790 ·
I maintainer di solito rispondono entro 1 giorno
-
P2 testing
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 90/100
I maintainer di solito rispondono entro 1 giorno
-
area-Infrastructure-coreclr os-ios os-maccatalyst os-tvos untriaged
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
dotnet/runtime#134766 · 3 commenti ·
I maintainer di solito rispondono entro 1 giorno
-
0 - Backlog Bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
BrighterCommand/Brighter#4444 ·
I maintainer di solito rispondono entro 1 giorno
-
bug
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 1 giorno