Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Python: text_chunker's was_split flag can never be True, so every separator tier always runs (dead early-exit)

Đang mở Phù hợp với người mới
#14,490 1 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
2/5
Thời gian dự kiến
1-3 giờ
Mức phù hợp với người mới
82/100
Loại issue
Lỗi
Độ rõ ràng
Đặc tả rõ ràng
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
ai

Hướng nghiên cứu

Bắt đầu trong python/semantic_kernel/text/text_chunker.py bằng cách đọc _split_str_lines(), _split_str() và _split_list(), sau đó chạy bản tái hiện split_plaintext_lines được cung cấp để quan sát số lần gọi ở mỗi tầng. Được xem là hoàn tất khi một lần tách thực sự đánh dấu flag, vòng lặp thoát sau tầng dấu phân cách thành công và đầu ra vẫn không thay đổi trong khi tránh các lần gọi dư thừa.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

python triage

What happens

text_chunker.py's was_split / input_was_split flag, used to short-circuit the separator-tier loop once a good split has been found, can never actually become True. It is initialized to False in every function that returns it, and is only ever combined with other values that are (by the same argument) always False. As a result, split_plaintext_lines(), split_markdown_lines(), split_plaintext_paragraph() and split_markdown_paragraph() always run through every separator tier (newline, ., ?!, ;, :, ,, brackets, space, -, and finally the unconditional half-length hard split) on every already-split segment, instead of stopping once an earlier, more meaningful tier has produced a valid split.

Where

python/semantic_kernel/text/text_chunker.py

  • _split_str_lines() (used by both split_plaintext_lines and split_markdown_lines):
lines: list[str] = []
was_split = False
for split_option in separators:
    if not lines:
        lines, was_split = _split_str(...)
    else:
        lines, was_split = _split_list(...)
    if was_split:
        break  # pragma: no cover

Note the maintainers' own # pragma: no cover on the break — their coverage tooling has already flagged that this branch never executes as True, which matches the root cause below.

  • _split_str():
def _split_str(text, max_tokens, separators, trim, token_counter=_token_counter):
    input_was_split = False
    if not text:
        return [], input_was_split
    ...
    if token_counter(text) <= max_tokens:
        return text_as_is, input_was_split          # False
    ...
    elif set(separators) & set(text) and len(text) > 2:
        ...
    else:
        return text_as_is, input_was_split          # False
    if 0 < cutpoint < len(text):
        lines = []
        for text_part in [text[:cutpoint], text[cutpoint:]]:
            split, has_split = _split_str(text_part, max_tokens, separators, trim, token_counter)
            lines.extend(split)
            input_was_split = input_was_split or has_split   # OR of two values that are False by induction
    else:
        return text_as_is, input_was_split          # False
    return lines, input_was_split

No line in this function (or in _split_list(), which OR's _split_str's always-False result over a list) ever assigns True to this flag. By induction over the recursion, every base case returns False, so every recursive case (which only ORs base-case results together) also returns False. The flag is dead: it is computed, threaded through three functions, and checked, but can never be True.

Why this matters

The whole point of TEXT_SPLIT_OPTIONS / MD_SPLIT_OPTIONS being an ordered list (prefer splitting on newlines, then sentence punctuation, then softer punctuation, then whitespace, then a hard character cut as last resort) is defeated: because was_split is always False, _split_str_lines never breaks early and always falls through every remaining tier for every segment already produced, even when a stronger separator already produced a fully within-budget split. This is a real perf/efficiency defect for a function whose entire purpose is chunking large documents for embeddings/RAG (SKPY-CHUNK), since it does up to ~9x more recursive splitting work than the algorithm is designed to do.

Repro

import semantic_kernel.text.text_chunker as tc

call_count = {"split_str": 0, "split_list": 0}
orig_split_str, orig_split_list = tc._split_str, tc._split_list

def counting_split_str(*a, **k):
    call_count["split_str"] += 1
    return orig_split_str(*a, **k)

def counting_split_list(*a, **k):
    call_count["split_list"] += 1
    return orig_split_list(*a, **k)

tc._split_str, tc._split_list = counting_split_str, counting_split_list

text = "\n".join(f"Line number {i} is short." for i in range(30))
lines = tc.split_plaintext_lines(text, max_token_per_line=50)
print("output lines:", len(lines))
print("call counts:", call_count)

Output on semantic-kernel==1.44.1 (byte-identical to current main):

output lines: 4
call counts: {'split_str': 43, 'split_list': 9}

The input already fully fits (30 short lines, each well under the 50-token budget) after the very first tier (["\n", "\r"]) splits it. If the early-exit worked as designed, only tier 1 would run. Instead, _split_list is invoked once for every one of the remaining 9 tiers (., ?!, ;, :, ,, brackets, space, -, None), each re-scanning the already-finished lines.

You can also confirm the flag is unconditionally False directly:

from semantic_kernel.text.text_chunker import _split_str
lines, was_split = _split_str("A sentence that is long enough to need splitting here.", max_tokens=5, separators=["."], trim=True)
print(was_split)  # False, even though a real split happened

Expected vs actual

  • Expected: once a separator tier fully resolves a segment (or the recursive split materially succeeds), _split_str_lines should stop trying weaker tiers for that segment, per the tiered design and the if was_split: break intent.
  • Actual: was_split can never be True, so every tier always runs against every already-split segment, regardless of how the input was chunked.

Suggested direction

In _split_str, the branch that actually performs a split (if 0 < cutpoint < len(text):) needs to record that a split happened at this level, e.g. input_was_split = True before/alongside the recursive OR, not only OR the recursive children's flags. The current code never marks "a split occurred here," only "a split occurred deeper," which can never be true from any base case.

Environment

  • semantic-kernel 1.44.1 (PyPI), verified byte-for-byte identical against python/semantic_kernel/text/text_chunker.py on the current main branch.
  • Python 3.13, Windows.
Ngôn ngữ chính
C#
Star
28.6k
Fork
4.8k
Merge trung bình
13 giờ 24 phút
Pull request đã merge (30 ngày)
11

Chuẩn bị môi trường

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/semantic-kernel

Tất cả issue của microsoft/semantic-kernel

Issue tương tự

Thêm issue về C#

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.