Hacktoberfest 2026: le issue che i maintainer hanno segnato per ottobre, aperte e adatte ai principianti. Sfoglia le issue Hacktoberfest

Python: Bug: split_plaintext_paragraph / split_markdown_paragraph can return a chunk larger than max_tokens

Aperta Adatta ai principianti
#14,566 0 commenti 0 reazioni 0 assegnatari Vedi su GitHub

I maintainer di solito rispondono entro 4 giorni

@xThreeh ci sta già lavorando.

Dal 7/10/2026.

  • #14567 di @xThreeh — aperta

Valutazione

Difficoltà
2/5
Tempo stimato
1-3 ore
Idoneità per principianti
75/100
Tipo di issue
Bug
Chiarezza
Specificata chiaramente
Stato di attività
Attiva
Stack tecnologico
python
Ambito
backend

Direzione di ricerca

Leggi python/semantic_kernel/text/text_chunker.py, soprattutto _split_text_paragraph in prossimità delle righe collegate, ed esamina tests/unit/text/test_text_chunker.py. Esegui prima i test del text chunker; aggiorna il comportamento di unione e i risultati attesi interessati, in modo che l’ultimo chunk venga unito solo quando token_counter conferma che ci sta. Il lavoro è completato quando nessun chunk restituito supera max_tokens e i test pertinenti passano.

Scritto dal modello di indicizzazione a partire dal testo della issue.

Descrizione

python triage

Describe the bug

split_plaintext_paragraph and split_markdown_paragraph can return a paragraph whose token count is above max_tokens.

At the end of _split_text_paragraph, a short last paragraph is merged into the previous one. The merge check counts words (len(paragraph.split(" "))) instead of calling token_counter:

https://github.com/microsoft/semantic-kernel/blob/cc8a15fa356f02dcb7bc64999392ca02f3167312/python/semantic_kernel/text/text_chunker.py#L139-L150

A word is usually more than one token, so the merged paragraph can go over the limit even though the word count is under it. This matters when max_tokens is the embedding model's input limit.

To Reproduce

With the default token counter, using the input of the existing test_split_text_paragraph_evenly test:

from semantic_kernel.text import split_plaintext_paragraph

text = [
    "This is a test of the emergency broadcast system. This is only a test.",
    "We repeat, this is only a test. A unit test.",
    "A small note. And another. And once again. Seriously, this is the end. We're finished. All set. Bye.",
    "Done.",
]
chunks = split_plaintext_paragraph(text, 15)
print([len(c) // 4 for c in chunks])  # [12, 5, 11, 10, 16] -> last chunk is 16 tokens, limit is 15

With a real tokenizer (tiktoken, cl100k_base) passed as token_counter, on generated multi-line English text and limits from 64 to 512, 955 of 8000 runs (about 12%) returned an over-limit chunk. The worst one was 314 tokens for max_tokens=256.

Expected behavior

No returned paragraph is above max_tokens. The last paragraph should only be merged into the previous one if the merged text fits, measured with token_counter.

Platform

  • Language: Python
  • Source: main at cc8a15fa3
  • OS: Windows 11, Python 3.12

Additional context

Eight existing unit tests in tests/unit/text/test_text_chunker.py expect an over-limit last chunk (for example 16 tokens with max_token_per_line = 15). A fix would change their expected output, so I wanted to raise it here first. I have a small fix with tests ready and I'm happy to open the PR.

Lingua principale
C#
Stelle
28.6k
Fork
4.8k
Merge medio
2g 1h
PR unite (30g)
20

Preparare l'ambiente

Apri in Codespaces

Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.

Come iniziare

  1. Leggi tutta la issue e poi la guida ai contributi del progetto.
  2. Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
  3. Fai un fork del repository e lavora su un branch.
  4. Apri una pull request che faccia riferimento al numero della issue.

Altre issue di microsoft/semantic-kernel

Tutte le issue di microsoft/semantic-kernel

Issue simili

Altre issue su C#

Ricevi le nuove issue nella tua casella

Un breve riepilogo di issue GitHub adatte ai principianti.