Python: Bug: split_plaintext_paragraph / split_markdown_paragraph can return a chunk larger than max_tokens
I maintainer di solito rispondono entro 4 giorni
Valutazione
- Difficoltà
- 2/5
- Tempo stimato
- 1-3 ore
- Idoneità per principianti
- 75/100
Direzione di ricerca
Leggi python/semantic_kernel/text/text_chunker.py, soprattutto _split_text_paragraph in prossimità delle righe collegate, ed esamina tests/unit/text/test_text_chunker.py. Esegui prima i test del text chunker; aggiorna il comportamento di unione e i risultati attesi interessati, in modo che l’ultimo chunk venga unito solo quando token_counter conferma che ci sta. Il lavoro è completato quando nessun chunk restituito supera max_tokens e i test pertinenti passano.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Describe the bug
split_plaintext_paragraph and split_markdown_paragraph can return a paragraph whose token count is above max_tokens.
At the end of _split_text_paragraph, a short last paragraph is merged into the previous one. The merge check counts words (len(paragraph.split(" "))) instead of calling token_counter:
A word is usually more than one token, so the merged paragraph can go over the limit even though the word count is under it. This matters when max_tokens is the embedding model's input limit.
To Reproduce
With the default token counter, using the input of the existing test_split_text_paragraph_evenly test:
from semantic_kernel.text import split_plaintext_paragraph
text = [
"This is a test of the emergency broadcast system. This is only a test.",
"We repeat, this is only a test. A unit test.",
"A small note. And another. And once again. Seriously, this is the end. We're finished. All set. Bye.",
"Done.",
]
chunks = split_plaintext_paragraph(text, 15)
print([len(c) // 4 for c in chunks]) # [12, 5, 11, 10, 16] -> last chunk is 16 tokens, limit is 15
With a real tokenizer (tiktoken, cl100k_base) passed as token_counter, on generated multi-line English text and limits from 64 to 512, 955 of 8000 runs (about 12%) returned an over-limit chunk. The worst one was 314 tokens for max_tokens=256.
Expected behavior
No returned paragraph is above max_tokens. The last paragraph should only be merged into the previous one if the merged text fits, measured with token_counter.
Platform
- Language: Python
- Source:
mainat cc8a15fa3 - OS: Windows 11, Python 3.12
Additional context
Eight existing unit tests in tests/unit/text/test_text_chunker.py expect an over-limit last chunk (for example 16 tokens with max_token_per_line = 15). A fix would change their expected output, so I wanted to raise it here first. I have a small fix with tests ready and I'm happy to open the PR.
- Lingua principale
- C#
- Stelle
- 28.6k
- Fork
- 4.8k
- Merge medio
- 2g 1h
- PR unite (30g)
- 20
Preparare l'ambiente
Avvia il container di sviluppo del progetto nel browser, con il tuo account GitHub.
- Nessun Dockerfile né file Docker Compose
- Ha un modello di pull request
- Leggi la guida per i contributori
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/semantic-kernel
-
.Net: gpt-image-1 is the default image model in the .NET OpenAI connector, and OpenAI shuts it down on October 23Forse già presa @nightcityblade l’ha presa 4 giorni fa. Aperta.NET triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 73/100
microsoft/semantic-kernel#14526 · 1 commento ·
I maintainer di solito rispondono entro 4 giorni
-
Python: VolatileMemoryStore.get_batch and get_nearest_matches ignore with_embeddings=False (deepcopy result is discarded)Forse già presa @VANDRANKI l’ha presa 6 giorni fa. Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
microsoft/semantic-kernel#14522 ·
I maintainer di solito rispondono entro 4 giorni
-
Python: VolatileMemoryStore.get_nearest_match returns an un-awaited coroutine instead of a (MemoryRecord, score) tupleForse già presa @VANDRANKI l’ha presa 6 giorni fa. Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
microsoft/semantic-kernel#14521 ·
I maintainer di solito rispondono entro 4 giorni
-
.Net: Bug: BinaryContent does not decode the %xx escapes of a non-base64 data URIForse già presa @Laurianti l’ha presa 6 giorni fa. Aperta.NET triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
microsoft/semantic-kernel#14518 ·
I maintainer di solito rispondono entro 4 giorni
-
Python: FunctionCallContent.combine_arguments drops a streamed "{}" chunk, producing invalid JSON argumentsForse già presa @VANDRANKI l’ha presa 7 giorni fa. Apertapython triage
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
microsoft/semantic-kernel#14512 · 2 commenti ·
I maintainer di solito rispondono entro 4 giorni
Tutte le issue di microsoft/semantic-kernel
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
activescott/lessmsi#306 ·
-
triage
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 82/100
rjmurillo/moq.analyzers#1384 ·
-
Variables passed to Compensated are not set on the routing slipForse di nuovo libera Una pull request per questa issue è stata chiusa senza essere unita. Aperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 75/100
MassTransit/MassTransit#6249 ·
-
security
Difficoltà 2/5 1-3 ore Idoneità per principianti 82/100
Sendspin/sendspin-dotnet#339 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 88/100
owasp-dep-scan/dosai#79 ·
I maintainer di solito rispondono entro 1 giorno