Hacktoberfest 2026: the issues maintainers tagged for October, open and beginner-friendly. Browse Hacktoberfest issues

Python: Bug: split_plaintext_paragraph / split_markdown_paragraph can return a chunk larger than max_tokens

Open Beginner friendly
#14,566 0 comments 0 reactions 0 assignees View on GitHub

Maintainers usually reply within 4 days

@xThreeh is already working on this.

Since Oct 7, 2026.

Assessment

Difficulty
2/5
Estimated time
1-3 hours
Newbie friendliness
75/100
Issue type
Bug
Clarity
Clearly specified
Activity status
Active
Tech stack
python
Domain
backend

Research direction

Read python/semantic_kernel/text/text_chunker.py, especially _split_text_paragraph around the linked lines, and inspect tests/unit/text/test_text_chunker.py. Run the text chunker tests first; update the merge behavior and affected expected outputs so the final chunk is merged only when token_counter confirms it fits. Done means no returned chunk exceeds max_tokens and the relevant tests pass.

Written by the indexing model from the issue text.

Description

python triage

Describe the bug

split_plaintext_paragraph and split_markdown_paragraph can return a paragraph whose token count is above max_tokens.

At the end of _split_text_paragraph, a short last paragraph is merged into the previous one. The merge check counts words (len(paragraph.split(" "))) instead of calling token_counter:

https://github.com/microsoft/semantic-kernel/blob/cc8a15fa356f02dcb7bc64999392ca02f3167312/python/semantic_kernel/text/text_chunker.py#L139-L150

A word is usually more than one token, so the merged paragraph can go over the limit even though the word count is under it. This matters when max_tokens is the embedding model's input limit.

To Reproduce

With the default token counter, using the input of the existing test_split_text_paragraph_evenly test:

from semantic_kernel.text import split_plaintext_paragraph

text = [
    "This is a test of the emergency broadcast system. This is only a test.",
    "We repeat, this is only a test. A unit test.",
    "A small note. And another. And once again. Seriously, this is the end. We're finished. All set. Bye.",
    "Done.",
]
chunks = split_plaintext_paragraph(text, 15)
print([len(c) // 4 for c in chunks])  # [12, 5, 11, 10, 16] -> last chunk is 16 tokens, limit is 15

With a real tokenizer (tiktoken, cl100k_base) passed as token_counter, on generated multi-line English text and limits from 64 to 512, 955 of 8000 runs (about 12%) returned an over-limit chunk. The worst one was 314 tokens for max_tokens=256.

Expected behavior

No returned paragraph is above max_tokens. The last paragraph should only be merged into the previous one if the merged text fits, measured with token_counter.

Platform

  • Language: Python
  • Source: main at cc8a15fa3
  • OS: Windows 11, Python 3.12

Additional context

Eight existing unit tests in tests/unit/text/test_text_chunker.py expect an over-limit last chunk (for example 16 tokens with max_token_per_line = 15). A fix would change their expected output, so I wanted to raise it here first. I have a small fix with tests ready and I'm happy to open the PR.

Dominant language
C#
Stars
28.6k
Forks
4.8k
Avg merge
2d 1h
Merged PRs (30d)
20

Getting set up

Open in Codespaces

Starts the project's dev container in your browser, under your own GitHub account.

First steps

  1. Read the whole issue, then the project's contributing guide.
  2. Comment on the issue to say you are picking it up — it saves two people doing the same work.
  3. Fork the repository and make your change on a branch.
  4. Open a pull request that references the issue number.

More from microsoft/semantic-kernel

All issues in microsoft/semantic-kernel

Similar issues

More C# issues

Get new issues in your inbox

A short digest of beginner-friendly GitHub issues.