Semantic Chunking Chunk Size Bug
@seankim658 is already working on this.
Since Jul 8, 2024.
Assessment
This issue has not been assessed yet.
Description
Llamaindex's SemanticSplitterNodeParser can sometimes produce chunks that are too large for the embedding model. Unfortunately there is no max length option for the semantic chunking to avoid this issue.
Will have to eventually subclass the SemanticSplitterNodeParser and create a two level safety net that will naively split large chunks into sub-chunks in order to stay under the embedding model input token limits.
Reference:
https://github.com/run-llama/llama_index/issues/12270
- Dominant language
- Python
- Stars
- 0
- Forks
- 0
- PR merge metrics
- No merged PRs in 30d
Contributor guide
No contributing guide indexed for this repository
First steps
- Read the whole issue, then the project's contributing guide.
- Comment on the issue to say you are picking it up — it saves two people doing the same work.
- Fork the repository and make your change on a branch.
- Open a pull request that references the issue number.
More from biocompute-objects/bco-rag
-
MongoDB Backend? Openenhancement question
biocompute-objects/bco-rag#12 · 1 assignee ·
-
enhancement
biocompute-objects/bco-rag#10 · 1 assignee ·
-
enhancement
biocompute-objects/bco-rag#9 · 1 assignee ·
-
enhancement
biocompute-objects/bco-rag#8 · 1 assignee ·
-
enhancement
biocompute-objects/bco-rag#7 · 1 assignee ·
All issues in biocompute-objects/bco-rag
Similar issues
-
Difficulty 2/5 1-3 hours Newbie friendliness 88/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 82/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 78/100
-
enhancement
Difficulty 2/5 1-3 hours Newbie friendliness 72/100
-
Difficulty 2/5 1-3 hours Newbie friendliness 74/100