Request to Modify Code to Enable TEXT_SPLITTER_EMBEDDING_MODEL Customization through Configuration File
@sumitkbh đang làm issue này rồi.
Từ ngày 18/1/2024.
Đánh giá
Issue này chưa được đánh giá.
Mô tả
I am looking to create a Chinese RAG demo service using RetrievalAugmentedGeneration.
However, I encountered an issue where the default SentenceTransformersTokenTextSplitter model used in the RetrievalAugmentedGeneration/common/utils.py file is hardcoded as 'intfloat/e5-large-v2'. This model generates a significant number of [UNK] tokens when processing Chinese text.
I would like the ability to specify a specific model for the text splitter, similar to how the embedding model can be specified through the config.yaml file.
Thank you for your assistance and support.
- Ngôn ngữ chính
- Jupyter Notebook
- Star
- 4.2k
- Fork
- 1.1k
- Merge trung bình
- 10 giờ 15 phút
- Pull request đã merge (30 ngày)
- 1
Hướng dẫn đóng góp
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của NVIDIA/GenerativeAIExamples
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 70/100
NVIDIA/GenerativeAIExamples#361 · 2 bình luận ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
NVIDIA/GenerativeAIExamples#299 ·
-
Độ khó 3/5 1-2 ngày Mức phù hợp với người mới 35/100
NVIDIA/GenerativeAIExamples#437 ·
-
hyperlink not working Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 45/100
NVIDIA/GenerativeAIExamples#400 ·
-
Độ khó 5/5 Hơn một tuần Mức phù hợp với người mới 10/100
NVIDIA/GenerativeAIExamples#399 ·