Regression: Falcon-E models cannot be loaded — "unknown pre-tokenizer type: 'falcon_e'" (support added in #268 was lost in the llama.cpp submodule update)
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 4/5
- Thời gian dự kiến
- 3-5 ngày
- Mức phù hợp với người mới
- 55/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Khá rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Lĩnh vực
- backend, machine-learning
Hướng nghiên cứu
Bắt đầu với src/llama-vocab.cpp trong submodule 3rdparty/llama.cpp và so sánh các trường hợp pre-tokenizer Falcon của nó với utils/convert-hf-to-gguf-bitnet.py, bao gồm cả mapping falcon_e hiện có. Tái hiện lỗi bằng lệnh hf download và build/bin/llama-cli được cung cấp, sau đó xác minh rằng Falcon-E GGUF được tải mà không cần --override-kv và metadata vừa được chuyển đổi được nhận diện; cập nhật README nếu việc hỗ trợ vẫn chưa khả dụng.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Summary
The Falcon-E family is listed as supported in the README and in setup_env.py (tiiuae/Falcon-E-1B/3B-Instruct, -Base), but no Falcon-E GGUF can be loaded with the current tree: the vocabulary loader rejects the falcon_e pre-tokenizer. This affects both the official GGUF published by TII and any GGUF freshly converted with this repo's converter, which still emits tokenizer.ggml.pre = "falcon_e".
Environment
microsoft/BitNetat0b341e5(currentmain), submodule3rdparty/llama.cppat390c3077- macOS 26.5.2, Apple M2 Pro; reproduced with two different builds of this tree (arm64 and x86_64), the error is raised before any backend is used
Steps to reproduce
hf download tiiuae/Falcon-E-3B-Instruct-GGUF ggml-model-i2_s.gguf --local-dir models/Falcon-E-3B-Instruct
build/bin/llama-cli -m models/Falcon-E-3B-Instruct/ggml-model-i2_s.gguf -p "Hello" -n 16
llama_model_load: error loading model: error loading model vocabulary: unknown pre-tokenizer type: 'falcon_e'
llama_model_load_from_file_impl: failed to load model
GGUF metadata: general.architecture = llama, tokenizer.ggml.pre = "falcon_e", 224 I2_S tensors.
Root cause
- PR #268 (May 2025, "Add falcon-e support") added the
falcon_epre-tokenizer: a hash→name mapping inutils/convert-hf-to-gguf-bitnet.py(still present today, lines 327-328:res = "falcon_e") plus the corresponding vocab support in the llama.cpp submodule of that time. - The current submodule (
isHuangXin/llama.cpp@390c3077, rebased on a much newer upstream) has nofalcon_ecase insrc/llama-vocab.cpp: the Falcon-related pre-tokenizers it knows arefalcon,falcon3andfalcon-h1(lines ~2129-2156). The C++ side of #268 was therefore lost, while the Python side kept producing the now-unknown name.
Workaround
--override-kv tokenizer.ggml.pre=str:falcon3 makes the model load and produce coherent output (in my tests Falcon-E-3B-Instruct answered general questions correctly and solved a classic river-crossing puzzle). I have not verified that the falcon3 regex set is identical to what falcon_e was meant to use, so tokenization may differ on edge cases.
Suggested fix
Re-add LLAMA_VOCAB_PRE_TYPE_FALCON_E in src/llama-vocab.cpp (or alias falcon_e to falcon3 if the pre-tokenizer rules are the same), and add the Falcon-E hash to the upstream-style pre-tokenizer table used by the converter. Until then, the README should not list Falcon-E as supported. Related: #508 (missing tokenizer.ggml.pre in the converters), #619 (conversion flow).
- Ngôn ngữ chính
- C++
- Star
- 40.4k
- Fork
- 3.7k
- Chỉ số merge pull request
- Không có pull request nào được merge trong 30 ngày
Chuẩn bị môi trường
Dự án này không cung cấp dev container, Dockerfile hay hướng dẫn đóng góp, nên bạn cần tự thiết lập môi trường: hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/BitNet
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed rowĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
-
Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.Có thể đã có người làm @ousamabenyounes đã nhận 54 ngày trước. Đang mở
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 76/100
Tất cả issue của microsoft/BitNet
Issue tương tự
-
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 78/100
espressif/esp-matter#1874 ·
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 78/100
opencv/opencv_contrib#4231 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
MiSTer-devel/Main_MiSTer#1341 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
`-static-libstdc++` breaks buildĐang mở
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
linux-test-project/lcov#552 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 1/5 1-3 giờ Mức phù hợp với người mới 94/100
Maintainer thường phản hồi trong vòng 1 ngày