Regression: Falcon-E models cannot be loaded — "unknown pre-tokenizer type: 'falcon_e'" (support added in #268 was lost in the llama.cpp submodule update)
まだ誰も着手していません。
評価
- 難易度
- 4/5
- 見積もり時間
- 3〜5日
- 初心者へのやさしさ
- 55/100
- issue の種類
- バグ
- 明瞭さ
- おおむね明確
- 活発さ
- 活発
調査の方向性
3rdparty/llama.cpp サブモジュール内の src/llama-vocab.cpp から始め、Falcon の pre-tokenizer ケースを、既存の falcon_e マッピングを含めて utils/convert-hf-to-gguf-bitnet.py と比較します。提供されている hf download および build/bin/llama-cli コマンドで失敗を再現し、その後、Falcon-E GGUF が --override-kv なしで読み込まれること、および新しく変換されたメタデータが認識されることを確認します。サポートが引き続き利用できない場合は README を更新します。
索引モデルが issue の本文から書いたものです。
説明
Summary
The Falcon-E family is listed as supported in the README and in setup_env.py (tiiuae/Falcon-E-1B/3B-Instruct, -Base), but no Falcon-E GGUF can be loaded with the current tree: the vocabulary loader rejects the falcon_e pre-tokenizer. This affects both the official GGUF published by TII and any GGUF freshly converted with this repo's converter, which still emits tokenizer.ggml.pre = "falcon_e".
Environment
microsoft/BitNetat0b341e5(currentmain), submodule3rdparty/llama.cppat390c3077- macOS 26.5.2, Apple M2 Pro; reproduced with two different builds of this tree (arm64 and x86_64), the error is raised before any backend is used
Steps to reproduce
hf download tiiuae/Falcon-E-3B-Instruct-GGUF ggml-model-i2_s.gguf --local-dir models/Falcon-E-3B-Instruct
build/bin/llama-cli -m models/Falcon-E-3B-Instruct/ggml-model-i2_s.gguf -p "Hello" -n 16
llama_model_load: error loading model: error loading model vocabulary: unknown pre-tokenizer type: 'falcon_e'
llama_model_load_from_file_impl: failed to load model
GGUF metadata: general.architecture = llama, tokenizer.ggml.pre = "falcon_e", 224 I2_S tensors.
Root cause
- PR #268 (May 2025, "Add falcon-e support") added the
falcon_epre-tokenizer: a hash→name mapping inutils/convert-hf-to-gguf-bitnet.py(still present today, lines 327-328:res = "falcon_e") plus the corresponding vocab support in the llama.cpp submodule of that time. - The current submodule (
isHuangXin/llama.cpp@390c3077, rebased on a much newer upstream) has nofalcon_ecase insrc/llama-vocab.cpp: the Falcon-related pre-tokenizers it knows arefalcon,falcon3andfalcon-h1(lines ~2129-2156). The C++ side of #268 was therefore lost, while the Python side kept producing the now-unknown name.
Workaround
--override-kv tokenizer.ggml.pre=str:falcon3 makes the model load and produce coherent output (in my tests Falcon-E-3B-Instruct answered general questions correctly and solved a classic river-crossing puzzle). I have not verified that the falcon3 regex set is identical to what falcon_e was meant to use, so tokenization may differ on edge cases.
Suggested fix
Re-add LLAMA_VOCAB_PRE_TYPE_FALCON_E in src/llama-vocab.cpp (or alias falcon_e to falcon3 if the pre-tokenizer rules are the same), and add the Falcon-E hash to the upstream-style pre-tokenizer table used by the converter. Until then, the README should not list Falcon-E as supported. Related: #508 (missing tokenizer.ggml.pre in the converters), #619 (conversion flow).
- 主要言語
- C++
- スター
- 40.4k
- フォーク
- 3.7k
- PR マージ指標
- 30日以内にマージされた PR はありません
環境構築
このプロジェクトには開発コンテナ、Dockerfile、コントリビューションガイドがありません。まず README を読み、一般的な手順ははじめてのコントリビューションガイドを参照してください。
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
microsoft/BitNet のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed rowオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
-
Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.対応中かも @ousamabenyounes が 56 日前に担当しました。 オープン
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
-
難易度 2/5 1〜3時間 初心者へのやさしさ 76/100
microsoft/BitNet の issue をすべて見る
似ている issue
-
enhancement
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信
-
iOS: hidden scale bar invalidates its intrinsic content size on every layout pass of MLNMapViewオープン
難易度 2/5 1〜3時間 初心者へのやさしさ 85/100
maplibre/maplibre-native#4723 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 78/100
HarbourMasters/Shipwright#7320 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
-
Make Catch2 optional when `RDK_BUILD_CPP_TESTS=OFF`対応中かも @pechersky が今日担当しました。 オープンbug
難易度 2/5 1〜3時間 初心者へのやさしさ 82/100
メンテナーはふだん 2 日以内に返信