Regression: Falcon-E models cannot be loaded — "unknown pre-tokenizer type: 'falcon_e'" (support added in #268 was lost in the llama.cpp submodule update)
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 4/5
- Tempo stimato
- 3-5 giorni
- Idoneità per principianti
- 55/100
- Tipo di issue
- Bug
- Chiarezza
- Abbastanza chiara
- Stato di attività
- Attiva
- Ambito
- backend, machine-learning
Direzione di ricerca
Inizia da src/llama-vocab.cpp nel submodule 3rdparty/llama.cpp e confronta i relativi casi del pre-tokenizer Falcon con utils/convert-hf-to-gguf-bitnet.py, inclusa la mappatura falcon_e esistente. Riproduci il problema con il comando hf download e build/bin/llama-cli fornito, quindi verifica che Falcon-E GGUF venga caricato senza --override-kv e che i metadati appena convertiti vengano riconosciuti; aggiorna il README se il supporto rimane non disponibile.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
The Falcon-E family is listed as supported in the README and in setup_env.py (tiiuae/Falcon-E-1B/3B-Instruct, -Base), but no Falcon-E GGUF can be loaded with the current tree: the vocabulary loader rejects the falcon_e pre-tokenizer. This affects both the official GGUF published by TII and any GGUF freshly converted with this repo's converter, which still emits tokenizer.ggml.pre = "falcon_e".
Environment
microsoft/BitNetat0b341e5(currentmain), submodule3rdparty/llama.cppat390c3077- macOS 26.5.2, Apple M2 Pro; reproduced with two different builds of this tree (arm64 and x86_64), the error is raised before any backend is used
Steps to reproduce
hf download tiiuae/Falcon-E-3B-Instruct-GGUF ggml-model-i2_s.gguf --local-dir models/Falcon-E-3B-Instruct
build/bin/llama-cli -m models/Falcon-E-3B-Instruct/ggml-model-i2_s.gguf -p "Hello" -n 16
llama_model_load: error loading model: error loading model vocabulary: unknown pre-tokenizer type: 'falcon_e'
llama_model_load_from_file_impl: failed to load model
GGUF metadata: general.architecture = llama, tokenizer.ggml.pre = "falcon_e", 224 I2_S tensors.
Root cause
- PR #268 (May 2025, "Add falcon-e support") added the
falcon_epre-tokenizer: a hash→name mapping inutils/convert-hf-to-gguf-bitnet.py(still present today, lines 327-328:res = "falcon_e") plus the corresponding vocab support in the llama.cpp submodule of that time. - The current submodule (
isHuangXin/llama.cpp@390c3077, rebased on a much newer upstream) has nofalcon_ecase insrc/llama-vocab.cpp: the Falcon-related pre-tokenizers it knows arefalcon,falcon3andfalcon-h1(lines ~2129-2156). The C++ side of #268 was therefore lost, while the Python side kept producing the now-unknown name.
Workaround
--override-kv tokenizer.ggml.pre=str:falcon3 makes the model load and produce coherent output (in my tests Falcon-E-3B-Instruct answered general questions correctly and solved a classic river-crossing puzzle). I have not verified that the falcon3 regex set is identical to what falcon_e was meant to use, so tokenization may differ on edge cases.
Suggested fix
Re-add LLAMA_VOCAB_PRE_TYPE_FALCON_E in src/llama-vocab.cpp (or alias falcon_e to falcon3 if the pre-tokenizer rules are the same), and add the Falcon-E hash to the upstream-style pre-tokenizer table used by the converter. Until then, the README should not list Falcon-E as supported. Related: #508 (missing tokenizer.ggml.pre in the converters), #619 (conversion flow).
- Lingua principale
- C++
- Stelle
- 40.4k
- Fork
- 3.7k
- Metriche di merge delle PR
- Nessuna PR unita negli ultimi 30g
Preparare l'ambiente
Questo progetto non fornisce container di sviluppo, Dockerfile né guida per i contributori, quindi l'ambiente è a tuo carico: parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/BitNet
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
i2_s SIGSEGVs at n_ubatch >= 32: BLAS backend dequantises by a row stride 4x the real packed rowAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
-
Warning: `huggingface-cli` is deprecated and no longer works. Use `hf` instead.Forse già presa @ousamabenyounes l’ha presa 61 giorni fa. Aperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 92/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 76/100
Tutte le issue di microsoft/BitNet
Issue simili
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 85/100
objectionary/eo-graphs#80 ·
-
bug C/C++ code
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 78/100
webarkit/WebARKitLib#85 ·
I maintainer di solito rispondono entro 1 giorno
-
SD Card Size correctionAperta
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
-
[request] vsg/1.1.16Apertaupstream update
Difficoltà 2/5 1-3 ore Idoneità per principianti 65/100
conan-io/conan-center-index#31142 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 72/100