Hacktoberfest 2026:维护者为十月标记出来的 issue,仍然开放、适合新手。 浏览 Hacktoberfest issue

Regression: Falcon-E models cannot be loaded — "unknown pre-tokenizer type: 'falcon_e'" (support added in #268 was lost in the llama.cpp submodule update)

未关闭
#622 0 条评论 0 个 reaction 已指派 0 人 在 GitHub 查看

还没有人认领这个 Issue。

评估

难度
4/5
预计耗时
3-5 天
新手友好度
55/100
Issue 类型
缺陷
描述清晰度
基本清楚
活跃度
活跃
技术栈
cpp, python

调研方向

从 3rdparty/llama.cpp 子模块中的 src/llama-vocab.cpp 开始,将其中的 Falcon 预分词器用例与 utils/convert-hf-to-gguf-bitnet.py 进行比较,包括现有的 falcon_e 映射。使用提供的 hf download 和 build/bin/llama-cli 命令重现该失败,然后验证 Falcon-E GGUF 无需 --override-kv 即可加载,并且新转换的元数据能够被识别;如果仍无法提供支持,请更新 README。

由索引模型根据 Issue 内容生成。

描述

Summary

The Falcon-E family is listed as supported in the README and in setup_env.py (tiiuae/Falcon-E-1B/3B-Instruct, -Base), but no Falcon-E GGUF can be loaded with the current tree: the vocabulary loader rejects the falcon_e pre-tokenizer. This affects both the official GGUF published by TII and any GGUF freshly converted with this repo's converter, which still emits tokenizer.ggml.pre = "falcon_e".

Environment
  • microsoft/BitNet at 0b341e5 (current main), submodule 3rdparty/llama.cpp at 390c3077
  • macOS 26.5.2, Apple M2 Pro; reproduced with two different builds of this tree (arm64 and x86_64), the error is raised before any backend is used
Steps to reproduce
hf download tiiuae/Falcon-E-3B-Instruct-GGUF ggml-model-i2_s.gguf --local-dir models/Falcon-E-3B-Instruct
build/bin/llama-cli -m models/Falcon-E-3B-Instruct/ggml-model-i2_s.gguf -p "Hello" -n 16
llama_model_load: error loading model: error loading model vocabulary: unknown pre-tokenizer type: 'falcon_e'
llama_model_load_from_file_impl: failed to load model

GGUF metadata: general.architecture = llama, tokenizer.ggml.pre = "falcon_e", 224 I2_S tensors.

Root cause
  • PR #268 (May 2025, "Add falcon-e support") added the falcon_e pre-tokenizer: a hash→name mapping in utils/convert-hf-to-gguf-bitnet.py (still present today, lines 327-328: res = "falcon_e") plus the corresponding vocab support in the llama.cpp submodule of that time.
  • The current submodule (isHuangXin/llama.cpp @ 390c3077, rebased on a much newer upstream) has no falcon_e case in src/llama-vocab.cpp: the Falcon-related pre-tokenizers it knows are falcon, falcon3 and falcon-h1 (lines ~2129-2156). The C++ side of #268 was therefore lost, while the Python side kept producing the now-unknown name.
Workaround

--override-kv tokenizer.ggml.pre=str:falcon3 makes the model load and produce coherent output (in my tests Falcon-E-3B-Instruct answered general questions correctly and solved a classic river-crossing puzzle). I have not verified that the falcon3 regex set is identical to what falcon_e was meant to use, so tokenization may differ on edge cases.

Suggested fix

Re-add LLAMA_VOCAB_PRE_TYPE_FALCON_E in src/llama-vocab.cpp (or alias falcon_e to falcon3 if the pre-tokenizer rules are the same), and add the Falcon-E hash to the upstream-style pre-tokenizer table used by the converter. Until then, the README should not list Falcon-E as supported. Related: #508 (missing tokenizer.ggml.pre in the converters), #619 (conversion flow).

主要语言
C++
星标
40.4k
派生
3.7k
PR 合并指标
30 天内没有已合并 PR

环境准备

这个项目没有提供开发容器、Dockerfile 或贡献指南,环境需要你自己搭建:先看它的 README,通用步骤见我们的新手贡献指南。

从这里开始

  1. 先读完整个 Issue,再读项目的贡献指南。
  2. 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
  3. Fork 仓库,在一个分支上完成修改。
  4. 提交 Pull Request,并在描述里引用这个 Issue 编号。

microsoft/BitNet 的其他 Issue

查看 microsoft/BitNet 的全部 Issue

相似的 Issue

更多 C++ Issue

把新 issue 发到你的邮箱

精选适合新手参与的 GitHub issue 摘要。