Scores are stored in a 32-bit NumPy array even when K and V are quantized
维护者通常 1 天内回复
还没有人认领这个 Issue。
评估
- 难度
- 2/5
- 预计耗时
- 1-3 小时
- 新手友好度
- 65/100
- Issue 类型
- 缺陷
- 描述清晰度
- 描述清楚
- 活跃度
- 停滞
- 技术栈
- numpy, python
- 领域
- backend, performance
调研方向
该 issue 指向 llama_cpp/llama.py 第 460 行,此处使用 dtype=np.single 创建了一个 NumPy 数组。修复方法是修改 dtype,使其与 K 和 V 张量的量化方式匹配(例如,16 位使用 np.half)。首先,检查周围的代码,以了解 scores 数组的使用方式,以及量化张量的 dtype。然后相应地修改 dtype,并通过使用较大上下文运行服务器命令进行测试,以确保不会发生内存错误。
由索引模型根据 Issue 内容生成。
描述
Prerequisites
Please answer the following questions for yourself before submitting an issue.
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Expected Behavior
It should load the scores into an array with the appropriate data type
Current Behavior
Instead, it loads them into 32 bit ndarray
Environment and Context
Windows 10
Python 3.11.9
Latest CUDA 12.1 wheel as of now
RTX 3090
RTX 3070 (Hidden to application in this test)
32GB RAM
Failure Information
Traceback (most recent call last):
File "<frozen runpy>", line 198, in _run_module_as_main
File "<frozen runpy>", line 88, in _run_code
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\server\__main__.py", line 100, in <module>
main()
app = create_app(
^^^^^^^^^^^
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\server\app.py", line 150, in create_app
set_llama_proxy(model_settings=model_settings)
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\server\app.py", line 70, in set_llama_proxy
_llama_proxy = LlamaProxy(models=model_settings)
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\server\model.py", line 31, in __init__
self._current_model = self.load_llama_from_model_settings(
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\server\model.py", line 236, in load_llama_from_model_settings
_model = create_fn(
^^^^^^^^^^
File "C:\Users\Ethan\miniconda3\envs\torch\Lib\site-packages\llama_cpp\llama.py", line 460, in __init__
self.scores: npt.NDArray[np.single] = np.ndarray(
^^^^^^^^^^^
numpy.core._exceptions._ArrayMemoryError: Unable to allocate 113. GiB for an array with shape (200000, 151552) and data type float32
Steps to Reproduce
Use the openai server with this command (or similar):
python -m llama_cpp.server --model glm-4-9b-chat-1m-Q4_0.gguf --flash_attn true --type_k 6 --type_v 6 --n_gpu_layers -1 --n_ctx 200000
Solution
The issue is on this line:
https://github.com/abetlen/llama-cpp-python/blob/main/llama_cpp/llama.py#L460C9-L462C10
I set it to dtype=np.half and it worked, but my kv quants are still at 6bit so I think it could go lower.
- 主要语言
- Python
- 星标
- 10.6k
- 派生
- 1.5k
- 平均合并
- 3 小时 57 分钟
- 30 天内合并 PR
- 4
环境准备
- 没有 Dockerfile 或 Docker Compose 文件
- 没有 Pull Request 模板
- 阅读贡献指南
从这里开始
- 先读完整个 Issue,再读项目的贡献指南。
- 在 Issue 下留言说明你要接手 —— 这能避免两个人做同样的事。
- Fork 仓库,在一个分支上完成修改。
- 提交 Pull Request,并在描述里引用这个 Issue 编号。
abetlen/llama-cpp-python 的其他 Issue
-
难度 2/5 1-3 小时 新手友好度 88/100
abetlen/llama-cpp-python#2371 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2352 · 1 条评论 · 2 个 reaction ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 75/100
abetlen/llama-cpp-python#2314 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2211 · 2 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 65/100
abetlen/llama-cpp-python#2210 ·
维护者通常 1 天内回复
查看 abetlen/llama-cpp-python 的全部 Issue
相似的 Issue
-
#bug
难度 1/5 1 小时以内 新手友好度 92/100
apache/superset#44923 · 1 条评论 ·
维护者通常 2 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
lawndoc/stack-back#123 ·
-
Add: Entune未关闭
难度 2/5 1-3 小时 新手友好度 76/100
AbdelStark/awesome-typesafe-jev#187 ·
维护者通常 1 天内回复
-
bug good first issue
难度 2/5 1-3 小时 新手友好度 88/100
repowise-dev/repowise#2966 · 1 条评论 ·
维护者通常 1 天内回复
-
难度 2/5 1-3 小时 新手友好度 84/100
维护者通常 2 天内回复