ggml-org/llama.cpp

server : add "token healing" support

Open

#5,765 创建于 2024年2月28日

在 GitHub 查看
 (9 评论) (14 反应) (0 负责人)C++ (18,202 fork)batch import
enhancementgood first issueroadmap

仓库指标

Star
 (110,169 star)
PR 合并指标
 (平均合并 6天 8小时) (30 天内合并 389 个 PR)

描述

Prerequisites

Please answer the following questions for yourself before submitting an issue.

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Feature Description

Hi! I am experimenting with using llama.cpp as a general-purpose code completion backend, similar to TabNine.

I am encountering a small problem: if the completion prompt ends mid-word, the results are not very accurate. For example, for a prompt such as Five, Four, Thre [sic], the model will often ignore the typo and suggest , Two (forming Thre, Two).

I think, as an option to the /completion server API, the following optional behavior would be useful:

  1. Tokenize the text
  2. Chop off the last token
  3. Run the prediction with the remaining tokens, but only consider those tokens whose bytes start with the bytes of the last token.

Thanks!

贡献者指南