ggml-org/llama.cpp
server: avoid full prompt eval when 'prompt >= ctx'
Geschlossen
#6.855 geöffnet am 23.04.2024
enhancementgood first issue
Repository-Metriken
- Stars
- (124.134 Sterne)
- PR-Merge-Metriken
- (Durchschn. Merge 6T 8h) (389 gemergte PRs in 30 T)
Beschreibung
When using the server for multi-turn chat, soon or later the prompt is going to surpass the context size, the current approach truncate the prompt by half of the context size excluding n_keep:
By doing that, common_part is going to match only n_keep tokens (when cache_prompt: true):
Technically, this is not a full prompt eval, n_keep is not revaluated, but it would be better to avoid this if possible, specially because prompt eval is slow on CPU.