ggml-org/llama.cpp

llama : refactor the llm.build_xxx functions

Fechada

#5.239 aberto em 31 de jan. de 2024

 (3 comentários) (3 reações) (0 responsável)C++ (21.764 forks)batch import
good first issuerefactoringroadmap

Métricas do repositório

Stars
 (124.134 estrelas)
Métricas de merge de PR
 (Mesclagem média 6d 8h) (389 fundiu PRs em 30d)

Description

Now that we support a large amount of architectures, we can clearly see the patterns when constructing the compute graphs - i.e. optional biases, different norm types, QKV vs Q+K+V, etc.

We should deduplicate the copy-paste portions in functions such as llm.build_llama(), llm.build_falcon(), etc.

The advantage of the current code is that it is easy to look into the graph of a specific architecture. When we refactor this, we will lose this convenience to some extend. So we should think about making this refactoring in such a way that we don't completely obscure which parts of the graph belong to which architectures

Open for ideas and suggestions how to do this best

Guia do colaborador