openai API `max_completion_tokens` argument is ignored
まだ誰も着手していません。
評価
- 難易度
- 2/5
- 見積もり時間
- 1〜3時間
- 初心者へのやさしさ
- 65/100
調査の方向性
OpenAI互換のチャット completions endpointを処理するサーバーコードを調べます。おそらく llama_cpp/server/app.py または同様のファイルにあります。リクエストパラメータが解析されている場所と、max_tokens が使用されている場所を見つけます。max_completion_tokens がどのように処理されるべきかと比較します。パラメータ名については OpenAI API 仕様を確認します。修正では、max_completion_tokens が読み取られ、生成ロジックに渡されることを保証します。サーバーを実行し、提供されているクライアントスクリプトを使って、トークン制限が守られることを確認してテストします。
索引モデルが issue の本文から書いたものです。
説明
Prerequisites
Please answer the following questions for yourself before submitting an issue.
- I am running the latest code. Development is very rapid so there are no tagged versions as of now.
- I carefully followed the README.md.
- I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
- I reviewed the Discussions, and have a new bug or useful enhancement to share.
Current Behavior
I'm running llama-server with following command:
python3 -m llama_cpp.server --model models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf --clip_model_path models/mys/ggml_llava-v1.5-13b/mmproj-model-f16.gguf --model_alias llava-v1.5-13b-q4_k --chat_format llava-1-5 --port 10322
(models downloaded from https://huggingface.co/mys/ggml_llava-v1.5-13b/tree/main)
When I call the server using openai python package:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:10322/v1", # "http://<Your api-server IP>:port"
api_key = "sk-no-key-required"
)
chat_completion = client.chat.completions.create(
model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
messages=[
{"role": "user", "content": "Write a limerick about python exceptions"}
],
max_tokens=3,
)
print(chat_completion.usage.completion_tokens) # returns 3, ok.
print(chat_completion.choices[0].finish_reason) # returns "length", ok.
chat_completion = client.chat.completions.create(
model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
messages=[
{"role": "user", "content": "Write a limerick about python exceptions"}
],
max_completion_tokens=3,
)
print(chat_completion.usage.completion_tokens) # returns much more than 3 (complete answer).
print(chat_completion.choices[0].finish_reason) # returns "stop".
According to OpenAI API, max_completion_tokens argument is replacing the deprecated max_tokens argument.
It's seems that only max_tokens is not ignored by the server.
Environment and Context
llama_cpp installed with pip install llama-cpp-python[server]
print(llama_cpp.__version__): 0.3.6
print(openai.__version__): 1.59.7
- 主要言語
- Python
- スター
- 10.6k
- フォーク
- 1.5k
- 平均マージ
- 6時間 43分
- マージ済み PR(30日)
- 2
環境構築
はじめの一歩
- issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
- 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
- リポジトリをフォークし、ブランチを切って変更します。
- issue 番号を参照したプルリクエストを送ります。
abetlen/llama-cpp-python のほかの issue
-
難易度 2/5 1〜3時間 初心者へのやさしさ 88/100
abetlen/llama-cpp-python#2371 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
abetlen/llama-cpp-python#2352 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 75/100
abetlen/llama-cpp-python#2314 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
abetlen/llama-cpp-python#2211 · コメント 2 件 ·
-
難易度 2/5 1〜3時間 初心者へのやさしさ 65/100
abetlen/llama-cpp-python#2210 ·
abetlen/llama-cpp-python の issue をすべて見る
似ている issue
-
bug status/needs-triage
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
prowler-cloud/prowler#12887 · コメント 1 件 ·
メンテナーはふだん 1 日以内に返信
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
難易度 1/5 1時間未満 初心者へのやさしさ 92/100
use-agent-os/agent-os#3484 ·
メンテナーはふだん 2 日以内に返信
-
bug
難易度 2/5 1〜3時間 初心者へのやさしさ 86/100
open-telemetry/opentelemetry-python-contrib#5113 · コメント 2 件 · リアクション 2 件 ·
メンテナーはふだん 1 日以内に返信
-
external
難易度 2/5 1〜3時間 初心者へのやさしさ 68/100
langchain-ai/docs#6255 ·
メンテナーはふだん 1 日以内に返信
-
難易度 2/5 1〜3時間 初心者へのやさしさ 72/100
メンテナーはふだん 1 日以内に返信