Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

openai API `max_completion_tokens` argument is ignored

オープン 初心者向け
#1,907 コメント 0 件 リアクション 3 件 担当者 0 名 GitHub で見る

まだ誰も着手していません。

評価

難易度
2/5
見積もり時間
1〜3時間
初心者へのやさしさ
65/100
issue の種類
バグ
明瞭さ
明確に書かれている
活発さ
停滞
技術スタック
python
領域
api, backend

調査の方向性

OpenAI互換のチャット completions endpointを処理するサーバーコードを調べます。おそらく llama_cpp/server/app.py または同様のファイルにあります。リクエストパラメータが解析されている場所と、max_tokens が使用されている場所を見つけます。max_completion_tokens がどのように処理されるべきかと比較します。パラメータ名については OpenAI API 仕様を確認します。修正では、max_completion_tokens が読み取られ、生成ロジックに渡されることを保証します。サーバーを実行し、提供されているクライアントスクリプトを使って、トークン制限が守られることを確認してテストします。

索引モデルが issue の本文から書いたものです。

説明

Prerequisites

Please answer the following questions for yourself before submitting an issue.

  • I am running the latest code. Development is very rapid so there are no tagged versions as of now.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new bug or useful enhancement to share.

Current Behavior

I'm running llama-server with following command:

python3 -m llama_cpp.server --model models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf --clip_model_path models/mys/ggml_llava-v1.5-13b/mmproj-model-f16.gguf --model_alias llava-v1.5-13b-q4_k --chat_format llava-1-5 --port 10322

(models downloaded from https://huggingface.co/mys/ggml_llava-v1.5-13b/tree/main)

When I call the server using openai python package:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:10322/v1", # "http://<Your api-server IP>:port"
    api_key = "sk-no-key-required"
)

chat_completion = client.chat.completions.create(
    model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
    messages=[
        {"role": "user", "content": "Write a limerick about python exceptions"}
    ],
    max_tokens=3,
)
print(chat_completion.usage.completion_tokens)  # returns 3, ok.
print(chat_completion.choices[0].finish_reason)  # returns "length", ok.

chat_completion = client.chat.completions.create(
    model="models/mys/ggml_llava-v1.5-13b/ggml-model-q4_k.gguf",
    messages=[
        {"role": "user", "content": "Write a limerick about python exceptions"}
    ],
    max_completion_tokens=3,
)
print(chat_completion.usage.completion_tokens)  # returns much more than 3 (complete answer).
print(chat_completion.choices[0].finish_reason)   # returns "stop".

According to OpenAI API, max_completion_tokens argument is replacing the deprecated max_tokens argument.
It's seems that only max_tokens is not ignored by the server.

Environment and Context

llama_cpp installed with pip install llama-cpp-python[server]
print(llama_cpp.__version__): 0.3.6
print(openai.__version__): 1.59.7

主要言語
Python
スター
10.6k
フォーク
1.5k
平均マージ
6時間 43分
マージ済み PR(30日)
2

環境構築

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

abetlen/llama-cpp-python のほかの issue

abetlen/llama-cpp-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。