Hacktoberfest 2026:メンテナが10月に向けて印を付けた、オープンで初心者向けの issue。 Hacktoberfest の issue を見る

Bug: server-side compaction is not emitted on Responses tool-call-only turns

オープン
#3,075 コメント 1 件 リアクション 0 件 担当者 0 名 GitHub で見る

メンテナーはふだん 1 日以内に返信

@nightcityblade がすでに取り組んでいます。

2026年5月2日 から。

  • #3344 @redactdeveloper による — マージされずにクローズ
  • #3417 @Sudhanwa-git による — マージされずにクローズ

評価

難易度
4/5
見積もり時間
3〜5日
初心者へのやさしさ
38/100
issue の種類
バグ
明瞭さ
説明が足りない
活発さ
静か
技術スタック
python
領域
api

調査の方向性

まず、提供されたツール呼び出しループと context_management 設定を使用し、responses.create と responses.parse で報告されたシーケンスを再現します。SDK が返す出力項目を、基盤となる Responses API の動作と比較します。Python ライブラリがツール呼び出しのみのターンで compaction items を破棄するかどうかを判定し、破棄する場合は修正後の動作に対する回帰テストを定義すれば完了です。

索引モデルが issue の本文から書いたものです。

説明

bug
Confirm this is an issue with the Python library and not an underlying OpenAI API
  • This is an issue with the Python library
Describe the bug

Describe the bug

I am using the Responses API through openai-python with:

  • context_management=[{"type": "compaction", "compact_threshold": 1000}]
  • store=False
  • gpt-5.4

I see different behavior depending on the output type of the turn:

  • For a long plain request, response.output contains:

    • message
    • compaction
  • For a long request that returns only function_call, response.output contains only:

    • function_call
  • If I continue the tool loop and the next turn is again only function_call, there is still no compaction.

  • Only when the model finally returns an assistant message does response.output include:

    • message
    • compaction

This means that in tool-heavy agent loops with several consecutive tool-call turns, context can continue growing without any emitted compaction item, and the loop can eventually hit
context_length_exceeded before compaction appears.

I reproduced this through openai-python using both client.responses.create(...) and client.responses.parse(...).

If this is expected backend/API behavior rather than a Python SDK issue, please let me know and I can move the report.

To Reproduce

  1. Create a long input that is clearly above the compaction threshold.
  2. Enable server-side compaction with a very low threshold, for example:
    context_management=[{"type": "compaction", "compact_threshold": 1000}]
  3. Force the first turn to produce a function_call.
  4. Send the corresponding function_call_output.
  5. If the model produces another function_call, observe that there is still no compaction item in response.output.
  6. Observe that compaction only appears once the model finally emits an assistant message.

Observed output from my repro:

R1 input_tokens 5084
R1 output_types ['function_call']

R2 input_tokens 5119
R2 output_types ['function_call']

R3 input_tokens 5154
R3 output_types ['message', 'compaction']

For comparison, a plain long request with the same threshold produces compaction immediately:

CREATE input_tokens 5007
CREATE output_types ['message', 'compaction']
To Reproduce
import asyncio
from openai import AsyncOpenAI
from azure.identity.aio import DefaultAzureCredential, get_bearer_token_provider

AZURE_ENDPOINT = "https://<your-resource>.openai.azure.com/openai/v1/"
MODEL = "gpt-5.4"

async def main():
    cred = DefaultAzureCredential()
    token_provider = get_bearer_token_provider(
        cred,
        "https://cognitiveservices.azure.com/.default"
    )

    client = AsyncOpenAI(
        base_url=AZURE_ENDPOINT,
        api_key=token_provider,
    )

    long_text = "context " * 5000

    tools = [{
        "type": "function",
        "name": "echo_tool",
        "description": "Echo a short string",
        "parameters": {
            "type": "object",
            "properties": {
                "text": {"type": "string"}
            },
            "required": ["text"],
            "additionalProperties": False
        }
    }]

    cm = [{"type": "compaction", "compact_threshold": 1000}]

    conversation = [{
        "role": "user",
        "content": (
            long_text +
            "\n\nCall echo_tool twice in sequence. "
            "First with text=first. After I return the tool result, "
            "call echo_tool again with text=second. "
            "Only after the second tool result, answer DONE."
        )
    }]

    for step in range(1, 5):
        response = await client.responses.create(
            model=MODEL,
            input=conversation,
            tools=tools,
            store=False,
            context_management=cm,
        )

        print(f"R{step} input_tokens:", response.usage.input_tokens)
        print(f"R{step} output_types:", [getattr(i, 'type', None) for i in response.output])

        conversation.extend(response.output)

        function_calls = [i for i in response.output if getattr(i, "type", None) == "function_call"]
        if function_calls:
            for idx, fc in enumerate(function_calls, start=1):
                conversation.append({
                    "type": "function_call_output",
                    "call_id": fc.call_id,
                    "output": f"tool-result-{step}-{idx}",
                })
        else:
            break

    await client.close()
    await cred.close()

asyncio.run(main())
Code snippets

OS

Windows

Python version

3.11.5

Library version

openai 2.21.0

主要言語
Python
スター
31.8k
フォーク
7.3k
平均マージ
1日 3時間
マージ済み PR(30日)
131

環境構築

Codespaces で開く

このプロジェクトの開発コンテナを、あなたの GitHub アカウントでブラウザ上に起動します。

はじめの一歩

  1. issue を最後まで読み、次にプロジェクトのコントリビューションガイドを読みます。
  2. 着手することを issue にコメントします — 二人が同じ作業をするのを防げます。
  3. リポジトリをフォークし、ブランチを切って変更します。
  4. issue 番号を参照したプルリクエストを送ります。

openai/openai-python のほかの issue

openai/openai-python の issue をすべて見る

似ている issue

Python の issue をもっと見る

新しい issue をメールで受け取る

初心者向けの GitHub issue を短くまとめたダイジェスト。