LiteLLMChatTarget does not flag or survive output-token truncation
Maintainer thường phản hồi trong vòng 1 ngày
Chưa có ai nhận issue này.
Đánh giá
- Độ khó
- 2/5
- Thời gian dự kiến
- 1-3 giờ
- Mức phù hợp với người mới
- 84/100
- Loại issue
- Lỗi
- Độ rõ ràng
- Đặc tả rõ ràng
- Mức độ hoạt động
- Sôi nổi
- Công nghệ
- python
- Lĩnh vực
- backend-api-design
Hướng nghiên cứu
Bắt đầu trong pyrit/prompt_target/litellm_chat_target.py và so sánh cách xử lý response của nó với pyrit/prompt_target/openai/_response_adapter.py và openai_chat_target.py. Chạy tests/unit/prompt_target/target/test_litellm_chat_target.py, bao phủ cả hai trường hợp đã tái hiện: một response có finish_reason="length" được đánh dấu là bị cắt ngắn, và một response bị cắt ngắn rỗng trả về một phần tử rỗng mà không phát sinh exception.
Do mô hình lập chỉ mục viết ra từ nội dung của issue.
Mô tả
Describe the bug
LiteLLMChatTarget does not handle output-token truncation, even though it speaks the same OpenAI-compatible Chat Completions shape that OpenAIChatTarget already handles.
With max_tokens set (pyrit/prompt_target/litellm_chat_target.py:138), a response that stops at the limit arrives with finish_reason == "length", and two things go wrong:
- A clipped answer is stored as a complete one.
MessagePiece.mark_as_truncatedexists for exactly this (pyrit/models/messages/message_piece.py:196): "A truncated piece may still carry a partial answer withresponse_error == "none", so without this flag a consumer cannot tell a complete answer from a clipped one."is_truncatedstaysFalsefor aLiteLLMChatTargetpiece, so a scorer or a red teamer reads a cut-off answer as the model's real answer. - A truncation that empties the response aborts the turn.
validate_chat_completion_responseraisesEmptyResponseExceptionwhen there is no content, audio or tool call (pyrit/prompt_target/common/chat_completions_response_parser.py:107), and it runs before the construct step. This target is deliberately not wrapped inpyrit_target_retry(LiteLLM owns retry vianum_retries), so nothing catches it — the conversation just stops.
How OpenAIChatTarget already handles the same case
pyrit/prompt_target/openai/openai_chat_target.py:314-331 warns, flags the first piece, and returns a graceful empty piece when truncation left nothing. pyrit/prompt_target/openai/_response_adapter.py:105-114 is the shape to mirror: validate() warns and returns early when is_truncated, otherwise calls validate_chat_completion_response; is_truncated() is get_finish_reason(response=response) == "length".
warn_truncated_response documents that its wording lives in common "to keep targets from drifting apart" (pyrit/prompt_target/common/utils.py:197), and build_empty_truncated_response is already shared for this case.
Steps/Code to Reproduce
Reproduced on main at 7b533109, through the public entry point (tests/unit/prompt_target/target/test_litellm_chat_target.py, litellm stubbed so no provider call is made). Two tests, one per failure mode:
truncated = _mock_response(content="The answer is", finish_reason="length")
litellm_stub.acompletion = AsyncMock(return_value=truncated)
result = await target.send_prompt_async(message=_user_message())
assert result[0].message_pieces[0].is_truncated is True # currently False
empty = _mock_response(content=None, finish_reason="length")
empty.choices[0].message.content = None
litellm_stub.acompletion = AsyncMock(return_value=empty)
result = await target.send_prompt_async(message=_user_message()) # currently raises
assert result[0].message_pieces[0].response_error == "empty"
Expected Results
The first assertion holds; the second call returns a graceful empty piece instead of raising.
Actual Results
FAILED test_token_limit_truncation_marks_the_piece
E AssertionError: assert False is True
... 'finish_reason': 'length'}).is_truncated
FAILED test_token_limit_truncation_with_no_content_does_not_raise
E pyrit.exceptions.exception_classes.EmptyResponseException: Status Code: 204,
Message: The chat returned an empty response (no content, audio, or tool_calls).
pyrit/prompt_target/common/chat_completions_response_parser.py:109: EmptyResponseException
Screenshots
Not applicable.
Versions
- OS: macOS
- Python version: 3.11+
- PyRIT version:
mainat7b533109
Proposed change
Confined to pyrit/prompt_target/litellm_chat_target.py: add _is_truncated_response on the existing get_finish_reason helper; warn via warn_truncated_response(..., limit_parameter="max_tokens") and skip strict validation when truncated; in _construct_message_from_response_async, mark_as_truncated() the first piece and return build_empty_truncated_response(...) when truncation left no pieces.
One question, since this changes error behaviour: should a truncated-but-empty LiteLLM response continue as error="empty" like the OpenAI target, or is aborting the turn preferred here? Happy to follow whichever the project prefers. A branch with the change and both tests is on feiiiiii5:fix/litellm-truncation-marker if you want to look at the diff.
- Ngôn ngữ chính
- Python
- Star
- 4.5k
- Fork
- 896
- Merge trung bình
- 3 ngày 2 giờ
- Pull request đã merge (30 ngày)
- 210
Chuẩn bị môi trường
Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.
Bắt đầu từ đâu
- Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
- Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
- Fork repository và làm thay đổi trên một nhánh.
- Mở pull request có tham chiếu số hiệu của issue.
Issue khác của microsoft/PyRIT
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
microsoft/PyRIT#2888 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 91/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 84/100
Maintainer thường phản hồi trong vòng 1 ngày
-
Bug: triage GUI help wanted
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
microsoft/PyRIT#2868 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 82/100
Maintainer thường phản hồi trong vòng 1 ngày
Tất cả issue của microsoft/PyRIT
Issue tương tự
-
bug status/needs-triage
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
prowler-cloud/prowler#12887 · 1 bình luận ·
Maintainer thường phản hồi trong vòng 1 ngày
-
area: desktop platform: macos priority: p3 status: ready type: enhancement
Độ khó 1/5 Dưới một giờ Mức phù hợp với người mới 92/100
use-agent-os/agent-os#3484 ·
Maintainer thường phản hồi trong vòng 2 ngày
-
bug
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 86/100
open-telemetry/opentelemetry-python-contrib#5113 · 2 bình luận · 2 reaction ·
Maintainer thường phản hồi trong vòng 1 ngày
-
external
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 68/100
langchain-ai/docs#6255 ·
Maintainer thường phản hồi trong vòng 1 ngày
-
Độ khó 2/5 1-3 giờ Mức phù hợp với người mới 72/100
Maintainer thường phản hồi trong vòng 1 ngày