Hacktoberfest 2026: những issue maintainer đã đánh dấu cho tháng Mười, đang mở và phù hợp người mới. Xem issue Hacktoberfest

Title: Feature Request: Attack strategy for Excessive Agency / unauthorized tool invocation

Đang mở
#2,903 0 bình luận 0 reaction 0 người được giao Xem trên GitHub

Maintainer thường phản hồi trong vòng 2 ngày

Chưa có ai nhận issue này.

Đánh giá

Độ khó
5/5
Thời gian dự kiến
Hơn một tuần
Mức phù hợp với người mới
30/100
Loại issue
Tính năng
Độ rõ ràng
Cần làm rõ
Mức độ hoạt động
Sôi nổi
Công nghệ
python
Lĩnh vực
ai, security

Hướng nghiên cứu

Start by reading openai_response_target.py and the existing Tool/ToolProvider interfaces, then compare the single-turn and multi-turn attack structures, including Crescendo. Confirm with maintainers how tool-call output can be inspected and scored. Done means an agreed attack design and implementation that evaluates actual unauthorized or excessive tool calls rather than only text output.

Do mô hình lập chỉ mục viết ra từ nội dung của issue.

Mô tả

feature-request

Summary

PyRIT already supports exposing tools to a target via Tool / ToolProvider
(see openai_response_target.py), but I don't see an attack strategy that
specifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.

This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.

Proposed approach (open to feedback before implementing)

  • A new attack, e.g. ExcessiveAgencyAttack, that:
    1. Takes a target configured with a defined set of "allowed" tools/scope
      (via the existing Tool/ToolProvider system).
    2. Attempts to elicit a tool call outside that scope — either a tool the
      agent has access to but shouldn't use for the stated task, or a chained
      sequence of legitimate calls that together exceed the intended
      permission boundary.
    3. Scores success based on the tool call actually made (inspecting the
      target's tool-call output), not on the text response — this is the
      distinct piece existing text/output scorers don't cover.
  • Could reuse the single-turn or multi-turn structure depending on whether
    the elicitation needs conversation history (multi-turn is likely more
    realistic here, similar to how Crescendo works for text escalation).

Questions for maintainers

  1. Is there existing tooling for asserting on tool-call output specifically
    (as opposed to text output) that I should reuse for scoring?
  2. Any prior discussion/PR on agentic risk testing I should be aware of
    before designing this?

Happy to take this on once the approach is validated.

Ngôn ngữ chính
Python
Star
4.5k
Fork
896
Merge trung bình
2 ngày 22 giờ
Pull request đã merge (30 ngày)
220

Chuẩn bị môi trường

Chúng tôi chưa kiểm tra các tệp thiết lập môi trường của dự án này. Hãy bắt đầu từ README và xem hướng dẫn đóng góp lần đầu của chúng tôi để biết các bước chung.

Bắt đầu từ đâu

  1. Đọc hết issue, rồi đọc hướng dẫn đóng góp của dự án.
  2. Bình luận trên issue rằng bạn sẽ nhận — tránh hai người làm cùng một việc.
  3. Fork repository và làm thay đổi trên một nhánh.
  4. Mở pull request có tham chiếu số hiệu của issue.

Issue khác của microsoft/PyRIT

Tất cả issue của microsoft/PyRIT

Issue tương tự

Thêm issue về Python

Nhận issue mới trong hộp thư của bạn

Bản tóm tắt ngắn những issue GitHub phù hợp với người mới.