Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Title: Feature Request: Attack strategy for Excessive Agency / unauthorized tool invocation

Abierto
#2,903 0 comentarios 0 reacciones 0 asignados Ver en GitHub

Los mantenedores suelen responder en 2 días

Nadie ha tomado este issue todavía.

Evaluación

Dificultad
5/5
Tiempo estimado
Más de una semana
Aptitud para principiantes
30/100
Tipo de issue
Nueva funcionalidad
Claridad
Necesita aclaración
Estado de actividad
Activo
Stack tecnológico
python
Área
ai, security

Línea de trabajo

Start by reading openai_response_target.py and the existing Tool/ToolProvider interfaces, then compare the single-turn and multi-turn attack structures, including Crescendo. Confirm with maintainers how tool-call output can be inspected and scored. Done means an agreed attack design and implementation that evaluates actual unauthorized or excessive tool calls rather than only text output.

Escrito por el modelo de indexación a partir del texto del issue.

Descripción

feature-request

Summary

PyRIT already supports exposing tools to a target via Tool / ToolProvider
(see openai_response_target.py), but I don't see an attack strategy that
specifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.

This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.

Proposed approach (open to feedback before implementing)

  • A new attack, e.g. ExcessiveAgencyAttack, that:
    1. Takes a target configured with a defined set of "allowed" tools/scope
      (via the existing Tool/ToolProvider system).
    2. Attempts to elicit a tool call outside that scope — either a tool the
      agent has access to but shouldn't use for the stated task, or a chained
      sequence of legitimate calls that together exceed the intended
      permission boundary.
    3. Scores success based on the tool call actually made (inspecting the
      target's tool-call output), not on the text response — this is the
      distinct piece existing text/output scorers don't cover.
  • Could reuse the single-turn or multi-turn structure depending on whether
    the elicitation needs conversation history (multi-turn is likely more
    realistic here, similar to how Crescendo works for text escalation).

Questions for maintainers

  1. Is there existing tooling for asserting on tool-call output specifically
    (as opposed to text output) that I should reuse for scoring?
  2. Any prior discussion/PR on agentic risk testing I should be aware of
    before designing this?

Happy to take this on once the approach is validated.

Lenguaje dominante
Python
Estrellas
4.5k
Forks
896
Merge medio
2 d 22 h
PR fusionados (30 d)
230

Preparar el entorno

Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de microsoft/PyRIT

Todos los issues de microsoft/PyRIT

Issues similares

Más issues de Python

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.