Title: Feature Request: Attack strategy for Excessive Agency / unauthorized tool invocation
Los mantenedores suelen responder en 2 días
Nadie ha tomado este issue todavía.
Evaluación
- Dificultad
- 5/5
- Tiempo estimado
- Más de una semana
- Aptitud para principiantes
- 30/100
Línea de trabajo
Start by reading openai_response_target.py and the existing Tool/ToolProvider interfaces, then compare the single-turn and multi-turn attack structures, including Crescendo. Confirm with maintainers how tool-call output can be inspected and scored. Done means an agreed attack design and implementation that evaluates actual unauthorized or excessive tool calls rather than only text output.
Escrito por el modelo de indexación a partir del texto del issue.
Descripción
Summary
PyRIT already supports exposing tools to a target via Tool / ToolProvider
(see openai_response_target.py), but I don't see an attack strategy that
specifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.
This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.
Proposed approach (open to feedback before implementing)
- A new attack, e.g.
ExcessiveAgencyAttack, that:- Takes a target configured with a defined set of "allowed" tools/scope
(via the existingTool/ToolProvidersystem). - Attempts to elicit a tool call outside that scope — either a tool the
agent has access to but shouldn't use for the stated task, or a chained
sequence of legitimate calls that together exceed the intended
permission boundary. - Scores success based on the tool call actually made (inspecting the
target's tool-call output), not on the text response — this is the
distinct piece existing text/output scorers don't cover.
- Takes a target configured with a defined set of "allowed" tools/scope
- Could reuse the single-turn or multi-turn structure depending on whether
the elicitation needs conversation history (multi-turn is likely more
realistic here, similar to how Crescendo works for text escalation).
Questions for maintainers
- Is there existing tooling for asserting on tool-call output specifically
(as opposed to text output) that I should reuse for scoring? - Any prior discussion/PR on agentic risk testing I should be aware of
before designing this?
Happy to take this on once the approach is validated.
- Lenguaje dominante
- Python
- Estrellas
- 4.5k
- Forks
- 896
- Merge medio
- 2 d 22 h
- PR fusionados (30 d)
- 230
Preparar el entorno
Aún no hemos revisado los archivos de configuración de este proyecto. Empieza por su README y consulta nuestra guía para la primera contribución para los pasos generales.
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/PyRIT
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
Los mantenedores suelen responder en 2 días
-
Bug: triage GUI help wanted
Dificultad 2/5 1-3 horas Aptitud para principiantes 86/100
microsoft/PyRIT#2868 · 1 comentario ·
Los mantenedores suelen responder en 2 días
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 88/100
Los mantenedores suelen responder en 2 días
-
Dificultad 3/5 1-2 días Aptitud para principiantes 76/100
Los mantenedores suelen responder en 2 días
Todos los issues de microsoft/PyRIT
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 85/100
kornia/kornia#5263 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Metadata correction for W16-5400Abiertoapproved correction metadata
Dificultad 1/5 Menos de una hora Aptitud para principiantes 88/100
acl-org/acl-anthology#10133 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
BasedHardware/omi#20084 ·
Los mantenedores suelen responder en 1 día
-
bug needs-acceptance wg/evaluation-quality
Dificultad 2/5 1-3 horas Aptitud para principiantes 76/100
vllm-project/semantic-router#4424 ·
Los mantenedores suelen responder en 1 día