Title: Feature Request: Attack strategy for Excessive Agency / unauthorized tool invocation
I maintainer di solito rispondono entro 2 giorni
Nessuno ha ancora preso questa issue.
Valutazione
- Difficoltà
- 5/5
- Tempo stimato
- Più di una settimana
- Idoneità per principianti
- 30/100
Direzione di ricerca
Start by reading openai_response_target.py and the existing Tool/ToolProvider interfaces, then compare the single-turn and multi-turn attack structures, including Crescendo. Confirm with maintainers how tool-call output can be inspected and scored. Done means an agreed attack design and implementation that evaluates actual unauthorized or excessive tool calls rather than only text output.
Scritto dal modello di indicizzazione a partire dal testo della issue.
Descrizione
Summary
PyRIT already supports exposing tools to a target via Tool / ToolProvider
(see openai_response_target.py), but I don't see an attack strategy that
specifically tests Excessive Agency — whether an agent can be manipulated
into invoking a tool outside its intended scope, chaining tool calls beyond
what a task requires, or calling a tool that its system prompt explicitly
restricts.
This maps to the OWASP Top 10 for LLM Applications (Excessive Agency /
Insecure Plugin Design) and is a distinct risk category from prompt injection
or jailbreaking the model's text output — it targets the agent's actions,
not just its words.
Proposed approach (open to feedback before implementing)
- A new attack, e.g.
ExcessiveAgencyAttack, that:- Takes a target configured with a defined set of "allowed" tools/scope
(via the existingTool/ToolProvidersystem). - Attempts to elicit a tool call outside that scope — either a tool the
agent has access to but shouldn't use for the stated task, or a chained
sequence of legitimate calls that together exceed the intended
permission boundary. - Scores success based on the tool call actually made (inspecting the
target's tool-call output), not on the text response — this is the
distinct piece existing text/output scorers don't cover.
- Takes a target configured with a defined set of "allowed" tools/scope
- Could reuse the single-turn or multi-turn structure depending on whether
the elicitation needs conversation history (multi-turn is likely more
realistic here, similar to how Crescendo works for text escalation).
Questions for maintainers
- Is there existing tooling for asserting on tool-call output specifically
(as opposed to text output) that I should reuse for scoring? - Any prior discussion/PR on agentic risk testing I should be aware of
before designing this?
Happy to take this on once the approach is validated.
- Lingua principale
- Python
- Stelle
- 4.5k
- Fork
- 896
- Merge medio
- 2g 22h
- PR unite (30g)
- 220
Preparare l'ambiente
Non abbiamo ancora controllato i file di configurazione di questo progetto. Parti dal suo README e consulta la nostra guida al primo contributo per i passaggi generali.
Come iniziare
- Leggi tutta la issue e poi la guida ai contributi del progetto.
- Commenta sulla issue per dire che te ne occupi tu — evita che due persone facciano lo stesso lavoro.
- Fai un fork del repository e lavora su un branch.
- Apri una pull request che faccia riferimento al numero della issue.
Altre issue di microsoft/PyRIT
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
microsoft/PyRIT#2905 · 3 commenti ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
I maintainer di solito rispondono entro 2 giorni
-
Bug: triage GUI help wanted
Difficoltà 2/5 1-3 ore Idoneità per principianti 86/100
microsoft/PyRIT#2868 · 1 commento ·
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 88/100
I maintainer di solito rispondono entro 2 giorni
-
Difficoltà 3/5 1-2 giorni Idoneità per principianti 72/100
I maintainer di solito rispondono entro 2 giorni
Tutte le issue di microsoft/PyRIT
Issue simili
-
Difficoltà 1/5 Meno di un'ora Idoneità per principianti 72/100
letsencrypt/cp-cps#353 ·
-
Marble Madness II is missingAperta
Difficoltà 2/5 1-3 ore Idoneità per principianti 68/100
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 84/100
PedestrianDynamics/pyFDS-Evac#394 ·
I maintainer di solito rispondono entro 1 giorno
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
DOI-USGS/pywatershed#421 ·
-
Difficoltà 2/5 1-3 ore Idoneità per principianti 78/100
python-pillow/Pillow#10087 · 1 commento ·
I maintainer di solito rispondono entro 1 giorno