Python: [Feature]: A deterministic pre-execution verification middleware for agent actions — AgentDojo v2.2 re-run: ASR=0 / FP=0 (open artifacts, full fix-cycle trajectory inside)
Los mantenedores suelen responder en 1 día
@eavanvalkenburg ya está trabajando en esto.
Desde el 7/10/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Description
Discussions like #8862 (budget enforcement in AgentLoopMiddleware) and #7853 (sandbox abstraction for tool execution) point at the same gap: agent frameworks today observe actions, but the decision to block a dangerous action before it executes is usually left ad hoc. We built that missing piece and would like to offer it to this community as a candidate middleware contract.
What it is
agent-action-verifier (https://github.com/Lsy1533133/agent-action-verifier) is a deterministic pre-execution enforcement layer: every pending action (tool call, network egress, file write) is checked against the agent's declared plan before execution. The judge is fixed-constant math — no LLM call in the enforcement path, no learned parameters — so it adds microsecond-scale cost per action (P99 = 0.13 ms in the synthetic stress suite; production latency not yet measured) and is deterministic and auditable by construction.
Closed-loop results (each reported number is backed by a checksummed artifact in the repo)
- AgentDojo official benchmark, real LLM in the loop (97 tasks × 2 rounds, official injection suite, official
security()judgment): ASR = 0, FP = 0 in the v2.2 full re-run (FP-rate 95% upper bound ≈3.8% at n=97 benign rounds) - Published fix-cycle trajectory FP 16 → 1 → 0 across v1.2 → v2.1 → v2.2 — same-source iteration, failures included, not independent stability trials
- Cross-model spot check: GLM subset (24 benign + 24 attack): FP = 0, ASR = 0
- Synthetic stress layer (100,000 seeded scenarios): FN = 0, FP = 0; interception Wilson-95 lower bound 99.99%+ on that synthetic distribution
- White-box adaptive attacks: 49 cases, 16 adaptive vector families — 43 hard-blocked, 0 bypass; 6 boundary cases documented and not counted as bypasses under the stated threat model
Integration shape (matches your middleware seam)
Wrap the action-dispatch point; the verifier receives (declared plan, pending action) and returns PASS / VETO + rule id. verifier_interface.pyi in the repo specifies the contract; typical adapter is ~10 lines around an AgentLoopMiddleware. Rule families: out-of-plan action, egress breach, scope escalation, tool-consent violation, dangerous value class, arithmetic guard, uninitialised state.
Honest boundaries
Synthetic scenarios are abstracted from publicly disclosed incident categories, not production traffic. Same-source fix cycles ≠ independent stability trials. GLM subset vs full run differ in model and sample size and are not directly comparable. This layer complements monitoring/alignment — it does not replace them.
Verify it yourself
python verify_artifacts.py in the repo recomputes every SHA-256 chain and the Wilson-95 bounds from raw counts — no trust required, stdlib only.
We'd genuinely value feedback on the metrics methodology, and if a deterministic enforcement seam fits the framework's roadmap, if maintainers see a fit, we can align the contract with the framework's middleware design.
Code Sample
Language/SDK
Both
- Lenguaje dominante
- Python
- Estrellas
- 13.9k
- Forks
- 2.4k
- Merge medio
- 1 d 18 h
- PR fusionados (30 d)
- 443
Preparar el entorno
Inicia el contenedor de desarrollo del proyecto en tu navegador, con tu propia cuenta de GitHub.
- Sin Dockerfile ni archivo de Docker Compose
- Tiene una plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de microsoft/agent-framework
-
.NET: Proposal: add an llms.txtAbierto.NET python triage
Dificultad 1/5 Menos de una hora Aptitud para principiantes 85/100
microsoft/agent-framework#9092 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Python: raw-data content mappings lose annotations and attachment metadataPosiblemente ocupada @moonbox3 la tomó hace 8 días. Abiertopython triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
microsoft/agent-framework#8632 · 2 comentarios ·
Los mantenedores suelen responder en 1 día
-
Python: Clarify when to use platformPosiblemente ocupada @eavanvalkenburg la tomó hace 8 días. Abiertopython triage
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
microsoft/agent-framework#8599 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
.NET Compaction - Update docs to refer to `AIContextProvider` deep divePosiblemente ocupada Un pull request vinculado a esta issue está abierto o ya se fusionó. Abierto.NET compaction documentation
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
microsoft/agent-framework#4629 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
Python: [Bug]: Media type detection documentation examples contain invalid base64Posiblemente ocupada @eavanvalkenburg la tomó hoy. Abiertoagents python reproduced
microsoft/agent-framework#9187 · 1 comentario · 1 asignado ·
Los mantenedores suelen responder en 1 día
Todos los issues de microsoft/agent-framework
Issues similares
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
UKGovernmentBEIS/inspect_ai#5781 ·
Los mantenedores suelen responder en 2 días
-
Bump .cicd to wamp-cicd 4c2f9ac: `just land` refuses open A18 decisions, `just where` lists themAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 82/100
crossbario/cfxdb#139 ·
-
Bump .cicd to wamp-cicd 4c2f9ac: `just land` refuses open A18 decisions, `just where` lists themAbierto
Dificultad 1/5 Menos de una hora Aptitud para principiantes 84/100
crossbario/txaio#241 ·
-
UX
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
mediajunkie/piper-morgan-product#1963 ·
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 64/100