Make the skill security scan gate deterministic
Los mantenedores suelen responder en 1 día
@danbarr ya está trabajando en esto.
Desde el 5/10/2026.
Evaluación
Este issue todavía no se ha evaluado.
Descripción
Make the skill security scan gate deterministic
Problem
The blocking skill security scan (scripts/skill-scan/run_scan.py) still flips between pass and fail on byte-identical content. We've traced the flips to the scanner's LLM meta-analyzer (--enable-meta). It runs after all the analyzers and decides which findings are false positives. Since the ATR rule pack produces thousands of raw HIGH/CRITICAL regex hits, the meta-analyzer is effectively making the blocking decision, and it doesn't make the same call twice. That's how a skill passes on its PR and then fails on the merge to main (the PR scan cache isn't visible to main pushes, so main always rescans).
What changed in skill-scanner 2.2.0
Upstream's new Recommended Settings page lines up with what we've been seeing:
- Avoid
--enable-meta. In their measurements it cost 16.4 points of recall for a 0.3-point FPR gain, and it's now off by default. - Use a policy preset plus the LLM judge, and block at HIGH. New
low-noiseandquietpresets demote noisy rules to LOW and cap the judge's low-confidence findings at LOW.quietalso caps "contextual risk" findings (a risky capability with no sign of malicious intent). - Opt-in rule packs aren't part of any measured gating setup. They only show up in the unmeasured "hunting" setup, which is never meant to gate.
The ATR pack's own README says the same thing: ATR's enforce lane loads only maturity: stable rules, and all 38 of its skill-targeted rules are maturity: test. The bundled pack also grew from 470 rules (scanner 2.0.13) to 712 (2.2.0), so every scanner bump can add new CRITICAL regexes.
What we measured
Scanner 2.2.0, run locally against all 196 skills at their pinned refs.
Rules only (no LLM, no meta-analyzer). These findings are deterministic. They're what reaches process_scan_results.py once the meta-analyzer stops filtering.
| Config | Skills with HIGH+ | HIGH+ findings | Skills still blocked after existing allowlists | Findings after allowlists |
|---|---|---|---|---|
| ATR + PromptGuard, default policy (current packs) | 163 | 4,719 | 142 | 3,045 |
ATR + PromptGuard, low-noise |
163 | 4,717 | 142 | 3,046 |
ATR + PromptGuard, quiet |
162 | 4,702 | 142 | 3,035 |
| PromptGuard only, default policy | 31 | 87 | 20 | 48 |
PromptGuard only, low-noise |
30 | 84 | 20 | 48 |
PromptGuard only, quiet |
27 | 65 | 18 | 33 |
ATR accounts for 98% of HIGH+ hits, and the presets don't touch ATR rules at all. ATR rules also make up 251 of the 553 allowlist entries in the catalog today. Without ATR, quiet leaves 33 findings across 18 skills. I spot-checked these and they're false positives (credential-storage advice, subprocess.run in helper scripts, and so on). Allowlisting them is one-time work, because the findings don't change between runs.
With the LLM judge. 6 skills × 3 configs × 3 repeats, openai/gpt-5.6-terra (our CI model), consensus off so raw run-to-run variance shows. Sample: the two skills that flipped before (mongodb-mcp-setup, vercel-cli-with-tokens), plus claude-api, gha-security-review, huggingface-paper-publisher, and clickhouse-best-practices.
| Config | Gate verdict flips | Blocking-set changes |
|---|---|---|
Current (ATR + PromptGuard, --enable-meta) |
1 of 6 skills | 2 of 6 skills |
PromptGuard only, low-noise, no meta |
0 | 0 |
PromptGuard only, quiet, no meta |
0 | 0 |
- Current config flip:
mongodb-mcp-setupwent pass, pass, BLOCK. On the third run the meta-analyzer kept four ATR CRITICALs that it had dropped in the first two. - Current config, unstable reasons:
claude-apiblocked every run, but on a different set of rules each time (5, 4, then 2). - No-meta configs: every skill gave the same verdict, for the same reasons, on all three runs.
- The judge itself never raised anything above MEDIUM in any of the 54 runs. All the variance came from the meta-analyzer. That also means the judge-side caps weren't exercised on this sample.
Proposal
In run_scan.py:
- Remove
--enable-meta. - Drop
atrfrom--rule-packs, keepingpromptguard. - Add
--policy quiet. - Keep
--use-llm,--use-trigger,--use-behavioral,--llm-consensus-runs 3, and the HIGH block threshold.
In the same PR, add per-skill allowlist entries for the 33 remaining findings so nothing that passes today starts failing on its next rescan.
This is close to upstream's "lowest FPR" recommended setup: quiet, judge on, block at HIGH. The differences are that we keep PromptGuard, trigger, and behavioral, and use a different judge model.
Tradeoffs
- Losing ATR's signal. ATR has agent-specific signatures that the core rules don't. If we want them, we could run ATR in a separate pass that reports but never blocks (a PR comment or job summary).
quietvslow-noise. Upstream measuredquiet+ judge at 50.3% recall / 7.2% FPR, andlow-noise+ judge at 63.2% / 13.4%, on their malicious-skill benchmark. On our catalog,low-noisewould need 48 allowlist entries across 20 skills instead of 33 across 18. The extras are mostly 12YARA_jailbreak_generichits in theskill-scannerskill.quietfits a gate on a vetted-publisher catalog better.low-noiseis the choice if we care more about recall.- The judge is still an LLM. Without the meta-analyzer it adds findings but can no longer remove deterministic ones. With the caps, a low-confidence judge finding can't reach HIGH. A confident HIGH from the judge could still flip a verdict. We didn't see one in this sample, but six benign skills is a small sample.
Rollout notes
run_scan.pyis part of the trusted-scan cache key, so changing it invalidates every cached verdict. Skills rescan as they're touched, not all at once. It doesn't touchcmd/dockhand/orinternal/skills/, so it won't trigger a full-catalog rebuild.- Follow-ups: prune the 251 ATR allowlist entries that would stop matching anything, and update
docs/security.mdandscripts/skill-scan/README.md(the README's rule counts are already out of date).
Open questions
quietorlow-noise?- Do we want the report-only ATR pass, or drop ATR entirely?
- Is
PG_PII_CREDENTIAL_HARVESTING(9 of the 33 remaining hits, across 6 skills, all false positives in the sample) better handled with per-skill allowlist entries or a policy override to MEDIUM?
- Lenguaje dominante
- Go
- Estrellas
- 8
- Forks
- 7
- Merge medio
- 1 d 23 h
- PR fusionados (30 d)
- 99
Preparar el entorno
- Sin Dockerfile ni archivo de Docker Compose
- Sin plantilla de pull request
- Leer la guía de contribución
Primeros pasos
- Lee el issue completo y luego la guía de contribución del proyecto.
- Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
- Haz un fork del repositorio y trabaja en una rama.
- Abre un pull request que haga referencia al número del issue.
Más de stacklok/dockyard
-
grype high security
Dificultad 2/5 1-3 horas Aptitud para principiantes 68/100
stacklok/dockyard#950 · 4 comentarios ·
Los mantenedores suelen responder en 1 día
-
needs-triage
Dificultad 5/5 Más de una semana Aptitud para principiantes 38/100
Los mantenedores suelen responder en 1 día
-
critical grype high security
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
Los mantenedores suelen responder en 1 día
-
critical grype high security
Dificultad 3/5 1-2 días Aptitud para principiantes 48/100
Los mantenedores suelen responder en 1 día
-
critical grype high security
Dificultad 3/5 1-2 días Aptitud para principiantes 55/100
stacklok/dockyard#1043 · 1 comentario ·
Los mantenedores suelen responder en 1 día
Todos los issues de stacklok/dockyard
Issues similares
-
agent-research-recommend agent-review-finding chore
Dificultad 2/5 1-3 horas Aptitud para principiantes 78/100
jordansmall/spindrift#4821 · 1 comentario ·
Los mantenedores suelen responder en 1 día
-
area:web
Dificultad 2/5 1-3 horas Aptitud para principiantes 82/100
praetorianer777/GoTome#178 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
oracle/go-oracledb#105 ·
Los mantenedores suelen responder en 1 día
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 84/100
-
Dificultad 2/5 1-3 horas Aptitud para principiantes 70/100
Los mantenedores suelen responder en 1 día