Hacktoberfest 2026: los issues que los mantenedores marcaron para octubre, abiertos y aptos para principiantes. Explorar issues de Hacktoberfest

Make the skill security scan gate deterministic

Cerrado
#1,047 0 comentarios 0 reacciones 1 asignado Ver en GitHub

Los mantenedores suelen responder en 1 día

@danbarr ya está trabajando en esto.

Desde el 5/10/2026.

Evaluación

Este issue todavía no se ha evaluado.

Descripción

needs-triage

Make the skill security scan gate deterministic

Problem

The blocking skill security scan (scripts/skill-scan/run_scan.py) still flips between pass and fail on byte-identical content. We've traced the flips to the scanner's LLM meta-analyzer (--enable-meta). It runs after all the analyzers and decides which findings are false positives. Since the ATR rule pack produces thousands of raw HIGH/CRITICAL regex hits, the meta-analyzer is effectively making the blocking decision, and it doesn't make the same call twice. That's how a skill passes on its PR and then fails on the merge to main (the PR scan cache isn't visible to main pushes, so main always rescans).

What changed in skill-scanner 2.2.0

Upstream's new Recommended Settings page lines up with what we've been seeing:

  • Avoid --enable-meta. In their measurements it cost 16.4 points of recall for a 0.3-point FPR gain, and it's now off by default.
  • Use a policy preset plus the LLM judge, and block at HIGH. New low-noise and quiet presets demote noisy rules to LOW and cap the judge's low-confidence findings at LOW. quiet also caps "contextual risk" findings (a risky capability with no sign of malicious intent).
  • Opt-in rule packs aren't part of any measured gating setup. They only show up in the unmeasured "hunting" setup, which is never meant to gate.

The ATR pack's own README says the same thing: ATR's enforce lane loads only maturity: stable rules, and all 38 of its skill-targeted rules are maturity: test. The bundled pack also grew from 470 rules (scanner 2.0.13) to 712 (2.2.0), so every scanner bump can add new CRITICAL regexes.

What we measured

Scanner 2.2.0, run locally against all 196 skills at their pinned refs.

Rules only (no LLM, no meta-analyzer). These findings are deterministic. They're what reaches process_scan_results.py once the meta-analyzer stops filtering.

Config Skills with HIGH+ HIGH+ findings Skills still blocked after existing allowlists Findings after allowlists
ATR + PromptGuard, default policy (current packs) 163 4,719 142 3,045
ATR + PromptGuard, low-noise 163 4,717 142 3,046
ATR + PromptGuard, quiet 162 4,702 142 3,035
PromptGuard only, default policy 31 87 20 48
PromptGuard only, low-noise 30 84 20 48
PromptGuard only, quiet 27 65 18 33

ATR accounts for 98% of HIGH+ hits, and the presets don't touch ATR rules at all. ATR rules also make up 251 of the 553 allowlist entries in the catalog today. Without ATR, quiet leaves 33 findings across 18 skills. I spot-checked these and they're false positives (credential-storage advice, subprocess.run in helper scripts, and so on). Allowlisting them is one-time work, because the findings don't change between runs.

With the LLM judge. 6 skills × 3 configs × 3 repeats, openai/gpt-5.6-terra (our CI model), consensus off so raw run-to-run variance shows. Sample: the two skills that flipped before (mongodb-mcp-setup, vercel-cli-with-tokens), plus claude-api, gha-security-review, huggingface-paper-publisher, and clickhouse-best-practices.

Config Gate verdict flips Blocking-set changes
Current (ATR + PromptGuard, --enable-meta) 1 of 6 skills 2 of 6 skills
PromptGuard only, low-noise, no meta 0 0
PromptGuard only, quiet, no meta 0 0
  • Current config flip: mongodb-mcp-setup went pass, pass, BLOCK. On the third run the meta-analyzer kept four ATR CRITICALs that it had dropped in the first two.
  • Current config, unstable reasons: claude-api blocked every run, but on a different set of rules each time (5, 4, then 2).
  • No-meta configs: every skill gave the same verdict, for the same reasons, on all three runs.
  • The judge itself never raised anything above MEDIUM in any of the 54 runs. All the variance came from the meta-analyzer. That also means the judge-side caps weren't exercised on this sample.
Proposal

In run_scan.py:

  1. Remove --enable-meta.
  2. Drop atr from --rule-packs, keeping promptguard.
  3. Add --policy quiet.
  4. Keep --use-llm, --use-trigger, --use-behavioral, --llm-consensus-runs 3, and the HIGH block threshold.

In the same PR, add per-skill allowlist entries for the 33 remaining findings so nothing that passes today starts failing on its next rescan.

This is close to upstream's "lowest FPR" recommended setup: quiet, judge on, block at HIGH. The differences are that we keep PromptGuard, trigger, and behavioral, and use a different judge model.

Tradeoffs
  • Losing ATR's signal. ATR has agent-specific signatures that the core rules don't. If we want them, we could run ATR in a separate pass that reports but never blocks (a PR comment or job summary).
  • quiet vs low-noise. Upstream measured quiet + judge at 50.3% recall / 7.2% FPR, and low-noise + judge at 63.2% / 13.4%, on their malicious-skill benchmark. On our catalog, low-noise would need 48 allowlist entries across 20 skills instead of 33 across 18. The extras are mostly 12 YARA_jailbreak_generic hits in the skill-scanner skill. quiet fits a gate on a vetted-publisher catalog better. low-noise is the choice if we care more about recall.
  • The judge is still an LLM. Without the meta-analyzer it adds findings but can no longer remove deterministic ones. With the caps, a low-confidence judge finding can't reach HIGH. A confident HIGH from the judge could still flip a verdict. We didn't see one in this sample, but six benign skills is a small sample.
Rollout notes
  • run_scan.py is part of the trusted-scan cache key, so changing it invalidates every cached verdict. Skills rescan as they're touched, not all at once. It doesn't touch cmd/dockhand/ or internal/skills/, so it won't trigger a full-catalog rebuild.
  • Follow-ups: prune the 251 ATR allowlist entries that would stop matching anything, and update docs/security.md and scripts/skill-scan/README.md (the README's rule counts are already out of date).
Open questions
  • quiet or low-noise?
  • Do we want the report-only ATR pass, or drop ATR entirely?
  • Is PG_PII_CREDENTIAL_HARVESTING (9 of the 33 remaining hits, across 6 skills, all false positives in the sample) better handled with per-skill allowlist entries or a policy override to MEDIUM?
Lenguaje dominante
Go
Estrellas
8
Forks
7
Merge medio
1 d 23 h
PR fusionados (30 d)
99

Preparar el entorno

Primeros pasos

  1. Lee el issue completo y luego la guía de contribución del proyecto.
  2. Comenta en el issue que vas a ocuparte — evita que dos personas hagan lo mismo.
  3. Haz un fork del repositorio y trabaja en una rama.
  4. Abre un pull request que haga referencia al número del issue.

Más de stacklok/dockyard

Todos los issues de stacklok/dockyard

Issues similares

Más issues de Go

Recibe los nuevos issues en tu correo

Un resumen breve de issues de GitHub para principiantes.