Download Latest Version 2.2.1 source code.zip (3.5 MB) Google Add to Preferred Sources
Home / 2.2.0
Name Modified Size InfoDownloads / Week
Parent folder
cel-go-dependencies.json 2026-10-02 4.3 kB
cel-go-helpers.json 2026-10-02 5.1 kB
requirements.txt 2026-10-02 169.4 kB
sbom.cdx.json 2026-10-02 546.2 kB
2.2.0 source code.tar.gz 2026-10-02 2.9 MB
2.2.0 source code.zip 2026-10-02 3.5 MB
README.md 2026-10-02 4.8 kB
Totals: 7 Items   7.2 MB 0

Highlights

  • A setup guide built from measurements. Recommended Settings lists setups by goal. Every recommended setup runs the LLM judge. Results on MaliciousSkillBench's held-out split (839 malicious, 545 benign), judge on Gemma 4 26B-A4B:
Goal Setup Recall FPR F1
Highest F1 balanced + judge, review MEDIUM+ 66.7% 15.4% 75.5%
Smaller review queue low-noise + judge, review MEDIUM+ 63.2% 13.4% 73.5%
Lowest FPR quiet + judge, review MEDIUM+ 50.3% 7.2% 64.9%
Lowest FPR, gate only quiet + judge, block HIGH+ 33.1% 5.5% 48.5%
  • New low-noise and quiet presets (#230). They demote the rules that most often flag harmless real skills, and cap the judge's low-confidence findings (low_confidence_max_severity) and contextual-risk findings (contextual_risk_max_severity). quiet requires the judge.
  • A stronger judge (#230). A new prompt separates a skill's purpose from misuse. Repair of a self-contradicting SAFE verdict is on by default, and --llm-decompose runs one focused pass per concern. Bedrock structured output and meta-analyzer schema wiring are fixed.
  • Current default model: claude-sonnet-5-5, or bedrock/us.anthropic.claude-sonnet-5-5 on Bedrock (#241). The old default, claude-3-5-sonnet-20241022, is retired. Temperature is left out for models that reject it. Sonnet 5.5, Opus 5.5 and Fable 5.1 get plain JSON output with the schema in the prompt.
  • Scoped suppressions for rules, by skill and path (#223).
  • New providers:
  • Adjudicator credentials now resolve for every provider (#233).
  • Experimental on-device Apple Foundation Model support (#235).
  • Any OpenAI-compatible gateway works through SKILL_SCANNER_LLM_PROVIDER=openai-compatible.
  • Detection:
  • Frontmatter description and when_to_use are scanned (#229).
  • New SUPPLY_CHAIN_REGISTRY_REDIRECT rule (HIGH) and UNDECLARED_NETWORK_DESTINATION rule (LOW) (#236).
  • OSV reads package.json (#228).
  • The judge reads bounded excerpts of oversized files instead of dropping them (#231).
  • Pre-commit can run the judge: set use_llm, llm_model and llm_provider in .skill_scannerrc (#241).
  • Reusable workflow (#241): new scanner_version, llm_provider, llm_base_url and aws_region inputs. A bedrock/ model installs the [bedrock] extra, and configuration errors are reported separately from findings.
  • Docs (#240):
  • grouped website navigation
  • new LLM Providers and Results & Tuning pages
  • every page updated to list all five presets and current model IDs

Behavior changes

  • --use-llm, --use-behavioral and --enable-meta fail closed. If the requested analyzer can't be built (no key, missing extra, unknown provider), the scan exits 2 with the fix in the message, and the REST API returns 400. Previously the scan ran rules only and exited 0.
  • Exit codes: 2 means a usage or configuration error, such as an unknown policy (previously 1). 1 still means findings at or above --fail-on-severity.
  • Verdict repair is on by default. A SAFE verdict that lists findings becomes SUSPICIOUS and keeps the findings. To restore the strict path, set SKILL_SCANNER_LLM_REPAIR_INCONSISTENT_VERDICT=0.
  • Severity changes:
  • Unpinned dependencies are reported at LOW instead of MEDIUM.
  • FILE_MAGIC_MISMATCH can be LOW when only a text label disagrees with the extension.
  • Several false-positive fixes landed in the static and YARA rules (#230).
  • Frontmatter description and when_to_use are now scanned by the core rules, YARA and the active-directive rules. Opt-in packs still scan the body only.
  • The wizard no longer turns on the meta-analyzer by default. With the meta-analyzer working, it cost 16.4 points of recall.

Fixes

  • [#234]: scan-all --recursive --lenient no longer splits a skill's Markdown subfolders into separate skills (#238).
  • [#232]: --adjudicate resolves credentials for every provider (#233).
  • [#220]: oversized code files and SKILL.md bodies get bounded excerpts instead of being skipped (#231).
  • [#225]: the live C2 IP in the ATR description text is defanged.
  • A malformed meta-analysis priority_order no longer fails the batch (#224).
  • upload-sarif has actions: read (#221).
  • Dependency floors were raised to clear pip-audit advisories (#239).

Measured figures and methodology: https://huggingface.co/spaces/Vineethsain/skill-scanner-vs-skillspector and Measured results.

Full Changelog: https://github.com/cisco-ai-defense/skill-scanner/compare/2.1.0...2.2.0

Source: README.md, updated 2026-10-02