← Browse

@aws-samples/skill-eval

B

Evaluate AI Agent Skills across safety, quality, reliability, and cost efficiency. Audit for security issues (secrets, injection, unsafe installs), test functional correctness with-skill vs without-skill, measure trigger precision, classify cost-efficiency tradeoffs, track version lifecycle, and generate unified grades. Use when evaluating a skill before installing, auditing marketplace skills, proving your skill works with automated tests, setting up CI/CD quality gates, or comparing two skill versions. NOT for: evaluating full agent systems, testing non-skill plugins, runtime performance benchmarking, or monitoring production agent behavior.

skillclaude

Install

agr install @aws-samples/skill-eval --target claude

Writes 9 files into .claude/skills/, pinned to git-75a05043.

  • .claude/skills/skill-eval/.gitignore
  • .claude/skills/skill-eval/AGENTS.md
  • .claude/skills/skill-eval/CODE_OF_CONDUCT.md
  • .claude/skills/skill-eval/CONTRIBUTING.md
  • .claude/skills/skill-eval/LICENSE
  • .claude/skills/skill-eval/README.md
  • .claude/skills/skill-eval/SKILL.md
  • .claude/skills/skill-eval/demo.sh
  • .claude/skills/skill-eval/pyproject.toml

Document


name: skill-eval description: "Evaluate AI Agent Skills across safety, quality, reliability, and cost efficiency. Audit for security issues (secrets, injection, unsafe installs), test functional correctness with-skill vs without-skill, measure trigger precision, classify cost-efficiency tradeoffs, track version lifecycle, and generate unified grades. Use when evaluating a skill before installing, auditing marketplace skills, proving your skill works with automated tests, setting up CI/CD quality gates, or comparing two skill versions. NOT for: evaluating full agent systems, testing non-skill plugins, runtime performance benchmarking, or monitoring production agent behavior."

Skill Eval — Agent Skill Evaluation Framework

Evaluate Agent Skills across four dimensions: safety (audit), quality (functional), reliability (trigger), and cost efficiency (Pareto classification).

Quick Start

skill-eval audit /path/to/skill          # Is it safe?
skill-eval report /path/to/skill         # Full grade (audit + functional + trigger)
skill-eval functional /path/to/skill     # Quality: with-skill vs without-skill
skill-eval trigger /path/to/skill        # Reliability: activation precision

Decision Tree

  • "Is this skill safe?"skill-eval audit <path>
  • "Full evaluation with grade"skill-eval report <path>
  • "Full repo security review"skill-eval audit <path> --include-all
  • "Write eval cases"skill-eval init <path>, then edit evals/
  • "Compare two versions"skill-eval compare <old> <new>
  • "Check for regressions"skill-eval snapshot <path>, then skill-eval regression <path>
  • "Track changes"skill-eval lifecycle <path> --save --label v1.0

Commands

CommandPurpose
auditSecurity & structure scan (secrets, permissions, spec compliance)
functionalQuality eval — runs prompts with and without skill, grades output
triggerReliability eval — tests activation precision for relevant/irrelevant queries
reportUnified grade combining audit (40%) + functional (40%) + trigger (20%)
compareSide-by-side comparison of two skills on the same eval cases
snapshotSave current audit as regression baseline
regressionCheck for score regressions against baseline
lifecycleVersion tracking and change detection
initGenerate eval scaffold from SKILL.md frontmatter

For detailed flags and examples, see references/cli-reference.md.

Eval File Format

Functional evals (evals/evals.json):

[{"id": "case-1", "prompt": "...", "assertions": ["contains 'expected'"], "files": ["files/input.csv"]}]

Trigger queries (evals/eval_queries.json):

[{"query": "relevant question", "should_trigger": true}, {"query": "unrelated question", "should_trigger": false}]

Scoring

Grades: A (90+), B (80-89), C (70-79), D (60-69), F (<60). Findings deduct: CRITICAL −25, WARNING −10, INFO −2.

For the full security check reference and OWASP mapping, see references/security-checks.md.

Trustgrade B

  • passBody integrity

    Whether the stored document is plausibly the kind of file the artifact declares, rather than something fetched by mistake.

  • passType matchnot applicable to this artifact type

    Whether the artifact is really the kind of thing its metadata claims it is.

  • passFreshness

    How long since the source repository was last pushed to.

  • passPrompt injection

    Scans the artifact's own text for instructions aimed at your agent rather than at you.

  • warnLicensecopyleft/unknown — index-and-link only

    Whether the source repository declares an SPDX license permissive enough to redistribute.

How the grade is calculated

Each check contributes 0 points when it passes, 1 when it warns, and 2 when it fails. The total maps to a letter:

  • Aevery check passed
  • Bone warning
  • Ctwo warnings
  • Dprompt injection or body integrity failed, or three warnings
  • Fone of those failed, and something else is wrong

These are automated hygiene checks, not a security audit, and not a dependency or vulnerability scan. A grade of A means nothing was flagged — not that the artifact is safe.

Versions

  • git-75a0504345b12026-07-31