Library · Skills: teach the agent your way of working

Evals: skills that improve themselves

Engineer65 minUpdated: October 2026
32 of 105 in the library

Module: Skills: reusable expertise | Time: ~25 min theory + 40 min practice


The gist

A skill without testing is like hiring a cook and never tasting the food. Maybe they cook brilliantly. Maybe they put salt where the sugar goes. An eval is a tasting system: it takes a skill, runs it through real scenarios, gives a precise score and shows exactly where things went wrong. A skill with a low pass rate can be noticeably improved in a few iterations, with no guessing.


Key concepts

  • eval.json: the file that tests a skill's output. It sets specific inputs and checks the quality of what comes out
  • trigger_eval.json: activation testing. When the skill should kick in and when it shouldn't
  • Pass rate: the percentage of test cases the skill passed
  • Assertions: specific checks such as exact count, character limit, format compliance, relevance
  • Improvement loop: write → eval → find weaknesses → improve → eval again
  • Skill Creator: the skill-creator plugin from Anthropic's official catalog, which helps you create tests and run them

About file names. The current Skill Creator (as of October 2026, per the Claude Code documentation) stores tests in evals/evals.json, writes the grades of the checks to grading.json, and the "with skill vs. without" comparison to benchmark.json. It checks triggering through description tuning on sets of "should / should not trigger" queries. Below, the files are called eval.json and trigger_eval.json, as in a teaching diagram: the principle is the same, but for the exact names and fields, look at what the plugin creates for you.


Theory

What an eval is and why you need it

Without an eval, you improve a skill by feel: "seems better now." With an eval, you see numbers.

🎨 Picture this: an eval is like hiring a test panel for a restaurant. Not "seems tasty," but 20 people with a questionnaire: food temperature, wait time, matches the menu, salt, presentation. The numbers tell you what to improve. Feelings lie.

An example from a teaching demo (the numbers are illustrative, yours will differ): a skill for generating YouTube titles was tested with an eval:

Metric With skill Without skill
Pass rate 100% 33–50%
Exact count (10 titles) 6/6 tests 2/6 tests
Character limit (60 characters) 100% compliance often broken
Variety of phrasing high low

One iteration and the difference is visible. Without an eval, this kind of analysis would take a lot of manual testing.


eval.json: testing output quality

🎨 Picture this: eval.json is the spec sheet for a product at a factory. Each test case is a specific line item: length 10 to 60 mm, load up to 500 kg, color RAL 3020. The product passes inspection and goes to the warehouse. If it fails, it goes back to the shop floor.

It's a file with 3–5 test cases. Each case has:

  1. Input: a specific request to the skill
  2. Expected: what should come out
  3. Assertions: specific checks you can measure

The structure of eval.json:

json
{
  "skill": "youtube-title-generation",
  "test_cases": [
    {
      "id": "beginner-coding-video",
      "input": {
        "topic": "Claude Code for beginners",
        "angle": "first steps with no programming knowledge",
        "count": 10
      },
      "assertions": [
        {
          "type": "exact_count",
          "value": 10,
          "description": "Must be exactly 10 titles"
        },
        {
          "type": "max_length",
          "value": 60,
          "description": "Each title is no more than 60 characters"
        },
        {
          "type": "framework_adherence",
          "description": "Frameworks used: curiosity, specificity, emotion"
        },
        {
          "type": "topic_relevance",
          "description": "All titles are relevant to Claude Code for beginners"
        },
        {
          "type": "variety",
          "description": "At least 5 different title structures (not the same template)"
        }
      ]
    },
    {
      "id": "tool-comparison-video",
      "input": {
        "topic": "VS Code vs Cursor vs Devin Desktop (formerly Windsurf) comparison",
        "angle": "what a developer should pick in 2026",
        "count": 10
      },
      "assertions": [
        {"type": "exact_count", "value": 10},
        {"type": "max_length", "value": 60},
        {"type": "includes_comparison", "description": "At least 3 titles contain an explicit comparison"},
        {"type": "framework_adherence"},
        {"type": "variety"}
      ]
    }
  ]
}

trigger_eval.json: testing activation

A skill should kick in on the right requests and stay quiet on unrelated ones. A trigger eval checks exactly that.

Structure: 20 tests, 10 "should trigger" + 10 "should not":

json
{
  "skill": "youtube-title-generation",
  "trigger_tests": {
    "should_trigger": [
      "come up with titles for a video about AI",
      "give me 10 ideas for a video name",
      "help me name a video about Claude Code",
      "brainstorm YouTube titles for my tutorial",
      "which titles will make the video clickable",
      "I want 15 title options to compare",
      "generate titles for my topic",
      "what should I put in the video title",
      "suggestions for video title",
      "come up with catchy titles"
    ],
    "should_not_trigger": [
      "write a script for a YouTube video",
      "make a thumbnail for the channel",
      "how do I improve my channel's SEO",
      "write a blog post on this topic",
      "how do I get more YouTube subscribers",
      "help with the video description",
      "channel growth strategy",
      "competitor analysis in my niche",
      "how do I edit a video",
      "come up with a content plan for the month"
    ]
  }
}

🎨 Picture this: trigger_eval is a test for a security guard. The guard should let in only the right people (should_trigger) and stop strangers (should_not_trigger). If the guard lets everyone in, the skill fires on unrelated requests. If the guard stops everyone, the skill never works at all.


The self-improvement loop

This isn't a one-time test. It's a loop you repeat until you're satisfied.

Code
Skill version 1
     ↓
Run eval → result: 60% pass rate
     ↓
Analysis: what exactly failed?
  - exact_count: 8 out of 10 (doesn't count correctly)
  - max_length: broken in 3 cases
     ↓
Improve the skill:
  - add an explicit rule "always exactly N titles"
  - add a rule "60 characters max, check each one"
     ↓
Skill version 2
     ↓
Run eval → result: 90% pass rate
     ↓
One more iteration → 100%

🎨 Picture this: the skill improvement loop is like an athlete's training. Run 5 km, time it. See where you ran out of breath. Work on your breathing. Run again. Without timing, it's just "feels faster." With timing, it's actual seconds of progress.

The idea of the loop: tell the skill to improve, run the eval again, look at the numbers. It's a repeating loop where the skill improves based on data, not feelings (provided you're the one approving the improvements).


Common mistakes with evals

  1. Not running the eval after changing the skill. Changed one line in the skill? Run the eval. A small change can break triggering or change the output format. The eval will show it fast.

  2. Too few test cases. 1-2 test cases don't cover edge cases. At least 3 cases for eval.json and 10+10 for trigger_eval.json (10 "should trigger" + 10 "should not").

  3. Testing only the happy path. "The skill works when everything's fine" isn't enough. Add cases with bad input, empty data, boundary values.

  4. Not saving eval results between iterations. If you don't record each version's pass rate, you can't see progress. Keep a log: v1 = 44%, v2 = 71%, v3 = 89%.


A real eval.json with 3 test cases

json
{
  "skill": "email-cold-outreach",
  "test_cases": [
    {
      "id": "saas-founder",
      "input": {
        "recipient_role": "CEO of a SaaS startup",
        "product": "AI support automation",
        "tone": "professional"
      },
      "assertions": [
        {"type": "max_length", "value": 200, "description": "No more than 200 words"},
        {"type": "includes", "value": "call-to-action", "description": "Has a specific CTA"},
        {"type": "excludes", "value": "spam words", "description": "No words like: free, urgent, exclusive offer"}
      ]
    },
    {
      "id": "ecommerce-manager",
      "input": {
        "recipient_role": "E-commerce marketer",
        "product": "SEO audit",
        "tone": "casual"
      },
      "assertions": [
        {"type": "max_length", "value": 150, "description": "Casual = shorter"},
        {"type": "tone_check", "description": "Informal tone, no corporate jargon"},
        {"type": "includes", "value": "personalization", "description": "Personalized to the role"}
      ]
    },
    {
      "id": "edge-case-empty",
      "input": {
        "recipient_role": "",
        "product": "AI tool",
        "tone": "professional"
      },
      "assertions": [
        {"type": "graceful_handling", "description": "The skill handles an empty field without an error"},
        {"type": "fallback", "description": "Uses a generic greeting if no role is given"}
      ]
    }
  ]
}

How to run an eval: commands

First install the plugin (the /plugin menu will show the catalog name):

Type this into the chat
/plugin install skill-creator@claude-plugins-official

Create an eval for an existing skill and run it:

Type this into the chat
Check my youtube-title-generation skill with skill-creator: write tests for it and run them

The Claude Code documentation gives this example phrasing: "evaluate my summarize-changes skill with skill-creator". The tests go into an evals/ folder inside the skill folder, and each run happens in a separate subagent.

The skill folder after creating an eval (a diagram for understanding; real file names may differ):

Code
.claude/skills/youtube-title-generation/
├── SKILL.md              ← the main skill file
├── evals/
│   ├── eval.json         ← output quality tests (in the plugin: evals.json)
│   └── trigger_eval.json ← activation tests
└── references/
    └── title-examples.md ← title examples (if any)

Metrics in the eval report

After a run, you get a report:

Type this into the chat
EVAL REPORT: youtube-title-generation
======================================
Test case 1: "Claude Code for beginners"
  WITH skill:    6/6 assertions PASSED ✓
  WITHOUT skill: 2/6 assertions passed ✗

Test case 2: "VS Code vs Cursor vs Devin Desktop"
  WITH skill:    5/6 assertions PASSED ✓
  WITHOUT skill: 3/6 assertions passed ✗

SUMMARY:
  With skill:    91.7% pass rate (11/12 assertions)
  Without skill: 41.7% pass rate (5/12 assertions)

ANALYSIS:
  Biggest advantage: format compliance (+100%)
  Weakness found: exact_count failed in test 2
  Recommendation: add explicit counting rule to skill

🎨 Picture this: an eval report is like a printout of lab results from your doctor. "Hemoglobin 11 g/dL, below normal. Recommended: iron." No need to guess what hurts. A specific number, a specific treatment.

When you see "Weakness found," that's a hint about exactly what to improve in the skill. No guessing. The numbers speak for themselves.


Practice

Assignment: create an eval for an existing skill

  1. Pick any skill you created in earlier lessons (or create a simple skill for this exercise)
  2. Give the command: Write tests for my [name] skill with skill-creator. Skill Creator will generate test cases and activation checks
  3. Look through the eval.json and trigger_eval.json it created: do you understand what they check?
  4. Give the command: Run the tests for my [name] skill
  5. Read the report: what's the pass rate? What failed?
  6. Improve the skill based on the weak spots the eval found
  7. Run the eval again and compare the pass rate before and after

Goal: reach at least an 80% pass rate (a rough benchmark) and understand the logic of iterating


Tools and resources

  • Skill Creator: install with /plugin install skill-creator@claude-plugins-official (or find it through the /plugin menu)
  • eval.json (in the plugin: evals/evals.json): quality tests, created in the evals/ folder inside the skill
  • trigger_eval.json: activation tests (in the plugin this is description tuning on "should / should not trigger" sets)
  • Claude Code Skills documentation: the official guide to skills, including the section on testing skills
  • Claude Code: requests like "check my skill with skill-creator" work right in the chat

Key takeaways

Evals move skill improvement from the realm of feelings to the realm of numbers. The pass rate grows through iterations, not guesswork (in the demo, from 60% to 90%+ in a couple of iterations; the numbers are illustrative).

eval.json tests output quality (what the skill produces). trigger_eval.json tests when the skill should activate. You need both.

The loop: skill → eval → find a weakness → improve → eval again. Repeat until you're happy with the result. Don't stop at the first version.



Next lesson

→ Use Don't Build: the Skills Ecosystem: when to grab something ready-made and when to build your own. After that: Hooks: automatic rules in Claude Code.

The mark stays in this browser only and is never sent anywhere. My progress