The gist
A skill without testing is like hiring a cook and never tasting the food. Maybe they cook brilliantly. Maybe they put salt where the sugar goes. An eval is a tasting system: it takes a skill, runs it through real scenarios, gives a precise score and shows exactly where things went wrong. A skill with a low pass rate can be noticeably improved in a few iterations, with no guessing.
Key concepts
- eval.json: the file that tests a skill's output. It sets specific inputs and checks the quality of what comes out
- trigger_eval.json: activation testing. When the skill should kick in and when it shouldn't
- Pass rate: the percentage of test cases the skill passed
- Assertions: specific checks such as exact count, character limit, format compliance, relevance
- Improvement loop: write → eval → find weaknesses → improve → eval again
- Skill Creator: the
skill-creatorplugin from Anthropic's official catalog, which helps you create tests and run them
About file names. The current Skill Creator (as of October 2026, per the Claude Code documentation) stores tests in evals/evals.json, writes the grades of the checks to grading.json, and the "with skill vs. without" comparison to benchmark.json. It checks triggering through description tuning on sets of "should / should not trigger" queries. Below, the files are called eval.json and trigger_eval.json, as in a teaching diagram: the principle is the same, but for the exact names and fields, look at what the plugin creates for you.
Theory
What an eval is and why you need it
Without an eval, you improve a skill by feel: "seems better now." With an eval, you see numbers.
An example from a teaching demo (the numbers are illustrative, yours will differ): a skill for generating YouTube titles was tested with an eval:
| Metric | With skill | Without skill |
|---|---|---|
| Pass rate | 100% | 33–50% |
| Exact count (10 titles) | 6/6 tests | 2/6 tests |
| Character limit (60 characters) | 100% compliance | often broken |
| Variety of phrasing | high | low |
One iteration and the difference is visible. Without an eval, this kind of analysis would take a lot of manual testing.
eval.json: testing output quality
It's a file with 3–5 test cases. Each case has:
- Input: a specific request to the skill
- Expected: what should come out
- Assertions: specific checks you can measure
The structure of eval.json:
{
"skill": "youtube-title-generation",
"test_cases": [
{
"id": "beginner-coding-video",
"input": {
"topic": "Claude Code for beginners",
"angle": "first steps with no programming knowledge",
"count": 10
},
"assertions": [
{
"type": "exact_count",
"value": 10,
"description": "Must be exactly 10 titles"
},
{
"type": "max_length",
"value": 60,
"description": "Each title is no more than 60 characters"
},
{
"type": "framework_adherence",
"description": "Frameworks used: curiosity, specificity, emotion"
},
{
"type": "topic_relevance",
"description": "All titles are relevant to Claude Code for beginners"
},
{
"type": "variety",
"description": "At least 5 different title structures (not the same template)"
}
]
},
{
"id": "tool-comparison-video",
"input": {
"topic": "VS Code vs Cursor vs Devin Desktop (formerly Windsurf) comparison",
"angle": "what a developer should pick in 2026",
"count": 10
},
"assertions": [
{"type": "exact_count", "value": 10},
{"type": "max_length", "value": 60},
{"type": "includes_comparison", "description": "At least 3 titles contain an explicit comparison"},
{"type": "framework_adherence"},
{"type": "variety"}
]
}
]
}trigger_eval.json: testing activation
A skill should kick in on the right requests and stay quiet on unrelated ones. A trigger eval checks exactly that.
Structure: 20 tests, 10 "should trigger" + 10 "should not":
{
"skill": "youtube-title-generation",
"trigger_tests": {
"should_trigger": [
"come up with titles for a video about AI",
"give me 10 ideas for a video name",
"help me name a video about Claude Code",
"brainstorm YouTube titles for my tutorial",
"which titles will make the video clickable",
"I want 15 title options to compare",
"generate titles for my topic",
"what should I put in the video title",
"suggestions for video title",
"come up with catchy titles"
],
"should_not_trigger": [
"write a script for a YouTube video",
"make a thumbnail for the channel",
"how do I improve my channel's SEO",
"write a blog post on this topic",
"how do I get more YouTube subscribers",
"help with the video description",
"channel growth strategy",
"competitor analysis in my niche",
"how do I edit a video",
"come up with a content plan for the month"
]
}
}The self-improvement loop
This isn't a one-time test. It's a loop you repeat until you're satisfied.
Skill version 1
↓
Run eval → result: 60% pass rate
↓
Analysis: what exactly failed?
- exact_count: 8 out of 10 (doesn't count correctly)
- max_length: broken in 3 cases
↓
Improve the skill:
- add an explicit rule "always exactly N titles"
- add a rule "60 characters max, check each one"
↓
Skill version 2
↓
Run eval → result: 90% pass rate
↓
One more iteration → 100%The idea of the loop: tell the skill to improve, run the eval again, look at the numbers. It's a repeating loop where the skill improves based on data, not feelings (provided you're the one approving the improvements).
Common mistakes with evals
Not running the eval after changing the skill. Changed one line in the skill? Run the eval. A small change can break triggering or change the output format. The eval will show it fast.
Too few test cases. 1-2 test cases don't cover edge cases. At least 3 cases for eval.json and 10+10 for trigger_eval.json (10 "should trigger" + 10 "should not").
Testing only the happy path. "The skill works when everything's fine" isn't enough. Add cases with bad input, empty data, boundary values.
Not saving eval results between iterations. If you don't record each version's pass rate, you can't see progress. Keep a log: v1 = 44%, v2 = 71%, v3 = 89%.
A real eval.json with 3 test cases
{
"skill": "email-cold-outreach",
"test_cases": [
{
"id": "saas-founder",
"input": {
"recipient_role": "CEO of a SaaS startup",
"product": "AI support automation",
"tone": "professional"
},
"assertions": [
{"type": "max_length", "value": 200, "description": "No more than 200 words"},
{"type": "includes", "value": "call-to-action", "description": "Has a specific CTA"},
{"type": "excludes", "value": "spam words", "description": "No words like: free, urgent, exclusive offer"}
]
},
{
"id": "ecommerce-manager",
"input": {
"recipient_role": "E-commerce marketer",
"product": "SEO audit",
"tone": "casual"
},
"assertions": [
{"type": "max_length", "value": 150, "description": "Casual = shorter"},
{"type": "tone_check", "description": "Informal tone, no corporate jargon"},
{"type": "includes", "value": "personalization", "description": "Personalized to the role"}
]
},
{
"id": "edge-case-empty",
"input": {
"recipient_role": "",
"product": "AI tool",
"tone": "professional"
},
"assertions": [
{"type": "graceful_handling", "description": "The skill handles an empty field without an error"},
{"type": "fallback", "description": "Uses a generic greeting if no role is given"}
]
}
]
}How to run an eval: commands
First install the plugin (the /plugin menu will show the catalog name):
/plugin install skill-creator@claude-plugins-official
Create an eval for an existing skill and run it:
Check my youtube-title-generation skill with skill-creator: write tests for it and run them
The Claude Code documentation gives this example phrasing: "evaluate my summarize-changes skill with skill-creator". The tests go into an evals/ folder inside the skill folder, and each run happens in a separate subagent.
The skill folder after creating an eval (a diagram for understanding; real file names may differ):
.claude/skills/youtube-title-generation/
├── SKILL.md ← the main skill file
├── evals/
│ ├── eval.json ← output quality tests (in the plugin: evals.json)
│ └── trigger_eval.json ← activation tests
└── references/
└── title-examples.md ← title examples (if any)Metrics in the eval report
After a run, you get a report:
EVAL REPORT: youtube-title-generation ====================================== Test case 1: "Claude Code for beginners" WITH skill: 6/6 assertions PASSED ✓ WITHOUT skill: 2/6 assertions passed ✗ Test case 2: "VS Code vs Cursor vs Devin Desktop" WITH skill: 5/6 assertions PASSED ✓ WITHOUT skill: 3/6 assertions passed ✗ SUMMARY: With skill: 91.7% pass rate (11/12 assertions) Without skill: 41.7% pass rate (5/12 assertions) ANALYSIS: Biggest advantage: format compliance (+100%) Weakness found: exact_count failed in test 2 Recommendation: add explicit counting rule to skill
When you see "Weakness found," that's a hint about exactly what to improve in the skill. No guessing. The numbers speak for themselves.
Practice
Assignment: create an eval for an existing skill
- Pick any skill you created in earlier lessons (or create a simple skill for this exercise)
- Give the command:
Write tests for my [name] skill with skill-creator. Skill Creator will generate test cases and activation checks - Look through the eval.json and trigger_eval.json it created: do you understand what they check?
- Give the command:
Run the tests for my [name] skill - Read the report: what's the pass rate? What failed?
- Improve the skill based on the weak spots the eval found
- Run the eval again and compare the pass rate before and after
Goal: reach at least an 80% pass rate (a rough benchmark) and understand the logic of iterating
Tools and resources
- Skill Creator: install with
/plugin install skill-creator@claude-plugins-official(or find it through the/pluginmenu) - eval.json (in the plugin:
evals/evals.json): quality tests, created in theevals/folder inside the skill - trigger_eval.json: activation tests (in the plugin this is description tuning on "should / should not trigger" sets)
- Claude Code Skills documentation: the official guide to skills, including the section on testing skills
- Claude Code: requests like "check my skill with skill-creator" work right in the chat
Key takeaways
Evals move skill improvement from the realm of feelings to the realm of numbers. The pass rate grows through iterations, not guesswork (in the demo, from 60% to 90%+ in a couple of iterations; the numbers are illustrative).
eval.json tests output quality (what the skill produces). trigger_eval.json tests when the skill should activate. You need both.
The loop: skill → eval → find a weakness → improve → eval again. Repeat until you're happy with the result. Don't stop at the first version.
Related lessons
- ← Building a skill from scratch LIVE: creating a skill with Skill Creator, a first look at evals
- ← Skills architecture: the two skill archetypes, the quarterly audit
- ← What Skills are: the basic concepts of skills and frontmatter
Next lesson
→ Use Don't Build: the Skills Ecosystem: when to grab something ready-made and when to build your own. After that: Hooks: automatic rules in Claude Code.
The mark stays in this browser only and is never sent anywhere. My progress