Files
SlipItIn/.github/skills/skill-evaluation/references/output-template.md
Tim Krampitz 01046b01e4 Neue Skills, Referenzen & OpenWiki-Doku integriert
Umfangreiche Erweiterung der Skill-Bibliothek: Neue Skills für Humanisierung (Englisch/PT-BR), Design-Validierung, AI-SEO und Coolify-Deployment inkl. Regelwerke, Presets, Pattern-Referenzen, Testfälle und Automatisierungsskripte. Zusätzliche Skills für Revenue-Centric Design, Pier Cloud, OKF, Lebenslauf- und LinkedIn-Optimierung sowie zahlreiche Referenzdateien, Checklisten und YAML/JSON/Markdown-Templates. Einführung einer vollständigen OpenWiki-Dokumentation mit Architektur-, Domain- und Workflow-Beschreibungen, zentralem Index und automatisierten Updates. Modularer Aufbau, restriktive Lizenzen und umfassende Qualitäts- und Evaluationsmechanismen für alle neuen Inhalte.
2026-07-26 14:00:58 +02:00

159 lines
5.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Output Template
Emit the scorecard exactly in this structure (step 9 of the workflow).
```markdown
# Skill Evaluation — {skill name}
> Evaluated: {date}
> Source: {path}
> Evaluator: skill-evaluation v2.1.0
> Framework: [Anthropic Skill Best Practices](https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills) + Matt Pocock's [writing-great-skills](https://www.youtube.com/watch?v=UNzCG3lw6O0)
## Summary
| Metric | Value |
|--------|-------|
| Overall Score | {weighted}/100 |
| Grade | {A/B/C/D/F} |
| Category | {category} |
| Invocation | {model-invoked / user-invoked} |
| Files | {count} |
| Criteria scored / N/A | {n} scored, {m} N/A |
## Scorecard
### Axis 1 — Trigger
| # | Criterion | Weight | Score | Notes |
|---|-----------|--------|-------|-------|
| 1 | Invocation design | 2x | {n}/100 | {evidence} |
| 2 | Description quality | 2x | {n}/100 | {evidence} |
### Axis 2 — Structure
| # | Criterion | Weight | Score | Notes |
|---|-----------|--------|-------|-------|
| 3 | Steps vs. reference clarity | 1x | {n}/100 | {evidence} |
| 4 | Branch-aware disclosure & pointers | 2x | {n}/100 | {evidence} |
| 5 | Conciseness | 2x | {n}/100 | {evidence} |
| 6 | Coherent scope | 1x | {n}/100 | {evidence} |
### Axis 3 — Steering
| # | Criterion | Weight | Score | Notes |
|---|-----------|--------|-------|-------|
| 7 | Leading words | 2x | {n}/100 | {evidence} |
| 8 | Completion criteria & legwork | 2x | {n/100 or N/A} | {evidence} |
| 9 | Gotchas section | 2x | {n}/100 | {evidence} |
| 10 | Grounded in expertise | 2x | {n}/100 | {evidence} |
| 11 | Avoids railroading | 1x | {n}/100 | {evidence} |
### Axis 4 — Pruning
| # | Criterion | Weight | Score | Notes |
|---|-----------|--------|-------|-------|
| 12 | No-ops (deletion test) | 2x | {n}/100 | {evidence with line citations} |
| 13 | Single source of truth | 1x | {n}/100 | {evidence} |
| 14 | Relevance & sediment | 1x | {n}/100 | {evidence} |
### Conditional criteria
| # | Criterion | Weight | Score | Notes |
|---|-----------|--------|-------|-------|
| 15 | Setup flow | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
| 16 | Memory mechanism | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
| 17 | Scripts & libraries | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
| 18 | On-demand hooks | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
## Trigger Eval
{For user-invoked skills, write: "N/A — user-invoked skill, no description to test."}
### Prompts tested
| # | Prompt | Expected | Triggered | Other skills |
|---|--------|----------|-----------|--------------|
| 1 | {prompt text} | should-trigger | yes/no | {list or none} |
| 2 | {prompt text} | should-trigger | yes/no | {list or none} |
| 3 | {prompt text} | should-trigger | yes/no | {list or none} |
| 4 | {prompt text} | should-trigger | yes/no | {list or none} |
| 5 | {prompt text} | should-trigger | yes/no | {list or none} |
| 6 | {prompt text} | should-not-trigger | yes/no | {list or none} |
| 7 | {prompt text} | should-not-trigger | yes/no | {list or none} |
| 8 | {prompt text} | should-not-trigger | yes/no | {list or none} |
| 9 | {prompt text} | should-not-trigger | yes/no | {list or none} |
| 10 | {prompt text} | should-not-trigger | yes/no | {list or none} |
### Results
| Metric | Value |
|--------|-------|
| Should-trigger hit rate | {X}/5 |
| Should-not-trigger leak rate | {X}/5 |
| Other skills observed | {list or none} |
### Observations
{Free-form notes: patterns in what triggered or didn't, description wording
gaps revealed, sibling skills that competed, etc.}
## Failure Modes Detected
| Mode | Evidence | Root cause | Defense |
|------|----------|------------|---------|
| {mode, or a single row "None detected"} | {file:line} | {cause} | {defense} |
## Prioritized Actions
### 1. {action}
**Evidence:** {file:line or section}
**Fix:** {specific recommendation}
### 2. {action}
**Evidence:** {file:line or section}
**Fix:** {specific recommendation}
(35 total, each tied to a detected failure mode)
## Bonus Patterns
| Pattern | Status | Notes |
|---------|--------|-------|
| Validation loops | {Present/Absent/N/A} | {detail} |
| Output templates | {Present/Absent/N/A} | {detail} |
| Procedures over declarations | {Present/Absent/N/A} | {detail} |
| Defaults over menus | {Present/Absent/N/A} | {detail} |
| Trace-checkable steering | {Present/Absent/N/A} | {detail} |
## Grade Scale
{copy the Grade Scale table from SKILL.md}
---
*Generated by [skill-evaluation](https://github.com/fabricioctelles/skills) v2.1.0, merging the [Anthropic skill quality framework](https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills) with Matt Pocock's [writing-great-skills](https://www.youtube.com/watch?v=UNzCG3lw6O0) methodology.*
```
## Comparison mode
When `compare` is set, add a side-by-side table across all 18 criteria.
Leave a cell N/A rather than scoring it 0, and exclude N/A rows from the
Overall row's weighted math for that skill.
```markdown
## Comparison: {skill A} vs {skill B}
| # | Criterion | {A} | {B} | Delta |
|---|-----------|-----|-----|-------|
| 1 | Invocation design | 60 | 85 | +25 |
| 2 | Description quality | 25 | 70 | +45 |
| ... | ... | ... | ... | ... |
| 15 | Setup flow | N/A | 80 | — |
| **Overall** | | **43** | **72** | **+29** |
```