Neue Skills, Referenzen & OpenWiki-Doku integriert
Umfangreiche Erweiterung der Skill-Bibliothek: Neue Skills für Humanisierung (Englisch/PT-BR), Design-Validierung, AI-SEO und Coolify-Deployment inkl. Regelwerke, Presets, Pattern-Referenzen, Testfälle und Automatisierungsskripte. Zusätzliche Skills für Revenue-Centric Design, Pier Cloud, OKF, Lebenslauf- und LinkedIn-Optimierung sowie zahlreiche Referenzdateien, Checklisten und YAML/JSON/Markdown-Templates. Einführung einer vollständigen OpenWiki-Dokumentation mit Architektur-, Domain- und Workflow-Beschreibungen, zentralem Index und automatisierten Updates. Modularer Aufbau, restriktive Lizenzen und umfassende Qualitäts- und Evaluationsmechanismen für alle neuen Inhalte.
This commit is contained in:
112
.github/skills/skill-evaluation/references/categories.md
vendored
Normal file
112
.github/skills/skill-evaluation/references/categories.md
vendored
Normal file
@@ -0,0 +1,112 @@
|
||||
# Skill Categories Reference
|
||||
|
||||
Sources:
|
||||
- [Lessons from building Claude Code: How we use skills](https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills) — Anthropic, Jun 2026
|
||||
- [Best practices for skill creators](https://agentskills.io/skill-creation/best-practices) — Agent Skills spec
|
||||
- [Extend Claude with skills](https://code.claude.com/docs/en/skills) — Claude Code docs
|
||||
|
||||
---
|
||||
|
||||
## 1. `library-and-api-reference`
|
||||
|
||||
Skills that explain how to correctly use a library, CLI, or SDK. Can be internal or public libraries that the model struggles with. Often include reference code snippets and gotchas lists.
|
||||
|
||||
**Signals:** Has API endpoint docs, CLI command reference, code examples, "how to call X" patterns.
|
||||
|
||||
**Examples:** billing-lib, internal-platform-cli, sandbox-proxy
|
||||
|
||||
---
|
||||
|
||||
## 2. `product-verification`
|
||||
|
||||
Skills that describe how to test or verify code is working. Often paired with Playwright, tmux, or other external tools. These have the most measurable impact on output quality — worth investing an engineer-week.
|
||||
|
||||
**Signals:** Has test scripts, assertion patterns, Playwright/Cypress flows, "verify that X" instructions.
|
||||
|
||||
**Examples:** signup-flow-driver, checkout-verifier, tmux-cli-driver
|
||||
|
||||
---
|
||||
|
||||
## 3. `data-fetching-and-analysis`
|
||||
|
||||
Skills that connect to data and monitoring stacks. Include libraries to fetch data with credentials, dashboard IDs, common query patterns.
|
||||
|
||||
**Signals:** Has database queries, dashboard references, metric/event schemas, "how to find X in our data" patterns.
|
||||
|
||||
**Examples:** funnel-query, cohort-compare, grafana, datadog
|
||||
|
||||
---
|
||||
|
||||
## 4. `business-process-automation`
|
||||
|
||||
Skills that automate repetitive workflows into one command. Usually simple instructions but may depend on other skills or MCPs. Saving results in log files helps consistency.
|
||||
|
||||
**Signals:** Has "do this weekly/daily" patterns, aggregates from multiple sources, posts to Slack/channels, formats structured output.
|
||||
|
||||
**Examples:** standup-post, create-ticket, weekly-recap
|
||||
|
||||
---
|
||||
|
||||
## 5. `code-scaffolding-and-templates`
|
||||
|
||||
Skills that generate framework boilerplates for a specific function. May combine with composable scripts. Especially useful when scaffolding has natural-language requirements beyond pure code.
|
||||
|
||||
**Signals:** Has templates, "new X" generators, boilerplate structures, asset files to copy.
|
||||
|
||||
**Examples:** new-workflow, new-migration, create-app
|
||||
|
||||
---
|
||||
|
||||
## 6. `code-quality-and-review`
|
||||
|
||||
Skills that enforce code quality and help review code. Can include deterministic scripts for robustness. May run as hooks or in GitHub Actions.
|
||||
|
||||
**Signals:** Has style rules, review checklists, linting patterns, "reject if X" logic, adversarial review patterns.
|
||||
|
||||
**Examples:** adversarial-review, code-style, testing-practices
|
||||
|
||||
---
|
||||
|
||||
## 7. `ci-cd-and-deployment`
|
||||
|
||||
Skills that help fetch, push, and deploy code. May reference other skills to collect data.
|
||||
|
||||
**Signals:** Has deploy commands, build pipelines, PR management, rollout/rollback logic, environment configs.
|
||||
|
||||
**Examples:** babysit-pr, deploy-service, cherry-pick-prod
|
||||
|
||||
---
|
||||
|
||||
## 8. `runbooks`
|
||||
|
||||
Skills that take a symptom (alert, error, Slack thread) and walk through multi-tool investigation producing a structured report.
|
||||
|
||||
**Signals:** Has symptom→tool→diagnosis flows, "if you see X check Y" decision trees, report templates.
|
||||
|
||||
**Examples:** service-debugging, oncall-runner, log-correlator
|
||||
|
||||
---
|
||||
|
||||
## 9. `infrastructure-operations`
|
||||
|
||||
Skills that perform routine maintenance and ops, some involving destructive actions with guardrails. Make it easier to follow best practices in critical operations.
|
||||
|
||||
**Signals:** Has cleanup/orphan detection, cost investigation, dependency approval, confirmation gates for destructive actions.
|
||||
|
||||
**Examples:** resource-orphans, dependency-management, cost-investigation
|
||||
|
||||
---
|
||||
|
||||
## Classification Decision Tree
|
||||
|
||||
1. Does it primarily teach how to **call an API/CLI/SDK**? → `library-and-api-reference`
|
||||
2. Does it **verify** that something works (test, assert, validate)? → `product-verification`
|
||||
3. Does it **query data** from monitoring/analytics/databases? → `data-fetching-and-analysis`
|
||||
4. Does it **automate a repeating team process** (standup, report, ticket)? → `business-process-automation`
|
||||
5. Does it **generate new code/files** from templates? → `code-scaffolding-and-templates`
|
||||
6. Does it **review/lint/enforce quality** on existing code? → `code-quality-and-review`
|
||||
7. Does it **build/deploy/ship** code to environments? → `ci-cd-and-deployment`
|
||||
8. Does it **diagnose problems** from symptoms to structured findings? → `runbooks`
|
||||
9. Does it perform **infrastructure maintenance/cleanup** with guardrails? → `infrastructure-operations`
|
||||
|
||||
If a skill spans multiple categories, pick the one that describes its **primary action** — what the user gets when they invoke it.
|
||||
116
.github/skills/skill-evaluation/references/mechanics.md
vendored
Normal file
116
.github/skills/skill-evaluation/references/mechanics.md
vendored
Normal file
@@ -0,0 +1,116 @@
|
||||
# Mechanics: Predictability, Invocation, Hierarchy, Steering, Failure Modes
|
||||
|
||||
Reference for scoring Axes 1, 3, and 4 of the skill-evaluation rubric — a
|
||||
deliberately self-contained condensation of Matt Pocock's `writing-great-skills`
|
||||
GLOSSARY, kept in-skill so the evaluator runs anywhere without that skill
|
||||
installed (sync manually if the upstream GLOSSARY changes). Not a tutorial:
|
||||
look a bolded term up here rather than re-deriving it.
|
||||
|
||||
## 1. Root virtue: Predictability
|
||||
|
||||
A skill exists to wrangle determinism out of a stochastic system.
|
||||
**Predictability** is the agent taking the same *process* every run, not
|
||||
producing the same output — a brainstorming skill should predictably diverge;
|
||||
its tokens vary, its behavior doesn't. Every criterion in the rubric is a lever
|
||||
on this one virtue: conciseness, steering, and pruning are symptoms of
|
||||
predictability, not separate virtues competing with it.
|
||||
|
||||
## 2. Invocation trade-off
|
||||
|
||||
Two invocation modes, each paying a different cost:
|
||||
|
||||
- **Model-invoked** (default; no `disable-model-invocation`): keeps a
|
||||
description the agent reads every turn. Pays permanent **context load** —
|
||||
tokens and attention spent on every turn — in exchange for autonomous
|
||||
firing and reachability by other skills.
|
||||
- **User-invoked** (`disable-model-invocation: true`): the description is
|
||||
stripped from the agent's reach; only a human typing the skill's name can
|
||||
fire it, and no other skill can reach it either. Zero context load, but
|
||||
spends **cognitive load** — the human becomes the index of which skills
|
||||
exist and when to reach for each.
|
||||
|
||||
Pick model-invocation only when the agent must reach the skill on its own, or
|
||||
another skill must reach it. A skill that only ever fires by hand should be
|
||||
user-invoked and carry no trigger scaffolding it doesn't need. When
|
||||
user-invoked skills multiply past what a human can remember, a **router
|
||||
skill** — one user-invoked skill naming the others and when to reach for each
|
||||
— cures the accumulated cognitive load. That fix operates at the portfolio
|
||||
level, not the single-skill level this evaluation scores.
|
||||
|
||||
## 3. Content types & hierarchy
|
||||
|
||||
A skill mixes two content types freely: **steps** (ordered actions, each
|
||||
ending on a **completion criterion**) and **reference** (definitions, rules,
|
||||
facts consulted on demand). All-steps, all-reference, and mixed skills are
|
||||
equally valid — neither shape is a smell.
|
||||
|
||||
The **information hierarchy** ranks material by how immediately the agent
|
||||
needs it: in-skill step, then in-skill reference, then reference disclosed
|
||||
behind a **context pointer** in a linked file. Material every **branch** (a
|
||||
distinct way the skill is invoked) needs belongs inline; material only some
|
||||
branches need belongs behind a pointer — branching is the disclosure test. A
|
||||
pointer's *wording*, not its target, decides whether the agent reaches it and
|
||||
how reliably; a must-have target behind weak wording is a variance bug, and
|
||||
the fix is sharper wording, tried before pulling the material back inline.
|
||||
|
||||
**Co-location** governs what sits beside a piece of content once placed: a
|
||||
concept's definition, rules, and caveats belong under one heading, not
|
||||
scattered, so reading one part brings its neighbors with it.
|
||||
|
||||
## 4. Steering
|
||||
|
||||
**Leading words** are compact, pretrained concepts (*tight*, *red*, *lesson*)
|
||||
the agent thinks with while executing. Repeated consistently, they recruit
|
||||
priors the model already holds and anchor a region of behavior in the fewest
|
||||
tokens — cheaper and stickier than spelling the same quality out in prose. A
|
||||
leading word works twice: in the body it anchors execution (the same behavior
|
||||
fires every time the word appears); in the description it anchors invocation.
|
||||
It is also **trace-checkable** — distinctive enough that its appearance in the
|
||||
agent's reasoning traces confirms the skill actually shaped behavior.
|
||||
|
||||
A completion criterion must be *checkable* (can the agent tell done from
|
||||
not-done?) and, where it matters, *exhaustive* ("every X accounted for", not
|
||||
"produce a list"). A vague criterion invites **premature completion** —
|
||||
attention slipping to being done rather than to the work. The exhaustiveness
|
||||
demand also binds flat reference with no steps: "every rule applied" drives
|
||||
thorough **legwork** over a checklist the same way a sharp step criterion
|
||||
drives it over an action.
|
||||
|
||||
Skills should **avoid railroading**: procedures the agent adapts, not
|
||||
declarations of exact output; defaults with brief alternatives, not
|
||||
exhaustive menus.
|
||||
|
||||
## 5. Failure modes
|
||||
|
||||
- **Premature completion** — ending a step before it's genuinely done.
|
||||
Defense, in order: sharpen the completion criterion first (cheap, local);
|
||||
only if it's irreducibly vague *and* the rush is actually observed, split
|
||||
the sequence so later steps are hidden.
|
||||
- **Duplication** — the same meaning in more than one place. Costs
|
||||
maintenance and tokens, and inflates that meaning's rank past its real
|
||||
weight. Fix: collapse to a **single source of truth**, often via a leading
|
||||
word.
|
||||
- **Sediment** — stale layers that accumulate because adding feels safe and
|
||||
removing feels risky. The default fate of any skill without a pruning
|
||||
discipline.
|
||||
- **Sprawl** — a skill simply too long, independent of whether lines are
|
||||
stale or duplicated. Cure: disclose reference behind pointers, split by
|
||||
branch or sequence so each path carries only what it needs.
|
||||
- **No-op** — a line that changes nothing because the model already does it
|
||||
by default. The test: does it change behavior versus the default? Apply
|
||||
the **deletion test** sentence by sentence, not paragraph by paragraph — if
|
||||
removing the sentence leaves behavior unchanged, delete the whole sentence,
|
||||
don't trim words from it. A weak leading word (*be thorough* when the agent
|
||||
is already thorough-ish) is a no-op; the fix is a stronger word
|
||||
(*relentless*), not a different technique.
|
||||
- **Weak steering** — an instruction is present but the agent doesn't
|
||||
reliably follow it. Usually a leading word too weak to beat the default, or
|
||||
no leading word at all where a verbose passage is trying to do its job.
|
||||
- **Buried steps** — in-file reference so heavy it soaks the steps beneath
|
||||
it, turning attention to them into a coin-flip. Defense: progressive
|
||||
disclosure — push the reference behind a pointer.
|
||||
|
||||
**Relevance vs. no-op**: relevance asks whether a line still bears on the
|
||||
task; no-op asks whether it changes behavior. A line can be relevant (right
|
||||
topic) and still be a no-op (the model would do it anyway) — run both checks,
|
||||
they don't imply each other.
|
||||
158
.github/skills/skill-evaluation/references/output-template.md
vendored
Normal file
158
.github/skills/skill-evaluation/references/output-template.md
vendored
Normal file
@@ -0,0 +1,158 @@
|
||||
# Output Template
|
||||
|
||||
Emit the scorecard exactly in this structure (step 9 of the workflow).
|
||||
|
||||
```markdown
|
||||
# Skill Evaluation — {skill name}
|
||||
|
||||
> Evaluated: {date}
|
||||
> Source: {path}
|
||||
> Evaluator: skill-evaluation v2.1.0
|
||||
> Framework: [Anthropic Skill Best Practices](https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills) + Matt Pocock's [writing-great-skills](https://www.youtube.com/watch?v=UNzCG3lw6O0)
|
||||
|
||||
## Summary
|
||||
|
||||
| Metric | Value |
|
||||
|--------|-------|
|
||||
| Overall Score | {weighted}/100 |
|
||||
| Grade | {A/B/C/D/F} |
|
||||
| Category | {category} |
|
||||
| Invocation | {model-invoked / user-invoked} |
|
||||
| Files | {count} |
|
||||
| Criteria scored / N/A | {n} scored, {m} N/A |
|
||||
|
||||
## Scorecard
|
||||
|
||||
### Axis 1 — Trigger
|
||||
|
||||
| # | Criterion | Weight | Score | Notes |
|
||||
|---|-----------|--------|-------|-------|
|
||||
| 1 | Invocation design | 2x | {n}/100 | {evidence} |
|
||||
| 2 | Description quality | 2x | {n}/100 | {evidence} |
|
||||
|
||||
### Axis 2 — Structure
|
||||
|
||||
| # | Criterion | Weight | Score | Notes |
|
||||
|---|-----------|--------|-------|-------|
|
||||
| 3 | Steps vs. reference clarity | 1x | {n}/100 | {evidence} |
|
||||
| 4 | Branch-aware disclosure & pointers | 2x | {n}/100 | {evidence} |
|
||||
| 5 | Conciseness | 2x | {n}/100 | {evidence} |
|
||||
| 6 | Coherent scope | 1x | {n}/100 | {evidence} |
|
||||
|
||||
### Axis 3 — Steering
|
||||
|
||||
| # | Criterion | Weight | Score | Notes |
|
||||
|---|-----------|--------|-------|-------|
|
||||
| 7 | Leading words | 2x | {n}/100 | {evidence} |
|
||||
| 8 | Completion criteria & legwork | 2x | {n/100 or N/A} | {evidence} |
|
||||
| 9 | Gotchas section | 2x | {n}/100 | {evidence} |
|
||||
| 10 | Grounded in expertise | 2x | {n}/100 | {evidence} |
|
||||
| 11 | Avoids railroading | 1x | {n}/100 | {evidence} |
|
||||
|
||||
### Axis 4 — Pruning
|
||||
|
||||
| # | Criterion | Weight | Score | Notes |
|
||||
|---|-----------|--------|-------|-------|
|
||||
| 12 | No-ops (deletion test) | 2x | {n}/100 | {evidence with line citations} |
|
||||
| 13 | Single source of truth | 1x | {n}/100 | {evidence} |
|
||||
| 14 | Relevance & sediment | 1x | {n}/100 | {evidence} |
|
||||
|
||||
### Conditional criteria
|
||||
|
||||
| # | Criterion | Weight | Score | Notes |
|
||||
|---|-----------|--------|-------|-------|
|
||||
| 15 | Setup flow | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
|
||||
| 16 | Memory mechanism | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
|
||||
| 17 | Scripts & libraries | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
|
||||
| 18 | On-demand hooks | 1x | {n/100 or N/A} | {evidence or reason for N/A} |
|
||||
|
||||
## Trigger Eval
|
||||
|
||||
{For user-invoked skills, write: "N/A — user-invoked skill, no description to test."}
|
||||
|
||||
### Prompts tested
|
||||
|
||||
| # | Prompt | Expected | Triggered | Other skills |
|
||||
|---|--------|----------|-----------|--------------|
|
||||
| 1 | {prompt text} | should-trigger | yes/no | {list or none} |
|
||||
| 2 | {prompt text} | should-trigger | yes/no | {list or none} |
|
||||
| 3 | {prompt text} | should-trigger | yes/no | {list or none} |
|
||||
| 4 | {prompt text} | should-trigger | yes/no | {list or none} |
|
||||
| 5 | {prompt text} | should-trigger | yes/no | {list or none} |
|
||||
| 6 | {prompt text} | should-not-trigger | yes/no | {list or none} |
|
||||
| 7 | {prompt text} | should-not-trigger | yes/no | {list or none} |
|
||||
| 8 | {prompt text} | should-not-trigger | yes/no | {list or none} |
|
||||
| 9 | {prompt text} | should-not-trigger | yes/no | {list or none} |
|
||||
| 10 | {prompt text} | should-not-trigger | yes/no | {list or none} |
|
||||
|
||||
### Results
|
||||
|
||||
| Metric | Value |
|
||||
|--------|-------|
|
||||
| Should-trigger hit rate | {X}/5 |
|
||||
| Should-not-trigger leak rate | {X}/5 |
|
||||
| Other skills observed | {list or none} |
|
||||
|
||||
### Observations
|
||||
|
||||
{Free-form notes: patterns in what triggered or didn't, description wording
|
||||
gaps revealed, sibling skills that competed, etc.}
|
||||
|
||||
## Failure Modes Detected
|
||||
|
||||
| Mode | Evidence | Root cause | Defense |
|
||||
|------|----------|------------|---------|
|
||||
| {mode, or a single row "None detected"} | {file:line} | {cause} | {defense} |
|
||||
|
||||
## Prioritized Actions
|
||||
|
||||
### 1. {action}
|
||||
|
||||
**Evidence:** {file:line or section}
|
||||
|
||||
**Fix:** {specific recommendation}
|
||||
|
||||
### 2. {action}
|
||||
|
||||
**Evidence:** {file:line or section}
|
||||
|
||||
**Fix:** {specific recommendation}
|
||||
|
||||
(3–5 total, each tied to a detected failure mode)
|
||||
|
||||
## Bonus Patterns
|
||||
|
||||
| Pattern | Status | Notes |
|
||||
|---------|--------|-------|
|
||||
| Validation loops | {Present/Absent/N/A} | {detail} |
|
||||
| Output templates | {Present/Absent/N/A} | {detail} |
|
||||
| Procedures over declarations | {Present/Absent/N/A} | {detail} |
|
||||
| Defaults over menus | {Present/Absent/N/A} | {detail} |
|
||||
| Trace-checkable steering | {Present/Absent/N/A} | {detail} |
|
||||
|
||||
## Grade Scale
|
||||
|
||||
{copy the Grade Scale table from SKILL.md}
|
||||
|
||||
---
|
||||
|
||||
*Generated by [skill-evaluation](https://github.com/fabricioctelles/skills) v2.1.0, merging the [Anthropic skill quality framework](https://claude.com/blog/lessons-from-building-claude-code-how-we-use-skills) with Matt Pocock's [writing-great-skills](https://www.youtube.com/watch?v=UNzCG3lw6O0) methodology.*
|
||||
```
|
||||
|
||||
## Comparison mode
|
||||
|
||||
When `compare` is set, add a side-by-side table across all 18 criteria.
|
||||
Leave a cell N/A rather than scoring it 0, and exclude N/A rows from the
|
||||
Overall row's weighted math for that skill.
|
||||
|
||||
```markdown
|
||||
## Comparison: {skill A} vs {skill B}
|
||||
|
||||
| # | Criterion | {A} | {B} | Delta |
|
||||
|---|-----------|-----|-----|-------|
|
||||
| 1 | Invocation design | 60 | 85 | +25 |
|
||||
| 2 | Description quality | 25 | 70 | +45 |
|
||||
| ... | ... | ... | ... | ... |
|
||||
| 15 | Setup flow | N/A | 80 | — |
|
||||
| **Overall** | | **43** | **72** | **+29** |
|
||||
```
|
||||
Reference in New Issue
Block a user