Skip to content

Declare and validate each skill's capabilities - #135

Merged
richardmhope merged 4 commits into
mainfrom
claude/codebase-review-o7y3i9
Sep 25, 2026
Merged

richardmhope merged 4 commits into
mainfrom
claude/codebase-review-o7y3i9

Conversation

@richardmhope

Copy link
Copy Markdown
Collaborator

What & why

Skills didn't declare what they ask an agent to do, so a reviewer or installer had to read every instruction to find out. This adds a required, versioned capabilities block to every skill's meta.yaml (schema 1). It covers:

  • files read and edited;
  • commands;
  • network use, and why;
  • credentials;
  • agent tools;
  • files created.

Anything not declared is not requested. The docs say plainly that this is a declaration for review, and that skilldeck can't enforce it inside an agent.

  • Validation (skilldeck.capabilities, registry) rejects:

    • malformed blocks;
    • commands given as script paths, with shell operators, or as an interpreter given a script or inline code;
    • artifact paths that could leave the project: absolute, a drive, ~, \, . or ...
  • Bundle rules. A skill directory is exactly a regular meta.yaml plus a regular skill.md. It rejects:

    • symlinks and Windows junctions, including a linked skill directory;
    • subdirectories;
    • undeclared executables, detected by script suffix, exec bit, #! or binary header;
    • any other file;
    • links in skill.md to files the skill can't ship. This covers Markdown links, images, reference definitions, autolinks, and HTML tag src/href. Code, prose and indented blocks aren't checked.

    OS and editor leftovers (.DS_Store, swap files, …) are ignored when loading, so one stray file doesn't break every command. provenance --verify and catalog stay strict and report them as unexpected files. tests/test_capabilities.py builds the adversarial bundles in tmp_path.

  • Preview. skilldeck show <skill> --summary and skilldeck install … --dry-run show:

    • the source and the build recorded at build time;
    • the digest checked against the shipped content manifest, with a pointer to provenance --verify for real verification;
    • the deprecation state and the declared capabilities.

    The dry run also shows what each install would do, or the error it would hit (including a destination folder that can't be created). It writes nothing.

  • Installed files. Adapters add a ## Declared capabilities section, addressed to the agent, only to skills that ask for more than a read-only review: an edit, a credential, an agent tool, an artifact, or a command beyond read-only git. The ten read-only review skills install byte for byte as before; only dependency-review, logging and test-review gain the section. The adapter contract skill declares its project test command, so the fixtures pin the section (contract sha256:7c21e0856472).

  • Catalog. catalog --json gains an additive capabilities field. schema_version stays 1; a future capability schema would be a breaking catalog change.

  • The 13 bundled skills each declare what their instructions ask for, with a version bump: patch for the declaration alone, and dependency-review goes to 0.5.0 (a minor bump) because its instructions changed.

    • test-review now runs a regression test against the pre-change code only in a temporary git worktree.
    • dependency-review runs pip-audit only as pip-audit --disable-pip on fully pinned files, following pip-audit's security model, and declares reading package registry pages.
  • Structure test. It keeps the declared commands and the skill bodies in sync, including commands from common CLIs that no skill declares yet. The plugin tree and manifests are regenerated.

  • Breaking for skill authors, per docs/lifecycle.md: meta.yaml now requires capabilities, and a skill directory may hold only meta.yaml and skill.md.

An independent review found two problems, both fixed here:

  • the notice originally went on every skill, mostly repeating its body;
  • a single stray .DS_Store broke every command.

It also found:

  • the pip-audit and test-review instruction issues above;
  • --dry-run reporting "would install" when a destination folder couldn't be created;
  • the link check matching prose;
  • gaps in the command checks.

Not done here: the golden-diff evals weren't re-run, because they're paid. The dependency-review and test-review bodies changed, and three skills gained the section, so a maintainer should run them by hand.

Closes #73

Tracking: #94

Type of change

  • Bug fix
  • New feature
  • New or updated skill
  • Docs only
  • Refactor / internal

Checklist

  • Ran uv run --extra dev ruff check . && uv run --extra dev ruff format --check . && uv run --extra dev mypy && uv run --extra dev pytest (1137 passed, also with -W error::EncodingWarning and --resolution lowest-direct)
  • Added or updated tests
  • Updated docs where relevant
  • Added a CHANGELOG.md entry under ## [Unreleased]
  • For a skill change: bumped that skill's version in meta.yaml

🤖 Generated with Claude Code

https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX


Generated by Claude Code

Every skill's meta.yaml now carries a required, versioned `capabilities`
block (schema 1): files read (none/diff/repo) and edited (none/repo),
commands it may run, network use and why, credentials, agent tools, and
files it may create. Anything undeclared is not requested; the declaration
is for review, not enforcement.

- skilldeck.capabilities validates the block (unknown or missing keys,
  other schema numbers, commands given as paths or with shell operators,
  artifact paths that are absolute, name a drive or home, use backslashes
  or have ./.. components).
- The registry enforces the bundle rules: exactly regular meta.yaml and
  skill.md, rejecting symlinks (including a symlinked skill directory),
  subdirectories, undeclared executables (script suffix, execute bit,
  #!, binary headers) and any other file, plus skill.md links to files the
  skill cannot ship. provenance --verify and catalog apply the same rules.
- Every adapter appends a "Declared capabilities" section to a skill that
  asks for more than reading files; the adapter contract skill declares its
  `git diff`, so the fixtures pin that section (contract sha256:d8b7d4463e84).
- `skilldeck show <skill> --summary` and `skilldeck install --dry-run`
  preview a skill's source, build, digest check, deprecation and declared
  capabilities; the dry run also reports what installing would do and
  writes nothing.
- `skilldeck catalog --json` reports `capabilities` (additive; schema_version
  stays 1).
- All 13 bundled skills declare what their instructions ask for, with patch
  version bumps; plugin tree and content manifests regenerated.

Closes #73

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
- Render the "Declared capabilities" notice only for skills that ask for
  more than a read-only review (an edit, a credential, an agent tool, an
  artifact, or a command beyond read-only git); the ten read-only review
  skills install byte-for-byte as before. Reword the notice for the agent
  ("Beyond reading the repository, this skill asks you to: ... It asks for
  nothing else.") and rename Capabilities.beyond_reading to
  beyond_review_baseline. The contract skill now declares
  <the project's test command>, so the fixtures still pin the notice
  (contract sha256:7c21e0856472).
- Loading ignores OS and editor leftovers (.DS_Store, Thumbs.db,
  desktop.ini, __pycache__, ._*, .#*, *~, #*#, Vim swap files) so one stray
  file no longer breaks every command; provenance --verify and catalog
  still report them as unexpected files, with main's wording. .DS_Store is
  gitignored.
- install --dry-run checks the nearest existing folder on the way to each
  destination and reports the error a real install would hit.
- test-review runs a regression test against pre-change code only in a
  temporary git worktree (declared); dependency-review runs pip-audit only
  as `pip-audit --disable-pip` on fully pinned files, per pip-audit's
  security model, and declares reading package registry pages.
- The link check matches src/href only inside HTML tags, skips indented
  code blocks, and checks autolinks.
- The structure test knows common CLIs, so a new undeclared scanner, curl
  or test run is caught; spans a skill only quotes are listed per skill.
- Declared commands reject interpreters given a script or inline code;
  Windows junctions count as links; the catalog docs say a new capability
  schema is a breaking catalog change; summary wording points to
  provenance --verify and docs/verifying-releases.md.

Part of #73

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
Conflicts resolved keeping both sides:
- CHANGELOG.md: #134's lifecycle Added bullets and this branch's capability
  and bundle-validation bullets; [Unreleased] ### Security still names no
  skill or adapter.
- docs/authoring-skills.md: #134's Deprecating paragraph (CHANGELOG
  ### Deprecated, notice period, lifecycle.md) ends "Deprecating a skill",
  followed by this branch's Capabilities and "What a skill directory may
  hold" sections.
- docs/catalog.md: #134's removed-skill sentence and this branch's
  capabilities bullet.
- tests/test_catalog.py: uses #134's shared tests/_schema.py validator; the
  `enum` keyword this branch's catalog schema needs is ported there.

tests/test_lifecycle.py's synthetic skills now carry the capability
declaration that meta.yaml requires.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
docs/lifecycle.md (#134) defines skill SemVer and what is breaking for skill
authors; apply it to this branch:

- dependency-review goes to 0.5.0, a minor bump: running pip-audit only on
  fully pinned files and adding registry-page lookups changes how it gathers
  evidence, which can change what it reports. test-review stays at 0.3.1, a
  patch: the temporary-worktree step only spells out a safe way to do what
  it already asked, and what it reports is unchanged. Both CHANGELOG entries
  say why.
- Requiring `capabilities` and limiting a skill directory to meta.yaml and
  skill.md tightens rules existing metadata could fail, which the policy
  calls breaking for skill authors: add a **Breaking:** Changed entry.
- docs/lifecycle.md's meta.yaml section mentions `capabilities`: its own
  schema number, and how a declaration change versions a skill.

Part of #73

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HtiCGzpikMrkDYBkfQG5CX
@richardmhope
richardmhope merged commit e2a88a1 into main Sep 25, 2026
17 checks passed
@richardmhope
richardmhope deleted the claude/codebase-review-o7y3i9 branch September 25, 2026 03:46
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Define and enforce a skill capability manifest

2 participants