A Floor, Not a Ceiling: Why My Agent Skills Use Everything Claude Code Offers
Most of my agent skills started life as Matt Pocock’s. grill-me, to-spec, to-tickets, tdd, diagnosing-bugs, triage: the names are his, and so is a lot of the thinking. If you haven’t used them, start there. They’re good, and they’re deliberately small.
They’re also deliberately harness-neutral. Matt’s README says they “work with any model”, and they install through skills.sh into Claude Code, Codex, and anything else that reads the Agent Skills format. One text, every agent.
For a while I held my fork to the same rule. My repo’s guide said to name a harness-specific tool “only where the skill genuinely cannot work without it.” This week I reversed it:
Every skill must run in any harness. That’s the floor, not the ceiling.
A skill can now name a Claude Code feature when that feature does the job better, as long as the same sentence says what to do without it. I went through all 49 skills with that rule and changed about 30 of them. This post is about what that changes for you, the person in the chair, because that’s where the difference shows up.
What “harness-neutral” costs you
A harness-neutral skill can only describe capabilities in words every agent understands: “ask the user”, “dispatch a sub-agent”, “wait for CI”. Each agent then does its best with what it has. In Claude Code, that usually means the generic version of something it has a dedicated tool for.
Some concrete cases from my own runs:
- Questions arrive as a wall of numbered text. The grilling protocol asks the whole frontier of open decisions in one round, each with a recommended answer. In chat, you scroll, then type “1 yes, 2 the second one, 3 actually no, because…” and hope the agent maps your reply back correctly.
- One debugging path can’t run at all. The bug-diagnosis skill’s last-resort loop is a bash script that walks a human through steps with
read -p. Claude Code’s shell has no terminal, so the agent can’t run it. You end up running it yourself and pasting the output back. - Waiting happens by polling. “Wait for the review bots” turns into the agent checking, sleeping, and checking again, burning turns and context while nothing changes.
- Long runs end silently. A two-hour autonomous build finishes, or stops on a question, and you find out whenever you next look at the terminal.
None of that is a bug in the skills. It’s the price of writing for the lowest common denominator. Every harness gets the same text, so no harness gets more than the text.
Questions you answer by clicking
Claude Code has AskUserQuestion: a real question UI with options, a recommended choice marked first, multi-select, and a preview pane for side-by-side code or mockups.
My grilling skill now puts each decision through it. When the options are things you compare by looking (two function signatures, two payload shapes, two screen layouts), each option carries its snippet in the preview, so you pick while looking at the code instead of reading three paragraphs describing it. design-it-twice does the same with competing interface designs. release shows the version, the rule that chose the bump, and the drafted notes in one confirm screen before it creates a tag that can’t be taken back.
Cleanup questions changed the most. They used to be a yes/no: “Can I delete everything I created?” Saying “yes, but keep the second worktree” meant free text the agent had to parse. Now it’s a multi-select of exactly what the run created, nothing pre-ticked. You tick what goes.
And the bug-diagnosis loop that couldn’t run? Each scripted step is now a question with Done and Couldn’t options, and each capture is a set of answers to pick from. The script still exists for every other harness.
Waiting that doesn’t cost you anything
Claude Code can run a command in the background and wake the agent when it exits, stream a log line by line with Monitor, schedule its own wake-up with ScheduleWakeup, and send a desktop or phone notification with PushNotification.
That turns a lot of babysitting into nothing:
autopilotstreams its milestones throughMonitorinstead of polling the run log, and pings you when the run needs input or fails. Not when it’s merely progressing.verifytails the app log while it clicks through the flow, so an exception shows up tied to the action that caused it, not in a log you grep afterwards.resolve-review-commentsused to push its fixes and declare victory. Bots re-review minutes later and often open new threads. Now it waits on CI in the background, schedules a wake-up for the bots’ second pass, and loops (three rounds at most).triage-github-prstarts its own independent review before the wait for bot reviews, so the two overlap instead of running back to back.
The notification rule is the one I care about most. It fires when a run stops on a decision or finishes while you’re likely away. A ping you didn’t need is worse than no ping.
Subagents with guardrails, not just “a sub-agent”
Matt’s skills already fan out work to sub-agents. A harness-neutral skill has to say “dispatch a sub-agent” and leave the rest to the agent. A Claude Code plugin can ship typed agents: a fixed tool set, a cheaper model where the thinking happens elsewhere, preloaded vocabulary, and memory across sessions.
The one that matters most is boring on purpose. Review comments from bots and strangers are untrusted text, and an agent that reads them is a prompt-injection target. My finding-verifier agent reads them for triage-github-pr, fix-github-issues, and now resolve-review-comments, and a plugin hook holds it to read-only commands. If a comment talks it into running anything that writes, deletes, or pushes, the hook blocks the command before it runs. The skill still says what to do without the plugin: any read-only sub-agent, or triage inline.
A few others from this pass:
implement-specnames each parallel implementer, so a failing test goes back to the same one withSendMessage. That implementer already knows its worktree and ticket, so there’s nothing to re-learn. A stopped run can cancel the ones you pick withTaskStop.panelruns each external CLI reviewer in its own throwaway worktree, so a CLI that decides to “help” by editing files never touches your tree.dependency-auditreads the actual changelog for every major version bump instead of relying on the model’s memory of it.
Things a skill can’t do alone: hooks and mods
Some improvements aren’t instructions at all.
A hook runs on harness events, whether or not any skill is active. The new one is small: when a session opens in a git worktree of a Laravel app with no vendor/ yet, it tells the agent to bootstrap the worktree before running anything. Every other session sees nothing. I wanted it to do the bootstrap itself, but WorktreeCreate replaces git’s worktree creation and plugins can’t register it, so the hook points the agent at the bootstrap steps instead.
A mod is code that runs inside Claude Code’s UI. My plugin ships two. /scratch opens a live pane of every in-flight effort in .scratch/: specs, tickets, blockers, what’s claimed, what’s next, without spending a turn. A band above the prompt lists every subagent a skill dispatched, with its tool count, last tool, and outcome. When implement-spec has four implementers running, you can see which one is stuck.

Neither changes what the skills do. They change how much of it you can see.
Pages instead of files
Some skills produce things meant for another person. A prototype for a product manager, a questionnaire for a domain expert, a lesson with a quiz. Handing over an HTML file worked, but nothing came back.
Where the Artifact tool is available, those now publish as private pages, and sharing is your call. teach lessons record quiz answers in the page’s database, and the next session reads them back and turns repeated misses into learning records, so the next lesson targets what you actually got wrong. A questionnaire’s recipient answers on the page, and to-questionnaire pulls the answers back into the Markdown record.
The rule that keeps this honest
All of this would be a trap if it made the skills Claude-Code-only. So every one of these changes follows the same shape:
Name the capability, then the tool, then the fallback — in the same sentence.
“Ask the user to choose (AskUserQuestion where available, otherwise in chat).” The fallback has to be a complete path on its own, because an instruction naming a tool the running agent doesn’t have isn’t just unused. It’s unfollowable. A CI check in my repo fails any skill that names one of my plugin’s agents without a fallback clause.
It has costs, and I’d rather name them:
- Skills get longer. Each feature adds a clause. I capped every change at about ten lines and pushed longer material into reference files.
- Harness features move. During this very pass I wrote that Claude Code’s Agent tool has no
run_in_backgroundparameter. It does, but only when fork mode is off. I caught it because I’d made a rule to check the docs before writing a tool name into a skill. Expect to maintain this. - You’re betting on one harness. If you switch agents weekly, Matt’s approach of one neutral text is simpler to live with. Mine still runs there; it just runs the plain way.
Which should you use?
If you want small skills you’ll read, fork, and bend to your own process across several agents, use Matt’s. They’re the better starting point, and they’re where mine came from.
If you live in Claude Code and want the skills to use what it offers (questions you click, waits you don’t babysit, agents that can’t be talked into running commands, and a view of what’s running), that’s what mine are for now.
The general principle holds beyond skills: portability is a floor. Write so that everything runs everywhere, then let each environment do its best work on top. The fallback is the complete path. The feature is the better one.