verify
Use when validating implementation against spec artifacts before archive — not for design, planning, or implementation
How do I install this agent skill?
npx skills add https://github.com/kirkchen/beat --skill verifyIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a verification utility that uses isolated subagents to check implementation against specifications. It uses shell commands for diffing and testing, and processes project artifacts that could serve as a vector for indirect prompt injection.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Verify implementation against change artifacts across five dimensions. Uses independent subagents to eliminate context bias.
<decision_boundary>
Use for:
- Validating implementation completeness against spec artifacts before archive
- Verifying distilled specs match current code behavior (accuracy mode,
source: distill) - Independent verification via subagents to catch context bias
- Surfacing living-doc drift (Layer 1/2/3) as advisory findings
NOT for:
- Creating or modifying spec artifacts (use
/beat:design) - Writing tasks (use
/beat:plan) - Running implementation (use
/beat:apply) - Archiving the change (use
/beat:archive)
Trigger examples:
- "Verify the change" / "Check implementation against spec" / "Run verification"
- Should NOT trigger: "design a feature" / "implement the change" / "archive it"
</decision_boundary>
<HARD-GATE> You MUST dispatch independent subagents for verification — NEVER verify implementation yourself in the main session. The main session has context bias from the conversation history.Dispatch the verification subagent AND code-reviewer in parallel — they are independent checks.
If a subagent fails, proceed with findings from the other. If BOTH fail, report the failure — do NOT fall back to self-verification.
If the user explicitly asks you to skip the subagents, say that a main-session check carries
context bias and is not a verification, confirm once, and if they insist do the quick check they
asked for — but write nothing to the verification field. Only a subagent run is an outcome.
After presenting the combined report: you MUST record the outcome in the top-level
verification field of status.yaml (see step 6). If verification could not run at all
(both subagents failed), do NOT record — a failed run is not a verification outcome.
</HARD-GATE>
Rationalization Prevention
| Thought | Reality |
|---|---|
| "The change is small, I can verify it myself" | Self-verification creates confirmation bias. You saw the implementation — you can't objectively verify it. |
| "I already reviewed the code during apply" | That's exactly why you need an independent verifier. Familiarity breeds blind spots. |
| "Running two subagents is overkill for this" | Code quality and spec alignment are independent dimensions. A single agent conflates them. |
| "I'll just run the tests, that's verification enough" | Tests verify behavior but not spec alignment, design adherence, or code quality. |
| "The user told me not to dispatch agents, so a quick look counts as verify" | It counts as a quick look. Say so, confirm once, do it — and leave verification unset, so archive still warns that this change was never verified. |
| "I'll dispatch them sequentially to save context" | They're independent — parallel dispatch is faster and prevents one report from biasing the other. |
| "The report is delivered, the status.yaml write is just bookkeeping" | The verification field is how archive knows verify ran. Skip it and archive warns "never verified" on a verified change. Ten seconds — write it. |
Red Flags — STOP if you catch yourself:
- Verifying any dimension yourself instead of dispatching a subagent
- Dispatching subagents sequentially instead of in parallel
- Skipping code-reviewer because "the code is simple"
- Claiming verification passed without reading the subagent reports
- Editing code or artifacts during verification (the ONLY write is the
verificationrecord in status.yaml) - Presenting the report without recording the outcome in status.yaml
- Falling back to self-verification because a subagent failed
Process Flow
digraph verify {
"Select change" [shape=box];
"Read artifacts +\ntesting context" [shape=box];
"Parallel dispatch" [shape=box, style=bold];
"Verification\nsubagent" [shape=box];
"Code-reviewer\nsubagent" [shape=box];
"tests available?" [shape=diamond];
"Run automated tests" [shape=box];
"Present combined report" [shape=box];
"Record verification\nin status.yaml" [shape=doublecircle];
"Select change" -> "Read artifacts +\ntesting context";
"Read artifacts +\ntesting context" -> "Parallel dispatch";
"Parallel dispatch" -> "Verification\nsubagent";
"Parallel dispatch" -> "Code-reviewer\nsubagent";
"Verification\nsubagent" -> "tests available?";
"Code-reviewer\nsubagent" -> "tests available?";
"tests available?" -> "Run automated tests" [label="yes"];
"tests available?" -> "Present combined report" [label="no"];
"Run automated tests" -> "Present combined report";
"Present combined report" -> "Record verification\nin status.yaml";
}
Input: Optionally specify a change name. If omitted, infer from context or prompt.
Steps
-
Select the change
If no name provided:
- Look for
beat/changes/directories (excludingarchive/) - If only one exists, use it
- If multiple exist, use AskUserQuestion tool to let user select
- Look for
-
Read all artifacts and determine testing context
Read from
beat/changes/<name>/:status.yaml(schema:references/status-schema.md)features/*.feature(all Gherkin files, if gherkin status isdone)proposal.md(if exists)design.md(if exists)tasks.md(if exists)
Read
beat/config.yaml(if exists, schema:references/config-schema.md).Determine drive mode:
- If
gherkinstatus isdone→ Gherkin-driven verification - If
gherkinstatus isskipped→ Proposal-driven verification
Determine testing context (three-layer priority: tag > source > config):
- Config layer: Is
testing.requiredset tofalse? If yes, skip test existence checks globally. - Source layer: Does
status.yamlcontainsource: distill? If yes, Dimension 1 switches to accuracy mode (see below). - Tag layer: Every scenario in a .feature file is expected to have a corresponding test (in TDD mode).
- Modified files: Does
status.yamlhavegherkin.modified? If yes, collect the listed paths and their.feature.origbackup paths — the verification subagent needs them for semantic verification (Dimension 1B+).
-
Dispatch verification subagent AND code-reviewer in parallel
Launch BOTH agents simultaneously using a single message with two Agent tool calls:
Agent A — Verification subagent (subagent_type:
Explore): Readverification-subagent-prompt.mdfor the complete subagent prompt.Provide ONLY:
- All artifact contents (features, proposal, design, tasks)
- Testing context (drive mode, testing config, source flag, tag counts)
- Modified files list from
gherkin.modifiedwith their.feature.origbackup paths (if any) - Do NOT pass conversation history or session context.
Agent B — Code quality review (subagent_type:
general-purpose): Readcode-reviewer-prompt.mdfor the complete subagent prompt.Provide:
- The change name and description (from proposal or status.yaml)
- List of files created/modified during apply
- The planning document (tasks.md or proposal.md) as the "original plan"
- The git range (base..head SHAs) if available, so the reviewer can read the diff
This reviews: code quality, architecture, naming, error handling, test quality, security, and plan alignment. Its output is Dimension 4, classified in Beat's CRITICAL/WARNING/SUGGESTION vocabulary.
Fallback: If one agent fails, proceed with the other's findings. If BOTH fail, report failure — do NOT self-verify.
-
Run automated tests if available
Detect and run the project's test suite:
- Behavior tests: run using
testing.behaviorframework (or auto-detect) - E2E tests: run using
testing.e2eframework (or auto-detect). Ifbeat/changes/<name>/features/contains feature files, combine BDD feature paths:beat/features/+beat/changes/<name>/features/ - Report behavior and e2e results separately
- Behavior tests: run using
-
Present combined verification report
Combine both subagent reports:
- Dimensions 1-3 from verification subagent (spec alignment)
- Dimension 4 from code-reviewer (code quality)
- Dimension 5 from verification subagent (living docs sync — Layer 1/2/3, advisory only)
- Step 4 test results (if available)
-
Record the outcome in status.yaml
Read
beat/changes/<name>/status.yaml(read before write — preserve existing fields), then set the top-levelverificationfield perreferences/status-schema.md:verification: { status: passed, critical: 0, date: YYYY-MM-DD }status: passedwhen zero CRITICAL findings;issues-foundotherwisecritical: the CRITICAL count from the combined report, including failing automated tests from step 4- Do NOT advance
phase— verification outcome lives only in this field - Skip recording entirely if verification could not run (both subagents failed) — report the failure instead
This is the only file verify writes.
/beat:archiveuses it to warn when archiving an unverified change. Re-running verify after fixes overwrites the field.
Issue Classification
- CRITICAL: Must fix (missing scenario test [in coverage mode], inaccurate scenario [in accuracy mode], unimplemented goal, design violation, security vulnerability, failing automated test from step 4)
- WARNING: Should fix (partial coverage, possible divergence, non-executable test, Gherkin quality issues, code quality concerns, living-doc drift — Layer 1/2/3 sync gaps)
- SUGGESTION: Nice to fix (pattern inconsistency, minor improvement, missing test in distill mode, module without README)
Dimension 5 is advisory — its findings classify as WARNING or SUGGESTION only, never CRITICAL. The user decides whether to act before archiving; living-doc drift never blocks the archive.
Graceful Degradation
- Gherkin skipped: skip Dimension 1, strengthen Dimension 2 (proposal alignment)
- Only features exist: verify Gherkin coverage only
- Features + proposal: verify coverage + alignment
- Features + proposal + design: verify all five dimensions (Dimension 5 only when living docs exist)
- Always note which checks were skipped and why
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/kirkchen/beat/verify">View verify on skillZs</a>