double-check
Use when performing a standalone PR double-check in a clean context — review, fix, and post.
How do I install this agent skill?
npx skills add https://github.com/fellowship-dev/dogfooded-skills --skill double-checkIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill implements a robust and secure Pull Request review and fix pipeline. It employs architectural safeguards such as isolated execution contexts for critical judgment, exact-head SHA verification to prevent race conditions, and clear boundaries between data collection and mutation stages.
- Socketwarn
1 alert: gptAnomaly
- Snykwarn
Risk: MEDIUM · 1 issue
What does this agent skill do?
double-check
Purpose
Standalone second-pass PR review, stage-partitioned into an ICM procedure. Fetch the PR + the
first review + the diff and check it out (setup), then run ONE cohesive critical-judgement
review in a clean context (verify the first review's claims, find missed edge cases, and check
tests/docs together → a single consolidated verdict), then apply fixes and re-run tests if
needed (fix), then post the curated comment, apply the double-checked label, and write the
report (post). Behaviorally equivalent to the original double-check skill, just isolated so the
judgement step is not polluted by the orchestrator's history.
Arguments
| Param | Required | Default | Notes |
|---|---|---|---|
pr | yes | — | PR number, e.g. 742 |
repo | yes | — | org/repo, e.g. fellowship-dev/booster-pack |
Parse from $ARGUMENTS: first token is pr, second is repo.
GitHub auth is ambient — no token env var. The pod's git-credential-pylot helper and the gh
shim mint short-lived App installation tokens per operation, so git URLs must stay plain
(https://github.com/<org>/<repo>.git); inline credentials bypass the helper and expire mid-run.
What it does
4-stage SEQUENTIAL ICM procedure (no parallel stages):
| Stage | Mode | Description |
|---|---|---|
| 01-setup | subagent | Fetch PR metadata including current HEAD, classify the incoming first-review receipt (its Head reviewed line) as current/stale/absent, capture comments + full diff, and checkout PR branch + merge base |
| 02-review | subagent | ONE cohesive critical review in clean context: reconcile the PR's claims against the diff, verify first review's claims, find missed edge cases, check tests/docs → consolidated verdict + curated findings |
| 03-fix | subagent | Apply MUST-FIX (and worthwhile NICE-TO-HAVE) fixes, re-run tests, push — only if fixes are needed |
| 04-post | inline | Re-fetch the live head, promote only when it equals the exact 40-hex head stage 02 reviewed, or perform one clean restart; then post curated review comment and apply labels only for the matching head |
Handoff locations
All handoffs live in the repo working directory:
.procedure-output/double-check/{stage}/handoff.md
The setup stage records the local checkout dir (REPO_DIR) in its handoff so the fix stage
operates on the same working tree. Each subagent stage receives only the handoffs it needs —
never the full orchestrator context.
Execution
Exact-head cycles (sequential subagents)
Run one Task per stage, one after another. Do NOT launch any stages in parallel. Do not start the
next stage until the current one completes. Start with restart_count=0. If stage 04 sees a
different full remote SHA, it stops without posting a verdict or touching labels, and the
orchestrator runs a complete conditional 01 → 02 → 03 → 04 cycle at the observed SHA with
restart_count=1. Stage 02 still receives only its new setup handoff: never orchestration
history. A second transition, or any unreadable live head, stops blocked.
Each Task prompt must be self-contained:
- Include only the stage's input handoff paths
- Include the path to the stage's CONTEXT.md
- Pass
prandrepovalues - Do NOT pass orchestrator history or prior reasoning
Task prompt template:
You are running stage {NN}-{name} of the double-check procedure.
PR: {pr} REPO: {repo}
Read your stage instructions:
skills/double-check/stages/{NN}-{name}/CONTEXT.md
Your inputs:
{list only the input handoff paths from that stage's CONTEXT.md}
Write your output to:
.procedure-output/double-check/{NN}-{name}/handoff.md
Execute all steps in CONTEXT.md. Write handoff.md before exiting.
Stage gating:
- Stage 02 is the isolated critical-judgement step. Its prompt MUST carry only the setup handoff (PR + first review + diff) — nothing else. This is the clean-context window the whole proc exists for.
- After stage 02, read its handoff. If
fixes_needed: false, SKIP stage 03 (no fixes to apply) and go straight to stage 04. Otherwise run stage 03.
Stage 04 (inline)
Run stage 04 yourself in the orchestrator context — do NOT spawn a Task. Read CONTEXT.md:
skills/double-check/stages/04-post/CONTEXT.md
Run the live claims-vs-diff and exact-head gates (gh pr view), then only for a matching head
post the comment and apply the label. Verify labels/comment actually landed, write the report
file, and emit the [pylot] outcome=... marker from the orchestrator (never from a subagent).
If stage 04 exits 3, set RESTART_COUNT=1 and run a new complete conditional
Stage 01 → 02 → 03 → 04 cycle at the current live head; do not pass the old
Stage 02 or Stage 03 handoff. If it exits 2, it is terminal blocked: do not run a promotion path.
Stage handoff chain
01-setup ─► 02-review ─► 03-fix ─► 04-post (inline, reads 01+02+03)
│ ▲
└── fixes_needed:false ┘ (skip 03)
Exit paths
- First-check fail closed (negative verdict or claims mismatch): stage 04 posts a
<!-- pylot:first-check-fail-closed -->comment, removes/withholdsdouble-checked, adds or retainsneeds-work, and creates no positive follow-on. It emits:[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-check {repo}#{pr} — verdict {verdict}, double-checked withheld, needs-work retained" status=success - Re-check PASS (PR had
needs-work, verdict=ready): stage 04 removesneeds-work, re-togglesdouble-checked(remove + re-add), and emits:[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-checked re-check PASS {repo}#{pr} — loop closed, cto-review re-fired" status=success - Re-check FAIL (PR had
needs-work, verdict=needs-work): stage 04 leavesneeds-workin place, does NOT re-toggledouble-checked, posts a structured verdict comment with a<!-- pylot:recheck-fail -->marker (idempotent — skipped if marker already present), and emits:[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-checked re-check FAIL {repo}#{pr} — needs-work retained" status=success - First-check success: only an explicit
readyverdict at the exact live 40-hex head may applydouble-checked; it emits:[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-checked {repo}#{pr} — verdict ready, {N} findings curated, {N} fixes pushed" status=success - Failure: failing stage emits
[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-check failed at stage NN: {reason}" status=failed - Blocked: setup cannot fetch/checkout the PR (e.g. merge conflict, missing PR), the live
head cannot be read, or a second head transition occurs →
[pylot:$PYLOT_OUTCOME_NONCE] outcome="double-check blocked: {reason}" status=blocked(a deliberate stop —blockedis its own terminal state, not a failure)
Rework Follow-up Mode
Whether — and how — a "produce the fix" follow-up dispatch exists at all is
repo policy, not protocol: some orgs wire a CTO-rework automation on top
of this skill, some don't. Read the repo playbook (GET /admin/playbooks/<org>/<repo>) for a rework-dispatch section; if it names a
contract, it looks like this:
A gate rule dispatches back into this skill when a prior review left
needs-work, with a task string identifying it as a PRODUCE follow-up (not a re-judge). The playbook names the exact trigger string/label to match and the automation's identity — treat those as resolved values below, not literal text to hardcode here.
When dispatched under that contract, do the following BEFORE running the stages above (the playbook is the source of truth for the trigger and rule name — keep this protocol in sync with it, don't let the two drift):
- Scan the PR's comments first. If a prior attempt already posted a
⚠️ NEEDS HUMANmarker, or a previous rework already tried and failed on this SAME gate, do NOT loop — leave it for a human and stop. Never re-check and re-applyneeds-workwithout producing anything. - Read the latest CTO review comment for its specific action items.
- Actually ADDRESS them. Code/test/doc fixes: make the changes and push.
Staging evidence requested: deploy the branch to staging, run the required
procedures, and update the PR body (not a comment) with the evidence
plus a
deployed_sha: <sha>line — the CTO gate scans the body. If the ask genuinely exceeds this skill's scope (needs a specialized runner, a human decision, or credentials you lack): do NOT silently re-applyneeds-work— post a comment beginning⚠️ NEEDS HUMANstating exactly what's blocking and what would unblock it, then STOP. - Only after the items are actually addressed:
gh pr edit <number> --repo <repo> --remove-label "double-checked,needs-work", then run the stages above fresh so the label chain re-fires.
No rework-dispatch section in the playbook? This mode does not apply — run the stages above as a normal review/fix/post pass.
Hard Rules
- SEQUENTIAL ONLY — one Task per stage, run one after another. NO parallel Task launches, ever.
- The review is ONE cohesive stage — do NOT split stage 02 into per-file or per-dimension subagents. Correctness, edge cases, tests, docs, deps, and security are judged together in a single verdict.
- Stage 02 gets a clean context — only the setup handoff and its artifact files (PR body, first review, changed files, diff). Never pass orchestrator history into it.
- Stage 04 runs inline — the
[pylot] outcome=...marker MUST come from the orchestrator. - Never pass full orchestrator context into subagent Task prompts — inputs only.
- Each stage writes handoff.md before the next stage reads it.
- Do not skip stages except stage 03 when
fixes_needed: false(an explicit, allowed skip). - NO Quest. Reporting is the local report file only — no Quest POST, no
127.0.0.1:4242, noquest.fellowship.dev, noQUEST_TOKEN. - Apply labels only after the comment posts successfully (stage 04). On re-check PASS,
remove
needs-workBEFORE re-addingdouble-checked— this is the structural loop-break. On re-check FAIL, do NOT touch labels or re-toggledouble-checked. - The diff is the only evidence; the PR body is a claim. Stage 02 reconciles every concrete
claim in the title/body against the changed files, and stage 04 re-checks it against the live
PR. A claim with no code behind it and no pointer to where it landed is
needs-work— "intentional", "the commit message explains it", and a LOW risk tier are NOT waivers.double-checkedis withheld until the body matches the diff. (pylot#2649, PR pylot#2782.) - Stage 04 verifies its own side effects — after labelling,
gh pr viewthe PR and confirm the expected labels/comment are actually there. Reporting success on unverified side effects is the failure this skill exists to catch in others. - Curate the first review, never re-derive it blind — setup captures the first review's
comments verbatim and its
Head reviewedreceipt line; stage 02 curates those findings at tier-scaled depth (escalate-only). No first review found → full-depth fresh review. - A stale first review is historical evidence, not a blocker or current coverage — compare
its
head_shawith the post-rebase PR HEAD. Continue the cohesive review against the complete current diff, re-check prior findings, and post a new current-head receipt. Never restart the whole pipeline merely because the incoming receipt is stale. - Promotion binds to the final full SHA — immediately before every verdict comment or label
mutation, fetch
headRefOid. The 40-character live SHA must equal the Stage 02reviewed_head_sha, fail closed. On the first mismatch restart cleanly; on a second mismatch or failed retrieval stop blocked. A delta inspection, file list, short SHA, or local HEAD never substitutes for equality. A stale or blocked run never mutatesdouble-checkedor triggers downstream automation. - First-check promotion is explicitly positive only — apply
double-checkedonly when the reviewer verdict is exactlyreadyand bound to the exact live head. Any negative, missing, malformed, stale, or conflicting signal fails closed: remove/withholddouble-checked, add or retainneeds-work, and do not create CTO, FlowChad, staging, or merge follow-ons. - Success markers come from stage 04's templates ONLY (pylot#3392). The four documented
status=successmarkers in stage 04's "Emit outcome marker" section are the complete set. Never synthesize a free-form success marker from the orchestrator, and never emit ANYstatus=successunless stage 04 ran to completion — receipt comment posted and post-action verification passed. On 2026-09-05 an orchestrator emitted "queued-worker result ok" without running stage 04: the mission terminalizeddonewith no receipt and the completed work was invisible for ~6h. If you cannot run stage 04, the outcome isstatus=failedorstatus=blockedwith the reason — a fabricated success is the worst possible exit. - Verbatim artifacts are written by shell, never retyped. Stage 01 redirects the PR body, first review, manifest and diff into files beside its handoff. Copying them through the Write tool was the largest setup cost and can silently drift from the source.
- Wait on processes, not clocks. A long command (tests, a hooked
git push) runs in the foreground with the tool's longest timeout, or in the background with its PID recorded and a blocking wait that returns the moment it exits (stage 03's "Waiting on a long command"). Never poll with a fixedsleep Nof 30 seconds or more: measured runs slept up to 10 minutes past completion. - The full suite is not this skill's job. Stage 03 tests its own fix delta once, with the repo's scoped gate; CI and the release gate stay the full-suite authority. Never run a test suite twice for one push (an explicit run and then the same run again in a pre-push hook).
Reference files
CONTEXT.md— architecture overviewstages/NN-name/CONTEXT.md— per-stage inputs, task, output contractshared/review-comment-template.md— curated PR comment template (stage 04)shared/report-template.md— local report file template (stage 04)
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/fellowship-dev/dogfooded-skills/double-check">View double-check on skillZs</a>