speckit-runner
Use when implementing a GitHub issue end-to-end — runs speckit with an independent advisory review before creating one PR.
How do I install this agent skill?
npx skills add https://github.com/fellowship-dev/dogfooded-skills --skill speckit-runnerIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill automates a complex GitHub issue workflow using secondary AI workers and external APIs. It implements several security best practices, including supervisor verification and fail-closed post-conditions. However, it presents an indirect prompt injection surface by processing untrusted data from GitHub issues and external URLs, and it executes shell commands discovered from the target repository.
- Socketwarn
1 alert: gptSecurity
- Snykwarn
Risk: MEDIUM · 1 issue
What does this agent skill do?
Run the speckit pipeline for issue $0 in repo $1.
You are an operator running in-session. Drive a producer through the speckit phases, then give its pushed checkpoint to a separate review worker before PR creation.
Gateway: $PYLOT_API (or $PYLOT_GATEWAY_URL). Token: $PYLOT_DISPATCH_TOKEN.
Mission: $PYLOT_JOB_ID. Repo: $1 (or $PYLOT_REPO).
Define this cleanup once at session start. The EXIT trap makes detached supervisor checkout removal executable on every terminal path, including an early failure.
cleanup_supervisor_checkout() {
if [ -n "${SUPERVISOR_CHECKOUT:-}" ] && [ -d "$SUPERVISOR_CHECKOUT" ]; then
git worktree remove --force "$SUPERVISOR_CHECKOUT" >/dev/null 2>&1 || true
fi
}
trap cleanup_supervisor_checkout EXIT
Drive the worker in the foreground by polling the worker API — in short chunks you re-run yourself. After queueing each phase prompt, run the Poll-to-idle snippet (Step P) with the Bash tool, using the default Bash timeout (do NOT pass a long
timeout). Each call returns in under 2 minutes with aPOLL_RESULT; while it printsPOLL_RESULT=running, run Step P again immediately. Every ~30 min it printsPOLL_RESULT=block_elapsedwith a decision packet (heartbeat age, output-changed flag, log tail) — review it and, if the worker is healthy, run Step P again to grant another block. Repeat untilPOLL_RESULT=done.The trap (read this): the harness auto-backgrounds any Bash command that runs past its tool
timeout(default ~120 s). A backgrounded poll is fatal. You may see a completion notification arrive for a background task — ignore that signal as a reason to wait. Those notifications only fire while your session is actively running tool calls; the instant you end your turn to "wait for it," the headlessclaude -psession exits and the mission is finalized as failed — while the worker is still healthy. So: never set a long Bashtimeouton Step P, never background it, never "wait for a notification," never end your turn while a worker turn is in flight. If Step P ever gets backgrounded, that is a bug — kill it and run it again. Each call is short synchronous shell you run, read, and re-run yourself.
Step 0: Reconcile Before Starting
Before spawning anything, check both terminal issue state and an already-open PR. This is the resume boundary: a rerun must report the existing PR instead of creating another one.
REPO="${1:-$PYLOT_REPO}"
ISSUE_STATE=$(gh issue view $0 --repo "$REPO" --json state --jq '.state' 2>/dev/null || echo "OPEN")
if [ "$ISSUE_STATE" = "CLOSED" ]; then
echo "[speckit-runner] already complete — issue $0 is CLOSED"
exit 0
fi
OPEN_PRS=$(gh pr list --repo "$REPO" --state open --limit 100 \
--json number,url,headRefName,closingIssuesReferences 2>/dev/null || echo '[]')
EXISTING_PR=$(printf '%s' "$OPEN_PRS" | jq -c --argjson issue "$0" --arg repo "$REPO" '
[.[] | select(any(.closingIssuesReferences[]?;
.number == $issue and ((.repository.owner.login + "/" + .repository.name) == $repo)))][0] // empty')
# A pushed checkpoint may not have a PR yet. If exactly one remote branch has
# an issue-number segment, also reconcile an open PR by its exact head ref.
BRANCH_CANDIDATES=$(gh api "repos/$REPO/branches" --paginate --jq '.[].name' 2>/dev/null \
| awk -v n="$0" '$0 ~ ("(^|[-_/])" n "($|[-_/])")')
if [ "$(printf '%s\n' "$BRANCH_CANDIDATES" | sed '/^$/d' | wc -l | tr -d ' ')" = "1" ]; then
RESUME_BRANCH=$(printf '%s\n' "$BRANCH_CANDIDATES" | sed '/^$/d')
HEAD_PR=$(printf '%s' "$OPEN_PRS" | jq -c --arg head "$RESUME_BRANCH" \
'[.[] | select(.headRefName == $head)][0] // empty')
[ -n "$EXISTING_PR" ] || EXISTING_PR="$HEAD_PR"
fi
if [ -n "$EXISTING_PR" ]; then
PR_URL=$(printf '%s' "$EXISTING_PR" | jq -r '.url')
echo "[speckit-runner] resume reconciliation found existing PR: $PR_URL"
exit 0
fi
If either reconciliation branch exits, emit its resolved marker as your final full assistant line (not from Bash and not inside a fence), then stop:
- Closed issue: [pylot:$PYLOT_OUTCOME_NONCE] outcome="already complete — issue $0 is CLOSED" status=success
- Existing PR: [pylot:$PYLOT_OUTCOME_NONCE] outcome="resume reconciliation found existing PR: $PR_URL" status=success
outcome="already complete" is only valid here — when the issue is genuinely CLOSED. Never emit it because of a timeout or missing notification.
An open PR is also terminal for this run, but report it as a reconciled existing
PR, not as "already complete." Exact closing linkage is preferred; exact head
matching is the fallback for a unique checkpoint branch. If no PR exists, pass
that unique RESUME_BRANCH to the worker and resume it only when its checkpoint
matches the issue. Ambiguous matches are advisory context: start from the default
branch rather than guessing.
Step 1: Spawn Worker
REPO="${1:-$PYLOT_REPO}"
SPAWN_RESP=$(curl -s --max-time 90 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
-H "Content-Type: application/json" \
-d "{\"repo\": \"$REPO\"}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers")
WID=$(echo "$SPAWN_RESP" | python3 -c 'import sys,json; print(json.load(sys.stdin).get("worker_id",""))' 2>/dev/null)
if [ -z "$WID" ]; then
echo "[speckit-runner] worker spawn failed: $(echo $SPAWN_RESP | head -c 200)"
exit 1
fi
echo "[speckit-runner] worker spawned: $WID"
If worker spawn fails, emit the following resolved marker as your final full assistant line (not from Bash and not inside a fence), then stop:
[pylot:$PYLOT_OUTCOME_NONCE] outcome="worker spawn failed: $(echo $SPAWN_RESP | head -c 200)" status=failed
This skill ships a poll-worker.sh helper next to this file — the boot-sync copies the whole skill dir, so it lands at ~/.claude/skills/speckit-runner/poll-worker.sh on the operator. It is the only way you poll a worker (see Step P) — never hand-roll a poll loop inline, never wait for a notification.
The same boot-sync also ships two validator executables used by the invariant-matrix
flow, both landing at ~/.claude/skills/speckit-runner/<name>.sh:
validate-invariant-matrix.sh MATRIX— the schema gate. Checks a single TSV matrix (header, 13 columns, stableINV-NNNIDs, allowedclass/statevalues, provenance/evidence completeness, and passed/failed rows carrying a realrepository/checkpoint/receipt). Used by Step 4.5's supervisor reconciliation and by the test suite to validate both the canonical schema fixture and the persisted planning matrix.validate-review-output.sh MATRIX REVIEW— the coverage gate, run after the schema gate passes. Checks the independent reviewer's output against a matrix that already satisfies the schema gate: every applicable row must get exactly oneROW INV-NNN verdict=...line (or aFINDINGchallenging it), coverage must be complete, and malformed matrices are rejected outright. Used in Step 5 to decideFIRST_REVIEW_STATUS=availablevsunavailable.
Together they form a two-stage gate: a matrix must pass the schema gate before its coverage can be validated against reviewer output — a structurally invalid matrix never reaches the coverage check.
Step P: Poll-to-idle (run the helper, loop while RUNNING)
After queueing a phase prompt, poll only by running the bundled helper with the Bash tool (default timeout — do not pass a long one):
bash ~/.claude/skills/speckit-runner/poll-worker.sh "$WID" "$TURN_SEQ"
Read its last line:
POLL_RESULT=done(exit 0) → worker finished this turn; the printed output carries the phase marker. Proceed.POLL_RESULT=running(exit 10) → worker is healthy and still working — run the exact same command again (re-inlineWID/TURN_SEQ; the script resumes its cumulative timer via a state file). The implement phase needs many of these — keep going.POLL_RESULT=block_elapsed(exit 20) → a poll block (default 30 min) elapsed; the worker is NOT stopped and is very likely still working. The script printed a decision packet:heartbeat_age,output_changed, and a tail of the worker's output. Decide — heartbeat_age is the primary signal:- heartbeat fresh (< ~5 min) → run the exact same command again. That grants another block. This is the DEFAULT for a healthy worker — long implement turns legitimately take multiple blocks; never fail a healthy worker just because time passed.
output_changed=emptymid-turn is NORMAL (worker output only lands at turn end) and is not a stuck signal. - heartbeat stale (> ~10 min, W_STATE still running) → the worker is wedged. Run Step C and emit a failed outcome quoting the packet.
output_changed=noacross consecutive blocks is corroborating evidence only, never sufficient by itself.
- heartbeat fresh (< ~5 min) → run the exact same command again. That grants another block. This is the DEFAULT for a healthy worker — long implement turns legitimately take multiple blocks; never fail a healthy worker just because time passed.
POLL_RESULT=ceiling_timeout(exit 1) → the hard ceiling (default 4 h per turn) was hit; the current worker has already been stopped by the script. Run Step C, then emit a failed outcome.
Each call returns in <2 min by design, so it never gets backgrounded. Never hand-roll a poll loop, set a long Bash timeout, background the call, wait for a notification, or end your turn while a worker turn is in flight — any of those abandons a healthy worker and fails the mission.
Step C: Always-Run Cleanup + Resume Receipt
Run this before every terminal outcome — success, partial, blocked, or failed. It stops every spawned worker while preserving pushed Git checkpoints. It also prints the exact remote head the next invocation can reconcile.
for WORKER_ID in "${WID:-}" "${RWID:-}"; do
[ -z "$WORKER_ID" ] || curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WORKER_ID}/stop" >/dev/null 2>&1 || true
done
if [ -n "${BRANCH:-}" ]; then
SAVED_HEAD=$(gh api "repos/$REPO/git/ref/heads/$BRANCH" --jq '.object.sha' 2>/dev/null || true)
[ -z "$SAVED_HEAD" ] || echo "[speckit-runner] resume branch=$BRANCH head=$SAVED_HEAD"
fi
Producer failures keep their normal failed/blocked outcome, but only after Step C. Reviewer failures are different: stop only the reviewer, record review as unavailable, and continue with the producer checkpoint.
Step 2: speckit.preflight — Pre-flight + Specify
Queue the prompt, then poll per Step P: run bash ~/.claude/skills/speckit-runner/poll-worker.sh "$WID" "$TURN_SEQ", re-running it while POLL_RESULT=running.
PROMPT=$(python3 -c "import json,sys; print(json.dumps('You are a worker running inside repo $REPO. Issue: #$0.\n\nPre-Flight (MANDATORY — do this FIRST):\n1. Fetch issue: gh issue view $0 --repo $REPO --json title,body,labels,comments\n2. Check if closed: if CLOSED, emit [pylot:$PYLOT_OUTCOME_NONCE] outcome=\"already complete\" status=success and exit.\n3. Verify required labels exist (create '\''in-progress'\'' if missing).\n4. Gather real data: read issue comments, fetch referenced URLs, read existing code patterns.\n\nResume or Specify:\n5. Fetch origin and inspect remote branches whose names identify issue $0. Resume only when exactly one candidate has an existing checkpoint/spec matching this issue; otherwise update the default branch and start clean. Never create or duplicate a PR in this phase.\n6. Bootstrap speckit scaffolding if absent: if [ ! -f \".specify/scripts/bash/create-new-feature.sh\" ]; then /setup-speckit; fi\n7. If the resumed branch already has a valid specification for this issue, continue it. Otherwise run: /speckit-specify $0\n8. Read the generated specification. If there are open questions, answer them from pre-flight data, then run /speckit-clarify.\n9. Set BRANCH=\$(git branch --show-current), commit any completed specification checkpoint, and push the branch.\n\nWhen done: emit [pylot] phase=preflight status=done branch=\$BRANCH head=\$(git rev-parse HEAD)'))")
PROMPT_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
-H "Content-Type: application/json" \
-d "{\"prompt\": $PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}/prompt")
TURN_SEQ=$(echo "$PROMPT_RESP" | python3 -c 'import sys,json; print(json.load(sys.stdin).get("turn_seq",""))' 2>/dev/null)
echo "[speckit-runner] preflight prompt queued (turn_seq=$TURN_SEQ)"
Poll per Step P now — run bash ~/.claude/skills/speckit-runner/poll-worker.sh "$WID" "$TURN_SEQ" and re-run it while POLL_RESULT=running. When it prints POLL_RESULT=done, read the printed worker output for phase=preflight status=done. If absent or status=failed, run Step C and emit a failed outcome.
After a successful preflight, fetch last_output, set BRANCH from its marker,
and verify the emitted head matches the remote head. Keep BRANCH for every later
Step C receipt.
Step 3: speckit.plan — Plan + Tasks
Queue the prompt, then poll per Step P: run bash ~/.claude/skills/speckit-runner/poll-worker.sh "$WID" "$TURN_SEQ", re-running it while POLL_RESULT=running.
PROMPT=$(python3 -c "import json; print(json.dumps('Continue on the feature branch from the previous phase.\nRun: /speckit-plan $0\nRead the generated plan and verify the approach.\nRun: /speckit-tasks $0\nRead the generated tasks and verify they are concrete.\n\nDerive the durable invariant matrix described by ~/.claude/skills/speckit-runner/references/invariant-matrix.md from issue, spec, tasks, and trusted repository evidence. Inspect every discovery class, emit every source-backed boundary and no row for an unsupported class, and persist specs/<feature>/invariant-matrix.tsv with stable IDs and initial states. Commit and push the completed planning checkpoint so a later invocation can resume it.\nWhen done: emit [pylot] phase=plan status=done branch=\$(git branch --show-current) head=\$(git rev-parse HEAD), followed by MATRIX RECEIPT path=<path> repository=$REPO checkpoint=\$(git rev-parse HEAD)'))")
PROMPT_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
-H "Content-Type: application/json" \
-d "{\"prompt\": $PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}/prompt")
TURN_SEQ=$(echo "$PROMPT_RESP" | python3 -c 'import sys,json; print(json.load(sys.stdin).get("turn_seq",""))' 2>/dev/null)
echo "[speckit-runner] plan prompt queued (turn_seq=$TURN_SEQ)"
Poll per Step P now — run bash ~/.claude/skills/speckit-runner/poll-worker.sh "$WID" "$TURN_SEQ" and re-run it while POLL_RESULT=running. When it prints POLL_RESULT=done, check the printed worker output for phase=plan status=done. If blocked or failed, run Step C before emitting the blocked/failed outcome.
After success, verify the plan marker's branch and head against the same remote branch. This proves the resume checkpoint advanced before implementation begins.
Step 4: speckit.implement — Implement + Verify + Checkpoint
The producer implements and pushes a reviewable checkpoint, but does not open a PR yet. Queue this prompt and poll per Step P.
PROMPT=$(python3 -c "import json; print(json.dumps('Continue on the feature branch. Run: /speckit-implement $0\nAfter implementation:\n- Run the repository-defined verification for the affected behavior and fix failures.\n- Exercise the changed behavior through the strongest practical interface available in this repository.\n- Run: /speckit-analyze $0 && /speckit-checklist $0; resolve concrete gaps they find.\n- Commit every intended implementation and specification change, then push the current branch.\n- Do NOT create a PR. This pushed head is the independent-review checkpoint.\nWhen done, emit one marker line: [pylot] phase=implement status=done branch=\$(git branch --show-current) head=\$(git rev-parse HEAD). Then provide a concise CHECKPOINT EVIDENCE block containing the exact verification commands, pass/fail results, behavior exercised, and evidence asset ids or receipts (or explicit N/A).'))")
PROMPT_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"prompt\": $PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}/prompt")
TURN_SEQ=$(echo "$PROMPT_RESP" | python3 -c 'import sys,json; print(json.load(sys.stdin).get("turn_seq",""))' 2>/dev/null)
echo "[speckit-runner] implement prompt queued (turn_seq=$TURN_SEQ)"
Poll to done, then re-fetch the worker output and extract the emitted branch.
Confirm that the remote branch points at the emitted head before continuing.
PRODUCER_STATE=$(curl -s --max-time 20 -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}")
PRODUCER_OUT=$(printf '%s' "$PRODUCER_STATE" | jq -r '.last_output // ""')
BRANCH=$(printf '%s' "$PRODUCER_OUT" | sed -n 's/.*phase=implement status=done branch=\([^ ]*\).*/\1/p' | tail -1)
HEAD_SHA=$(printf '%s' "$PRODUCER_OUT" | sed -n 's/.*phase=implement status=done.* head=\([0-9a-f]*\).*/\1/p' | tail -1)
REMOTE_SHA=$(gh api "repos/$REPO/git/ref/heads/$BRANCH" --jq '.object.sha')
test -n "$BRANCH" && test "$HEAD_SHA" = "$REMOTE_SHA"
If the marker or remote head match is absent, run Step C and report a producer
failure. CHECKPOINT EVIDENCE is useful context but is never a verification
receipt: a producer may report it, but it cannot establish a passing check.
Step 4.5: Supervisor-Owned Verification
The operator, not the producer, discovers and executes the expected checks at
the immutable pushed HEAD_SHA, in a detached checkout the producer cannot edit:
SUPERVISOR_CHECKOUT=$(mktemp -d)
git fetch origin "$BRANCH"
git worktree add --detach "$SUPERVISOR_CHECKOUT" "$REMOTE_SHA"
Perform repository-owned verification discovery in that detached checkout.
Read the task/spec and the repository's own contributor, agent, build, CI, and
test instructions. Identify every check those sources require for this change.
Do not hardcode a language, framework, conventional filename, or command set in
this shared skill; do not manufacture checks from a producer's tests pass
claim. Execute each discovered command yourself from the supervisor checkout —
never by prompting the producer.
Load the checkpoint's specs/<feature>/invariant-matrix.tsv and validate it
against references/invariant-matrix.md. The supervisor alone reconciles row
states: mark every executed receipt whose repository or full checkpoint differs
from the reviewed checkpoint stale, then bind each fresh observed check to its
row as passed, failed, or unavailable. Preserve prior receipt text and
finding IDs. A producer summary, reviewer statement, or narrative-only result
can never set passed. Include the reconciled matrix path and row states in
SUPERVISOR_VERIFICATION.
The reconciled matrix is a supervisor-owned serialized artifact, not an edit in
the detached checkout. Set SUPERVISOR_MATRIX to its complete TSV content and
SUPERVISOR_MATRIX_RECEIPT to repository=$REPO checkpoint=$REMOTE_SHA; verify
both identities before every use. Embed both values verbatim in the review,
correction, final-review, resume, and PR prompts. The branch copy remains the
planning record: no producer may reconcile it or replace this artifact with a
summary. On resume, rebuild the artifact at the reconciled remote head before
any consumer uses prior evidence.
Then write a plain-markdown verification summary — one line per expected check:
the check identity, the verbatim command, and its observed result (passed only
on exit status 0; otherwise failed with the real exit status, not-run with
the reason, or unavailable naming the specific missing runtime/credential):
SUPERVISOR_VERIFICATION="verified at head $REMOTE_SHA:
- <check id>: <command> — <passed|failed (exit N)|not-run: reason|unavailable: reason>
..."
Keep the summary available to the review, correction, and PR steps. The session EXIT trap removes the detached worktree on every terminal path.
Step 5: Independent Advisory Review
Spawn a second LLM worker in the same repo. It receives the issue, branch, diff, supervisor verification summary, and proposed PR claims — never the producer's rationale. It starts from a clean checkout; its findings are advisory.
FIRST_REVIEW_STATUS="unavailable"
REVIEW_OUT="Independent review unavailable: reviewer did not start."
REVIEW_SPAWN=$(curl -s --max-time 90 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"repo\": \"$REPO\"}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers")
RWID=$(printf '%s' "$REVIEW_SPAWN" | jq -r '.worker_id // empty')
if [ -n "$RWID" ]; then
REVIEW_PROMPT=$(BRANCH="$BRANCH" SUPERVISOR_VERIFICATION="$SUPERVISOR_VERIFICATION" SUPERVISOR_MATRIX="$SUPERVISOR_MATRIX" SUPERVISOR_MATRIX_RECEIPT="$SUPERVISOR_MATRIX_RECEIPT" python3 -c "import json,os; print(json.dumps('Independently review issue #$0 in $REPO and pushed branch ' + os.environ['BRANCH'] + ' from a clean checkout. Read the issue, spec, tasks, trusted repository guidance, merge-base diff, and supervisor-owned matrix and receipts below. Do not request or use producer rationale. Do not edit, commit, push, or create a PR.\n\nAttempt to falsify every applicable invariant row by ID. For each surviving row emit exactly: ROW INV-NNN verdict=no-findings. For each challenge emit exactly one line: FINDING F-NNN row_id=INV-NNN priority=P0..P3 status=open evidence=<concrete evidence> suggested_action=<action>. A source-backed omission uses row_id=OMITTED. Cover every applicable row exactly once; unsupported classes need no row. Findings are advisory.\n\nSUPERVISOR MATRIX RECEIPT:\n' + os.environ['SUPERVISOR_MATRIX_RECEIPT'] + '\nSUPERVISOR MATRIX:\n' + os.environ['SUPERVISOR_MATRIX'] + '\nSUPERVISOR VERIFICATION:\n' + os.environ['SUPERVISOR_VERIFICATION'] + '\n\nEnd with [pylot] phase=independent-review status=done actionable=yes|no.'))")
REVIEW_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"prompt\": $REVIEW_PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/prompt")
REVIEW_SEQ=$(printf '%s' "$REVIEW_RESP" | jq -r '.turn_seq // empty')
if [ -z "$REVIEW_SEQ" ]; then
REVIEW_OUT="Independent review unavailable: reviewer prompt was not accepted."
curl -s --max-time 20 -X POST -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/stop" >/dev/null 2>&1 || true
RWID=""
fi
fi
When RWID exists, poll with the dedicated bounded helper — not Step P:
bash ~/.claude/skills/speckit-runner/poll-reviewer.sh "$RWID" "$REVIEW_SEQ"
Re-run only while its last line is REVIEW_POLL_RESULT=running. Each call is
short and the cumulative review ceiling is 15 minutes. done permits the output
check below. unavailable means the helper already stopped the reviewer: set
REVIEW_OUT to an unavailable reason, keep FIRST_REVIEW_STATUS=unavailable,
set RWID="" because the helper stopped it, and continue to Step 6. Never route a reviewer result through Step P's producer
failure/mission-stop behavior.
if [ -n "$RWID" ]; then
REVIEW_STATE=$(curl -s --max-time 20 -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}")
REVIEW_OUT=$(printf '%s' "$REVIEW_STATE" | jq -r '.last_output // ""')
REVIEW_FILE=$(mktemp)
MATRIX_FILE=$(mktemp)
printf '%s\n' "$REVIEW_OUT" >"$REVIEW_FILE"
printf '%s\n' "$SUPERVISOR_MATRIX" >"$MATRIX_FILE"
if bash ~/.claude/skills/speckit-runner/validate-review-output.sh "$MATRIX_FILE" "$REVIEW_FILE"; then
FIRST_REVIEW_STATUS="available"
else
FIRST_REVIEW_STATUS="unavailable"
REVIEW_OUT=$(printf 'Independent review unavailable: reviewer output was incomplete or malformed.\nPreserved raw output below is advisory only; do not treat it as complete coverage.\n\n%s\n' "$REVIEW_OUT")
curl -s --max-time 20 -X POST -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/stop" >/dev/null 2>&1 || true
RWID=""
fi
rm -f "$REVIEW_FILE" "$MATRIX_FILE"
fi
Step 6: One Producer Correction Pass
Send the independent suggestions to the producer exactly once. The producer must record what it changed and why it declined anything; it must not loop indefinitely.
CORRECTION_PROMPT=$(REVIEW_OUT="$REVIEW_OUT" SUPERVISOR_MATRIX="$SUPERVISOR_MATRIX" SUPERVISOR_MATRIX_RECEIPT="$SUPERVISOR_MATRIX_RECEIPT" python3 -c "import json,os; print(json.dumps('''This is the one bounded correction pass before PR creation. Independently assess the review suggestions below against the issue and repository. Implement the high-value valid corrections; decline inapplicable or disproportionate suggestions with a concrete reason. Preserve every invariant row ID, evidence state, receipt, and finding ID; record each finding as resolved or declined with its disposition. The supervisor-owned matrix below is read-only producer context. If the pushed head changes, treat all unmatched executed matrix receipts as stale pending fresh supervisor verification; producer prose must not restore passed. Re-run the repository-defined verification plus /speckit-analyze $0 and /speckit-checklist $0, commit any changes, and push the branch. Do not create a PR yet. Emit [pylot] phase=correction status=done branch=\$(git branch --show-current) head=\$(git rev-parse HEAD), followed by a concise CORRECTION SUMMARY that lists each suggestion as resolved or declined with rationale and records the latest verification results.
SUPERVISOR MATRIX RECEIPT:
''' + os.environ['SUPERVISOR_MATRIX_RECEIPT'] + '''
SUPERVISOR MATRIX:
''' + os.environ['SUPERVISOR_MATRIX'] + '''
INDEPENDENT REVIEW:
''' + (os.environ.get('REVIEW_OUT') or 'Review unavailable; verify the checkpoint yourself and report that independent review evidence was unavailable.'))) ")
CORRECTION_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"prompt\": $CORRECTION_PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}/prompt")
TURN_SEQ=$(printf '%s' "$CORRECTION_RESP" | jq -r '.turn_seq // empty')
Poll the producer to done. This is the only correction pass even if suggestions
remain. Capture and validate its receipt:
CORRECTION_STATE=$(curl -s --max-time 20 -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}")
CORRECTION_OUT=$(printf '%s' "$CORRECTION_STATE" | jq -r '.last_output // ""')
if ! printf '%s' "$CORRECTION_OUT" | grep -q 'phase=correction status=done' \
|| ! printf '%s' "$CORRECTION_OUT" | grep -q 'CORRECTION SUMMARY'; then
CORRECTION_VALID="no"
else
CORRECTION_HEAD=$(printf '%s' "$CORRECTION_OUT" | sed -n 's/.*phase=correction status=done.* head=\([0-9a-f]*\).*/\1/p' | tail -1)
CORRECTION_REMOTE_HEAD=$(gh api "repos/$REPO/git/ref/heads/$BRANCH" --jq '.object.sha' 2>/dev/null || true)
if [ -n "$CORRECTION_HEAD" ] && [ "$CORRECTION_HEAD" = "$CORRECTION_REMOTE_HEAD" ]; then
CORRECTION_VALID="yes"
else
CORRECTION_VALID="no"
fi
fi
If CORRECTION_VALID=no, run Step C, then emit the producer failure with the
saved branch receipt. Do not proceed to review or PR creation.
Step 6.5: Re-verify After Correction
The correction may have pushed a new HEAD. Never carry a passed result across
that boundary. If CORRECTION_HEAD differs from the head Step 4.5 verified,
discard the old summary and repeat Step 4.5's repository-owned discovery and
supervisor execution at CORRECTION_HEAD:
if [ "$CORRECTION_HEAD" != "$REMOTE_SHA" ]; then
echo '[speckit-runner] correction moved HEAD; re-verifying at the new head'
cleanup_supervisor_checkout
SUPERVISOR_CHECKOUT=$(mktemp -d)
git fetch origin "$BRANCH"
git worktree add --detach "$SUPERVISOR_CHECKOUT" "$CORRECTION_HEAD"
cd "$SUPERVISOR_CHECKOUT"
# Repeat Step 4.5 discovery and execution here, then rebuild
# SUPERVISOR_VERIFICATION with "verified at head $CORRECTION_HEAD:".
# Rebuild SUPERVISOR_MATRIX from those observations and set
# SUPERVISOR_MATRIX_RECEIPT="repository=$REPO checkpoint=$CORRECTION_HEAD".
fi
If fresh checks are failed, not-run, or unavailable, continue. Their exact states must remain in the final summary; none are PR-creation blockers.
Step 7: Optional Final Advisory Review
If the first review reported actionable=yes or the correction changed the
pushed head, prompt the reviewer once more to inspect the updated remote head.
Ask only which original concerns are resolved and which suggestions remain. Do
not start another producer correction pass. If the first review had no actionable
suggestions, set RESIDUAL_REVIEW="No residual suggestions." and skip this turn.
The final review remains advisory. Persist residual suggestions for disclosure; never convert them into a mission blocker merely because they remain.
RUN_FINAL_REVIEW="no"
if printf '%s' "$REVIEW_OUT" | grep -q 'actionable=yes' \
|| { [ -n "${CORRECTION_HEAD:-}" ] && [ "$CORRECTION_HEAD" != "$HEAD_SHA" ]; }; then
RUN_FINAL_REVIEW="yes"
fi
if [ "$FIRST_REVIEW_STATUS" = "available" ] && [ -n "$RWID" ] \
&& [ "$RUN_FINAL_REVIEW" = "yes" ]; then
FINAL_REVIEW_PROMPT=$(CORRECTION_OUT="$CORRECTION_OUT" SUPERVISOR_VERIFICATION="$SUPERVISOR_VERIFICATION" SUPERVISOR_MATRIX="$SUPERVISOR_MATRIX" SUPERVISOR_MATRIX_RECEIPT="$SUPERVISOR_MATRIX_RECEIPT" python3 -c "import json,os; print(json.dumps('Re-fetch branch $BRANCH and review its updated diff for issue #$0. Reassess only your original F-NNN findings against the correction summary and the supervisor-owned matrix at its exact receipt. Preserve finding IDs and report each as resolved, declined, or still open with evidence; do not erase unavailable review or non-passing row states. Do not edit, push, or create a PR. Remaining suggestions are advisory. End with [pylot] phase=final-review status=done.\n\nSUPERVISOR MATRIX RECEIPT:\n' + os.environ['SUPERVISOR_MATRIX_RECEIPT'] + '\nSUPERVISOR MATRIX:\n' + os.environ['SUPERVISOR_MATRIX'] + '\nPRODUCER CORRECTION SUMMARY:\n' + os.environ['CORRECTION_OUT'] + '\n\nCURRENT SUPERVISOR VERIFICATION:\n' + os.environ['SUPERVISOR_VERIFICATION']))")
FINAL_REVIEW_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"prompt\": $FINAL_REVIEW_PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/prompt")
REVIEW_SEQ=$(printf '%s' "$FINAL_REVIEW_RESP" | jq -r '.turn_seq // empty')
if [ -z "$REVIEW_SEQ" ]; then
RESIDUAL_REVIEW="Final advisory review unavailable: prompt was not accepted."
curl -s --max-time 20 -X POST -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/stop" >/dev/null 2>&1 || true
RWID=""
fi
else
RESIDUAL_REVIEW="No final review run: first review unavailable or had no actionable suggestions."
if [ -n "${RWID:-}" ]; then
curl -s --max-time 20 -X POST -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/stop" >/dev/null 2>&1 || true
RWID=""
fi
fi
When REVIEW_SEQ exists, use poll-reviewer.sh again and re-run it only while
REVIEW_POLL_RESULT=running. On unavailable, set a concise unavailable reason,
clear RWID, and continue. On done, fetch output and accept it only when it
contains phase=final-review status=done; otherwise mark it unavailable. Stop the
reviewer immediately after capturing a valid final output. This wait has the same
15-minute cumulative ceiling and can never fail the mission.
if [ -n "${RWID:-}" ] && [ -n "${REVIEW_SEQ:-}" ]; then
FINAL_REVIEW_STATE=$(curl -s --max-time 20 -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}")
FINAL_REVIEW_OUT=$(printf '%s' "$FINAL_REVIEW_STATE" | jq -r '.last_output // ""')
if printf '%s' "$FINAL_REVIEW_OUT" | grep -q 'phase=final-review status=done'; then
RESIDUAL_REVIEW="$FINAL_REVIEW_OUT"
else
RESIDUAL_REVIEW="Final advisory review unavailable: no valid completion marker."
fi
curl -s --max-time 20 -X POST -H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${RWID}/stop" >/dev/null 2>&1 || true
RWID=""
fi
Step 8: Create the Sole PR
Send the producer the final advisory record. This is the only PR creation boundary in the pipeline.
PR_PROMPT=$(SUPERVISOR_VERIFICATION="$SUPERVISOR_VERIFICATION" SUPERVISOR_MATRIX="$SUPERVISOR_MATRIX" SUPERVISOR_MATRIX_RECEIPT="$SUPERVISOR_MATRIX_RECEIPT" FIRST_REVIEW_STATUS="$FIRST_REVIEW_STATUS" REVIEW_OUT="$REVIEW_OUT" CORRECTION_OUT="$CORRECTION_OUT" RESIDUAL_REVIEW="$RESIDUAL_REVIEW" python3 -c "import json,os; print(json.dumps('''Create the single PR for issue #$0 from the already-pushed branch. First confirm the worktree is clean and the remote head matches local HEAD. Invoke /create-compelling-prs and use that skill to compose and open the PR — do not substitute a placeholder template.
Base the PR's Verification section on the supervisor verification summary below. It is authoritative over any producer claim. Consume the supervisor-owned matrix only when its receipt matches the exact current repository/head. Never reconcile or restore execution states yourself. If that identity does not match, disclose executed rows as stale or the matrix as unavailable; do not claim passed evidence for the current head. Disclose every applicable row whose state is failed, not-run, unavailable, or stale; preserve not-applicable rows without presenting them as gaps. Never state or imply readiness when any current check is non-passing. Include an Independent review section summarizing first-review availability, every unresolved or declined F-NNN finding, correction decisions, and residual or unavailable final review. Review suggestions and verification failures are transparent advisory context, not a reason to suppress the PR or create another PR boundary. Emit [pylot] phase=pr status=done pr=<PR_URL>.
CURRENT SUPERVISOR VERIFICATION:
''' + os.environ['SUPERVISOR_VERIFICATION'] + '''
SUPERVISOR MATRIX RECEIPT:
''' + os.environ['SUPERVISOR_MATRIX_RECEIPT'] + '''
SUPERVISOR MATRIX:
''' + os.environ['SUPERVISOR_MATRIX'] + '''
FIRST REVIEW STATUS:
''' + os.environ['FIRST_REVIEW_STATUS'] + '''
FIRST REVIEW:
''' + os.environ['REVIEW_OUT'] + '''
PRODUCER CORRECTION SUMMARY:
''' + os.environ['CORRECTION_OUT'] + '''
FINAL ADVISORY REVIEW:
''' + (os.environ.get('RESIDUAL_REVIEW') or 'Final advisory review unavailable.'))) ")
PR_RESP=$(curl -s --max-time 30 -X POST \
-H "Authorization: Bearer $PYLOT_DISPATCH_TOKEN" -H "Content-Type: application/json" \
-d "{\"prompt\": $PR_PROMPT}" \
"${PYLOT_API}/missions/${PYLOT_JOB_ID}/workers/${WID}/prompt")
TURN_SEQ=$(printf '%s' "$PR_RESP" | jq -r '.turn_seq // empty')
Poll the producer to done, find the PR URL in its output, and confirm with
gh pr view that its head branch and issue linkage match this run.
Any producer prompt, poll, marker, or reconciliation failure runs Step C before
emitting its terminal outcome.
ISSUE_NUMBER="$0"
source ~/.claude/skills/speckit-runner/shared/pr-postcondition.sh || {
echo '[speckit-runner] fatal: pr-postcondition.sh not found — boot sync may be incomplete' >&2
exit 1
}
verify_pr_postcondition || { echo '[speckit-runner] PR postcondition failed' >&2; exit 1; }
Step 9: Cleanup + Emit Outcome
Run Step C unconditionally, then report the PR and independent-review availability.
If the PR exists, residual or unavailable suggestions produce status=success,
not failed or blocked.
# Run the Step C snippet first. The session EXIT trap removes
# SUPERVISOR_CHECKOUT on this and every other terminal path.
if [ -n "$PR_NUM" ]; then
echo "[speckit-runner] complete: PR #$PR_NUM opened; independent_review=$FIRST_REVIEW_STATUS"
else
echo "[speckit-runner] verified checkpoint exists but no PR URL was confirmed — check worker output"
fi
After the reporting command, emit exactly one resolved marker as your final full assistant line (not from Bash and not inside a fence):
- PR confirmed: [pylot:$PYLOT_OUTCOME_NONCE] outcome="speckit complete: PR #$PR_NUM opened; independent_review=$FIRST_REVIEW_STATUS" status=success
- No PR URL: [pylot:$PYLOT_OUTCOME_NONCE] outcome="verified checkpoint exists but no PR URL was confirmed — check worker output" status=partial
Hard Rules
- Pre-flight is mandatory — the worker must gather real data before speckit phases
- Poll only via the scripts — producer turns use
poll-worker.sh; reviewer turns use boundedpoll-reviewer.sh. Both return in <2 min per call. Never hand-roll a loop, pass a long Bash timeout, background a poll, wait for a notification, or end the operator turn while a worker turn is in flight. - Never stop a healthy worker on a timer —
block_elapsedis a checkpoint, not a failure. Only stop a worker when its heartbeat is stale (> ~10 min) or it reported a failure; never on elapsed time or empty mid-turn output alone. The script alone enforces the hard ceiling. - Cleanup always runs — every terminal path runs Step C; reviewer unavailability stops that reviewer immediately and continues
- External dispatch remains fire-and-forget — foreground polling is internal mission execution, not a reason for the dispatcher to babysit the mission
- Review in clean context — the producer never reviews its own checkpoint; the reviewer never edits it
- Supervisor owns verification — discover commands only from the task and repository-owned instructions, execute them in the detached supervisor checkout at the exact pushed head, and report their observed results honestly. Producer prose can never create a passed result.
- Repository trust boundary —
/speckit-runnermust only be used against repositories whose contributor, agent, and CI instruction files are trusted to the same degree as the dispatch token. The supervisor executes repository-discovered commands inside its own session (holds$PYLOT_DISPATCH_TOKEN,$PYLOT_API, andghcredentials); a hostile or compromised instruction file in the target repository can exfiltrate those credentials. - Suggestions never gate — allow one producer correction pass, disclose anything residual, and continue to the PR boundary
- PR creation happens once and last — verification, analyze/checklist, checkpoint push, and advisory review all precede
/create-compelling-prs - Emit the outcome marker —
[pylot:$PYLOT_OUTCOME_NONCE] outcome=... status=is mandatory as your final full assistant line before exiting - "already complete" only at the dedup gate — only emit this when the issue is genuinely CLOSED (Step 0); never for timeouts or missing notifications
- One task, one PR — do not scope-creep into adjacent issues
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/fellowship-dev/dogfooded-skills/speckit-runner">View speckit-runner on skillZs</a>