agent-output-audit
Independent audit of AI-implemented work — certifies a completed task actually did what it claims, checking files, diffs, tests, and CI rather than the agent's self-report. Flags skipped or weakened tests, mock-hidden integration, snapshot drift, happy-path-only coverage, flaky retries, and status/evidence mismatches. Use when validating completed Compozy tasks, AI-authored PRs, or codex-loop iterations. Not for real-user, persona, or journey QA — use qa-execution for those.
How do I install this agent skill?
npx skills add https://github.com/pedronauck/skills --skill agent-output-auditIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a tool for auditing AI-generated code by running the project's own tests and build scripts. It analyzes the repository to find these commands and then executes them. It includes detailed protocols for identifying 'red flags' in AI-written tests, such as skipped tests or weakened assertions. No malicious behavior was detected.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Agent Output Audit
You are the independent evaluator. Answer one question — "Did the implementing agent actually do what task_NN.md says it did?" — from files, public behavior, tests, and CI. A self-report is not evidence. (Whether a real user can succeed at the product is qa-execution; run both on a Compozy slug and keep their outputs separate.)
Step 1: Discover the Repository Verification Contract
- Read root instructions, repository docs, and CI/build files before running commands.
- Run
python3 scripts/discover-project-contract.py --root .to surface candidate install/verify/build/test/lint/start commands and E2E signals. - Prefer repository-defined umbrella commands (
make verify,just verify, CI entrypoints) over language defaults. When discovery surfaces more than one plausible gate or mixes ecosystems, readreferences/project-signals.mdbefore choosing, and state the tie-breaker. - Read
references/e2e-coverage.mdbefore classifying any flow's coverage. - Resolve the audit artifact directory: the
audit-output-pathargument if given, else repository conventions, else/tmp/agent-output-audit-<slug>. Create itsaudit/subdirectory; store all bugs and reports under<audit-output-path>/audit/. - Detect Compozy mode. If
.compozy/tasks/<slug>/exists, record the slug and:- Read
state.yamlread-only —scripts/update-state.pyowns its mutation per the cy-codex-loop contract. - Read
_techspec.md(deliverable source of truth) and_tasks.md(task roster) when present. - List every
task_NN.mdand capture its frontmatterstatus:(pending|in_progress|completed). When frontmatter disagrees withstate.yaml, frontmatter is the source of truth. - Note the memory slot
.compozy/tasks/<slug>/memory/qa-execution.md— Step 5 writes it before any status flip.
- Read
Step 2: Run the Baseline Verification Gate
- Install dependencies with the repository-preferred command.
- Run the canonical gate once before any audit work, fastest-first: lint and type-check → build → unit tests → integration tests.
- If the E2E command is separate from the umbrella gate, decide whether to run it now or after runtime prerequisites are ready, and record that plan.
- On a baseline failure, read the first failing output and determine whether it is pre-existing or introduced by current work. Exclude a failure from audit scope only after a clean reproduction proves it unrelated.
- Flaky-failure protocol. Before classifying any baseline failure, run the failing test in isolation 3-5 times on the same SHA. If it passes at least once without a code change, record it as
flaky-suspectin theSUITE HEALTH SNAPSHOT(test name, attempts, retry outcome, suspected category) rather than promoting it to PASS. Readreferences/flaky-triage.mdbefore assigning a suspected category or proposing a quarantine.
Step 3: Audit Task Implementations
Skip this step only when no task, phase, PRD, tech spec, or implementation-plan artifacts exist.
- Read
references/independent-evaluator-protocol.mdin full before forming any verdict — it owns what does and does not count as evidence, and the transcript classification (genuine-failure/grader-bug/ambiguous-task/bypass-exploit). In Compozy mode, read the implementer'smemory/<phase>.mdartifacts and record anomaly classifications inmemory/qa-execution.md→Errors / Correctionsbefore judging the task. - Summarize each
task_NN.mdand its body into a Task Implementation Matrix (columns mirror cy-codex-loop frontmatter):task_path,declared_status(literal frontmatterstatus:)title,type,complexity,dependencies— mirrored from frontmattertechspec_deliverable— linked_techspec.mdsection when present- Requirements, subtasks, checklist items, success criteria, dependent files
implementation_evidence— files, modules, routes, commands, migrations, seeds, testsverification_evidence— commands executed, exit codes, output summariesqa_verdict—PASS|PARTIAL|FAIL|REOPEN|BLOCKED(distinct fromdeclared_status)ai_audit_findings— red flag IDs that fired in Step 4 with verdictaction—none|fixed|reopened-frontmatter|BUG-NNN.md filedlinked_bugs— BUG IDs
- Verify every completed or claimed-complete task against actual files, public behavior, automated tests, and acceptance criteria. Re-execute the smallest public proof against the current repository state.
- Assign
qa_verdict:PASS: every material requirement and success criterion has implementation and fresh verification evidence.PARTIAL: implementation exists but one or more non-critical requirements, tests, or evidence are missing.FAIL: claimed behavior does not work or a critical requirement is absent.REOPEN: frontmatter saysstatus: completedbut the QA verdict isPARTIALorFAIL.BLOCKED: a concrete prerequisite is missing. Validate every local boundary that does not need the missing dependency and report the blocked live validation separately.
Step 4: AI Test-Hygiene Scan (RF-1..RF-6)
- Read
references/ai-implementation-audit.mdin full before scanning the test diff of any task withdeclared_status: completed— it owns the RF-1..RF-6 scanners, the Requirement→Test mapping, and the verdict matrix. - Run the scans against the diff since the task baseline (
git log --follow <test_file>,git diff <baseline_sha>..HEAD). - Emit the verdict the matrix assigns. RF-1 (skip/only/xit/t.Skip inserted), RF-2 on a P0/P1 criterion (weakened assertion), RF-3 (mock on a dependency the TC declared Integration/E2E), and RF-4 on P0/P1 (unjustified snapshot drift) are automatic
FAIL. - Record findings in the matrix column
ai_audit_findingsand in the per-task block ofaudit-report.md. - Apply the Requirement→Test mapping: for every Success Criterion in
task_NN.mdand every linked_techspec.mdbullet, mark the matching testcovers/weak/missing. A checked item orstatus: completedwithout acoversrow is an audit failure.
Step 5: Reopen, File Bugs, Write Memory
- Mark every incomplete completed task
REOPEN. - In Compozy mode, write
memory/qa-execution.mdwith the cy-codex-loop canonical sections (Objective Snapshot,Important Decisions,Learnings,Files / Surfaces,Errors / Corrections,Ready for Next Run) before flipping anytask_NN.mdfrontmatter (memory-precedes-status invariant). - Edit the offending
task_NN.mdfrontmatterstatus:back topending(orin_progressif salvageable). Leavestate.yamlalone —update-state.pyowns it, and the next iteration reconciles from frontmatter. - File
BUG-<num>.mdunder<audit-output-path>/audit/issues/usingassets/issue-template.md, including: the task path (Reopens task:), the failed Success Criterion (Summary:), the original strict assertion when RF-2 fired (Root cause:), the red flag ID and verdict (Automation Follow-up:), and any transcript anomaly classification (Related:). - When the gap is a bounded root-cause fix inside the audit scope, implement it, add regression coverage, and rerun the task proof. Otherwise reopen the task.
Step 6: Quality Gates Verdict
- Re-run the canonical verification gate from scratch after the last code change made during the audit.
- Compile the Quality Gates section of
audit-report.md, eachPASS/FAIL/N/A:- Flaky rate <2% in the canonical suite.
- Zero
FAILfrom the AI test-hygiene scan on P0/P1 tasks. - Zero
Critical/Highissues open. - Coverage delta ≥ baseline (no regression).
- Zero unresolved
flaky-suspecton P0 flows.
- A
FAILon any gate blocks an unconditional PASS verdict for the run.
Step 7: Write the Audit Report
- Write the report to
<audit-output-path>/audit/audit-report.mdusingassets/audit-report-template.md, with all mandatory sections:- Claim / Command / Exit code / Verdict per command executed in Steps 2 and 6.
- AUTOMATED COVERAGE — support detected, harness, canonical command, required flows with classification, specs added or updated.
- TASK IMPLEMENTATION AUDIT — Compozy slug, plan sources, matrix totals, per-task verdicts, reopened/fixed/blocked tasks, links to bugs.
- SUITE HEALTH SNAPSHOT — flaky rate, flaky events, mutation score (when a harness exists), coverage delta vs baseline, blocked count, manual-only count, AI audit findings count.
- QUALITY GATES — PASS/FAIL/N/A per gate.
- ISSUES FILED — total, by severity, with
Reopens task:annotations. - Report each blocked scenario, missing credential, or environment gap with the exact command or prerequisite that stopped execution.
- In a Compozy slug, a final PASS feeds cy-codex-loop's
verify.last_status=PASSprecondition for Phase E — leaveupdate-state.pyto cy-codex-loop. - Before declaring the audit complete, confirm every item in
references/checklist.md— it is the exhaustive completion criterion across all steps.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/pedronauck/skills/agent-output-audit">View agent-output-audit on skillZs</a>