ax-audit
Audits agentic products for tool parity, authority, approval payloads, recovery, and trust using 24 rules and a ship verdict. Use when asked for an "AX audit", to review an agent approval flow, or whether an agent can operate the product.
How do I install this agent skill?
npx skills add https://github.com/mblode/agent-skills --skill ax-auditIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a comprehensive framework for auditing agentic products. It operates by performing static analysis (grep) on source code and PR diffs using a set of defined rules. No malicious patterns, data exfiltration, or obfuscation were detected. The skill's use of ripgrep for code analysis is consistent with its stated purpose.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
AX Audit
Feature-level reviewer for apps where an agent acts for the user. One question: does it earn trust, and where does it break?
- IS: rules-based audit of agentic surfaces (chat, tool execution, config, dashboards) across architecture (
rules-arch/) and trust (rules-ax/), ending in a ship-readiness verdict plus an AX Relationship Summary. - IS NOT: traditional frontend UX (use
ui-designAudit mode); developer-facing API, CLI, or type ergonomics (usedx-audit); public site or docs agent scores (useagent-ready); agent instruction files (useagents-md); what the product should do before it exists (useproduct-design).
No agentic features in scope? Stop. AX rules against forms and lists are noise.
Contents
- Audit workflow
- Two rule layers
- Tiers and verdict
- AX Relationship Summary
- Reference files
- Gotchas
- Audit self-check
- Related skills
Audit workflow
AX Audit progress:
- [ ] Step 1: Scope, via the diff against the PR base merge-base (PR mode) or explicit path (full sweep)
- [ ] Step 2: Detect agentic features per references/feature-playbooks.md
- [ ] Step 3: Run each detected feature's playbook in order, plus the diff-wide checks (PR mode only)
- [ ] Step 4: For each check, load the rule file and follow its detection recipe
- [ ] Step 5: Tier each finding per references/ship-readiness.md (rule override table wins)
- [ ] Step 6: Render verdict + findings + AX Relationship Summary per references/output-format.md
- [ ] Step 7: Run the audit self-check and report its evidence counts
PR-mode scope is the diff plus the tool definitions and orchestrator it touches. Findings in untouched files belong in a full sweep, not this verdict. Playbook annotations are a scan copy; the rule file is authoritative. parity-orphan-ui-action runs on every PR-mode audit and never in a full sweep, where there is no diff for it to read.
Rule greps name the most common identifiers, not every framework's spelling. When a grep misses in code that plainly does the thing (a gate, a stream, a tool result), check references/framework-signals.md for the stack's name for it before recording unknown.
Two rule layers
| Layer | Folder | Rules | Load when a playbook names |
|---|---|---|---|
| 1: Architecture | rules-arch/ | 11 | rules-arch/<category>-<slug>.md |
| 2: Experience | rules-ax/ | 13 | rules-ax/<category>-<slug>.md |
Categories: arch = parity, granularity, context, comm; ax = trust, control, context, comm. Shared prefixes are different rules: rules-arch/comm-no-approval-gate.md (no gate on the execution path) is not rules-ax/control-no-approval-gate.md (gate exists, stakes are wrong).
Run Layer 1 comm/parity and Layer 2 control/trust first. They hold the blockers. Category map: rules-arch/_sections.md, rules-ax/_sections.md.
| Priority | Layer | Category | Prefix | Rules |
|---|---|---|---|---|
| 1 | arch | Communication | comm- | 3 |
| 2 | arch | Parity | parity- | 4 |
| 3 | ax | Control | control- | 4 |
| 4 | ax | Trust | trust- | 3 |
| 5 | arch | Context | context- | 2 |
| 6 | ax | Communication | comm- | 4 |
| 7 | ax | Context | context- | 2 |
| 8 | arch | Granularity | granularity- | 2 |
Tiers and verdict
Three tiers: release-blocker, fix-this-sprint, backlog. Definitions, the generic surface bump, and verdict logic live in references/ship-readiness.md.
Precedence: the rule's own surface-override table > the generic bump > defaultTier. Apply at most one adjustment.
Verdict: ✅ READY (0 blockers, ≤3 sprint) · ⚠️ READY WITH FOLLOW-UP (0 blockers, ≥4 sprint) · ❌ NOT READY (≥1 blocker) · 🚫 INCOMPLETE (self-check failed).
Blockers outrank an incomplete audit. With ≥1 release-blocker and a failed self-check, report ❌ NOT READY and note the self-check failure beneath it: the blockers are established findings and stay actionable, while 🚫 reads as "nothing was learned" and sends the reader away. Reserve 🚫 for an audit with no blockers whose coverage you cannot vouch for.
AX Relationship Summary
Render after findings when any agentic feature was detected. Findings serve engineers; this serves designers and PMs. Four fields: evolution stage (behavior, not a label), trust signal (high/moderate/low plus one-line reason), key gap (one actionable sentence), trust question (one question only research can answer).
Reference files
| File | Read when |
|---|---|
references/feature-playbooks.md | Steps 2-3: detection heuristics, per-feature ordered checks, diff-wide checks |
references/framework-signals.md | Step 4, when the code uses AI SDK, MCP, the Claude Agent SDK, or AG-UI: where the gate, the stream, the completion signal, and the structured result live in each, with the spec defaults the rules lean on |
references/ship-readiness.md | Step 5: tier definitions, precedence, generic surface bump, verdict logic |
references/output-format.md | Step 6: findings JSON schema, summary schema, terminal rendering |
references/ax-evolution-curve.md | Writing the AX Relationship Summary: stage, action depth, costume vs intelligence, and the arguments with no rule that land in keyGap |
rules-arch/_sections.md | Layer 1 categories, default tiers, co-firing pairs |
rules-ax/_sections.md | Layer 2 categories, default tiers, co-firing pairs |
Gotchas
- Scope before rules. Running all 24 rules repo-wide on a 3-file PR buries a new release-blocker under pre-existing backlog noise; the verdict stops meaning "can this PR merge."
- The rule's override table is authoritative.
comm-no-intent-handshakedefaults tofix-this-sprintbut its table saysrelease-blockeron tool execution. Stacking the generic "+1 tier on tool execution" bump on an explicit override double-upgrades backlog findings into blockers. - A stop button not wired to
AbortController.abort()is a false affordance.control-no-escape-hatchstill fails: verify theabort()call, not the button label, or the audit passes a UI that lies to users. - A client
stop()that only closes the stream leaves the executor running.useChat().stop()aborts the fetch. Unless the route passesreq.signalintostreamText({ abortSignal })and toolexecutereads it, the server finishes every remaining tool call after the user pressed Stop. Trace the signal to the loop, not to the button. - Tool annotations are hints, not stakes. MCP tells clients to treat
annotationsfrom untrusted servers as untrusted; a gate that auto-approves on a third-party server'sreadOnlyHint: truehas handed the gate to that server.comm-no-approval-gatefails it. The spec defaults (destructiveHint: true,readOnlyHint: false) are the fail-closed baseline. - A framework approval flag is the gate's input, not the gate. AI SDK
toolApproval: "user-approval"emits atool-approval-requestpart and waits. A UI that never rendersstate === "approval-requested", or answers it withaddToolApprovalResponse({ approved: true })on arrival, has a gate in the type system and none for the user. Check the renderer and the response call, not the option. - Absence checks need a recorded file list. "Find components lacking X" greps return nothing both when everything passes and when nothing was scanned. List candidate files first (
rg -l <feature-pattern>), check each for the counter-pattern, and cite the file list as evidence. detection: observationalrules cannot fail on grep evidence alone.granularity-static-api-mapping,control-over-conversational,comm-no-generative-momentum, and the uncertainty-gradient half oftrust-no-confidence-cuesneed interaction-flow judgment; on static evidence alone, returnunknownwith a reason, notfail.- Gates fail in three separate places. Absent from the path (
comm-no-approval-gate), present but mismatched to the stakes (control-no-approval-gate), or correct and unreadable (control-thin-approval-payload). Report the first that holds and fix in that order. - Interactive gates do not cover unattended runs. Cron, webhook, and queue entry points reach the same executor with nobody to prompt.
comm-unrequested-action-no-consentaudits that path; evidence names the entry point, not the executor. ax-audit-ignore:<slug>comments count assuppressed, notpass. Report the count in the verdict block; a suppression with no reason is itself awarn.- Don't inflate tiers.
comm-no-generative-momentumandgranularity-static-api-mappingdefault tobacklog. One finding promoted torelease-blockerflips the whole PR to ❌ NOT READY, so promoting cosmetic ones trains the team to ignore the verdict entirely. - Don't duplicate
ui-designAudit mode findings. "Missing loading state" and "form clears on error" are its territory; duplicating them trains engineers to dismiss the whole AX report. - A Personally Intelligent agent that only ever suggests has plateaued. Memory stage is not trust. Name the highest action rung in
evolutionStage.behavioror the summary flatters a polite chatbot.
Audit self-check
Flag the audit INCOMPLETE if any of these hold, and include the counts as evidence (planned vs. run rules per playbook, unknown rate, suppressed count):
- Fewer rules ran than the playbooks planned
- More than 30% of rules returned
unknown. Count onlyunknownhere, neverout-of-scope: a rule whose layer is absent from the scope you were given was answered correctly, and a narrow diff is the scope Step 1 asks for. Marking a correctly scoped audit INCOMPLETE buries its real blockers under a verdict that reads as "we learned nothing". - Any
fail/warnfinding lacksfile:lineevidence or a fix snippet - Every finding landed in the same tier (suspect blanket assignment)
- AX Relationship Summary is missing despite detected agentic features
Related skills
ui-designAudit mode: traditional frontend UX around agentic surfaces; run both on agentic feature PRsdx-audit: same files, different reader. This skill asks whether an agent can operate and recover;dx-auditasks whether a human adopting the API, CLI, or types finds it ergonomicagent-ready: whether public docs and HTTP APIs are discoverable to coding agents; this skill audits in-product agent UXproduct-design: what the agentic feature should do, before this auditagents-md: CLAUDE.md / AGENTS.md instruction files
Maintenance only: evals/evals.json and evals/evaluation-scenarios.md hold regression scenarios for changes to this skill; neither loads during a user task.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/mblode/agent-skills/ax-audit">View ax-audit on skillZs</a>