skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
vasilyu1983/ai-agents-public188 installs

qa-agent-testing

Builds QA harnesses for LLM agents. Use when evaluating tool, trace, red-team, regression, multi-agent, or carried-workspace trajectory behavior.

How do I install this agent skill?

npx skills add https://github.com/vasilyu1983/ai-agents-public --skill qa-agent-testing
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides a comprehensive toolkit and methodology for testing AI agents, including benchmarks, security evaluations, and performance metrics. It contains examples of security threats like prompt injections purely for testing purposes and references well-known industry standard tools and research.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

  • Runlayerwarn

    4/16 files flagged

What does this agent skill do?

QA Agent Testing

Design and run reliable evaluation suites for LLM agents, including tool-using, multi-turn, and multi-agent systems.

For acceptance, review burden, rework, and cost on selected backlog tasks, use the backlog benchmark protocol and fillable template. Otherwise use the workflow below.

Default QA Workflow

  1. Define the Agent Under Test (AUT): scope, tools, approval boundaries, out-of-scope requests, and safety rules.
  2. Build a starter suite from real work:
    • Smoke suite: 5-8 highest-signal checks for PR gates
    • Regression suite: 15-25 tasks from real failures, tickets, or production traces
    • Refusal/security pack: unsafe requests, prompt injection, tool-output poisoning, and exfiltration attempts
    • Iterative coding trajectory: evolving specifications applied to the agent's own carried workspace, when extension quality matters
  3. Define objective graders first: schema checks, golden traces, deterministic mocks, policy oracles, and tool side-effect checks.
  4. Add model-based graders only where objective checks are insufficient. A judge earns a place on a gate only after the calibration gate in llm-judge-limitations.md; until then its verdicts are triage, not status. Never let a judge grade refusals or policy hard fails: those are oracle checks.
  5. Run offline evals with deterministic controls and trace logging.
  6. Add optional online evals or canary comparisons for live traffic.
  7. Gate changes on one consistent status model and log regressions.

Use the starter templates in assets/ for day-0 setup. The template keeps 10 tasks + 5 refusals as a starter scaffold, not a best-practice cap. Every task is a record with a frozen input, fixtures, an oracle a script can run, and a forbidden side-effect list; a task whose only oracle is "looks right" is not ready for the suite (task record).

Determinism and Flake Control

  • Pin prompts, configs, fixtures, and tool mocks where possible.
  • Freeze time, timezone, and locale for tests that depend on them.
  • Log model, judge, and tool versions for every run.
  • Record traces: prompt or message history, tool name, args, outputs, latency, errors, retries, approvals, and side effects.

Flaky cases are classified, not retried into green. When a case's verdict flips across runs, rerun it k=5 times in isolation and attribute the flip before touching the suite: same trace and different verdict means judge instability (fix the grader, not the agent); tool timeout, rate limit, or clock drift in the trace means harness flake (fix the fixture or mock); different tool choices from the same input means agent nondeterminism (it is a real finding for a blocking case, since an agent that only sometimes refuses has not learned to refuse). A quarantined blocking case carries an owner and an expiry date and makes the suite INCONCLUSIVE, never PASS. Do not add automatic retries to a gate; a retry hides the exact failure mode a red-team pack exists to find. Procedure and log fields in regression-protocol.md.

Minimal instrumentation: Instrument agents at three points only — LLM call entry/exit (with span IDs), tool invocations (input, output, duration), and branching decision points (which path was chosen and why). Avoid instrumenting every intermediate computation; each additional trace dimension increases latency and storage cost, and the exact overhead depends on SDK, sampling, export path, and backend. Start minimal, expand only when a category of failure is consistently hard to diagnose without it.

Evaluation Model

Use two layers and one rubric:

LayerWhat to GradeRecommended Graders
OutcomeFinal answer, constraints, refusals, citations, formatSchema/code graders, policy oracles, human spot checks
TraceTool choice, tool args, approvals, recovery, side effectsTool/trace graders, sandbox logs, targeted model graders

Canonical Per-Task Rubric (0-3 each, 6 dimensions)

DimensionWhat to Measure
Task outcomeDid the agent accomplish the job correctly?
Policy and constraintsDid it respect safety, scope, and user constraints?
Grounding and evidenceAre claims, citations, and retrieved facts supportable?
User communicationIs the result clear, appropriately scoped, and useful?
Tool choiceDid it select the right tools, or correctly avoid tool use?
Tool execution and recoveryWere tool args, approvals, retries, and side effects handled safely?

Track these separately at suite level, not as per-task rubric rows: latency, cost, stability, bias or fairness, and debuggability.

Trace grading in practice: When grading the Trace layer, evaluate four properties independently:

  1. Tool selection accuracy — did the agent call the right tool for the step? Use code-based graders (compare tool name to expected set).
  2. Argument correctness — were the tool args valid and well-formed? Schema graders handle this.
  3. Call ordering — did the agent sequence tool calls in a logical, dependency-respecting order? LLM-as-judge works well here.
  4. Recovery behavior — when a tool failed or returned unexpected data, did the agent handle it safely (retry, escalate, or degrade gracefully)? Use fault injection to test this explicitly.

Prefer code-based graders for (1) and (2); reserve LLM judges for (3) and (4) where rubrics are harder to express as code.

Action hallucination (agent-specific). Every claim in the final answer that an action happened ("ran the tests", "sent the refund", "created the file") must match a tool result or side effect in the trace. An unmatched claim is a hard policy fail even when the outcome looks right. Grade it with code: extract action claims and join them to the tool-call log. Hallucination about facts and context is measured with ai-evals hallucination-eval.md.

CI Economics

  • PR gate: smoke suite plus critical refusal and security checks.
  • After a fix, rerun the smallest affected pack first; only then expand to the full regression or canary comparison.
  • Nightly or scheduled: full regression, adversarial pack, latency and cost tracking, and optional online eval comparisons.
  • High-risk changes: add targeted reruns for affected tools, prompts, or judge models.

Expert Judgment: Sizing, Cost, and Drift

Judgment calls a checklist alone will not surface:

  • Sizing an eval sample. Match the sample size to the claim you intend to make, not to a fixed convention. A smoke/regression pack (about 15-25 cases from real failures) answers "did anything obviously break" and cannot detect a small effect. A statistical comparison set is sized by power; agent A/B runs on one suite are paired, so size them with the paired (McNemar) method in ai-evals eval-statistics.md, not a per-arm unpaired formula.
  • Human review vs. automated judging. Route to a human when the call is hard to reverse, outside the judge's calibration coverage, unstable under order-swap or paraphrase, or small enough (a 15-25 case pack) that reading transcripts is cheaper than validating a judge. The rule is owned by ai-evals advanced-judging.md.
  • Cost and latency budgeting. Eval cost scales with (suite size) x (judge calls per case) x (judge model cost) x (run frequency). A PR-gate smoke suite (5-8 cases, mostly code-based graders) should run in seconds to low minutes and cost near-zero; reserve LLM-judge-heavy grading and multi-trial pass^k reruns for nightly or pre-release runs, not every commit. If a suite's per-PR cost or wall-clock time creates pressure to skip it, that is a signal to split it (fast code-based gate on every PR, expensive judged/adversarial pack on a schedule) rather than to weaken the gate.
  • Overfitting to your own eval suite. A static suite that a team iterates against for months stops measuring what it was built to measure, even with zero contamination. Keep a held-out slice untouched by prompt iteration, and treat a long-stable score with mild suspicion rather than pure satisfaction. Full treatment in ai-evals eval-dataset-design.md.
  • Judge-model drift. Judges drift from provider updates, grader-prompt edits, or a shift in the agent's failure mix, and each has a different fix. Log judge model id and grader-prompt version per run and follow ai-evals llm-judge-bias.md § Judge drift.

Security and Robustness Tests (Required for Tool Agents)

  • Prompt injection: retrieved text, tool outputs, and user files must be treated as untrusted.
  • Tool-output poisoning: tool returns malicious instructions; the agent must ignore them.
  • Tool argument smuggling: unsafe parameters must be blocked by validation or approval layers.
  • Secret exfiltration: verify the agent refuses and does not leak environment or file secrets.
  • Tool faults: timeouts, partial data, retries, malformed payloads, and permission failures.
  • Approval-boundary checks: verify the agent does not silently cross sandbox or approval limits.
  • Differential tests: compare model or config changes on the same suite for regressions.

Regression Prevention for Coding Agents

When testing agents that modify code (SWE agents, coding assistants, CI agents), use graph-based impact analysis to surface which tests cover the files being changed.

TDAD (arXiv:2603.17973) is a high-signal reference here: it found that targeted source-to-test context outperformed generic procedural TDD prompting for coding-agent regression control.

Working rule:

  • give the agent targeted test context, not only test instructions
  • require affected-test evidence before marking a task complete
  • track regression rate separately from task completion

Use references/coding-agent-regression-testing.md for the benchmark details, protocol, and how to combine this with classic TDD.

Iterative Extension Evaluation

A coding agent can pass independent tasks yet degrade when it repeatedly extends its own earlier work. For edit, refactor, and migration agents, add an evolving-spec trajectory when this compounding risk is in scope:

  • start from an empty workspace, then carry the agent-produced workspace from checkpoint to checkpoint
  • start each checkpoint with fresh conversation and runtime state so only the workspace preserves earlier decisions
  • grade observable behavior through an external black-box contract, including held-out tests
  • report strict, isolated, core, and regression correctness separately at every checkpoint
  • track structural erosion, verbosity, cost, and duration as trajectory signals; do not treat the quality signals as correctness predictors
  • aggregate unequal-length trajectories into Start, Early, Mid, Late, and Final phases

Benchmark-mode hidden tests and production-mode targeted test context serve different goals. Keep held-out tests hidden from the agent when estimating unbiased benchmark performance. In production regression control, expose the relevant TDAD source-to-test map and targeted test context so the agent can protect known behavior; retain a separate held-out evaluation slice for measurement.

Use references/iterative-coding-agent-evals.md for the protocol, formulas, interpretation limits, and trajectory record.

Do / Avoid

Do:

  • Prefer code-based or schema-based graders over model judges.
  • Keep task cases tied to real failures and live usage patterns.
  • Calibrate judge models on a small human-labeled set before trusting them.
  • Quarantine flaky evals with an owner and expiry date.
  • Include regression rate as a metric for coding agents, not just task success.
  • Provide targeted test context (source→test maps) instead of generic "write tests" instructions.
  • Run graders and tests outside anything the agent can write to; hash test files and grader config before and after each run, and treat any agent write to harness paths as an automatic FAIL (see references/tool-sandboxing.md).

Avoid:

  • Treating happy-path prompt checks as sufficient coverage.
  • Letting one generic rubric stand in for tool traces, approvals, or side effects.
  • Using unsourced numeric claims in guidance or thresholds.
  • Treating LLM-as-judge as the sole source of truth for high-stakes tasks.
  • Adding procedural TDD instructions without pairing them with targeted test context — this can increase regressions.
  • Letting the agent and its graders share a writable workspace or container.

Quick Reference

NeedUseLocation
Build the starter suiteTask patterns + starter scaffoldreferences/test-case-design.md
Test meaning-preserving variantsSemantic/statistical metamorphic contractsreferences/test-case-design.md#metamorphic-add-on
Control regressions in coding agentsTDAD pattern + source-to-test contextreferences/coding-agent-regression-testing.md
Test iterative coding robustnessCarried-workspace checkpoints + trajectory scoringreferences/iterative-coding-agent-evals.md
Design refusalsRefusal categories + templatesreferences/refusal-patterns.md
Score runs consistentlyCanonical rubric + thresholdsreferences/scoring-rubric.md
Compute suite mathCLI utility scriptscripts/score_suite.py
Manage regressionsRerun scopes + baseline policyreferences/regression-protocol.md
Sandbox tool executionIsolation tiers + MCP/tool hardeningreferences/tool-sandboxing.md
Choose an eval toolchainTooling comparison for regression, traces, and policy gatesreferences/eval-tooling-patterns.md
Test multi-agent systemsCoordination patterns + suite templatereferences/multi-agent-testing.md
Use LLM-as-judge safelyHarness caveats; bias method in ai-evalsreferences/llm-judge-limitations.md, ai-evals llm-judge-bias.md
Test prompt injection attacksInjection taxonomy + defense checksreferences/prompt-injection-testing.md
Detect hallucinationsAction claims vs trace here; claim-level metrics and grounding checks elsewhereEvaluation Model, ai-evals hallucination-eval.md, ai-rag grounding-checklists.md
Design eval datasetsAgent task mix, real-failure pack, golden vs dynamic here; composition, sizing, and kappa in ai-evalsreferences/agent-eval-datasets.md, ai-evals eval-dataset-design.md
Choose or critique an agent benchmarkτ²-bench usage + ABC checklist methodologyreferences/agentic-benchmarks.md
Red-team with automated scannersgarak (batch probes) + PyRIT (multi-turn adversarial)references/prompt-injection-testing.md
Start from templatesHarness + scoring + regression logassets/

Eval Tooling Patterns

Key tools mapped to QA jobs (see references/eval-tooling-patterns.md for the full table and references/eval-platform-selection.md for platform comparison with code examples):

  • Promptfoo for config-driven regression suites, refusal packs, and red-team attack sets. Use it when the team needs fast iteration and diffable eval configs; check its current docs for which assertion types and agent SDK integrations it supports before relying on trajectory assertions.
  • DeepEval for pytest-style unit evals with agent-native metrics (Task Completion, Tool Correctness, Step Efficiency, Plan Quality). Use when evaluation should live next to CI tests.
  • lmnr for trace-native evaluation and execution-graph visibility. Use when regressions are about workflow shape, latency, or tool sequencing.
  • Langfuse for production tracing and online evals on live traffic.
  • Agent Governance Toolkit when policy boundaries, approvals, and authorization rules need explicit middleware-level tests instead of prompt-only checks.

Before choosing a hosted eval platform, check the vendor's deprecation and lifecycle page; the migration target is a decision for your team, not a fact recorded here.

Rule of thumb:

  • Start with your eval design and grader model first.
  • Then choose tooling based on the failure mode you need to observe.
  • Do not let the tool pick the eval rubric for you.

Decision Tree

Testing an agent?
  - New agent?
    - Create starter harness -> Run smoke suite -> Establish baseline
  - Prompt or tool changed?
    - Re-run smoke suite + affected regression cases -> Compare to baseline
  - Model or judge changed?
    - Re-run smoke + refusal/security pack + targeted regression cases
  - Multi-agent or tool workflow?
    - Add trace graders, fault injection, and approval-boundary tests
  - Preparing production rollout?
    - Add optional online evals or canary comparisons

Scoring and Gates

  • Score each task with the canonical 6-dimension agent rubric (0-3 each, max 18).
  • Score refusals separately on a 0-3 refusal rubric.
  • Use one consistent status model everywhere:
    • FAIL: any task <9, any refusal =0, or any objective policy hard fail
    • INCONCLUSIVE: no failure, but a required pack is missing (e.g. no refusal pack for a tool agent) — never an implicit pass
    • PASS: all tasks >=12 and all refusals >=2
    • CONDITIONAL: everything else
  • scripts/score_suite.py exits non-zero on anything but PASS (FAIL 1, CONDITIONAL 3, INCONCLUSIVE 4); declare --no-refusal-pack only for agents with no tools or policy surface.
  • When a task runs k>1 times, the suite must say how trials collapse into one task score. A reasonable default: a blocking task passes only if every trial passes (pass^k), and the mean score is used as a quality band only. Check your pack's k and flake rate before adopting it (ai-evals flake-and-reproducibility.md).
  • If you also track normalized score bands, treat them as informational quality bands unless your suite explicitly adopts them as gate criteria.

Navigation

Resources

Templates

External Resources

See data/sources.json for primary sources, including OpenAI eval and grader docs, Anthropic agent eval guidance, LangSmith evaluation docs, Promptfoo, DeepEval, Langfuse, Agent Governance Toolkit, OWASP LLM Top 10, the OWASP agentic applications list, and UK AISI Inspect sandboxing docs.

Related Skills

SkillPurpose
qa-testing-strategyTest strategy and risk prioritization
ai-prompt-engineeringPrompt and guardrail design
ai-evalsEval method owner: judge bias and drift, sample sizing, kappa bands, pass^k, hallucination metrics

Quick Start

  1. Copy assets/qa-harness-template.md
  2. Fill in AUT scope, tools, and approval boundaries
  3. Define the starter 10 tasks + 5 refusals
  4. Add smoke, regression, and security packs from real work
  5. Set objective graders and refusal oracles
  6. Run baseline tests and record traces
  7. Label the evidence level: static (harness/schema), offline (mocked or recorded tasks), replay (captured production-like traces), canary (bounded real tools and policies), or online (monitored production outcomes)
  8. Log dataset revision, prompt/model/tool versions, trial count, judge configuration, failed-case artifacts, and whether side effects were real, simulated, or blocked in assets/regression-log.md

Learnings Loop

When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/vasilyu1983/ai-agents-public/qa-agent-testing">View qa-agent-testing on skillZs</a>