evals-analyze
Use when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.
How do I install this agent skill?
npx skills add https://github.com/tikalk/adlc-team-skills --skill evals-analyzeIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill provides automated analysis of AI evaluation results and helps close development loops by routing failures to either automated checks or rule updates. It includes setup scripts for resolving project paths and configuration.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
evals-analyze
What this skill does
Provides cross-functional team elevation and closed-loop feedback following EDD Principle VIII (Close the Production Loop) by deep-analyzing trajectory failure traces and routing them to correct resolution pathways.
Output:
- Trajectory Analysis - Full multi-turn trace analysis with tool calls and context preservation (EDD Principle V)
- Failure Routing:
- Specification Failures (agent logic missing/ambiguous) → subclassified (EVAL-010):
- Mechanical (fixed, checkable pattern) → fix/extend the existing grader or add a unit test — a deterministic check, not a CDR
- Judgment gap (missing intent, ambiguity, needs context) → automatically triggers a local call to
team-levelupto propose new context rules inadlc branch drafts/cdr/to fix agent behavior
- Generalization Failures (grader flawed or lacks edge-case coverage) → Appends evaluator backlog items to the project backlog for ongoing monitoring.
- Specification Failures (agent logic missing/ambiguous) → subclassified (EVAL-010):
- Cross-Functional PR - Creates a team-ai-directives PR with insights and rule updates (EDD Principle X)
Key EDD Principles Applied:
- Principle VIII: Close Production Loop - Spec failures → fix directives; Gen failures → evaluator backlog
- Principle V: Trajectory Observability - Full multi-turn traces, not just outputs
- Principle X: Cross-Functional Observability - PMs, domain experts, and AI engineers collaborate
When to use
- After
/evals-validate: Analyze failures and resolve them - Closing a development loop: Translate evaluation failure insights into rule or evaluator fixes
- Reporting to stakeholders: Generate readable summaries for PMs and domain experts
When NOT to use
- Evals not yet executed: Run
/evals-validatefirst to generate results inevals/results/ - Trivial tasks: Closed-loop analysis is overhead for simple features
Process
User Input
$ARGUMENTS
--focus AREA— Focus analysis on specific areas (e.g., security, quality, performance)--dry-run— Analyze results and print report, but skip PR creation and local skill triggers
Execution Steps
Phase 1: Load Evaluation Results
- Reads results JSON from
evals/results/. - Extracts failure cases and full multi-turn conversation traces (including tool calls).
Phase 2: Failure Classification
Categorizes each failure trace:
- Specification Failure: The agent was correct relative to its context, but the rule/directive was missing, ambiguous, or incorrect. Subclassify (EVAL-010):
- Mechanical: the gap is a fixed, checkable pattern (banned API, import shape, file-location, syntactic shape) → deterministic check territory
- Judgment gap: the gap needs intent, context, or cross-file judgement → context rule territory
- Generalization Failure: The rule was correct, but the agent made a mistake anyway (hallucinated, missed a constraint, or grader lacked edge-case coverage).
Phase 3: Action Routing (Close the Loop)
- For Mechanical Specification Failures: Fix/extend the existing binary grader or add a unit test that enforces the pattern. Do NOT propose a CDR for a mechanically-checkable gap — pay once for the check instead of re-deriving it per session.
- For Judgment-gap Specification Failures: Automatically triggers local skill
/team-levelupwith the failure trace as input. This creates new rule/persona/example CDRs inadlc branch drafts/cdr/to fix the agent's behavior. - For Generalization Failures: Appends an evaluator backlog item to
evals/results/evaluator_backlog.mddetailing the needed grader edge-case updates.
Phase 4: Cross-Functional Insights & PR
- Generates a stakeholder-specific report in
evals/results/team_insights.md(tailored for PMs, domain experts, and AI engineers). - If git remote and gh CLI are available, commits rule/eval changes in
team-ai-directivesand opens a draft PR (usesteam-leveluplogic under the hood).
Verification
- Trajectory failure traces analyzed and classified
- Mechanical specification failures routed to grader/unit-test fixes (deterministic checks); judgment-gap failures routed to
/team-levelup(proposes CDRs inadlc branch drafts/cdr/) - Generalization failures written to
evals/results/evaluator_backlog.md - Stakeholder report
evals/results/team_insights.mdgenerated - Draft PR created in team-ai-directives (if applicable)
- Final report summary presented with PR link and backlog details
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/tikalk/adlc-team-skills/evals-analyze">View evals-analyze on skillZs</a>