skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
tikalk/adlc-team-skills99 installs

evals-analyze

Use when evaluation results need triage and loop-closing — spec failures route to deterministic checks or context rules, generalization failures to the evaluator backlog.

How do I install this agent skill?

npx skills add https://github.com/tikalk/adlc-team-skills --skill evals-analyze
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides automated analysis of AI evaluation results and helps close development loops by routing failures to either automated checks or rule updates. It includes setup scripts for resolving project paths and configuration.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

evals-analyze

What this skill does

Provides cross-functional team elevation and closed-loop feedback following EDD Principle VIII (Close the Production Loop) by deep-analyzing trajectory failure traces and routing them to correct resolution pathways.

Output:

  1. Trajectory Analysis - Full multi-turn trace analysis with tool calls and context preservation (EDD Principle V)
  2. Failure Routing:
    • Specification Failures (agent logic missing/ambiguous) → subclassified (EVAL-010):
      • Mechanical (fixed, checkable pattern) → fix/extend the existing grader or add a unit test — a deterministic check, not a CDR
      • Judgment gap (missing intent, ambiguity, needs context) → automatically triggers a local call to team-levelup to propose new context rules in adlc branch drafts/cdr/ to fix agent behavior
    • Generalization Failures (grader flawed or lacks edge-case coverage) → Appends evaluator backlog items to the project backlog for ongoing monitoring.
  3. Cross-Functional PR - Creates a team-ai-directives PR with insights and rule updates (EDD Principle X)

Key EDD Principles Applied:

  • Principle VIII: Close Production Loop - Spec failures → fix directives; Gen failures → evaluator backlog
  • Principle V: Trajectory Observability - Full multi-turn traces, not just outputs
  • Principle X: Cross-Functional Observability - PMs, domain experts, and AI engineers collaborate

When to use

  • After /evals-validate: Analyze failures and resolve them
  • Closing a development loop: Translate evaluation failure insights into rule or evaluator fixes
  • Reporting to stakeholders: Generate readable summaries for PMs and domain experts

When NOT to use

  • Evals not yet executed: Run /evals-validate first to generate results in evals/results/
  • Trivial tasks: Closed-loop analysis is overhead for simple features

Process

User Input

$ARGUMENTS
  • --focus AREA — Focus analysis on specific areas (e.g., security, quality, performance)
  • --dry-run — Analyze results and print report, but skip PR creation and local skill triggers

Execution Steps

Phase 1: Load Evaluation Results

  • Reads results JSON from evals/results/.
  • Extracts failure cases and full multi-turn conversation traces (including tool calls).

Phase 2: Failure Classification

Categorizes each failure trace:

  • Specification Failure: The agent was correct relative to its context, but the rule/directive was missing, ambiguous, or incorrect. Subclassify (EVAL-010):
    • Mechanical: the gap is a fixed, checkable pattern (banned API, import shape, file-location, syntactic shape) → deterministic check territory
    • Judgment gap: the gap needs intent, context, or cross-file judgement → context rule territory
  • Generalization Failure: The rule was correct, but the agent made a mistake anyway (hallucinated, missed a constraint, or grader lacked edge-case coverage).

Phase 3: Action Routing (Close the Loop)

  • For Mechanical Specification Failures: Fix/extend the existing binary grader or add a unit test that enforces the pattern. Do NOT propose a CDR for a mechanically-checkable gap — pay once for the check instead of re-deriving it per session.
  • For Judgment-gap Specification Failures: Automatically triggers local skill /team-levelup with the failure trace as input. This creates new rule/persona/example CDRs in adlc branch drafts/cdr/ to fix the agent's behavior.
  • For Generalization Failures: Appends an evaluator backlog item to evals/results/evaluator_backlog.md detailing the needed grader edge-case updates.

Phase 4: Cross-Functional Insights & PR

  • Generates a stakeholder-specific report in evals/results/team_insights.md (tailored for PMs, domain experts, and AI engineers).
  • If git remote and gh CLI are available, commits rule/eval changes in team-ai-directives and opens a draft PR (uses team-levelup logic under the hood).

Verification

  • Trajectory failure traces analyzed and classified
  • Mechanical specification failures routed to grader/unit-test fixes (deterministic checks); judgment-gap failures routed to /team-levelup (proposes CDRs in adlc branch drafts/cdr/)
  • Generalization failures written to evals/results/evaluator_backlog.md
  • Stakeholder report evals/results/team_insights.md generated
  • Draft PR created in team-ai-directives (if applicable)
  • Final report summary presented with PR link and backlog details

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/tikalk/adlc-team-skills/evals-analyze">View evals-analyze on skillZs</a>