skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
ai-evals-course/evals-skills184 installs

write-code-eval

Write code evaluators for known failure modes with objective rules. Use when code can check the rule from a trace, with or without a reference answer. Use `write-judge-prompt` when the rule requires interpretation.

How do I install this agent skill?

npx skills add https://github.com/ai-evals-course/evals-skills --skill write-code-eval
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides instructional guidelines for creating deterministic code evaluators to detect specific failure modes. It consists entirely of documentation and configuration without any executable code or external dependencies.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Write a code evaluator

Start with a failure mode found through error analysis. Write one check for that failure mode, much like a unit test that asserts what should hold for each trace.

  1. State the rule and identify the trace fields or reference data the check needs. If the rule requires interpretation, use write-judge-prompt.
  2. Implement the check in the project's language and eval framework. Return a result and a reason in the format that framework expects.
  3. Test known passes and failures, including borderline cases. Run the check on available traces and inspect mistakes. If the rule uses a proxy for human judgment, compare its results with human labels.

Examples

Failure modePossible check
Invalid output structureParse the output and check required fields
Missing or forbidden textMatch a string or pattern
Citation not in retrieved documentsCompare cited IDs with retrieved IDs
Bad tool callCheck arguments against the tool schema or run the call in a safe test environment
Wrong valueCompare the output with a reference value

Choose the check from the failure rule. For a failure with both objective and interpretive parts, check the objective part with code and use a judge for the rest.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/ai-evals-course/evals-skills/write-code-eval">View write-code-eval on skillZs</a>