skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
yonatangross/orchestkit201 installs

assess

Assesses and rates quality 0-10 across multiple dimensions (correctness, maintainability, security, performance, testability, simplicity) with pros/cons analysis. Compares against project conventions and prior decisions from memory. Produces structured evaluation reports with actionable improvement suggestions. Use when evaluating code, designs, architectures, or comparing alternative approaches.

How do I install this agent skill?

npx skills add https://github.com/yonatangross/orchestkit --skill assess
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill is a comprehensive code and design assessment tool that uses structured evaluation, scoring, and automated refutation. It employs background agents for multi-dimensional analysis and specialized local scripts for data persistence and report rendering. Security controls such as scope boundaries and egress-guarded commands are integrated into the workflow.

  • Socketpass

    No alerts

  • Snykwarn

    Risk: MEDIUM · 1 issue

  • Runlayerfail

    5/18 files flagged

What does this agent skill do?

Assess

Host-neutral workflow. Invoke by skill name (assess). Claude Code slash routing, YAML hook loaders, and .claude/chain live in references/claude-code.md.

Comprehensive assessment skill for answering "is this good?" with structured evaluation, scoring, and actionable recommendations.

🎯 Quick Start

assess backend/app/services/auth.py
assess our caching strategy
assess --model=opus the current database schema
assess frontend/src/components/Dashboard

Effort levels (CC 2.1.111+ adds xhigh)

EffortBehavior
low / mediumSubset of dimensions, faster turnaround
high (default)All six dimensions with pros/cons
xhighAll six dimensions + one additional assessor pass focused on uncertainty/caveats; emits confidence per dimension

xhigh silently falls back to high on a model that does not implement it: no error, no log line. doctor Category 14 reports this, and only when it can positively prove the active model lacks the tier.


Argument Resolution

Step 0: resolve a conversational reference first

$ARGUMENTS is often not a path. For a bare pronoun or deictic (them, this, that, these, they, same, the above, the last one, what we just did) or an empty target after flags are stripped, the subject is in the conversation. Read back for the NEAREST concrete one (a file just discussed, a diff or PR just opened, a component just investigated) and announce the resolution in one line, so a wrong guess costs a correction rather than a turn: "Reading 'them' as the 3 pretool guards we just probed; say otherwise and I'll switch."

Refusing is the bug, not the safe option. Asking "what does this refer to?" when the previous turn named the subject burns a round-trip re-deriving what is already on screen. Measured 2026-08-28: the operator sent assess them throguhly one message after "bug in orchestkit hooks", mid-investigation of pretool/bash/dangerous-command-blocker, and this skill replied that "them" had "no antecedent anywhere in this conversation". It had two.

Ask only when the conversation is genuinely empty (a fresh session opening with a bare pronoun). Every other case: resolve and announce.

Not unique to this skill: verify, cover, fix-issue, review-pr and implement all read $ARGUMENTS as a literal path or topic, and no skill mentions resolving a reference. Tracked separately; this one fixes its own door.

TARGET = "$ARGUMENTS"  # Full argument string, e.g., "backend/app/services/auth.py"
# $ARGUMENTS[0] is the first token (CC 2.1.59 indexed access)

# Model override detection (CC 2.1.72)
MODEL_OVERRIDE = None
for token in "$ARGUMENTS".split():
    if token.startswith("--model="):
        MODEL_OVERRIDE = token.split("=", 1)[1]  # "opus", "sonnet", "haiku", "fable"
        TARGET = TARGET.replace(token, "").strip()

Pass MODEL_OVERRIDE to all Agent() calls via model=MODEL_OVERRIDE when set. Accepts symbolic names (opus, sonnet, haiku, fable on harnesses whose Agent tool lists it; note fable is premium API spend after 2026-07-12) or full IDs (claude-opus-5-5) per CC 2.1.74.

Switching to Opus via /model (CC 2.1.144+): /model now changes the model for the current session only, so picking Opus for an assess run no longer persists past it. Press d in the picker only to set a default for new sessions.

Effort detection (CC 2.1.120+)

$CLAUDE_EFFORT is the primary signal. CC 2.1.120 sets this env var from /effort or the model picker. --effort= token in $ARGUMENTS is the explicit override fallback (also covers older CC).

# Read env first (CC 2.1.120+), then check explicit override
EFFORT = os.environ.get("CLAUDE_EFFORT")  # "low" | "medium" | "high" | "xhigh" | None
for token in "$ARGUMENTS".split():
    if token.startswith("--effort="):
        EFFORT = token.split("=", 1)[1]   # explicit override wins
        TARGET = TARGET.replace(token, "").strip()
EFFORT = EFFORT or "high"  # default when CC < 2.1.120 and no flag

Use EFFORT to gate dimension count, agent count, and the optional xhigh uncertainty pass — see "Effort levels" table above. On CC < 2.1.120 the env var is unset; the explicit --effort= override is the only path. doctor Category 14 reports a provably unsupported xhigh request.


STEP -1: MCP Probe + Resume Check

Load: Read("../chain-patterns/references/mcp-detection.md")

# 1. Probe MCP servers (once at skill start)
# memory is alwaysLoad in .mcp.json (CC 2.1.121+, #1541) — probe below kept as fallback for older CC:
ToolSearch(query="select:mcp__memory__search_nodes")

# 2. Store capabilities
Write(".claude/chain/capabilities.json", {
  "memory": probe_memory.found,
  "skill": "assess",
  "timestamp": now()
})

# 3. Check for resume
state = Read(".claude/chain/state.json")  # may not exist
if state.skill == "assess" and state.status == "in_progress":
    last_handoff = Read(f".claude/chain/{state.last_handoff}")

Phase Handoffs

PhaseHandoff FileContents
000-intent.jsonDimensions, target, mode
101-baseline.jsonInitial codebase scan results
202-evaluation.jsonPer-dimension scores + evidence
303-report.jsonFinal report, grade, recommendations

STEP 0: Verify User Intent with AskUserQuestion

BEFORE creating tasks, clarify assessment dimensions:

AskUserQuestion(
  questions=[{
    "question": "What dimensions to assess?",
    "header": "Dimensions",
    "options": [
      {"label": "Full assessment (Recommended)", "description": "All dimensions: quality, maintainability, security, performance"},
      {"label": "Code quality only", "description": "Readability, complexity, best practices"},
      {"label": "Security focus", "description": "Vulnerabilities, attack surface, compliance"},
      {"label": "Quick score", "description": "Just give me a 0-10 score with brief notes"}
    ],
    "multiSelect": false
  }]
)

Based on answer, adjust workflow (passed to Phase 2 as focus):

  • Full assessment (full): All 7 phases, parallel agents, dimension subset scaled by effort
  • Code quality only (quality): Skip security and performance phases
  • Security focus (security): Prioritize security-auditor agent
  • Quick score (quick): Single pass, brief output

STEP 0b: Select Orchestration Mode

Load details: Read("references/orchestration-mode.md") for env var check logic, Agent Teams vs Agent Tool comparison, and mode selection rules.


🚨 Task Management (CC 2.1.16)

# 1. Create main task IMMEDIATELY
TaskCreate(
  subject="Assess: {target}",
  description="Comprehensive evaluation with quality scores and recommendations",
  activeForm="Assessing {target}"
)

# 2. Create subtasks for each assessment phase
TaskCreate(subject="Understand target and gather context", activeForm="Understanding target")   # id=2
TaskCreate(subject="Discover scope and build file list", activeForm="Discovering scope")        # id=3
TaskCreate(subject="Rate quality across 6 dimensions", activeForm="Rating quality")             # id=4
TaskCreate(subject="Analyze pros and cons", activeForm="Analyzing pros/cons")                   # id=5
TaskCreate(subject="Compare alternatives", activeForm="Comparing alternatives")                 # id=6
TaskCreate(subject="Generate improvement suggestions", activeForm="Generating suggestions")     # id=7
TaskCreate(subject="Compile assessment report", activeForm="Compiling report")                  # id=8

# 3. Set dependencies for sequential phases
TaskUpdate(taskId="3", addBlockedBy=["2"])  # Scope needs target understanding
TaskUpdate(taskId="4", addBlockedBy=["3"])  # Rating needs scoped file list
TaskUpdate(taskId="5", addBlockedBy=["4"])  # Pros/cons needs quality scores
TaskUpdate(taskId="6", addBlockedBy=["4"])  # Alternatives need quality scores
TaskUpdate(taskId="7", addBlockedBy=["5", "6"])  # Suggestions need analysis
TaskUpdate(taskId="8", addBlockedBy=["7"])  # Report needs suggestions

# 4. Update status as you progress
TaskUpdate(taskId="2", status="in_progress")  # When starting
TaskUpdate(taskId="2", status="completed")    # When done — repeat for each subtask

🔄 Workflow Overview

PhaseActivitiesOutput
1. Target UnderstandingRead code/design, identify scopeContext summary
1.5. Scope DiscoveryBuild bounded file listScoped file list
2. Quality Rating6-dimension scoring (0-10)Scores with reasoning
3. Pros/Cons AnalysisStrengths and weaknessesBalanced evaluation
4. Alternative ComparisonScore alternativesComparison matrix
5. Improvement SuggestionsActionable recommendationsPrioritized list
6. Effort EstimationTime and complexity estimatesEffort breakdown
7. Assessment ReportCompile findingsFinal report

Phase 1: Target Understanding

Identify what's being assessed and gather context. TARGET here is the value Step 0 already resolved, which is not necessarily what the user typed.

# PARALLEL - Gather context
Read(file_path=TARGET)                                   # only when TARGET is a path
Grep(pattern=TARGET, output_mode="files_with_matches")   # topic or symbol
mcp__memory__search_nodes(query=TARGET)                  # past decisions

Read failing is NOT a reason to stop. A target resolved from the conversation is usually a subject rather than a filename ("the three pretool guards", "today's hook fixes"), so the Read misses and the Grep plus the conversation carry the context. Treat a failed Read as "this is a topic, not a path" and continue to Phase 1.5, which discovers the real file list anyway.


Phase 1.5: Scope Discovery

Load Read("references/scope-discovery.md") for the full file discovery, limit application (MAX 30 files), and sampling priority logic. Always include the scoped file list in every agent prompt.

Progressive Output (CC 2.1.76)

Output results incrementally as each evaluation phase completes:

After PhaseShow User
1. Target UnderstandingScope summary, file list, context
1.5. Scope DiscoveryBounded file list (max 30 files)
2. Quality RatingEach dimension's score as the evaluating agent returns
3. Pros/ConsBalanced evaluation summary

The Phase 2 workflow returns once, so show every dimension's score from its result and lead with priorityConcerns (any dimension below 4/10) as a concern needing user attention. Per-agent streaming applies only on the Agent tool fallback.


Phase 2: Quality Rating (6 Dimensions)

Rate each dimension 0-10 with weighted composite score. Load Read("../quality-gates/references/unified-scoring-framework.md") for dimensions, weights, grade interpretation, and per-dimension criteria. Load Read("references/quality-model.md") for assess-specific overrides.

Do NOT hand-roll the assessors. Run the executor, which owns Phases 2 and 2.5:

result = Workflow(
  scriptPath="${CLAUDE_SKILL_DIR}/workflows/assess-fanout.js",
  args={"target": TARGET, "effort": EFFORT, "focus": FOCUS,   # FOCUS from STEP 0
        "mode": "comparison" if COMPARING else "default",     # quality-model.md
        "domain": "frontend" or "backend",                     # picks the performance engineer
        "scopeFiles": SCOPE_FILES,                             # Phase 1.5 list
        "projectContext": MEMORY_CONTEXT,                      # Phase 1 memory search
        "rubric": Read("rubric.json"), "modelOverride": MODEL_OVERRIDE,
        "feature": FEATURE})
Write(".claude/chain/02-evaluation.json", result)

The script owns the mechanics, not the prose. It picks the assessors from focus and effort (security first), gives each a score schema that demands file:line evidence, sends every decision-bearing score to blind refuters (Phase 2.5), and computes the weighted composite, grade, rubric verdict and blockers. A score with no file:line evidence counts as unscored, a repeated dimension keeps only its first entry, and a selected dimension with a min_blocker that nobody scored is a blocker. It returns composite, grade, verdict, blockers (producer basis), postRefutation, chainVerdict, chainVerdictIfConfirmed, revisions, confirmationNeeded, manualReview, advisory, priorityConcerns, quickWins, unscored, rejectedDimensions, unassessed, dimensions, ledger and reasons. It never asks and never writes: those stay in this shell.

Fallback (no Workflow tool, ORCHESTKIT_FORCE_TASK_TOOL=1, or the cross-model lane below): Read("references/agent-spawn-definitions.md") for Agent tool and Agent Teams spawns, then run Phase 2.5 by hand.

Composite Score: Weighted average of the scored dimensions (see quality-model.md).


Phase 2.5: Adversarial Refutation (effort-gated)

The assessor that scores a dimension is also its only judge, a self-preferential bias. A separate blind refuter forms its own band for each decision-bearing score. Effort gate: low/medium skip it; high runs up to 4 single advisory refuters (no auto-swing); xhigh runs a 3-refuter majority that revises to the near band edge. On the Workflow path the script already ran it; the shell finishes it:

  1. Write the returned ledger to .claude/chain/02b-refutation.json (engine section 10).
  2. Re-open every cited file:line in revisions (engine section 3). A citation that does not hold reverts that dimension to its producer score; recompute the composite with the returned weights.
  3. If confirmationNeeded is non-empty, AskUserQuestion before using chainVerdictIfConfirmed; otherwise use chainVerdict. Refutation alone never raises a score or flips fail to pass (engine section 7).
  4. List manualReview dimensions as "not independently refuted", and surface every advisory overturn at high effort.

Protocol and assess bindings: Read("references/adversarial-refutation.md") (loads the shared engine ../../shared/rules/adversarial-refutation.md). Producer findings must first pass the evidence-replay gate before entering any score or verdict: Read("../../shared/rules/evidence-replay.md").

Cross-model refuter (optional, provenance-labeled, cost-gated)

When ORK_ALT_MODEL_CMD is configured and effort is high/xhigh, one quorum slot per high-weight or boundary-adjacent dimension score can route to a non-Claude model (Codex/GPT) for diverse failure modes. Off by default; substitutes one same-model slot, stamps refuter_model for provenance, cannot silently raise the grade (engine §7), owns no credentials/egress (shells out via ORK_ALT_MODEL_CMD, matches the egress guard #2533), and degrades to same-model on an absent command. Shares the review-pr operational doc: Read("../review-pr/references/cross-model-refuter.md").

The workflow does not run this lane (a script cannot shell out). When the user wants it, choose the Agent tool fallback before Phase 2 and run Phases 2 and 2.5 there.

Refuters are ALWAYS isolated spawns with no team_name, fed only the dimension and the scoped files: no producer score, identity, or prose. Keep the producer-basis score AND a labeled post-refutation score.


Phases 3-7: Analysis, Comparison & Report

Load Read("references/phase-templates.md") for output templates for pros/cons, alternatives, improvements, effort, and the final report.

See also: Read("references/alternative-analysis.md") | Read("references/improvement-prioritization.md")


Phase 7b: Emit Dashboard Spec (json-render)

Parse --render= from $ARGUMENTS. Default is both.

ModeBehavior
markdownCurrent behavior — markdown assessment report only. No spec emitted.
json-renderEmit .claude/chain/assess-dashboard.json only. Skip markdown report.
bothEmit spec and markdown. Default — human reads the report, downstream skills parse the spec.

When emitting a spec:

  1. Load format and catalog: Read("references/dashboard-spec.md"). Example: references/dashboard-example.json.
  2. Build the spec using only catalog types: Card, StatGrid, DataTable, StatusBadge, BarMeter, Markdown. Top-level fields composite (number) and grade (string) are required for assess specs.
  3. One BarMeter per dimension scored. The verdict element is a StatusBadge with status success/warning/error mapped from grade (A/B → success, C → warning, D/F → error).
  4. Write to .claude/chain/assess-dashboard.json with compact JSON.
  5. Validate before declaring success:
node "${CLAUDE_SKILL_DIR}/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json --check

If validation fails, fall back to markdown-only and surface the error. Never write a partial spec.

  1. For --render=both, render the markdown view from the spec:
node "${CLAUDE_SKILL_DIR}/scripts/render-spec.mjs" .claude/chain/assess-dashboard.json

This guarantees JSON spec and markdown report stay in sync.

xhigh effort: when effort=xhigh is active, add a sibling Markdown element per dimension containing confidence and caveats from the uncertainty pass. Reference list it in the dimensions Card's children alongside the BarMeter. See references/dashboard-spec.md for the exact pattern.

Downstream consumption: implement reads .claude/chain/assess-dashboard.json and pulls the lowest-scoring dimension and high-priority improvements (effort ≤ 2 AND impact ≥ 4) without parsing markdown tables. Measured: assess spec ≈ 830 tokens vs ~3500 token markdown for the same content.


Phase 7c: Memory Writeback (signal-fired, optional)

When the assessment lands with a composite score, optionally persist scores + summary to the memory MCP knowledge graph as a typed entity. Future memory queries can then surface assessment lineage (which decisions did this codebase score 9/10 on testability? when did security regress below 7.0?).

python3 ${CLAUDE_SKILL_DIR}/scripts/memory_writeback.py "<assessment-dir>"

<assessment-dir> is the dir containing assessment.json (typically the session's .claude/chain/). The script writes a memory-writeback.json handoff alongside it.

Auto-skip conditions (all exit 0, all WARN-logged):

Skip reasonTrigger
no composite scoreassessment.json has no top-level composite numeric field
yg-mcp-core not importableyg-mcp-core>=0.3.0 not installed (orchestkit is public; yg-mcp-core lives on private pypi.yonyon.ai — HQ-only)
memory MCP unreachablememory MCP server down OR .mcp.json doesn't define memory

The created entity has:

  • name: <slug-or-dir>@<timestamp> (stable across re-runs — re-runs create new entities)
  • entityType: assessment (override with --entity-type <type>)
  • observations: composite=X.XX, one <dim>=X.XX per scored dimension, optional summary: ... and topic: ...

Mirrors Yonatan-HQ/hq-ext-plugin#194 (audio_podcast handler) and orchestkit#1886 (post-synthesis podcast) pattern. Unblocked by Yonatan-HQ/core#993 (yg-mcp-core 0.3.0).


Phase 7d: Emit Chain Verdict (stop-gating)

After the composite and grade are final (post-refutation, Phase 2.5), ALWAYS write the machine-readable verdict: this is the stop-gate implement reads before Phase 1. On the Workflow path it is the returned chainVerdict (or chainVerdictIfConfirmed after a yes in Phase 2.5). Mirror the Phase 7b spec-emit pattern: write compact JSON, never a partial file.

// .claude/chain/assess-verdict.json
{
  "rubric": "ork-rubric/1.0",
  "skill": "assess",
  "verdict": "fail",
  "composite": 5.1,
  "dimension_scores": {"correctness": 7.0, "maintainability": 6.5, "performance": 5.5, "security": 3.2, "scalability": 6.0, "testability": 4.8, "compliance": 6.2},
  "blockers": [
    {"dimension": "security", "score": 3.2, "reason": "Unparameterized SQL in auth path (src/api/auth.ts:42)"}
  ],
  "feature": "<assessment topic, e.g. first non-flag token of $ARGUMENTS>"
}

Verdict rules — thresholds come from rubric.json (schema: ../../shared/rubric.schema.json):

  • verdict = "fail" when composite < min_pass (5.5) OR any dimension scores below its min_blocker. Otherwise "pass".
  • Every dimension below its min_blocker gets a blockers[] entry — dimension, score, one evidence-backed reason. blockers is [] on pass.
  • Scores are the post-refutation numbers — the same ones in the report. Refutation never silently flips a fail to pass.

Consumers: implement Step -0.5 blocks Phase 1 on verdict == "fail" (user must fix-first or explicitly override); Phase 7c memory writeback persists the verdict + dimension scores to the memory graph (add a verdict=pass|fail observation) for cross-session learning.


Self-Reported Uncertainty (xhigh effort)

Current-generation models report their own limits far better than older tiers did. When xhigh effort is active, enrich each dimension's rating with a confidence level and a list of caveats — things the model couldn't verify, assumptions it relied on, or cases it didn't test.

Output schema per dimension (JSON):

{
  "dimension": "security",
  "score": 7.2,
  "confidence": "medium",              // "low" | "medium" | "high"
  "caveats": [
    "Didn't execute the SQL queries against a real DB to confirm parameterization",
    "Assumed NODE_ENV=production in deployment; didn't verify CI config",
    "Reviewed 12 of 15 handlers; remaining 3 deferred by scope filter"
  ],
  "evidence": ["src/api/auth.ts:42", "src/middleware/guard.ts:88"]
}

Rules:

  • Do not use confidence as an auto-gate. It's a signal for the human reader, not a pass/fail threshold.
  • caveats must be specific. "Didn't check X" with file paths beats "uncertainty about security".
  • If a caveat is cheap to resolve, resolve it instead of recording it. Caveats are for things that genuinely can't be verified within the skill's scope (e.g., production runtime behavior, future input patterns).
  • Composite score still computes from score only — not weighted by confidence — to keep the number comparable across runs.

💡 Grade Interpretation

Load Read("../quality-gates/references/unified-scoring-framework.md") for grade thresholds and scoring criteria.


Key Decisions

DecisionChoiceRationale
6 dimensionsComprehensive coverageAll quality aspects without overwhelming
0-10 scaleIndustry standardEasy to understand and compare
Parallel assessmentworkflows/assess-fanout.js, up to 4 assessorsScores, refutation and verdict held in code, not prose
Effort/Impact scoring1-5 scaleSimple prioritization math

Rules Quick Reference

RuleImpactWhat It Covers
complexity-metrics (load rules/complexity-metrics.md)HIGH7-criterion scoring (1-5), complexity levels, thresholds
complexity-breakdown (load rules/complexity-breakdown.md)HIGHTask decomposition strategies, risk assessment

Quality Bar

Done means all of these hold:

  • Every in-scope dimension scored 0-10 with evidence (file:line) backing the score, not vibes
  • Composite is the weighted average of the scored dimensions and the grade maps from that composite
  • At high/xhigh effort, decision-bearing scores passed the adversarial refutation lane before entering the composite
  • .claude/chain/assess-verdict.json written with verdict pass/fail and a blockers[] entry for every dimension below its min_blocker
  • Any dimension scoring below 4/10 is flagged immediately as a priority concern
  • If a json-render spec is emitted, it passes render-spec.mjs --check and carries the required composite + grade fields

📜 Related Skills

  • ork:verify - Post-implementation verification
  • ork:code-review-playbook - Code review patterns
  • ork:quality-gates - Task complexity assessment, gate patterns

Version: 1.9.0 (September 2026): Phases 2 and 2.5 run as a Workflow script (workflows/assess-fanout.js)

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/yonatangross/orchestkit/assess">View assess on skillZs</a>