skill-architect
Architect and audit portable agent skills for retrieval, predictability, progressive disclosure, evals, and repository gates. Use when asked to "create a skill", "author a skill", "improve a skill", "review a skill", "refactor a skill", "compare skill rubrics", "distill a skill from sessions", "reconcile skill plans", or "teach skill authoring"; when paths include `skills/<name>/SKILL.md`, `references/*.md`, helper scripts, evals, lockfiles, security manifests, or vendor adapters; or when mentions include skill-creator, Skill Development, skills.sh, trigger descriptions, completion criteria, leading words, or model/user invocation.
How do I install this agent skill?
npx skills add https://github.com/t4sh/skills4sh --skill skill-architectIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill provides a robust framework and set of tools for architecting and auditing AI agent skills. It includes well-documented Python scripts for validation, inspection, and scaffolding that use secure practices such as SafeLoader for YAML parsing and environment isolation for external tool execution. No malicious patterns or security risks were identified.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Skill Architect
Architect portable, high-quality agent skills from a vendor-neutral perspective. Use this skill as an independent planning, creation, review, refactoring, comparison, distillation, reconciliation, and teaching layer for any agent skill repository or runtime.
Command modes
Use command-like modes when the user names one explicitly. If no mode is named, infer the closest mode from the request and state the assumption.
| Mode | Use when | Primary reference | Output |
|---|---|---|---|
skill-architect: plan | Designing a skill, repo standard, gate, release, or migration before editing | Portable rubric, Naming and packaging | Scope, affected files, standards, verification path, rollback notes |
skill-architect: create | Creating a new portable skill or scaffolding a skill folder | Portable rubric | SKILL.md plus linked references/assets/scripts and repo metadata plan |
skill-architect: audit | Reviewing an existing skill or skill repo against the rubric | Portable rubric | Source-cited findings table, mechanical gate status, patch recommendations |
skill-architect: fix | Applying narrow deterministic fixes that validate_skill.py can prove | Portable rubric | Dry-run patch list or applied safe fixes, then validation output |
skill-architect: refactor | Improving an existing skill without changing its purpose | Portable rubric, Eval methodology | Trigger rewrite, reference split, routing cleanup, verification updates |
skill-architect: compare | Comparing upstream skill-development, authoring, or creator rubrics | Comparative study | Evidence matrix, reusable rules, adoption/rejection decisions |
skill-architect: distill | Reviewing one session, a session ID/path, or a group of sessions to extract a reusable skill | Eval methodology, Portable rubric | Session evidence map, reusable workflow, skill proposal or draft, eval plan |
skill-architect: reconcile | Refreshing stale skill plans, audits, session-distilled proposals, or blocked skill backlogs | Eval methodology, Portable rubric | Updated status, drift decisions, retired/reopened items, next executable plan |
skill-architect: teach | Explaining skill authoring standards or coaching a user/agent | Portable rubric, Vendor adapters | Concise lesson, examples, and next-step exercise or checklist |
Mode rules
- Plan before non-trivial multi-file changes, new gates, release work, or cross-vendor adapter changes.
- Create should start from concrete use cases, then draft the portable core before vendor metadata.
- Audit is read-only by default; present findings and wait unless the user also asks to patch.
- Fix is deterministic and dry-run-first; only apply
fix_skill.py --writefor safe mechanical edits, then runvalidate_skill.py. - Refactor preserves intent; do not rename, split, or change behavior beyond the requested scope without calling it out.
- Compare must cite inspected sources and separate source facts from adoption decisions.
- Distill must separate session-specific events from reusable patterns, cite the session source, and propose a skill only from repeated or high-value behavior.
- Reconcile must verify whether prior plans, findings, and proposals still match the current context; retire fixed or rejected items instead of re-reporting them.
- Teach should explain the smallest useful rule, then give one concrete example rather than dumping the whole rubric.
Operating mode
Start with the current working context, not an upstream rubric. Check the CWD, project files, installed skill location, or repository convention first. Use the Agent Skills specification for portable baseline syntax; external rubrics inform content quality, while local rules win on accepted frontmatter fields, manifests, lockfiles, CI, and release gates. Do not add experimental allowed-tools unless the target runtime and repository schema both accept and enforce it.
Use this order:
- Classify the request — choose
plan,create,audit,fix,refactor,compare,distill,reconcile, orteachbefore deciding files or checks. - Read the local standard — find the CWD/project/repository authoring standard, agent instructions, check commands, and target skill before editing.
- Select the lens — structure, quality/evals, vendor compatibility, distribution governance, or predictability.
- Run the predictability pass — check invocation fit, branch uniqueness, information hierarchy, completion criteria, leading words, duplication, sediment, sprawl, no-op lines, and premature-completion risk before proposing edits.
- Plan the artifact — decide the portable core, references, optional assets/scripts, vendor metadata, and validation path.
- Patch narrowly — avoid churn-only rewrites; improve only the requested surface and directly related standard violations.
- Verify mechanically — run repo checks or the closest local equivalent before handing back.
Skill architecture workflow
1. Gather concrete use cases
Capture only the details that determine structure:
- exact request phrases that should trigger the skill
- file paths, tools, APIs, or error messages that imply the skill
- repeated task shape: quick reference, workflow, router, review rubric, or deterministic helper
- target runtime capabilities: check required tools, permissions, and instruction loading; use vendor-adapters.md when host-specific execution or packaging details matter
- whether the skill needs references, assets, examples, or helper scripts
Ask one focused question if a structural choice is ambiguous. Otherwise infer from the CWD/project/repository pattern and state the assumption.
2. Separate portable core from adapters
The portable core lives in SKILL.md and linked references. Vendor details live in adapter sections or metadata files.
Use this split:
| Layer | Contains | Avoid |
|---|---|---|
| Portable core | trigger intent, workflow, checks, references, scripts/assets | one vendor's CLI assumptions as universal truth |
| Local convention | required frontmatter, versions, hashes, security files, README/plugin registry updates, or platform conventions | upstream defaults that fail local checks |
| Vendor adapter | OpenAI YAML, Claude plugin notes, Antigravity/Azure conventions, installation constraints | changing the portable workflow to fit one agent |
| Quality harness | eval prompts, baseline comparisons, review packets, pressure tests | claiming quality from prose review alone |
3. Draft with progressive disclosure
Keep SKILL.md focused on routing and immediate behavior. Move long comparisons, advanced examples, vendor matrices, and eval details to references/.
Minimum portable skill shape:
skills/<skill-name>/
├── SKILL.md
├── LICENSE
└── references/ # only when detail would bloat SKILL.md
Add assets/ or helper scripts only when they materially improve repeatability. Treat scripts as inert skill assets unless the local project or runtime explicitly allows executable helpers; document, link, hash, and security-scan them according to local conventions when available.
4. Review before patching
Prepare a compact audit packet before making non-trivial edits:
| Check | Evidence | Decision |
|---|---|---|
| Trigger description is concrete | SKILL.md frontmatter line | keep / patch |
| Invocation fit is justified | description, expected caller, context/cognitive load tradeoff | keep model-invoked / make user-invoked / add router |
| Branches are unique | description, command table, mode rules | keep / collapse synonym-only trigger branches |
| Information hierarchy is clean | section inventory | keep in SKILL.md / move to reference / split sequence |
| Completion criteria are checkable | every ordered step or command workflow | keep / sharpen / add done criteria |
| Leading words carry repeated concepts | repeated behavior terms and trigger language | keep prose / collapse to leading word |
| Pruning pass finds no sediment | duplicated, stale, no-op, or sprawling lines | keep / delete / disclose |
SKILL.md is lean enough | word count | keep / split |
| References are linked | link check | keep / patch |
| Vendor assumptions are isolated | relevant section/file | keep / patch |
| Mechanical gates are available | repository scripts, CI, evals, or fixtures | run / note skipped |
Run a predictability and pruning pass before judging style or patching content:
- Invocation fit: decide whether the skill should be model-invoked or user-invoked. Pay context load only when autonomous agent reach or cross-skill reach is necessary; recommend a router skill when many user-invoked skills create cognitive load.
- Description branches: keep one trigger per real branch, remove synonym-only trigger duplication, and front-load strong leading words that also appear in prompts, docs, or code.
- Information hierarchy: classify each major block as step, in-skill reference, or disclosed reference. Inline what every branch needs; move branch-only detail behind a context pointer.
- Completion criteria: every ordered step should end with a checkable done condition. Sharpen vague criteria before splitting a sequence.
- Premature-completion risk: if visible post-completion steps pull attention forward and the current criterion cannot be sharpened further, recommend a real context-boundary split.
- Leading words: replace repeated explanatory prose with a compact pretrained concept only when it improves predictable execution or invocation.
- Pruning: apply sentence-level relevance and no-op tests; remove duplication, sediment, and sprawl instead of polishing dead prose.
For behavior-shaping skills, add at least one forward test or prompt scenario. For any shipped CI, shell, or code — helper scripts, snippets, or workflow examples — review it for functional correctness and execute it against a fixture rather than vouching for it by reading. Reading alone routinely misses logic defects (wrong regex, broken cascade handling, crash-after-success) that a fixture run surfaces immediately; security-scanning and shape-checking do not cover correctness. An audit is incomplete until each embedded-code finding or pass includes a fixture command plus observed output, or an explicit skipped-risk note.
Use a bounded executable-surface triage before invoking a full code-review lens. Check whether the artifact has: fixture coverage for shipped scripts/snippets, fail-closed error handling, shell/runtime portability, pinned or cached external tools, path and quoting safety, network/credential boundaries, parser edge-case tests, and CI visibility. Report these as skill-architecture risks with evidence and a verification path; route algorithmic correctness, exploitability analysis, performance tuning, or large refactors to a dedicated code-review skill instead of expanding this skill's scope.
For eval and adapter claims, calibrate the finding to local policy and evidence quality: prompt-only eval catalogs are useful retrieval vectors but incomplete evidence for high-risk behavior claims, and missing vendor adapters are packaging defects only when the local repo requires them. Weak descriptions are deterministic failures only when the rule is portable and low-false-positive; in skills4sh, generic trigger-only wording such as Use when creating skills. fails mechanically.
Authors/reviewers evaluating Skill Architect itself only: skip this catalog when using Skill Architect for another task. The mode scenario catalog supplies self-contained fixture files for all nine modes, isolation instructions, grading criteria, and description-only routing prompts. Materialize a fresh case directory and withhold grading from the executor; the catalog is test input, not evidence that those modes passed.
For test vectors, harness setup, enumeration consistency, and severity calibration, use Eval methodology and Portable rubric rather than expanding SKILL.md.
5. Make handoffs executable
When producing a plan, proposal, audit packet, or distilled skill spec for another agent or a future session, write for the weakest plausible executor. Reference accessible authoritative artifacts with exact paths or URLs; include essential excerpts only when the executor needs them. Preserve local conventions, scope boundaries, verification commands with expected results, drift checks, and STOP conditions. For dependent tasks, specify consumed/produced interfaces and requirement coverage using the handoff plan rubric.
Record rejected findings or non-adopted patterns with one-line rationales so they do not return in the next audit or reconciliation pass.
6. Verify and hand off
Run the narrowest meaningful checks for the current working context. Prefer documented validation commands, CI gates, evals, fixture tests, hash checks, and security scans when present.
For skills4sh examples, the usual checks are:
npm run check:skill-standard
npm run check:drift
node bin/hash-check.mjs
npm test
Run a focused or full security scan when new files contain scripts, shell examples, credentials guidance, network examples, or other security-sensitive patterns.
Report exactly what passed, failed, or was skipped. Do not claim an eval, scan, or fixture test that was not run.
Helper scripts
Optional helper scripts live under assets/scripts/ so they ship as inert skill assets:
| Script | Use |
|---|---|
assets/scripts/inspect_skill.py | Summarize a skill folder: files, frontmatter, body word count, linked references, and obvious missing pieces |
assets/scripts/validate_skill.py | Run a local portable-skill validation pass against one skill folder, with warning-only progressive-disclosure hints |
assets/scripts/fix_skill.py | Dry-run or apply conservative deterministic fixes to SKILL.md, then rerun validation |
assets/scripts/scaffold_skill.py | Create a starter skill folder using the portable rubric |
assets/scripts/run_uv.py | Start uv after removing inherited UV_* host configuration |
inspect_skill.py and validate_skill.py each accept one skill folder per invocation. For multiple skills, invoke each helper separately for each folder; reuse the same Python environment.
Validation and inspection require Python >=3.10 plus the version-and-hash-pinned Python dependency file. Use an isolated runner; never install validation dependencies into the global Python environment.
Prefer uv when it is available. It is the preferred adapter, not a portable requirement. Invoke it through run_uv.py, which removes the entire inherited UV_* namespace before uv starts; --no-config only ignores persistent files and does not neutralize variables such as UV_INDEX or UV_NO_VERIFY_HASHES. --isolated alone still installs a surrounding project: add --no-project so a nearby pyproject.toml is not built or imported, --no-build so only wheels are used (the same pin as pip install --only-binary=:all:), and --no-config so host uv.toml / uv.ini is ignored. Run helpers with python -E so PYTHONPATH cannot replace pinned PyYAML. Do not use python -I; it drops the script directory from sys.path and breaks inspect_skill.py's sibling import of validate_skill.py.
python3 -E "/path/to/skill-architect/assets/scripts/run_uv.py" run --isolated --no-project --no-build --no-config --with-requirements "/path/to/skill-architect/assets/scripts/requirements.txt" python -E \
"/path/to/skill-architect/assets/scripts/inspect_skill.py" "/path/to/target-skill"
python3 -E "/path/to/skill-architect/assets/scripts/run_uv.py" run --isolated --no-project --no-build --no-config --with-requirements "/path/to/skill-architect/assets/scripts/requirements.txt" python -E \
"/path/to/skill-architect/assets/scripts/validate_skill.py" "/path/to/target-skill"
On Windows, use py -3 -E instead of python3 -E for the launcher command.
If uv is unavailable or cannot prepare its isolated environment, fall back to a temporary virtual environment. On macOS or Linux:
VALIDATION_VENV="$(mktemp -d)/venv"
python3 -E -m venv "$VALIDATION_VENV"
"$VALIDATION_VENV/bin/python" -E -m pip --isolated install --require-hashes --only-binary=:all: \
-r "/path/to/skill-architect/assets/scripts/requirements.txt"
"$VALIDATION_VENV/bin/python" -E "/path/to/skill-architect/assets/scripts/inspect_skill.py" "/path/to/target-skill"
"$VALIDATION_VENV/bin/python" -E "/path/to/skill-architect/assets/scripts/validate_skill.py" "/path/to/target-skill"
On Windows PowerShell:
$ValidationVenv = Join-Path ([System.IO.Path]::GetTempPath()) ([guid]::NewGuid().ToString())
py -3 -E -m venv $ValidationVenv
& "$ValidationVenv\Scripts\python.exe" -E -m pip --isolated install --require-hashes --only-binary=:all: `
-r "/path/to/skill-architect/assets/scripts/requirements.txt"
& "$ValidationVenv\Scripts\python.exe" -E "/path/to/skill-architect/assets/scripts/inspect_skill.py" "/path/to/target-skill"
& "$ValidationVenv\Scripts\python.exe" -E "/path/to/skill-architect/assets/scripts/validate_skill.py" "/path/to/target-skill"
A helper that starts and reports findings produced an available result, even when validation exits non-zero. Retry with the temporary-environment fallback only when uv, dependency resolution, or helper startup fails. Report validation or inspection as unavailable only after both isolated runner paths are unavailable or fail; name the attempted commands and setup errors. Do not silently substitute shape checks for YAML validation.
Scaffold and fix helpers need only Python. The fixer accepts column-zero literal mapping keys and simple single-line literal names; it refuses unsupported key or name syntax before any edits. Use manual review for those cases, then validate.
Read or run scripts only when the task needs deterministic inspection or scaffolding. Local validation commands and CI, when present, remain the source of truth.
Deterministic checks vs judgment
Skill work splits across four layers — keep them separate and route each check to the layer that can decide it:
| Layer | Owns | Examples |
|---|---|---|
Portable validator (validate_skill.py) | Minimal deterministic checks that hold inside one skill folder, no repo metadata or network | valid YAML and typed, non-empty fields; specification length limits + name/directory match + kebab-case, description has concrete trigger detail, body within the hard size cap, references/*.md linked from SKILL.md, in-skill relative Markdown link targets, and same/cross-file Markdown heading anchors |
Portable fixer (fix_skill.py) | Safe mechanical edits that require no taste or domain judgment; dry-run unless --write is explicit | frontmatter name normalization and insertion of missing references/*.md links in SKILL.md |
| Local binding gate (project CI) | Deterministic checks that depend on repo conventions; binding and a superset | file hashes, skills-lock.json, security manifests, doc-sync, semver — in skills4sh: check:drift, check:guardskills, hash-check, npm test |
| This skill's rubric (judgment) | What no script can decide | trigger quality, executable-surface triage, embedded-code correctness (review + fixture run), body altitude beyond the cap, vendor-isolation, narrative bloat |
The litmus test: if a script can decide it, automate it; if it needs reasoning or taste, it stays in the rubric. When a local binding gate exists, it wins over the portable validator on any overlap. The portable validator is deliberately narrow: it checks only facts visible inside one skill folder. Codifying a rule (for example a numeric size band) moves it from judgment toward the validator; treat that as the goal, not a loss.
Reference files
| File | Load when |
|---|---|
| references/comparative-study.md | Comparing Anthropic, OpenAI, Matt Pocock, Obra, Azure, Antigravity, or other skill-development/authoring/creator patterns |
| references/house-rubric.md | Creating or reviewing a portable rubric; checking predictability, invocation fit, completion criteria, pruning, enumeration consistency, executable-surface triage, and severity calibration |
| references/vendor-adapters.md | Runtime install roots, pointer files, and vendor metadata that must stay out of the portable workflow |
| references/eval-methodology.md | Testing triggers, building test-vector catalogs, setting up lightweight harnesses, comparing baseline behavior, or planning pressure tests |
| references/naming-and-packaging.md | Avoiding slug collisions, deciding names, updating lockfiles/security manifests, and packaging repo updates |
Non-negotiables
- Do not make a vendor-specific skill the portable rubric verbatim.
- Do not reinstall upstream
skill-creatorover an existing skill without checking provenance and collision risk. - Do not ship loose summary-only descriptions; the description must include a concise capability clause and concrete trigger/use conditions.
- Do not add unlinked reference files.
- Do not add helper scripts without a concrete repeatability benefit and a validation path.
- Do not treat
skills.shpopularity as quality proof; inspect source, tests, security posture, and portability.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/t4sh/skills4sh/skill-architect">View skill-architect on skillZs</a>