skill-retrieval
BM25-based skill retrieval plugin for Hermes Agent. Replaces the full skill list in the system prompt with a names-only compact view (~2K tokens) and injects top-K relevant skill descriptions per turn via BM25 retrieval (~300 tokens). Saves ~9K tokens/turn. Use when system prompt token overhead from skills is a concern, or when skill discovery quality matters.
How do I install this agent skill?
npx skills add https://github.com/moonlight-lupin/agent-skills --skill skill-retrievalIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a token-optimization plugin for the Hermes Agent that implements BM25 retrieval to dynamically disclose skill descriptions. It compacts the system prompt and includes an optional, opt-in semantic reranking feature via an external API. The implementation follows security best practices, including environment-based secret handling and query neutralization to prevent injection.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Skill Retrieval
This is a Hermes Agent plugin. It is not a Claude Code plugin and will
not load in Claude Code — that runtime has no pre_llm_call event, no Python
register() entry point, and reads .claude-plugin/plugin.json rather than
plugin.yaml. Requires Hermes Agent >=0.21.1 (requires_hermes in
plugin.yaml).
BM25-based progressive disclosure for Hermes Agent skills. Instead of dumping every skill description into the system prompt (~11.5K tokens), this plugin keeps a compact names-only index and injects only the top-K relevant descriptions per turn.
What it does
Two-phase progressive disclosure:
-
Phase 1 — System prompt compaction (an
llm_requestmiddleware): Rewrites the outgoing request's<available_skills>block so it lists skill names only (descriptions stripped). All skills remain discoverable by name (~2K tokens instead of ~11.5K). The system prompt Hermes builds is never modified — prompt caching is unaffected. -
Phase 2 — Per-turn BM25 retrieval (
pre_llm_callhook): Tokenizes the user message, ranks active skill descriptions with BM25 Okapi, and injects the top-K matches (~300 tokens) as context appended after the user message.
Architecture
Each turn (llm_request middleware)
│
▼
Phase 1: rewrite <available_skills> → names only (~2K tokens, request only)
Each turn (pre_llm_call)
│
▼
Phase 2: BM25Index.retrieve(user_message, top_k)
└── inject "## Retrieved Skills ..." into user message (~300 tokens)
The BM25 index is built lazily on the first turn from standalone skills
(~/.hermes/skills) and plugin-bundled skills (~/.hermes/plugins/*/skills),
then cached per Hermes home, tool capability snapshot, session platform and
resolved disabled-skill set (the same inputs Hermes keys its own skills prompt
cache on). The cache is rebuilt without a restart when Hermes clears its skills
prompt cache (skill_manage create/patch/delete, hub install, skill toggles)
or when a cheap on-disk manifest changes (skill root and category dir mtimes,
top-level SKILL.md files, config.yaml).
Retrieval uses a pure-stdlib inverted index (term → posting list of
precomputed BM25 weights) and is sub-millisecond for ~200 skills.
Token savings
| Stage | Tokens (approx.) |
|---|---|
| Before (full skill list in system prompt) | ~11.5K |
| After — names-only system prompt | ~2.0K |
| After — per-turn top-K descriptions | ~0.3K |
| Net per turn | ~2.3K (~9K saved) |
Measured on a Hermes install with ~300 skills; savings scale with skill count.
Installation
Copy or symlink this directory into the Hermes plugins folder:
# From this repo
ln -s "$(pwd)/plugins/skill-retrieval" ~/.hermes/plugins/skill-retrieval
# Or copy
cp -r plugins/skill-retrieval ~/.hermes/plugins/skill-retrieval
User plugins are opt-in: Hermes discovers the directory but does not load it
until it is enabled (this adds skill-retrieval to plugins.enabled in
~/.hermes/config.yaml):
hermes plugins enable skill-retrieval
Restart the agent session so register() runs — it registers the
llm_request middleware and the pre_llm_call hook.
Dependencies (install into the Hermes Python env if missing):
pip install pyyaml
Configuration
| Setting | Default | How to set |
|---|---|---|
TOP_K | 6 | Env var SKILL_RETRIEVAL_TOP_K |
| System prompt compaction | enabled | Set SKILL_RETRIEVAL_COMPACT=0 to disable compaction while keeping BM25 retrieval injection |
BM25 k1 | 1.5 | Constant in scripts/bm25_retriever.py |
BM25 b | 0.75 | Constant in scripts/bm25_retriever.py |
| Jev rerank | off | Set SKILL_RETRIEVAL_RERANK=jev to enable; key in TYPESAFE_API_KEY (legacy TYPESAFE_KEY accepted), read from os.environ |
| Rerank log path | $HERMES_HOME/data/jev-trial/rerank_ab_log.jsonl (rotates past 5 MB) | Env var SKILL_RETRIEVAL_RERANK_LOG |
export SKILL_RETRIEVAL_TOP_K=8
export SKILL_RETRIEVAL_COMPACT=0
export SKILL_RETRIEVAL_RERANK=jev
Jev rerank (optional, off by default)
When SKILL_RETRIEVAL_RERANK=jev, the BM25 top-N shortlist is reranked by the
Jev system-one judgment model (POST https://api.typesafe.ai/v1/systemone,
model jev-latest). Behavior:
- Fail-soft — 3.0 s timeout, no retries; on any error the request logs a warning and BM25's original order stands. Retrieval never blocks on Jev.
- Env-only config — all settings come from
os.environ(never a file read):TYPESAFE_API_KEY(legacyTYPESAFE_KEYaccepted). Hermes loads the active profile's.envinto environ at startup, so no cross-profile file is ever read. - A/B log — every query records the timestamp, session id, a 200-char
query excerpt, both orders, probabilities, tokens and latency to the log
path above; the log rotates past 5 MB;
scripts/rerank_ab_report.pysummarizes it locally. - Outbound cap — at most
RERANK_CANDIDATES(default 12) descriptions are sent per call, even when the caller's shortlist is longer; the remainder keeps BM25 order at the tail and never leaves the machine.
Verify it's working
Phase 1 silently no-ops outside a full Hermes runtime, and the BM25 index can silently empty. After restart, check the Hermes logs.
Healthy start — look for these log lines:
BM25 index built: N docs …Skill retrieval plugin registered (top_k=…, compact=true)
Degraded — these warnings mean it's not working:
Cannot locate prompt_builder — compaction skipped(Phase 1 failed; Phase 2 still runs for anonymous sessions, but named sessions skip injection because no capability snapshot is recorded)No active skills found for BM25 index(index is empty — zero retrieval injection)
How it works
- Tokenizer — lowercases text, strips punctuation, splits on whitespace.
- Corpus — each skill becomes
"name: description"plus its frontmattertriggers:list andmetadata.hermes.tags(when present) from SKILL.md YAML. Trigger phrases and tags participate in retrieval scoring but only the description is injected into the prompt. Disabled skills from~/.hermes/config.yamlare skipped. - Index — BM25 Okapi TF saturation + Lucene IDF
log(1 + (N-df+0.5)/(df+0.5))(always positive, so small corpora and common terms still score), stored as an inverted index:dict[str, list[tuple[int, float]]]mapping each term to a posting list of (doc_index, precomputed BM25 weight). - Retrieve — for each unique query token present in the index, walk its posting list and accumulate scores; sort by descending score (score > 0 only).
Performance
- Index built on first use and cached (~8 ms for 200 skills on a CPU-only
VM). Each later turn only stats the skill roots, their immediate
subdirectories and
config.yamlto validate the cache. A zero-skill install caches the empty result too (one warning, no rescans). - Retrieval is sub-millisecond (~0.03 ms mean for 200 skills). The inverted index touches only documents that share a query term — no full-corpus scan.
- No compiled dependencies. The plugin uses only the Python standard library
(plus pyyaml for config/frontmatter parsing). This removes a 154 MB
numpy/scipy install and a ~573 ms import cost, which matters for subprocess
spawning paths (e.g. a Claude Code
UserPromptSubmitvariant). - Failures in the hook return
None(no injection) so the agent keeps working.
Dependencies
pyyaml
Limitations
- BM25 is lexical, not semantic. Paraphrased queries that share few tokens with a skill's description may rank poorly even when the intent matches.
- Descriptions longer than 200 characters are truncated in the injected block;
use
skill_view(name)for the full skill body. - Compaction uses Hermes's read-only
agent.prompt_builderhelpers; if they cannot be imported, Phase 1 is skipped and retrieval falls back to the standalone loader. - A named session is only injected once its system prompt has been built in this process (that build records the session's tool capabilities). A session restored after a restart without a rebuild gets no injection rather than risking skills Hermes hides from it.
- A content-only edit of a nested
SKILL.md(category/skill/SKILL.md) made outside Hermes (e.g. in an editor) changes no directory mtime, so it is picked up only after Hermes clears its skills prompt cache or the agent restarts. Edits throughskill_manage, and any added/removed skill or config change, are picked up on the next turn. - Phase 1 depends on Hermes internals (
agent.prompt_builder) and can break on a Hermes upgrade. - BM25 top-1 precision is soft: the best-matching skill is often not rank 1,
though it usually lands within the first few results. Ranking depends entirely
on your own corpus and how its descriptions are worded, so
TOP_Kbelow ~5 is not recommended. - Compaction rewrites the request's
<available_skills>block using read-only imports fromagent.prompt_builderandagent.skill_utils(all ten disclosed in the README); if those change signature, any mid-scan exception falls back to the standalone loader, so the corpus is never silently empty. - The stdlib index computes in float64 (the previous scipy version used float32). Equal-scoring skills may order differently than before. This is harmless — the scores are genuine ties (~1e-6 difference) — but it is a real behaviour delta from the scipy version.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/moonlight-lupin/agent-skills/skill-retrieval">View skill-retrieval on skillZs</a>