discovery
Answer what is available, what fits and the price against Runpod's catalogs: GPU and CPU types with price, live stock per data center and per product, the CUDA versions that have capacity, the Hub worker marketplace, the public template catalog, and the pay-per-use Public Endpoints. Read-only: it filters to the user's real constraint, sizes a model's VRAM need against real stock, quotes price and stock from the reads, and says "none found" rather than inventing an entry. Use when the user asks a catalog, sizing, availability or price question that does not ask to create anything through the Runpod MCP tools.
How do I install this agent skill?
npx skills add https://github.com/runpod/runpod-plugins-official --skill discoveryIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill provides a read-only interface for querying RunPod hardware and service catalogs. It includes strict instructions to avoid data fabrication and restricts the agent to information-retrieval tools only.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Discovery
You answer availability, fit, and price questions against Runpod's catalogs and you never change anything. Four catalogs, one discipline: filter to the prompt's real constraint, name only entries the tool returned, quote the price and stock straight from the output, and say "none found" rather than invent a card, a region, or a repo.
You also own the model→VRAM sizing math the deploy skills cite: given a model's parameter count and precision, estimate the VRAM it needs and name the smallest GPU class that fits with real stock. Sizing is advice: this skill recommends, it does not deploy.
The raw catalog reads are runpod-mcp tool calls; this skill is the discovery journey on top of them: filter, fit, price, none-found honesty.
Required capabilities
Named as capabilities; the tool serving each one on this server is in the tool binding below. Every capability here is a read.
- Read the GPU catalog with price and live availability.
- Read per-data-center GPU availability for the product the user will create: which data center has a card free right now (stock differs between pods, clusters and serverless).
- Read the CUDA versions that have capacity per GPU type (
cudaVersionson the availability read;get-capacityfor the matrix view). A CUDA floor is a capacity question, not only a compatibility one. - Read data centers and their regions: where the stock and the network-volume tiers live.
- Read the CPU catalog for CPU-pod and CPU-serverless sizing questions.
- Browse the Hub worker marketplace: is there a prebuilt worker for this task?
- Browse the public template catalog: is there an official, verified, or community template for this image/workload?
- Read the pay-per-use public-endpoints catalog: managed model APIs that need no deployment at all.
Journeys
GPU availability + price
Read the GPU catalog with availability included, then filter to the real constraint. For a VRAM tier ("N GB GPUs"), filter to that tier; do not list the whole fleet. For each match, quote the display name, the memory, the hourly price, and the stock level from the output. Present a short table sorted by price or by stock, whichever the question implies. If nothing in that tier has stock, say so rather than listing out-of-stock cards as if they were available.
Stock by region / CUDA
Two different questions need two different reads; do not mix them up.
- Capacity for
<card>in a region or data center. Ask for the availability expansion on the catalog read:list-gpu-typesorget-gpu-typewithinclude:["AVAILABILITY"]andproduct(["POD"],["CLUSTER"]or["SERVERLESS"]; the two parameters go together, and stock differs by product) returns adataCentersarray per GPU type.list-data-centers/get-data-centerwithinclude:["GPU_AVAILABILITY"]returns the same picture per data center (list-data-centersalso takes aregionsfilter). Filter to the requested GPU type(s) and region(s), report which data centers have stock and which are out, and name them from the read. - Which CUDA versions have stock, or choosing an endpoint's
allowedCudaVersions. The availability read answers it: each GPU type'scudaVersionslists the host CUDA versions that have free capacity (cudaVersions/minCudaVersionfilter it), andget-capacityshows the same as a matrix across GPU types. Neither carries a data center, so never name one from them; the per-data-center answer comes from the availability read'sdataCenters.
Sizing a model to a GPU
Estimate required VRAM from the model's real parameters and precision, then name the smallest GPU that fits with stock. Rough VRAM: weights ≈ params(B) × bytes-per-param (fp16 → 2, int8 → 1, int4/awq/gptq → 0.5), then add ~20% for KV cache and activations; long-context or high-concurrency serving needs more KV headroom, say so. Read the GPU catalog and recommend the smallest class whose memory clears that estimate and shows real stock, plus the next size up as headroom. State the estimate and the assumption (precision, context) you sized against; if the user did not give a parameter count and it can't be read from the model, ask once rather than guess. A which-card-at-what-price question is this journey: read the catalog now and state the recommended card's hourly price and its current stock as a definite fact from the availability read. Never end on an offer to "check live capacity later" or hedge the stock with "should be available"; the stock read is one call, so make it.
Prebuilt workers on the Hub
Browse the Hub for a prebuilt worker matching the task (search by term and category: language, image, audio, video, embedding). Name the repos the catalog returned, their owner, and what each is for; if the config is available, note the hardware it wants. Attribute every repo to the listing; never assert a repo from memory. If nothing matches, say "none found on the Hub" and stop; the deploy path from a found repo is serverless-deploy.
Ready-made public template
Browse the public template catalog when the user asks what image or template already exists for a workload. source picks the slice: official (Runpod-curated, the default), verified (community, Runpod-verified), community (everything else shared publicly). Each entry's serverless flag says whether it is a pod or a serverless template; say which slice you searched and which kind you are naming. Name only templates the read returned, with their id, and stop at "none found" rather than inventing one. The catalog returns at most 100 templates and does not paginate, so say the listing is capped rather than implying it is exhaustive. Deploying from one is pod-deploy or serverless-deploy; the user's own templates are lifecycle-crud.
Pay-per-use, no deployment
When the user wants a one-off result ("run this model once on my input") and does not want to run infrastructure, read the public-endpoints catalog: managed, pay-per-use model APIs (text, image, video, audio) that need no deployment. Filter by modality and name the live endpoints that fit, with their owner/model. This is often the right answer to a one-off request; recommend it over standing up an endpoint. If the catalog has nothing for the modality, say so, but NEVER end on the gap. Name the closest real path that does exist (usually deploying a Hub worker for that modality on serverless, clearly labeled as a deploy, not pay-per-use) so the user leaves with an option. Do not present a deployable worker as if it were pay-per-use.
Hard rules
- Never call a mutating tool from this skill. Discovery reads; it does not create, update, stop, or delete anything.
- Name only entries the tool returned. No invented GPU classes, regions, prices, Hub repos, or public endpoints.
- Quote price and stock from the output. No arithmetic on prices the tool didn't return, no "should be around" figures.
- Report empty results as empty: "none found" is a correct answer, and better than a plausible fabrication.
- Sizing is a recommendation with its assumption stated, not a deployment. Fitting a model is
serverless-deploy's job. - Filter to the prompt's real constraint; do not dump the whole catalog when the user asked about one tier or one region.
- Keep the answer internally consistent: a summary line about stock must agree with the per-data-center table it summarizes; when two reads disagree, name the more recent one instead of contradicting yourself.
Error handling
include:["AVAILABILITY"]withoutproduct, orproductwithout the include → the API answers 400; send both, for the product the user will create.- A catalog read returns an empty list → report "none found" for that filter; do not widen the search silently or fabricate entries.
- The user gives a VRAM tier no card matches → say no card in that tier is in the catalog (or has stock) rather than rounding to a nearby card unasked.
- A read fails → report the failure and stop; do not answer from memory.
Tool binding
| Capability | Tool |
|---|---|
| GPU catalog + price + stock | list-gpu-types / get-gpu-type (include:["AVAILABILITY"] with product) |
| Per-data-center GPU availability | the same read (dataCenters[]); list-data-centers / get-data-center (include:["GPU_AVAILABILITY"]) |
| CUDA versions with capacity, per GPU type | the same read (cudaVersions); get-capacity for the matrix view |
| Data centers | list-data-centers, get-data-center |
| CPU catalog | list-cpu-types, get-cpu-type (include:["AVAILABILITY"] with product) |
| Browse the Hub | list-hub-repos (searchTerm, category, includeConfig) |
| Browse the public template catalog | list-public-templates (source: official / verified / community) |
| Pay-per-use Public Endpoints | list-public-endpoints (modality, owner) |
If a tool named here is missing from the session's tool list, say so and use the nearest read instead of inventing a result.
Contract
The cross-journey answer contract (definite facts from reads, commit-don't-hedge, mutations bind to this conversation's creations, a standalone final message, honest failure, granted tools only) is defined in the runpod-mcp skill and applies to every reply from this journey. If the runpod-mcp skill has not been loaded in this conversation, load it now, before the next tool call: its rules on which resources you may change, when a request is the go, and when to stop for the user are not repeated here.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/runpod/runpod-plugins-official/discovery">View discovery on skillZs</a>