skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
roboflow/computer-vision-skills271 installs

roboflow-training-and-evaluation

Use when training Roboflow models, diagnosing why a model underperforms, improving accuracy, or setting up a production feedback loop — covers architecture selection, model IDs, checkpoints, Model Evaluation (confusion matrix, per-class metrics, MCP eval tools), the diagnosis-first improvement playbook, and Active Learning through the Project Model Workflow block.

How do I install this agent skill?

npx skills add https://github.com/roboflow/computer-vision-skills --skill roboflow-training-and-evaluation
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides instructions for managing the training and evaluation of computer vision models on the Roboflow platform. It guides agents through architecture selection, Neural Architecture Search (NAS), and diagnostic workflows for model improvement. The skill identifies an attack surface for indirect prompt injection via the processing of untrusted project metadata and evaluation results.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

For agents — source-of-truth: This skill is authored in roboflow/computer-vision-skills and shipped with the Roboflow plugin. If your client has loaded the plugin (you'll see roboflow:<name> skills in your available skills list), use those local skills — they're read fresh from disk every session. The same content served as MCP resources at roboflow://skills/<name>/... is a fallback for clients without the plugin and may lag this repo. Don't call ReadMcpResourceTool for roboflow://skills/... URIs when a local roboflow:<name> skill is available.

Training & Evaluation on Roboflow

Training Flow

Upload/Annotate Images
  → Generate Dataset Version (preprocessing + augmentation + train/val/test split)
    → Pick Model Architecture + Size
      → Pick Checkpoint (COCO, Universe model, or previous version)
        → Train
          → Evaluate (auto-runs for paid users)

Version = frozen snapshot. Changes to the project after version creation do not affect it. Configure preprocessing (resize, contrast, etc.) and augmentation (flip, rotate, mosaic, etc.) during version generation.

Available Model Architectures

Object Detection

ArchitectureSizesDefault ResolutionNotes
RF-DETRPico, Nano, Small, Base, Medium, Large, XL, 2XL384-880 (varies by size)Best accuracy among named sizes; see RF-DETR NAS
Roboflow 3.0Fast, Accurate, Medium, Large, XL640x640YOLOv8-based. Medium+ require paid plan
YOLO26n/s/m/l/x640x640Also supports seg + pose
YOLOv12n/s/m/l/x640x640OD only
YOLOv11n/s/m/l/x640x640Also supports seg + pose
YOLOv8n/s/m/l/x640x640Also supports seg + pose
YOLO-NASSmall, Medium640x640
YOLOLite CPUn/s/m/l/x640x640Edge-optimized, beta
YOLOLite GPUn/s/m/l/x640x640Edge-optimized, beta
Roboflow InstantsingleN/A (no resize)Few-shot, free, OD only

Instance Segmentation

ArchitectureSizesDefault Resolution
RF-DETR SegNano, Small, Medium, Large, XL, 2XL312-768 (varies)
Roboflow 3.0 SegFast, Accurate, Medium, Large, XL640x640
YOLO-segv8/v11/v26 (n/s/m/l/x each)640x640
SAM 3 (Segment Anything 3)Large1008x1008

Semantic Segmentation

ArchitectureSizesDefault Resolution
DeepLabV3+Base>=512x512

Classification

ArchitectureSizesDefault Resolution
ViTBase224x224
ResNet18/34/50/101224x224
DINOv3Base, Small224x224

Keypoint / Pose

ArchitectureSizesDefault Resolution
YOLO-posev8/v11/v26 (n/s/m/l/x each)640x640

Multimodal / VLM

ArchitectureSizesDefault Resolution
Qwen3.50.8B, 2B448x448
Qwen3 VL2B448x448
SmolVLM256M, 2B384x384
Florence 2Base, Large768x768
PaliGemma 23B448x448
Qwen2.5 VL7B448x448

Model Selection Decision Tree

Follow this flowchart to pick the right model. Start at Step 1.

  1. Task type? OD / Instance Seg / Keypoint → Step 2. Classification / Semantic Seg / VLM → use specialized block directly.
  2. Target classes in COCO 80? Yes → Step 3. No → Step 6.
  3. Real-time? No (images/recorded video) → Step 4. Yes (live video) → Step 5.
  4. Non-real-time, COCO — OD / Inst Seg → prefer RF-DETR NAS; if unavailable, RF-DETR (detection) or RF-DETR Seg (instance segmentation) at Medium, Small for constrained HW, XL for accuracy-first. Keypoint → YOLO26 pose. Done.
  5. Real-time, COCO — Same, and NAS is the strongest option here because it reports measured latency per target hardware. Without it, same families at Nano–Small. Done.
  6. Non-COCO, which sub-task? OD → Step 7. Inst Seg → Step 8. Keypoint → Step 9.
  7. OD, non-COCO — Check Rapid exclusions (see below). If excluded → Step 13. Otherwise → recommend Roboflow Rapid (default) or SAM3 zero-shot as secondary option → Step 10.
  8. Inst Seg, non-COCO — SAM3 zero-shot (sam3/sam3_final, set class_names). Rapid does not support segmentation → Step 10.
  9. Keypoint, non-COCO — Not real-time → YOLO26 pose Medium–XL. Real-time → YOLO26 pose Nano–Small. Done.
  10. Real-time? No → try model (Step 11). Yes → warn about latency, try model (Step 12).
  11. Non-real-time trial — User confirms works → Done. Poor results → Step 13.
  12. Real-time trial — User confirms works → Done. Poor results → Step 13.
  13. Universe Model Search — search community models on Roboflow Universe. Good match → Done. No match → Step 14.
  14. Custom Training — OD / Inst Seg → prefer RF-DETR NAS. If NAS is unavailable, fine-tune RF-DETR (detection) or RF-DETR Seg (instance segmentation), sized by HW constraints. Done.

Model ID Reference

Read model-ids.md before calling a training tool. It contains the exact supported model_id values and the COCO class reference. Do not guess model IDs.

Model Selection Quick Guide

When you are training a model on the user's data for detection or instance segmentation, start with NAS. Neural Architecture Search searches the RF-DETR space against that data and reports a speed/accuracy frontier, so it is the option most likely to land on the best model for it. Reach for a single named architecture when NAS is unavailable (see prerequisites below), when the user asks for a specific one, or when a quick throwaway baseline is all that's wanted.

This summarises the training branches of the decision tree above; it does not override it. The tree may route a non-COCO request to Roboflow Rapid or SAM3 zero-shot first, neither of which trains a model — NAS only applies once custom training is the chosen path.

GoalRecommended
Object detection — no strong priorRF-DETR NAS (rfdetr-nas-parent)
Instance segmentation — no strong priorRF-DETR NAS Seg (rfdetr-nas-seg-parent)
Best accuracy, object detection, NAS unavailableRF-DETR (Large or XL)
Fast inference, object detection, NAS unavailableRF-DETR Nano or YOLOv11n
Best accuracy, instance segmentation, NAS unavailableRF-DETR Seg
Quick proof-of-concept (<1000 images)Roboflow Instant
ClassificationViT or DINOv3
Multimodal / text promptsQwen3.5 or SmolVLM

NAS parents exist only for object detection and instance segmentation. Keypoint, classification, semantic segmentation, and VLM tasks have no NAS option — use the named models above.

Comparing architectures (sweeps)

If you are comparing architectures rather than picking one, include a NAS parent as one of the candidates whenever the prerequisites are met — e.g. rfdetr-medium vs yolo26m vs rfdetr-nas-parent. Launch one trainings_create per candidate and keep each trainingId.

Before comparing results, read NAS results. It explains why a recommended child is not necessarily the highest-accuracy child and how to choose a fair comparison.

The NAS arm also takes longer than a single fine-tune, so report the named-model arms as they finish rather than blocking on NAS.

RF-DETR NAS (Neural Architecture Search)

Instead of picking a single RF-DETR size manually, NAS trains one parent model and mines many architectures out of it, reporting the speed/accuracy frontier so you can pick the one that fits your hardware budget.

  • What: A NAS run trains a single parent model, then searches the RF-DETR architecture space within that trained parent to identify frontier candidates, reporting each one's mAP and measured latency on target hardware (e.g., Jetson, T4 GPU). The output is a set of models on a Pareto frontier, plus a winner auto-selected per (metric, hardware) bucket using Roboflow's current ranking heuristic to balance validation accuracy against measured latency. Those per-bucket winners are not exposed individually: models_list flattens them into one recommended boolean per child (see Picking for a specific hardware target).
  • Tasks: Object Detection (rfdetr-nas) and Instance Segmentation (rfdetr-nas-seg).
  • When to use: When you want the best speed/accuracy tradeoff for a specific deployment target and don't want to A/B-test sizes manually. Especially valuable for edge hardware where latency budgets are tight.
  • Phases:
    1. Parent training — trains the one parent model the search draws from. This is the bulk of the wall-clock.
    2. Mining — searches architectures inside the trained parent and evaluates candidates to build the Pareto frontier (latency vs mAP). Candidates are derived from the parent rather than each being trained from scratch, so they appear in a burst near the end of the run. Each becomes a regular model you can deploy.
  • Model IDs: rfdetr-nas-parent (Standard — use this by default), rfdetr-nas-pecoret-parent (Fast), rfdetr-nas-base-parent (Plus) for object detection; rfdetr-nas-seg-parent for instance segmentation.
  • Prerequisites — check both before offering NAS:
    1. ≥15 validation images. versions_get returns splits.valid; below 15 the train call fails with insufficient_validation_images_for_nas, and waiting will not help — the user must generate a version with a larger validation split.
    2. Plan entitlement. NAS is included on Core and Growth plans; any other plan needs it granted on the workspace. Basic/starter/sandbox/research/trial need to upgrade; enterprise/legacy need to contact sales. Entitlement is not readable from the MCP, so this cannot be checked up front — a non-entitled workspace finds out when trainings_create rejects the run with code nas_not_available_for_plan. Treat that as a plan limit, not a transient error: do not retry it. Fall back to the named model for the task — rfdetr-medium (detection) or rfdetr-seg-medium (segmentation) — and say NAS is unavailable on the current plan and may need an upgrade or workspace enablement, using the plan on the error to tell which. Do not fall back to a hyperparameter sweep.
  • Start a run: trainings_create(project_id, version_number, model_type="rfdetr-nas-parent"). NAS launches through the normal training tool; there is no separate engine parameter. (The UI equivalent is the Train page with ?engine=nas.) Results land at /{workspace}/{project}/nas-runs/{versionId}.
  • Nothing to hand-tune: a NAS parent's whole hyperparameter surface is epochs (default 200, range 100–300). There are no learning-rate or loss-weight knobs, because the architecture search is the sweep.
  • Reading a run and selecting a hardware target: Read NAS results before interpreting child metrics or choosing a model. It covers paging, latency formats, and the limits of recommendation flags.
  • Deploy: Each NAS-produced model deploys like any other — pick one and use it as a normal Roboflow model. Call it the hardware-recommended child only when an authoritative recommendedByHardware entry actually says so; otherwise it is the child the user chose from the frontier. Inference type is rfdetr-nas / rfdetr-nas-seg, but it's served through the standard inference paths.
  • References: RF-DETR paper (arxiv), ICLR 2026, What is NAS? (blog).

Roboflow Instant / Rapid

Roboflow Instant

  • What: Few-shot model, trains in minutes, free
  • Task: Object Detection only
  • When to use: PoC, <1000 images, quick iteration
  • Auto-trains when you approve a batch and no Instant model exists yet
  • No preprocessing/augmentation -- uses images as-is
  • Deploy: Available in Workflows like any trained model
  • Manual trigger: Project > Models > Train Model > Roboflow Instant Model

Roboflow Rapid

  • What: Interactive annotation-and-training workflow — SAM3 pre-annotates a small image set, user reviews/corrects, a fast custom OD model trains automatically. Model keeps improving as it captures more production data.
  • Task: Object Detection only, non-COCO classes
  • When to use: Default path for non-COCO object detection when exclusions don't apply

Do NOT use Rapid when:

ExclusionWhy
OCR / text detection (characters, serial numbers, labels, receipts, license plates)SAM3 cannot reliably segment individual characters
Blueprints, floor plans, schematics, technical drawingsAbstract symbols and line-based elements not handled by SAM3 text prompting
More than 5 target classesSAM3 text prompting accuracy degrades significantly with many classes
Fine-grained visual distinctions (correct vs incorrect orientation, pass/fail, subtle defects)SAM3 cannot differentiate nearly identical objects; fine-tuned model needed
High-precision measurement / metrology (distances, dimensions, tolerances)SAM3 auto-labeling annotation precision insufficient for calibrated measurement

When Rapid is excluded → resume the decision tree at Step 13: Universe model search first, then custom training, which starts with RF-DETR NAS and falls back to named RF-DETR when its prerequisites are not met.

Checkpoint Training

OptionWhen to use
Public Checkpoint (COCO)First model version, default recommended
Universe CheckpointStar a Universe project first, then it appears as checkpoint option. Good for domain-specific transfer learning
Previous VersionAlready have a good model, want to improve with more data (all types except classification and SAM3)
Random InitializationAdvanced users only, usually worse results

Training Controls

  • Cancel Training: Stops job, no weights saved. Refund if early in training.
  • Early Stopping: Stops job, saves weights. Use when graphs show convergence with many epochs remaining. Charges for used credits.
  • NAS Training: Shows paired charts (mining progress + Pareto curve, then per-model training curves). May auto-stop on convergence. See RF-DETR NAS section below.

Post-Training Metrics

Metrics vary by project type:

Project TypeMetrics Shown
Object DetectionmAP@50, Precision, Recall, F1
ClassificationAccuracy
Instance Segmentation / KeypointmAP@50, Precision, Recall
Semantic SegmentationmIoU
MultimodalPerplexity

Model Evaluation (Paid Plans)

Auto-runs after training. Access: Models > click model version > View Evaluation.

FeatureWhat it shows
Production Metrics ExplorerPrecision/Recall/F1 at all confidence thresholds; recommends optimal confidence
Model Improvement RecommendationsActionable suggestions (false negatives, false positives, confused classes, insufficient data)
Performance by ClassCorrect predictions, misclassifications, false negatives, false positives per class; filterable
Confusion MatrixGround truth vs predictions grid; click cells to see specific images; adjustable confidence threshold
Vector ExplorerInteractive embedding clusters showing where model succeeds/fails

Deep link: https://app.roboflow.com/{workspace}/{project}/evaluation/{versionId}.

When a user asks why a model is bad or how to improve it, start with roboflow://skills/roboflow-training-and-evaluation/model-diagnosis. It maps every evaluation panel to a root cause (taxonomy, mislabeled data, inconsistent label standards, coverage gaps, too little data, bad data) and says which data to add next. The improvement playbook holds the compact decision tree and the training-side fixes (architecture, size, augmentation, overfitting).

The MCP server exposes every evaluation panel. All require the model-eval:read scope and return 409 model_eval_not_done while an evaluation is still running.

PanelTool
Find an evaluationmodel_evals_list (filter by project_id, version_number, or model_id; one at a time)
Headline mAP / precision / recall + app_urlmodel_evals_get
Model Improvement Recommendationsmodel_evals_get_recommendations ({"generated": false} when never produced)
Performance by Classmodel_evals_get_performance_by_class (split = train/valid/test)
mAP@50 / @50-95 / @75 per split, by object size and per classmodel_evals_get_map_results
Confusion Matrixmodel_evals_get_confusion_matrix (split, confidence 0-100; defaults to the optimal threshold)
Production Metrics Explorermodel_evals_get_confidence_sweep
Vector Explorermodel_evals_get_vector_analysis, then model_evals_get_image_predictions for per-image TP/FP/FN and cluster ids
Dataset Health Checkprojects_health (regenerate=True to recompute; first run can take minutes)

Viewing & Comparing Models

  • Models page: Project sidebar > Models. Shows all Instant + fine-tuned models with metrics, architecture, license, dataset version used.
  • Universe tab: Starred Universe models available for transfer learning.
  • Visualize: Quick test on test-set images, uploaded images, or webcam. Works for OD, segmentation, classification, keypoint. Not supported for multimodal.

MCP Tools Reference

ActionTool
Generate versionversions_generate
Start trainingtrainings_create
Find a training runtrainings_list(project_id, version_number) returns each run's trainingId and status
Check a known training runtrainings_get(project_id, version_number, training_id); use the paginated NAS path above for child model details
Get model infomodels_get
List modelsmodels_list
Find evaluationsmodel_evals_list
Evaluation summarymodel_evals_get
Evaluation recommendationsmodel_evals_get_recommendations
Per-class metricsmodel_evals_get_performance_by_class
mAP by split / object size / classmodel_evals_get_map_results
Confusion matrixmodel_evals_get_confusion_matrix
Confidence sweepmodel_evals_get_confidence_sweep
Vector analysis / per-image predictionsmodel_evals_get_vector_analysis, model_evals_get_image_predictions
Dataset health checkprojects_health

Related Pages

  • model-ids.md — exact training model IDs and COCO class coverage
  • roboflow://skills/roboflow-training-and-evaluation/model-diagnosis — start here for "why is my model bad / how do I improve it": run Model Evaluation, read each panel, map symptoms to root causes, decide which data to add
  • roboflow://skills/roboflow-training-and-evaluation/improvement-playbook — diagnostic decision tree, confusion matrix and per-class metric guide, recommendation types, architecture switching, augmentation, overfitting, iterative checklist
  • roboflow://skills/roboflow-training-and-evaluation/active-learning — production feedback loop: Project Model Workflow block, Active Learning, review, and retraining

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/roboflow/computer-vision-skills/roboflow-training-and-evaluation">View roboflow-training-and-evaluation on skillZs</a>