skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
ultralytics/skills98 installs

yolo-export

Use when exporting or deploying Ultralytics YOLO models in Platform or code — the Platform Export tab and yolo export/model.export() for ONNX, TensorRT, CoreML, Core AI, OpenVINO, LiteRT, NCNN, ExecuTorch, and NPUs (RKNN, QNN, Hailo, Ascend, IMX, Axelera, DeepX, Xilinx), FP16/INT8 quantization, benchmarking, and non-Python runtimes. For inference with .pt weights or Platform endpoints, see yolo-inference.

How do I install this agent skill?

npx skills add https://github.com/ultralytics/skills --skill yolo-export
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill provides documentation and examples for exporting YOLO models using the Ultralytics library. It describes standard library features, including the automatic installation of specific export toolchains when needed. No security issues were detected.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Export, quantization & deployment

Fastest route: export in Platform

Open a completed model's Export tab, select a format, configure its arguments, and click Start Export. Platform runs CPU exports directly and asks for a target GPU where the format requires one (notably TensorRT); download the artifact when the job completes. For TensorRT, select the deployment device's GPU family or Jetson target; if its TensorRT/CUDA versions differ from the device's, export locally on that device instead.

Use Platform when you do not want to install each exporter toolchain locally. Use the Python/CLI path below for custom calibration, repeatable automation, local hardware builds, or immediate parity validation. See Platform model export.

Quickstart

from ultralytics import YOLO

model = YOLO("runs/detect/train/weights/best.pt")
path = model.export(format="onnx")  # returns the exported file/dir path
yolo export model=best.pt format=onnx

Exports load straight back into YOLO() for predict/val — same API:

model = YOLO("best.onnx")  # or best.engine, best_openvino_model/, ...

Choose format by target hardware

Targetformat=Why
NVIDIA GPU / Jetsonengine (TensorRT)fastest on NVIDIA; build on the deployment device — engines are not portable across GPUs/TRT versions
Intel CPU/iGPU/NPUopenvinoup to 3× CPU speedup on Intel
Apple iOS/macOScoremlbroad OS coverage, Vision/iOS/Flutter support
Apple iOS 27+/macOS 27+coreainative .aimodel; iOS/Flutter SDKs default to coreml; export on Apple silicon macOS 26+ or x86_64 Linux
Androidlitert (renamed from tflite) or ncnnNCNN strong on ARM
Raspberry Pincnnbest ARM CPU latency
PyTorch Edgeexecutorch
Cross-platform / unsureonnxruns everywhere; start here, specialize when latency demands
NPUs (Rockchip/Qualcomm/Hailo/Huawei/Sony/Axelera/DeepX/AMD)rknn / qnn / hailo / ascend / imx / axelera / deepx / xilinxname= selects the exact chip for rknn/qnn/hailo/ascend/xilinx

Full 22-target local export matrix with per-format supported args: format-matrix.md (this folder).

Key arguments

ArgDefaultNotes
imgszmodelinherited from the loaded model; set explicitly to the deployment shape
quantizeNoneprecision request: 16/fp16, 8/int8/w8a8, w8a16, w8a32, or 32/fp32; support, speed, size, and accuracy are backend-dependent — see format-matrix.md and benchmark the target
dataNonerepresentative calibration data when required; use >300 images generally and 500+ for TensorRT. Omission selects a small task default, so pass deployment-representative data explicitly
dynamicFalsevariable input shape/batch where supported; check format-matrix.md and benchmark the target
batch1max batch baked into the export
simplifyTruesimplify ONNX graph
opsetNonecompatible ONNX opset selected automatically when unset; pin lower if the consumer runtime complains
nmsNoneNone exports raw outputs for external NMS (YOLO26 too); True embeds NMS where format-matrix.md lists nms; False selects the YOLO26/YOLOv10 NMS-free head, else raw outputs
workspaceNoneTensorRT builder GiB — lower if the build OOMs
deviceNoneTensorRT needs a GPU (auto-set to device=0 when unset); a GPU also speeds INT8 calibration
fraction1.0fraction of calibration data used

Verify parity after export (always)

yolo val model=best.pt data=data.yaml   # baseline
yolo val model=best.onnx data=data.yaml # compare the same task metric with the baseline

Acceptable differences depend on the task, model, backend, precision, and calibration data. Investigate unexpected gaps by matching imgsz and pre/post-processing and, where required, using representative calibration data. Also compare one prediction with .pt.

Benchmark all formats empirically

yolo benchmark model=best.pt data=data.yaml imgsz=640                                    # all formats at default precision
yolo benchmark model=best.pt data=data.yaml format=engine quantize=16 device=0 imgsz=640 # targeted FP16

Produces the task metric + latency per exportable format on this machine. Repeat for each supported precision and benchmark on deployment hardware, not your dev box.

Consuming exports outside Python

  • In raw runtimes (C++, mobile, JS) you own preprocessing (letterbox resize, BGR→RGB, /255) and output decoding.
  • Detect output layout follows nms: the default nms=None emits raw [4+nc, anchors] heads (xywh boxes + class scores) for every model, YOLO26 included, so you run NMS (imx instead always embeds NMS). nms=False on YOLO26/YOLOv10 emits NMS-free [max_det, 6] rows [x1,y1,x2,y2,conf,cls], and nms=True embeds NMS with the same row layout where supported; some formats fall back to raw heads (format-matrix.md). Segment, pose, and OBB add task-specific outputs. Check export warnings and shapes.
  • Class names travel in export metadata where supported; otherwise ship the names map alongside the model.
  • Serving: ultralytics.utils.triton.TritonRemoteModel for Triton; examples/ in the ultralytics repo has ONNXRuntime C++/Rust/Python references.

Troubleshooting

SymptomFix
Export crashes on missing packagemost backends auto-install on first export; rerun. On Linux, TensorRT auto-installs as tensorrt-cu<torch CUDA major>; elsewhere install it per NVIDIA docs. Hailo (Dataflow Compiler) and Ascend (CANN atc) need their vendor SDK first
Unsupported ONNX opset downstreamexport with lower opset=, or upgrade the runtime
TensorRT build OOM/slowlower workspace, batch=1, dynamic=False
Export much less accurateimgsz mismatch; too little/unrepresentative calibration data; wrong custom pre/post-processing; use a supported higher precision or backend
Engine fails on another machineTensorRT engines are device+version specific — rebuild on target
CoreML export fails on Windowsexport on macOS or Linux
Core AI export is unavailablerequires macOS 26+ on Apple silicon or x86_64 Linux (glibc 2.34+), torch>=2.8, and Python 3.11–3.14; use CoreML for broader production support
Deprecation warnings for half/int8/end2end/tfliteauto-forwarded (half→quantize=16, int8→quantize=8, end2end=True→nms=False, end2end=False→nms=None unless nms=True is also passed, tflite→litert) — switch to the new names

Related pages

  • format-matrix.md — all 22 local export targets, artifacts produced, and supported args. Read when using any format beyond onnx/engine/openvino/coreml.

If the installed version rejects an argument, trust the error text (format and quantize errors list valid values) and yolo cfg over this file.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/ultralytics/skills/yolo-export">View yolo-export on skillZs</a>