Enterprise-ready security: ISO 27001 + SOC 2 Type I compliant.

Ultralytics Benchmarks

How YOLO26 performs everywhere you run it — measured training throughput and inference latency on Ultralytics Platform GPUs, plus real on-device CPU, GPU, and NPU speeds on flagship phones, with memory, power, accuracy, and cost-efficiency.

GPU Training Throughput

YOLO26 detection trained on COCO at 640px with auto-batch across every NVIDIA GPU on the Platform. Compare across devices or by model size, switch the chart metric, or filter by generation; the table is fully sortable.

Training methodology

We use training throughput — images processed per second during training — as the yardstick; it correlates directly with time-to-solution. Every result is measured on Ultralytics Platform GPUs — the same NVIDIA hardware you rent for cloud training in one click, from entry-level workstation cards up to flagship data-center GPUs. We train YOLO26 at all five sizes (n/s/m/l/x) so you can match a model to your hardware budget.

Settings. 2 epochs on 25% of COCO at 640px, AMP mixed precision, single GPU, with auto-batch (batch=-1) selecting the largest batch that fits in memory. We report the steady-state second epoch, which excludes first-epoch warmup (dataset caching, CUDA graph capture) and the end-of-run validation pass. Resolved batch size, peak VRAM (including the CUDA context), and peak board power are recorded directly from each GPU.

Cost-efficiency. Images per dollar = throughput × 3600 ÷ hourly price, using Ultralytics Platform on-demand pricing — it often reorders the ranking dramatically, as value cards out-earn flagship GPUs per dollar. See also the Train and Benchmark mode docs.

Ready to build your next vision AI project?

Built on Ultralytics open source. Start training models in minutes.