Ultralytics Benchmarks
How YOLO26 performs everywhere you run it — measured training throughput and inference latency on Ultralytics Platform GPUs, plus real on-device CPU, GPU, and NPU speeds on flagship phones, with memory, power, accuracy, and cost-efficiency.
GPU Training Throughput
YOLO26 detection trained on COCO at 640px with auto-batch across every NVIDIA GPU on the Platform. Compare across devices or by model size, switch the chart metric, or filter by generation; the table is fully sortable.
Training methodology
We use training throughput — images processed per second during training — as the yardstick; it correlates directly with time-to-solution. Every result is measured on Ultralytics Platform GPUs — the same NVIDIA hardware you rent for cloud training in one click, from entry-level workstation cards up to flagship data-center GPUs. We train YOLO26 at all five sizes (n/s/m/l/x) so you can match a model to your hardware budget.
Settings. 2 epochs on 25% of COCO at 640px, AMP mixed precision, single GPU, with auto-batch (batch=-1) selecting the largest batch that fits in memory. We report the steady-state second epoch, which excludes first-epoch warmup (dataset caching, CUDA graph capture) and the end-of-run validation pass. Resolved batch size, peak VRAM (including the CUDA context), and peak board power are recorded directly from each GPU.
Cost-efficiency. Images per dollar = throughput × 3600 ÷ hourly price, using Ultralytics Platform on-demand pricing — it often reorders the ranking dramatically, as value cards out-earn flagship GPUs per dollar. See also the Train and Benchmark mode docs.
Train on the Best GPUs for Less
26 NVIDIA GPUs starting at $0.24/hr — from Ampere to Blackwell. No markup, no minimums, no commitment.
Ready to build your next vision AI project?
Built on Ultralytics open source. Start training models in minutes.