Where Ultralytics is headed
From the models that brought real-time vision AI to millions, to the releases on the horizon. Here's what we've shipped and what's coming next.
The release that brought real-time object detection to millions — PyTorch-native, fast, and remarkably easy to train.
One framework for detection, segmentation, classification, pose estimation, and oriented bounding boxes.
A refined architecture delivering higher accuracy with fewer parameters, unveiled at Ultralytics YOLO Vision 2024.
Our current recommended model — faster, more accurate, and production-ready across every vision task.
The end-to-end platform to annotate data, train YOLO models, and deploy across 42 global regions — all in one place.
Our first semantic segmentation models — dense per-pixel class labels for full-scene understanding.
A high-performance YOLO inference library and CLI written in Rust, running models on ONNX Runtime behind an API designed to match the Ultralytics Python package.
Run YOLO models directly in the browser, with no server and no Python. It runs on WebGPU with an automatic CPU and wasm fallback, behind a small TypeScript API powered by the Rust crate compiled to WebAssembly.
The research paper behind YOLO26 — detailing its NMS-free end-to-end design, the new MuSGD optimizer, and a state-of-the-art accuracy-latency tradeoff across all five model scales.
Faster, more accurate multi-object tracking — more stable identities through occlusions and crowded scenes for real-world video.
- Re-ID — re-identification keeps object identities consistent across cameras and after occlusions
Built-in knowledge distillation compresses large teacher models into smaller, faster students while helping preserve accuracy for efficient edge and real-time deployment.
YOLO26 depth models predict dense per-pixel distance maps from a single RGB camera, adding 3D spatial awareness without specialized depth sensors or lidar.
Ultralytics YOLO Vision 2026
The next flagship generation of YOLO, unveiled live at Ultralytics YOLO Vision 2026, expanding the family into vision-language:
- YOLO-VLM — a lightweight YOLO front-end feeding a deeper LLM layer for efficient vision-language pipelines
Export YOLO models to Apple's new .aimodel format for the Core AI runtime after iOS 27 and macOS 27 become generally available, while retaining Core ML for broader device compatibility.
Three product focus areas drive the Platform for the rest of the year:
- Auto-Training — iterative, LLM-driven training analysis that automatically diagnoses each run and refines the configuration over successive rounds to push accuracy higher
- On-Premise — run the Platform inside your own infrastructure, keeping data and training fully under your control
- Monitoring — production model monitoring to track performance, catch drift, and keep deployments healthy
New capabilities join the YOLO family throughout 2027:
- YOLO-OCR — fast, accurate text recognition
- YOLO-Face — facial recognition and analysis
- YOLO-Stereo3D — stereo 3D detection from binocular camera pairs for robotics, a camera-native alternative to lidar
Build on the latest YOLO
Start training and deploying with YOLO26 today — and be ready for everything that's next.