
Where Ultralytics is headed
From the models that brought real-time vision AI to millions, to the releases on the horizon. Here's what we've shipped and what's coming next.
The release that brought real-time object detection to millions — PyTorch-native, fast, and remarkably easy to train.
One framework for detection, segmentation, classification, pose estimation, and oriented bounding boxes.
A refined architecture delivering higher accuracy with fewer parameters, unveiled at Ultralytics YOLO Vision 2024.
Our current recommended model — faster, more accurate, and production-ready across every vision task.
The end-to-end platform to annotate data, train YOLO models, and deploy across 42 global regions — all in one place.
Our first semantic segmentation models — dense per-pixel class labels for full-scene understanding.
A high-performance YOLO inference library and CLI written in Rust, running models on ONNX Runtime behind an API designed to match the Ultralytics Python package.
Run YOLO models directly in the browser, with no server and no Python. It runs on WebGPU with an automatic CPU and wasm fallback, behind a small TypeScript API powered by the Rust crate compiled to WebAssembly.
The research paper behind YOLO26 — detailing its NMS-free end-to-end design, the new MuSGD optimizer, and a state-of-the-art accuracy-latency tradeoff across all five model scales.
Faster, more accurate multi-object tracking — more stable identities through occlusions and crowded scenes for real-world video.
- Re-ID — re-identification keeps object identities consistent across cameras and after occlusions
Built-in knowledge distillation compresses large teacher models into smaller, faster students while helping preserve accuracy for efficient edge and real-time deployment.
YOLO26 depth models predict dense per-pixel distance maps from a single RGB camera, adding 3D spatial awareness without specialized depth sensors or lidar.
Ask AI helps you train, compare, and improve models through a conversation, from a baseline run to the next experiment.
A visual workflow builder that connects YOLO models, conditions, and vision-language models to actions such as dataset collection and Slack alerts.
Track endpoint requests, latency, errors, and logs, plus live prediction statistics and temporary examples on dedicated endpoints.
The next flagship generation of YOLO, unveiled live at Ultralytics YOLO Vision 2026, expanding the family into vision-language. Join the waitlist.
YOLOE-27 is the next generation of real-time, open-vocabulary detection and segmentation, built to see anything through text, visual, or prompt-free inference.
Built for retail, warehouse, security, and aerial/drone use cases, available exclusively to Enterprise customers.
Export YOLO models to Apple's new .aimodel format for the Core AI runtime after iOS 27 and macOS 27 become generally available, while retaining Core ML for broader device compatibility.
Run the Platform inside your own infrastructure, keeping data and training fully under your control.
New capabilities join the YOLO family throughout 2027:
- YOLO-OCR — fast, accurate text recognition
- YOLO-Face — facial recognition and analysis
- YOLO-Stereo3D — stereo 3D detection from binocular camera pairs for robotics, a camera-native alternative to lidar
Build on the latest YOLO
Start training and deploying with YOLO26 today — and be ready for everything that's next.