NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1924 most downloaded on PyPI
Ultralytics YOLO 🚀 for SOTA object detection, instance segmentation, semantic segmentation, depth estimation, classification, pose estimation, oriented object detection, and multi-object tracking.
Last release today
04 Oct 2026
Ships on a steady schedule
a new release about every 8 days
Nearly every release is documented
notes for 60 of the last 60 stable releases
11 versions withdrawn
withdrawn after publishing
5 years old
848 releases · first in 2022
Version 8.4.173 adds native AMD Xilinx export for Versal AI Edge Series Gen 2 NPUs , opening a new path to deploy Ultralytics models on AMD edge hardw
Version 8.4.173 adds native AMD Xilinx export for Versal AI Edge Series Gen 2 NPUs, opening a new path to deploy Ultralytics models on AMD edge hardware.
xilinx export format — Export models through AMD Quark to a quantized ONNX model and Vitis AI configuration. The format also accepts vitis, vitisai, and versal as aliases..rai model for the NPU.ultralytics 8.4.173 Add AMD Xilinx Vitis AI export for Versal AI Edge Gen 2 NPUs by @glenn-jocher in #26519Full Changelog: v8.4.172...v8.4.173
One column per quarter.
v8.4.173 - ultralytics 8.4.173 Add AMD Xilinx Vitis AI export for Versal AI Edge Gen 2 NPUs (#26519) Latest
Latest
Compare
v8.4.172 makes oriented-box cropping work as expected and improves reliability across model training, export, dataset handling, and deployment.
v8.4.172 makes oriented-box cropping work as expected and improves reliability across model training, export, dataset handling, and deployment.
save_crop=True and ObjectCropper now save crops upright and aligned to each oriented bounding box. Areas extending beyond the image are filled with black.cloudpickle and filelock base dependencies and disallow new ones by @glenn-jocher in #26471Full Changelog: v8.4.171...v8.4.172
v8.4.172 - ultralytics 8.4.172 Add rotation-aligned OBB crops to save_crop (#26486)
Compare
v8.4.171 adds end-to-end AMD GPU support for Ultralytics workflows, alongside fixes to dataset handling, model export, and several task-specific behav
v8.4.171 adds end-to-end AMD GPU support for Ultralytics workflows, alongside fixes to dataset handling, model export, and several task-specific behaviors.
latest-amd Docker image, AMD GPU CI testing, and setup and usage guidance. MIGraphX support targets Linux x86_64 with Python 3.11 or newer.Full Changelog: v8.4.170...v8.4.171
v8.4.171 - Add AMD ROCm and MIGraphX support (#24137)
Compare
Ultralytics 8.4.170 restores the established disk-cache naming, so existing caches are reused instead of rebuilt. It also documents video frame skippi
Ultralytics 8.4.170 restores the established disk-cache naming, so existing caches are reused instead of rebuilt. It also documents video frame skipping for Platform inference.
image.npy, rather than image.jpg.npy.vid_stride for Platform inference: The parameter controls how often video frames are processed; its default of 1 processes every frame, and images are unaffected.cache=disk datasets can reuse their original cache files without rebuilding a second copy, saving time and disk space.vid_stride to process fewer video frames when the API supports it, potentially reducing inference work. No model architecture changes are included.Full Changelog: v8.4.169...v8.4.170
v8.4.170 - ultralytics 8.4.170 Revert disk cache naming from #26457 (#26460)
Compare
Ultralytics v8.4.169 improves dataset reliability, preserves more COCO annotations, and fixes Edge TPU export setup when Google’s old package reposito
Ultralytics v8.4.169 improves dataset reliability, preserves more COCO annotations, and fixes Edge TPU export setup when Google’s old package repository is unavailable.
sample.jpg and sample.png no longer share a cache file. This prevents one image’s pixels from being mistakenly used for another, in both detection-family and classification datasets.PATH is still used when available. The bundled compiler supports Linux x86_64.Full Changelog: v8.4.168...v8.4.169
v8.4.169 - ultralytics 8.4.169 Avoid disk cache collisions for same-stem images (#26457)
Compare
YOLO 8.4.168 delivers practical fixes for image and mask handling, prediction outputs, video timing, and exported models, alongside clearer training g
YOLO 8.4.168 delivers practical fixes for image and mask handling, prediction outputs, video timing, and exported models, alongside clearer training guidance. No model architecture changes are included.
bus.jpg and bus.png, get distinct output names.mask_ratio affects training masks, while optimizer=auto ignores lr0; it also clarifies deployment rename behavior. CLI examples were adjusted accordingly.SearchApp is tested using existing test images rather than triggering an extra large image download.lr0 that have no effect with the default automatic optimizer.mask_ratio does not change predicted mask resolution by @Y-T-G in #26441Full Changelog: v8.4.167...v8.4.168
v8.4.168 - ultralytics 8.4.168 Fix palette mask conversion, stale disk cache, AVIF orientation, same-stem predict outputs, and static-batch short inputs (#26443)
Compare
Ultralytics 8.4.167 reduces repeated downloads and storage for Platform dataset versions by reusing unchanged, content-addressed assets. No model arch
Ultralytics 8.4.167 reduces repeated downloads and storage for Platform dataset versions by reusing unchanged, content-addressed assets. No model architecture or training behavior changes are included.
v1.4.41.Full Changelog: v8.4.166...v8.4.167
v8.4.167 - ultralytics 8.4.167 Share content-addressed NDJSON assets across dataset versions (#26442)
Compare
Ultralytics 8.4.166 focuses on more reliable dataset handling and prediction outputs, alongside updated Platform and inference documentation. No model
Ultralytics 8.4.166 focuses on more reliable dataset handling and prediction outputs, alongside updated Platform and inference documentation. No model architecture changes are included.
Full Changelog: v8.4.165...v8.4.166
v8.4.166 - ultralytics 8.4.166 Support Pascal VOC-style palette and 16-bit PNG masks (#26426)
Compare
v8.4.165 improves package compatibility, export and inference reliability, and dataset handling, with several targeted fixes for model outputs and vis
v8.4.165 improves package compatibility, export and inference reliability, and dataset handling, with several targeted fixes for model outputs and vision workflows.
tests package, avoiding conflicts with other projects. Tests remain in the source distribution so Conda builds can still run them.tests imports, while source-based Conda builds retain the tests they need.Results.plot instance mode by @MohammadHijjawi97 in #26389tests package from the wheel by @Y-T-G in #26408Full Changelog: v8.4.164...v8.4.165
v8.4.165 - ultralytics 8.4.165 Stop installing top-level tests package from the wheel (#26408)
Compare
v8.4.164 makes FLOPs reporting more efficient and improves export, video inference, dataset validation, and reliability—without changing model archite
v8.4.164 makes FLOPs reporting more efficient and improves export, video inference, dataset validation, and reliability—without changing model architectures or predictions.
ultralytics-thop 2.2.0.model.info() and training startup logging lighter.Annotator(pil=True) from a font reference cycle by @Y-T-G in #26346labels autolabel path from non_max_suppression by @cainiao33 in #26356batched_mask_to_box on CPU by @Y-T-G in #26364Full Changelog: v8.4.163...v8.4.164
v8.4.164 - Delegate FLOPs counting to THOP 2.2.0 stride profiling (#26373)
Compare
YOLO models can now be exported to Apple Core AI format from x86_64 Linux, making Linux servers and CI environments useful for preparing models for Ap
YOLO models can now be exported to Apple Core AI format from x86_64 Linux, making Linux servers and CI environments useful for preparing models for Apple devices.
.aimodel files on x86_64 Linux as well as Apple silicon Macs running macOS 26 or later. The exported models target iOS 27 and macOS 27; running them still requires Apple hardware.coreai-torch 0.4.3, which performs optimization during conversion. The dependency now supports Python 3.11–3.14..aimodel files, and Core ML remains a separate option.AIProgram.optimize() by @cdeil in #26327ultralytics 8.4.163 Export Apple Core AI models on Linux by @glenn-jocher in #26333Full Changelog: v8.4.162...v8.4.163
v8.4.163 - ultralytics 8.4.163 Export Apple Core AI models on Linux (#26333)
Compare
YOLO prediction can now load the next batch of images while CUDA processes the current one, alongside performance, export, dataset, and Ultralytics Pl
YOLO prediction can now load the next batch of images while CUDA processes the current one, alongside performance, export, dataset, and Ultralytics Platform improvements. No model architectures or weights changed.
.yml support, OOM recovery, and tracking when source files share a name.Full Changelog: v8.4.161...v8.4.162
v8.4.162 - ultralytics 8.4.162 Prefetch the next image batch during CUDA prediction (#26319)
Compare
Ultralytics 8.4.161 is a reliability and usability release: it fixes several training, inference, plotting, and upload bugs, while improving dataset h
Ultralytics 8.4.161 is a reliability and usability release: it fixes several training, inference, plotting, and upload bugs, while improving dataset handling and clarifying export and Platform documentation.
last.pt when save=False means no best.pt was created..tgz archives now extract correctly, semantic segmentation can skip unnecessary polygon processing, and fuzz testing better recognizes expected missing-split errors.device option for Hailo and Ascend, and explained why static exports may produce different predictions on non-square images.device export argument for OpenVINO, Hailo and Ascend by @raimbekovm in #26286Full Changelog: v8.4.160...v8.4.161
v8.4.161 - Fix RT-DETR device moves, fused resume, empty images, bracket paths, odd plot grids and save=False uploads (#26301)
Compare
Ultralytics 8.4.160 makes dataset edits less likely to go unnoticed and improves memory use, training reliability, and everyday API behavior—without i
Ultralytics 8.4.160 makes dataset edits less likely to go unnoticed and improves memory use, training reliability, and everyday API behavior—without introducing a new model architecture.
data to choose a dataset explicitly. Depth calibration now requires data.Results.numpy() for CUDA-backed results by @aswanth-07 in #26282safe_download of single-file ZIP archives by @Nikhi00718 in #26273non_max_suppression() from mutating input predictions by @Nikhi00718 in #26278save_one_box() for list, tuple and NumPy boxes by @MohammadHijjawi97 in #26279pathlib.Path inputs in img2label_paths() by @MohammadHijjawi97 in #26280ultralytics 8.4.160 Invalidate dataset caches when files are edited in place by @aswanth-07 in #26283Full Changelog: v8.4.159...v8.4.160
v8.4.160 - ultralytics 8.4.160 Invalidate dataset caches when files are edited in place (#26283)
Compare
v8.4.159 improves dataset split handling, startup monitoring, and INT8 SavedModel export efficiency.
v8.4.159 improves dataset split handling, startup monitoring, and INT8 SavedModel export efficiency.
📥 Avoids unnecessary dataset downloads during training
split=val skips unused test images, reducing download time, storage, and preparation overhead.split=test retains test images and makes the trainer validate against that selected split.📊 Reports system metrics as soon as training starts
training_started Platform request now includes an initial system snapshot.🧠 Reduces peak memory during INT8 SavedModel export
🧪 Adds and updates tests and documentation
split setting now consistently controls which data is prepared, loaded, and evaluated across supported tasks.Full Changelog: v8.4.158...v8.4.159
v8.4.159 - Skip unused test downloads and report initial training system metrics (#26276)
Compare
Ultralytics 8.4.158 improves YOLO26 training reliability, model export compatibility, data handling, inference backends, and Ultralytics Platform work
Ultralytics 8.4.158 improves YOLO26 training reliability, model export compatibility, data handling, inference backends, and Ultralytics Platform workflows. 🚀
🔄 Mosaic augmentation now stays closed after OOM recovery
⚙️ More reliable AutoBatch sizing
🧠 Improved data loading and caching
spawn and forkserver DataLoader workers.📦 Stronger download and dataset handling
safe_download() can resume interrupted downloads with HTTP Range requests, handle encoded responses, and better tolerate temporary server errors.📤 More accurate model exports
🎭 Better SAM2 video masks
🍎 Expanded Apple deployment support
🛠️ Ultralytics Platform improvements
📚 Documentation and compatibility updates
safe_download retries with Range requests by @glenn-jocher in #26250safe_download robust to encoded responses, curl resumes, and 5xx bursts by @glenn-jocher in #26251ultralytics 8.4.158 Keep mosaic closed after an OOM batch auto-reduction by @rahultechenable in #26257Full Changelog: v8.4.157...v8.4.158
v8.4.158 - ultralytics 8.4.158 Keep mosaic closed after an OOM batch auto-reduction (#26257)
Compare
Ultralytics v8.4.157 makes TensorRT inference faster—up to 20% for FP16 and INT8 engines—while improving YOLOE prompt-free support, Apple Silicon perf
Ultralytics v8.4.157 makes TensorRT inference faster—up to 20% for FP16 and INT8 engines—while improving YOLOE prompt-free support, Apple Silicon performance, validation reliability, and training stability. 🚀
⚡ Faster TensorRT engines (PR #26223, @Y-T-G)
🎯 Expanded YOLOE-26 prompt-free inference (PR #26201)
nms setting can now select between standard NMS inference and NMS-free inference.set_vocab() can regenerate both branches for custom prompt-free YOLOE models.🍎 Improved Apple Silicon CPU and MPS performance
bincount operations with MPS-friendly alternatives, addressing severe slowdowns and buffer-size failures.🧪 More reliable training and validation
📐 Improved export and dataset behavior
predict() calls.nc=1, preserving foreground pixels in binary mask training.📦 Smoother dependency installation
uv lock and uv sync reliability.uv sync and uv lock extra resolution by @hylreg in #26219Full Changelog: v8.4.156...v8.4.157
v8.4.157 - Speed up TensorRT FP16 and INT8 engines up to 20% (#26223)
Compare
v8.4.156 improves remote NDJSON dataset reliability and makes INT8 TensorRT exports faster while preserving accuracy. 🚀
v8.4.156 improves remote NDJSON dataset reliability and makes INT8 TensorRT exports faster while preserving accuracy. 🚀
🔄 Reliable remote NDJSON conversion — PR #26242 by @cainiao33
check_file utility, making file handling more predictable.⚡ Faster and more accurate INT8 TensorRT exports — PR #26171 by @Y-T-G
🧪 More robust testing — PR #26217 by @glenn-jocher
🛠️ Maintenance and documentation
v1.4.39 to v1.4.40.Full Changelog: v8.4.155...v8.4.156
v8.4.156 - Refresh remote NDJSON at the conversion owner (#26242)
Compare
…option for newer PyTorch versions, reducing deprecation warnings while preserving behavior.
Ultralytics 8.4.155 improves dataset-cache safety, export reliability, training validation, and platform compatibility—helping YOLO26 workflows fail less often and provide clearer feedback. 🚀
🧹 Prevented stale labels.cache reuse after dataset configuration changes (PR #26207, the primary release change)
Dataset cache validation now includes important scan settings such as:
data.yaml trigger a fresh label scan instead of reusing incompatible cached labels.⛔ Invalid epoch values are now rejected early
Training configurations with epochs=0 or negative values now raise a validation error instead of silently running 100 epochs or producing invalid training outputs.
🖥️ Improved Windows OpenVINO inference
Windows CPU inference now explicitly requests FP32 precision to avoid reduced-precision kernel failures on affected systems.
📦 Better control over automatic dependency installation
Setting YOLO_AUTOINSTALL=False now also prevents automatic apt installations and Edge TPU compiler setup. Missing dependencies produce a warning or actionable error instead of modifying the environment unexpectedly.
📱 Fixed YOLO26 pose training on Apple MPS
The RLE pose-loss weights now use the correct float32 type, allowing YOLO26 pose training to run on MPS devices.
⚡ Faster font checking
Font lookup uses Matplotlib’s cached font list before rescanning the operating system, significantly reducing startup time on systems such as macOS.
📤 Improved export and inference paths
[0, 1], fixing collapsed bounding boxes and producing correct detections.🔧 Updated distributed training compatibility
DDP now uses the appropriate buffer-synchronization option for newer PyTorch versions, reducing deprecation warnings while preserving behavior.
☁️ Updated Platform SDK documentation
Documentation now recommends ultralytics-platform>=0.1.45 and explains how training metrics, checkpoints, arguments, and host information synchronize with the Ultralytics Platform.
🌐 Expanded repository mirroring infrastructure
A scheduled workflow now mirrors public Ultralytics repositories to GitLab daily, including branches, tags, Git LFS objects, metadata, and repository avatars.
📚 Refreshed heatmaps documentation
The heatmaps guide now features an updated YOLO26 tutorial video.
YOLO_AUTOINSTALL in apt and Edge TPU install paths by @Y-T-G in #26214Full Changelog: v8.4.154...v8.4.155
v8.4.155 - Fix stale labels.cache reuse after data.yaml semantic edits (#26207)
Compare
v8.4.154 improves CoreML export and inference reliability, restores accurate RT-DETR INT8 deployment, reduces training overhead, and strengthens datas
v8.4.154 improves CoreML export and inference reliability, restores accurate RT-DETR INT8 deployment, reduces training overhead, and strengthens dataset and Platform workflows. 🚀
🛠️ CoreML dynamic export fixed — PR #26199
dynamic=True without triggering a coremltools arange conversion error.⚡ RT-DETR OpenVINO INT8 export repaired
🏎️ Faster training, especially on GPUs
🧠 More memory-efficient SAM3 mask processing
🎯 Classification validation now prevents class-index mistakes
📡 Platform training callbacks now use the Platform SDK
ultralytics-platform>=0.1.45.✅ Clearer dataset and validation checks
kpt_shape is missing, including when stale label caches are present.save_json=True now reports small-, medium-, and large-object mAP on detection datasets using faster-coco-eval.data= override is provided.🌐 Platform workflow and documentation updates
Full Changelog: v8.4.153...v8.4.154
v8.4.154 - Fix CoreML dynamic anchor export and static multi-image inference, release 8.4.154 (#26199)
Compare
Ultralytics 8.4.153 delivers a critical SAM3 initialization fix, improves validation reliability, clarifies INT8 export behavior, and expands Ultralyt
Ultralytics 8.4.153 delivers a critical SAM3 initialization fix, improves validation reliability, clarifies INT8 export behavior, and expands Ultralytics Platform documentation and annotation capabilities. 🚀
🛠️ Fixed SAM3 compile=False handling — PR #26189 by @glenn-jocher
False disables compilation.True enables the default compilation mode.torch.compile activation and the resulting PyTorch 2.14 Dynamo error during SAM3 initialization.✅ Improved standalone semantic segmentation validation
🖼️ Fixed standalone classification validation with extra classes
IndexError crashes when the validation dataset contains more classes than the checkpoint supports.🔄 Preserved explicit datasets when resuming training
train(resume=True, data="...") now correctly honors the user-provided dataset instead of silently using the checkpoint’s original dataset.⚙️ Improved INT8 export behavior and documentation
dynamic is automatically enabled for INT8 exports.🎯 Preserved conf=0.0 during IMX export
0.001.conf=0.0 and the default conf=0.25 behavior.🧩 Expanded Platform annotation and billing documentation
📚 Documentation and links updated
conf=0.0.Full Changelog: v8.4.152...v8.4.153
v8.4.153 - Fix SAM3 compile=False handling and release 8.4.153 (#26189)
Compare
v8.4.152 improves system monitoring by adding cached NVIDIA driver and driver-supported CUDA version information to SystemLogger .
v8.4.152 improves system monitoring by adding cached NVIDIA driver and driver-supported CUDA version information to SystemLogger.
driver_version and cuda_version fields to system metrics collected on NVIDIA-enabled systems.SystemLogger starts, avoiding repeated NVML calls during metric collection.cuda_version represents the CUDA version supported by the installed NVIDIA driver, matching the value reported by nvidia-smi—not the installed CUDA toolkit or PyTorch build version.8.4.151 to 8.4.152.SystemLogger, without collecting them separately.Full Changelog: v8.4.151...v8.4.152
v8.4.152 - Report NVIDIA driver and CUDA versions in SystemLogger (#26173)
Compare
Ultralytics 8.4.151 improves export correctness, TensorRT QAT performance, tracking behavior, model diagnostics, and documentation—while introducing n
Ultralytics 8.4.151 improves export correctness, TensorRT QAT performance, tracking behavior, model diagnostics, and documentation—while introducing no new runtime arguments. 🚀
Preserved conf=0.0 in NMS exports 🎯
0.25 is now applied only when conf is omitted or set to None.Faster QAT TensorRT engines ⚡
Explicit zero confidence preserved in tracking 🎥
model.track(conf=0.0) now keeps the requested value instead of replacing it with the default 0.1.0.1, maintaining existing tracking behavior.Warnings for blank class labels ⚠️
AutoBackend now warns when model class names are empty or contain only whitespace.Expanded tracking and validation documentation 📚
imgsz, max_det, and quantize.fraction option for selecting subsets of training, validation, or test data.Improved PyTorch hook guidance 🧩
Updated YOLO27 preview documentation 🔭
Documentation rendering and layout fixes 🖼️
conf=0.0 without the value being silently changed to 0.25. ✅Full Changelog: v8.4.150...v8.4.151
v8.4.151 - Preserve conf=0.0 in nms=True exports instead of baking 0.25 into the graph (#26163)
Compare
v8.4.150 improves secure checkpoint loading, restores compatibility with fused YOLOE models, and adds performance and documentation updates across Ult
v8.4.150 improves secure checkpoint loading, restores compatibility with fused YOLOE models, and adds performance and documentation updates across Ultralytics. 🚀
⚡ Faster restricted checkpoint loading — PR #26155 by @glenn-jocher
forward and forward_fuse bindings required by fused YOLOE checkpoints.repr() on partially reconstructed modules.🛡️ Safer diagnostics and Ray Tune compatibility — PR #26151
2.41.0 or newer and removes obsolete compatibility code for older Ray releases.📈 More efficient RT-DETR training on dense datasets — PR #26150
📚 YOLO27 preview documentation
☁️ Expanded Ultralytics Platform deployment documentation
🏆 Clearer model benchmark tables
yolo login or other local variables.Full Changelog: v8.4.149...v8.4.150
v8.4.150 - Speed up restricted loading and restore fused checkpoints (#26155)
Compare
Ultralytics 8.4.149 improves training reliability, export compatibility, and Platform documentation—especially for resumed runs and third-party integr
Ultralytics 8.4.149 improves training reliability, export compatibility, and Platform documentation—especially for resumed runs and third-party integrations. 🚀
More reliable training resume behavior 🔄
data= argument still takes priority.Safer ONNX CUDA fallback ⚙️
Ray Tune compatibility improvements 📈
More robust Weights & Biases runs 📊
Improved invalid-argument errors 🛠️
RT-DETR LiteRT export and inference fixes 📱
SAM 3.1 becomes the default Platform smart annotation model ✨
Platform AutoTrain documentation reorganized 📚
Package version updated 📦
8.4.149.Full Changelog: v8.4.148...v8.4.149
v8.4.149 - Fix resume, ONNX CUDA fallback, Ray Tune and W&B training crashes (#26148)
Compare
Ultralytics 8.4.148 makes SAM 3.1 image prediction checkpoints load correctly, enabling reliable point, box, text, and exemplar-based segmentation. 🎯
Ultralytics 8.4.148 makes SAM 3.1 image prediction checkpoints load correctly, enabling reliable point, box, text, and exemplar-based segmentation. 🎯
SAM 3.1 checkpoint support 🧠
sam3.1_multiplex.pt works correctly with SAM and SAM3Predictor.sam3.pt checkpoints remain compatible; the key mapping is effectively a no-op for SAM 3.Supported SAM 3.1 prediction workflows 🖼️
SAM and SAM3Predictor.SAM3SemanticPredictor.sam3.pt with the SAM 3 video predictors for video workflows.Documentation updates 📚
sam3.1_multiplex.pt.Repository guidance improvements 🛠️
AGENTS.md and moved documentation-specific instructions into a new docs/AGENTS.md.Version update 📦
8.4.147 to 8.4.148.ultralytics 8.4.148 Load SAM 3.1 checkpoints into the SAM 3 image predictors by @glenn-jocher in #26144Full Changelog: v8.4.147...v8.4.148
v8.4.148 - Load SAM 3.1 checkpoints into the SAM 3 image predictors (#26144)
Compare
Ultralytics v8.4.147 improves preprocessing for grayscale images, strengthens SAM segmentation behavior, and refreshes Platform, installation, and CI
Ultralytics v8.4.147 improves preprocessing for grayscale images, strengthens SAM segmentation behavior, and refreshes Platform, installation, and CI documentation. 🚀
🖼️ Albumentations now supports grayscale images
🎯 SAM generation improvements
min_mask_region_area to SAM generation to remove small mask regions and holes.generate() calls behave consistently with normal prediction.min_mask_region_area.📚 Expanded Platform and agent documentation
platform-cli agent skill and ul cloud workflows.⚡ Improved installation and quickstart guidance
🛡️ More reliable downloads
safe_download() now makes curl fail on HTTP errors instead of saving an error page as if it were a valid file.🧪 Stronger CI and runner environments
gh) and jq to CPU and GPU runner images.min_mask_region_area to SAM generate() and scale crop point prompts to the resized crop by @raimbekovm in #26129Full Changelog: v8.4.146...v8.4.147
v8.4.147 - Apply Albumentations to grayscale images (#26133)
Compare
Ultralytics v8.4.146 improves RT-DETR reliability across small inputs, dynamic exports, and tracking, while updating export portability, Windows train
Ultralytics v8.4.146 improves RT-DETR reliability across small inputs, dynamic exports, and tracking, while updating export portability, Windows training behavior, data downloads, and GPU container environments. 🚀
🛠️ RT-DETR inference fixes
📦 More device-agnostic TorchScript exports
🎯 Improved FP16 embedded-NMS inference
🪟 Better Windows CUDA training defaults
channels_last memory format on Windows, where it could significantly reduce performance.channels_last=True.🌐 More reliable NDJSON image downloads
🐳 Updated GPU and export environments
g++ for compiled training workflows and TensorRT for CUDA 13.🧪 CI and maintenance updates
channels_last auto-enable on Windows CUDA training by @Y-T-G in #26105Full Changelog: v8.4.145...v8.4.146
v8.4.146 - Fix RT-DETR inference for small and dynamic inputs (#26120)
Compare
Improves Platform login reliability in Ultralytics 8.4.145 by reporting authentication and connection failures with a proper non-zero exit status, mak
Improves Platform login reliability in Ultralytics 8.4.145 by reporting authentication and connection failures with a proper non-zero exit status, making command-line automation more dependable. ✅
yolo login and delegated ul login flows.8.4.144 to 8.4.145 so the login fix is available through PyPI.ultralytics==8.4.145 receive the corrected login behavior.Full Changelog: v8.4.144...v8.4.145
v8.4.145 - Bump to 8.4.145 for login failure exit status (#26117)
Compare
v8.4.144 improves model reliability, numerical stability, inference precision, and deployment workflows without changing model architectures. 🚀
v8.4.144 improves model reliability, numerical stability, inference precision, and deployment workflows without changing model architectures. 🚀
More reliable model YAML loading 🧩
Explicitly requested files such as custom26n.yaml are now loaded before attempting a scale-unified fallback like custom26.yaml. This prevents custom model definitions from being silently replaced.
Safer training and assignment edge cases 🛡️
Empty-label batches now return correctly typed boolean foreground masks, keeping task-aligned assigners consistent with their normal output contract.
Stable CIoU calculations for reduced precision 🔢
The overlap calculation now avoids NaN values and invalid gradients for identical FP16 and BF16 boxes, improving training stability on lower-precision hardware.
More efficient repeated model.to() calls ⚡
Cached predictors are preserved when a model is already on the requested device or precision. This avoids unnecessary rebuilding, reducing latency and memory use during repeated inference setup.
Improved Triton FP32 behavior 🎯
Triton client tensors now remain in FP32 even when quantize=16 is requested. This prevents unintended input rounding and output conversion on the client while leaving the server-side model precision unchanged.
Faster semantic segmentation mosaic loading 🖼️
Semantic masks for buffered mosaic images are now retained in RAM instead of being decoded repeatedly. Benchmarks showed approximately 1.26× to 1.77× faster data loading, depending on the configuration.
More flexible CLI configuration 🛠️
Arguments provided before cfg=<file> are now preserved, and blank data values in copied configuration files correctly fall back to the task’s default dataset.
Broader OpenVINO/NNCF compatibility 📦
The general upper version limit for NNCF has been removed, allowing newer NNCF 3.x releases with modern PyTorch and OpenVINO installations while retaining legacy restrictions where required.
More informative quantization documentation 📚
QAT versus post-training quantization results are now documented for YOLO26 models. The examples show that QAT provides little benefit for the smallest model but can recover substantially more INT8 accuracy for larger models.
More robust CI and test infrastructure ✅
CLA permissions were corrected, DEEPX export environments are skipped on machines with less than 15 GiB of RAM, and checkpoint corruption tests now work correctly with channels-last tensors.
Documentation maintenance 🔗
The DL Streamer system requirements link was fixed, and the documentation license banner now uses a CDN-hosted asset.
NaN values that can interrupt training or corrupt gradients.Full Changelog: v8.4.143...v8.4.144
v8.4.144 - Fix model loading, numerical edge cases, and CI setup (#26106)
Compare
v8.4.143 introduces INT8 quantization-aware training for YOLO26, improves deployment and evaluation workflows, and delivers a broad documentation and
v8.4.143 introduces INT8 quantization-aware training for YOLO26, improves deployment and evaluation workflows, and delivers a broad documentation and integration refresh. 🚀
🧠 INT8 quantization-aware training (QAT) — PR #26083
quantize=8 during train mode.quantize argument and adds no new public API or dependency.📈 More complete validation metrics — PR #24489
🚀 Improved YOLO26 deployment workflows
nms=False and updates Triton, DALI, SAM, and C++ examples accordingly.🛠️ Training and model lifecycle fixes
reset_weights(), during resume operations, and in multi-dataset or distributed training.Results objects.🌍 Platform and authentication updates
📚 Extensive documentation improvements
Full Changelog: v8.4.142...v8.4.143
v8.4.143 - Add INT8 quantization-aware training via quantize=8 in train mode (#26083)
Compare
The legacy end2end argument is deprecated:
Ultralytics v8.4.142 unifies YOLO inference and export behavior under the nms option, making it easier to choose between maximum accuracy and NMS-free speed while improving benchmark consistency and deployment documentation.
🔄 Unified end2end into nms
nms=None (default) uses the one-to-many head with Ultralytics-managed NMS.nms=True uses the one-to-many head and embeds NMS in supported exports.nms=False selects the one-to-one NMS-free head when available.end2end argument is deprecated:
end2end=True maps to nms=False.end2end=False maps to nms=None, unless explicitly combined with nms=True.📦 Improved model fusion and export handling
📏 More reliable cross-format benchmarks
🐛 Training resume fix
🌐 Expanded deployment documentation
🧪 Broader test coverage and maintenance
nms behavior.8.4.142.nms setting now controls head selection and export-time NMS across prediction, validation, tracking, benchmarking, and export.nms=False when you want NMS-free predictions and a simpler post-processing pipeline.end2end configurations continue to work through compatibility mapping, but new code should use nms.end2end under nms=True|False|None by @glenn-jocher in #26066Full Changelog: v8.4.141...v8.4.142
Ultralytics 8.4.141 updates Axelera support to Voyager SDK 1.8.0, adds batched inference, and improves model compatibility, dependency management, and
Ultralytics 8.4.141 updates Axelera support to Voyager SDK 1.8.0, adds batched inference, and improves model compatibility, dependency management, and CI reliability. 🚀
Axelera integration upgraded to Voyager SDK 1.8.0 🔧
axelera-devkit and axelera-rt 1.8.0 packages.PATH and safely restores environment variables afterward.Batched Axelera inference is now supported 📦
batch>1 are routed through the Voyager scheduler.Axelera deployment requirements are documented more clearly 📚
Fixed legacy SPPF model reconstruction 🛠️
SPPF layers with the wrong activation behavior.OpenVINO dependencies are now opt-in 📦
export-base to a dedicated export-openvino extra.export extra still includes OpenVINO support.Conda CI dependency resolution improved ⚡
conda-forge with strict channel priority and removes inherited default channels.Full Changelog: v8.4.140...v8.4.141
v8.4.140 improves training reliability and model-state preservation, led by a fix for grayscale TIFF datasets that could previously crash YOLO26 segme
v8.4.140 improves training reliability and model-state preservation, led by a fix for grayscale TIFF datasets that could previously crash YOLO26 segmentation training.
🖼️ Fixed grayscale TIFF channel handling (PR #26061, @glenn-jocher)
Single-frame grayscale TIFF images now respect the requested color format. This prevents three-channel models from receiving one-channel batches and failing at the first convolution. Existing support for color TIFFs, multipage files, four-channel images, and multispectral stacking is preserved.
🧪 Expanded regression coverage for TIFF training
The existing multichannel test now also trains, validates, predicts, and exports with grayscale TIFF data.
🧠 Preserved model state during inference and validation (PR #26062)
Prediction and standalone validation now operate on independent model copies, preventing inference-time fusion, FP16 conversion, or end-to-end configuration changes from permanently modifying the caller’s model.
🚀 More reliable distributed training (PR #26050)
DDP workers now receive the parent trainer’s prepared model, arguments, and callbacks. In-memory weight changes, custom class names, and application callbacks are therefore retained during multi-GPU training. The release adds cloudpickle to support this state transfer.
🎯 Improved weight loading and provenance (PR #26039)
Loading weights from modules, checkpoint dictionaries, or files now preserves the correct training source and prefers EMA weights when available. Predictor caches are refreshed after model changes, reducing the risk of stale inference behavior.
🔧 Safer fusion detection and calibration (PR #26045)
Fused pretrained weights now generate a warning when loaded into an unfused model. Fusion detection is more accurate for convolutional, reparameterized, and end-to-end YOLO26 components. Depth calibration also preserves a trainable, saveable model state.
📦 Safer exports (PR #26052)
Exporting works on an isolated model copy before modifying names or head settings. This keeps the original model unchanged and avoids copying YOLOWorld’s large cached CLIP encoder unnecessarily.
🧬 Tuning now uses the caller’s loaded model (PR #26053)
Model.tune() preserves weights that were loaded or modified in memory instead of rebuilding each tuning iteration from the original model path.
📈 Small dataset fractions no longer discard all data (PR #26049)
Positive sampling fractions now retain at least one image, avoiding dataset-loading failures caused by rounding very small fractions down to zero. Explicit zero splits remain supported.
⚡ Faster single-image preprocessing (PR #25989)
Single-frame NumPy inputs bypass an unnecessary stacking operation, improving preprocessing speed while retaining batched-input behavior.
🧹 Updated documentation and package version
Inference and validation documentation now reflects the corrected model-state behavior, and the package version is bumped to 8.4.140.
Full Changelog: v8.4.139...v8.4.140
v8.4.139 improves training efficiency and reliability—especially by reducing validation memory usage—while refreshing model statistics, dataset handli
v8.4.139 improves training efficiency and reliability—especially by reducing validation memory usage—while refreshing model statistics, dataset handling, optimizer performance, and documentation. 🚀
Lower validation dataloader memory usage by @glenn-jocher:
Faster Muon/MuSGD optimizer updates:
Muon optimizer and related dead code were removed; MuSGD remains available. ⚡More reliable EMA checkpoints:
torch.compile.Correct semantic segmentation dataset detection:
masks/ directory is now recognized as PNG-mask semantic segmentation when no masks_dir entry is specified.Improved FP16 behavior documentation:
quantize=16 can cast a retained PyTorch model in place.Refreshed FLOPs and parameter reporting:
More robust and maintainable internals:
Expanded task visibility and documentation:
masks/ folder now behave as documented, avoiding incorrect output channels and class definitions.Full Changelog: v8.4.138...v8.4.139
Ultralytics v8.4.138 is a stability-focused release that restores compatibility with older checkpoints, fixes YOLO-World loading, and improves trainin
Ultralytics v8.4.138 is a stability-focused release that restores compatibility with older checkpoints, fixes YOLO-World loading, and improves training, inference, tracking, tuning, and documentation reliability. 🛠️
Legacy checkpoint loading fixed 🎯
Restricted checkpoint loading now recognizes loss and assignment classes stored in checkpoints created before 8.4.95, including detection, classification, pose, segmentation, rotated-box, and keypoint-related components.
YOLO-World loading fixed 🌍
Corrects package loading issues affecting YOLO-World models and related checkpoints.
MuSGD with channels_last no longer crashes ⚡
Replaces an incompatible tensor flattening operation with a layout-safe one, allowing CUDA training with the MuSGD optimizer and automatic channels_last support to run correctly.
SAM feature extraction optimized 🧠
SAM, SAM2, and SAM3 image embeddings are computed in inference mode, preventing unnecessary autograd graphs from being retained and reducing memory overhead during repeated inference.
Distributed training made more robust 🔧
Tuning results corrected 📈
Multi-dataset tuning now preserves dataset names and iteration order across distributed workers. Failed datasets are recorded with zero metrics instead of incomplete results, keeping tuning histories and fitness plots consistent.
Classification inference preprocessing improved 🚀
More preprocessing work is moved to the inference device and performed in batches, which can reduce CPU overhead and improve classification throughput.
BoT-SORT tracking made faster 🏃
Global motion compensation now caps corner detection at 400 points instead of 1,000, reducing optical-flow computation while retaining sufficient information for motion estimation.
Precision and quantization documentation clarified 📚
Documentation now explains that quantize may select or request different runtime precisions depending on the export format. This avoids implying that every backend supports the same FP16, FP32, or INT8 behavior.
Documentation quality updates ✨
Markdown tables were consistently formatted, the OBB navigation label was cleaned up, the TrackZone video was updated, and the YOLO26 CPU speed comparison now clearly identifies its YOLO26n-versus-YOLO11n ONNX baseline and hardware.
channels_last.quantize=16 or quantize=32 produces the same computation precision across PyTorch, ONNX, OpenVINO, NCNN, MNN, Triton, and other backends. 🔍OBB Dataset in the docs navigation bar by @Laughing-q in #26019quantize selects for each format by @raimbekovm in #26022Full Changelog: v8.4.137...v8.4.138
v8.4.137 automatically enables the faster channels-last memory layout for CUDA training on PyTorch 1.11+, improving GPU training performance while pre
v8.4.137 automatically enables the faster channels-last memory layout for CUDA training on PyTorch 1.11+, improving GPU training performance while preserving clear opt-out and compatibility options. 🚀
channels_last=None setting now automatically uses the NHWC memory format for CUDA training with PyTorch 1.11 and newer.channels_last=None: Automatically selects channels-last when supported.channels_last=False: Explicitly keeps the traditional NCHW format.channels_last=True: Explicitly requests channels-last, preserving previous behavior.channels_last setting is now included among the training options that can be updated when resuming a run.channels_last=False.Full Changelog: v8.4.136...v8.4.137
Version 8.4.136 improves hyperparameter tuning, inference performance, backend compatibility, and data handling—making YOLO workflows more reliable an
Version 8.4.136 improves hyperparameter tuning, inference performance, backend compatibility, and data handling—making YOLO workflows more reliable and efficient. 🚀
🎯 Smarter hyperparameter tuning — Current PR #25996 by @glenn-jocher
🧠 More robust channels-last inference
AutoBackend now owns memory-layout selection during backend construction, avoiding duplicated or unsafe conversions.⚡ Faster image preprocessing
🔍 More reliable prediction filtering
classes filter for YOLOE and World models when class IDs are supplied numerically.📷 Improved image and visualization handling
🏃 Tracking and distributed tuning improvements
🧪 Better validation and project maintenance
Full Changelog: v8.4.135...v8.4.136
v8.4.135 improves detection reliability by adapting max_det to dataset object counts and standardizing dataset fraction handling.
v8.4.135 improves detection reliability by adapting max_det to dataset object counts and standardizing dataset fraction handling.
🚀 Smarter max_det selection for detection, segmentation, pose, and OBB tasks
max_det is too low, it is automatically increased to match the observed dataset maximum.max_det values are preserved, but a warning is shown when they may limit validation recall.⚠️ Clearer warnings for object-count mismatches
max_det allows.max_det may increase validation cost, and cannot exceed the model or export format’s own capacity.📏 Consistent fraction boundary behavior
fraction=1 and fraction=1.0 now both mean “use the full dataset.”1 continue to represent an image count.0 and 0.0 remain available for skipping an optional test split.fraction=True are now rejected instead of being interpreted ambiguously.📚 Documentation and validation updates
max_det to observed data helps prevent artificially low validation recall.1 as an integer or 1.0 as a float.max_det settings continue to work, with warnings when they may restrict results.Full Changelog: v8.4.134...v8.4.135
v8.4.134 makes crowded-object training more resilient to GPU memory limits and increases the default hyperparameter search budget from 10 to 300 trial
v8.4.134 makes crowded-object training more resilient to GPU memory limits and increases the default hyperparameter search budget from 10 to 300 trials. 🚀
Faster TaskAlignedAssigner OOM recovery 🧠
Lower memory usage during target assignment 💾
Improved geometric and assignment processing ⚙️
Hyperparameter tuning now defaults to 300 trials 🔍
Tuner usage, and Ray Tune all now use 300 trials by default, up from 10.Version update 📦
8.4.134.Much better recovery from GPU out-of-memory errors ✅
Large or highly crowded batches can continue training without falling back to a very slow full-CPU assignment. In the reported xView benchmark, optimized single-image GPU assignment used about 1.56 GB of peak memory and completed in 0.376 seconds, compared with 194.86 seconds and 85.62 GB for the previous GPU-to-CPU fallback.
More practical training on dense datasets 🏙️
The changes are particularly valuable for aerial imagery, crowd analysis, and other datasets containing many objects per image.
No need to reduce the model’s forward batch size after recovery 📈
The fallback is isolated to target assignment, helping retain the intended training configuration while handling temporary memory pressure.
More effective automatic hyperparameter tuning 🎯
A 300-trial default gives the tuner substantially more opportunities to explore configurations and find stronger settings, especially for the broad YOLO26 search space.
Higher tuning cost unless overridden ⏱️
Users who rely on defaults should expect tuning jobs to run considerably longer and consume more compute. Smaller runs can still be requested explicitly when time or budget is limited.
Full Changelog: v8.4.133...v8.4.134
Ultralytics 8.4.133 improves hyperparameter tuning convergence, speeds up inference preprocessing, expands detection metrics, and simplifies edge-devi
Ultralytics 8.4.133 improves hyperparameter tuning convergence, speeds up inference preprocessing, expands detection metrics, and simplifies edge-device setup. 🚀
Smarter hyperparameter tuning — PR #25984 by @glenn-jocher
degrees or shear—to evolve more effectively.Faster predictor preprocessing — PR #25982 by @jahsef ⚡
Automatic channels-last CPU inference — PR #25983 by @JESUSROYETH
channels_last=True remains available for supported CPU and CUDA paths.More accurate INT8 calibration subsets — PR #25978 by @JESUSROYETH
fraction handling during classification and detection INT8 export calibration.Size-specific mAP for custom detection datasets — PR #25981 by @fcakyon 📈
save_json=True.Simpler edge-device installation
ultralytics package instead of the larger [export] extra.Improved Weights & Biases artifact control — PR #25985 by @glenn-jocher
save argument.save=False skips uploading the best checkpoint while retaining metrics and plots.save=True.Package update
save setting.[export] from edge-device install guides by @Y-T-G in #25977Full Changelog: v8.4.132...v8.4.133
🚀 Ultralytics v8.4.132 improves dataset efficiency, hardware compatibility, export workflows, and experiment-tracking documentation, with the headline
🚀 Ultralytics v8.4.132 improves dataset efficiency, hardware compatibility, export workflows, and experiment-tracking documentation, with the headline feature being finer control over test-split downloads.
🎯 Test-split control with fraction (PR #25966 — @fcakyon)
fraction now supports train, validation, and test splits.0 to skip test-image downloads entirely.⚡ More efficient NDJSON workflows
🩹 Corrected NMS and validation on Ascend NPU
📚 Improved YOLO26 end-to-end and export guidance
🔄 Broader support for exported non-YOLO models
YOLO() when task and imgsz are supplied explicitly.🧪 Better classification subset sampling
📈 Updated experiment-tracking integrations
🧩 Additional training and export corrections
fraction usage.Full Changelog: v8.4.131...v8.4.132
Ultralytics v8.4.131 adds Apple Core AI export and inference support for YOLO26, alongside important validation, training, model-configuration, and do
Ultralytics v8.4.131 adds Apple Core AI export and inference support for YOLO26, alongside important validation, training, model-configuration, and documentation improvements. 🚀
🍎 Apple Core AI export and inference
model.export(format="coreai") or the equivalent CLI command..aimodel asset format, which can be loaded again with YOLO("yolo26n.aimodel").export-coreai dependency group.⚡ Core AI deployment options
end2end=False produces raw predictions and can significantly reduce inference latency when post-processing is handled on the host..aimodel asset.✅ More reliable validation with split=train
🧮 Correct YOLO26 loss terminology
l1_loss from dfl_loss used by models with distribution-based box regression.🎯 Improved model configuration handling
parse_model now match exact scale letters, preventing unscaled or dictionary-based configurations from taking the wrong architecture branch.C3k2 configurations without an explicitly provided optional argument no longer fail for medium, large, or extra-large variants.🔤 YOLOE class reordering fixes
YOLOE.set_classes() now recognizes class-order changes and regenerates prompt embeddings when necessary.⚖️ Training robustness improvements
📟 Better progress bars in notebooks and narrow terminals
📚 Documentation and presentation updates
increment_path behavior.split=train, because training-time augmentation is no longer applied accidentally.dfl loss description for DFL-free YOLO26 by @raimbekovm in #25941Full Changelog: v8.4.130...v8.4.131
Version v8.4.130 makes dataset subset selection far more flexible and efficient, while improving tuning, tracking guidance, and dataset metadata. 🚀
Version v8.4.130 makes dataset subset selection far more flexible and efficient, while improving tuning, tracking guidance, and dataset metadata. 🚀
Count-based dataset limits 🎯
fraction now accepts a positive image count, such as fraction=1000, to train on exactly 1,000 images.fraction=[1000, 100] to limit the training and validation splits independently.fraction=0.1 still uses 10% of the dataset.1 means one image, while float 1.0 means the complete split.More efficient NDJSON and Platform dataset downloads ⚡
Improved hyperparameter tuning 🧠
Model.tune() now defaults to AdamW unless another optimizer is explicitly selected.Clearer tuning fitness plots 📈
tune_fitness.png now shows overall fitness progression, the best result achieved so far, and initial-versus-best fitness for each dataset.Expanded and clarified tracking documentation 🎥
More complete dataset license metadata 📚
fraction behavior.Full Changelog: v8.4.129...v8.4.130
v8.4.129 improves multi-dataset hyperparameter tuning, training precision, model export acceleration, data loading, and platform reliability—without i
v8.4.129 improves multi-dataset hyperparameter tuning, training precision, model export acceleration, data loading, and platform reliability—without introducing a new model architecture. 🚀
Multi-dataset tuning is now managed by MultiTrainer (PR #25937, @glenn-jocher):
Added BF16 mixed-precision training (PR #25931, @artest08):
amp now accepts True, False, "fp16", "bf16", and "fp32".Improved YOLO26 LiteRT exports for GPU delegates (PR #25914, @Y-T-G):
Faster image and FastSAM preprocessing (PRs #25935 and #25938, @JESUSROYETH):
More consistent ONNX CPU benchmarking (PR #23924, @Laughing-q):
ONNXBackend.Stronger export and validation coverage:
Improved Windows image compatibility (PR #21070, @Laughing-q):
imread_unicode for image paths containing non-ASCII characters.More reliable progress reporting for Ultralytics Platform (PR #25905, @Y-T-G):
Documentation and training guidance corrections:
l1_loss, pretrained-weight behavior, AutoBatch rules, AMP behavior, freezing, fine-tuning, K-Fold workflows, model YAML construction, and tuning output paths.Build reliability improvements for Jetson (PR #25930, @glenn-jocher):
More dependable distributed tuning: Multi-dataset experiments are now simpler to maintain, better isolated, and easier to analyze because results and output paths are tracked per dataset. MongoDB workers are also less likely to race when initializing defaults. 📈
More training options: Users with compatible CUDA hardware can choose BF16 for a practical balance of speed, memory use, and numerical stability. Existing FP16 and FP32 behavior remains available.
Faster edge deployment: LiteRT exports of YOLO26 end-to-end models should make better use of WebGPU and other GPU delegates, reducing unnecessary CPU fallback and potentially improving inference latency. ⚡
Better throughput for large workloads: Parallel image decoding and FastSAM preprocessing can reduce time spent waiting for CPU preprocessing, especially with large batches or many candidate crops.
Improved compatibility and confidence: Non-ASCII Windows paths, multi-input ONNX models, TensorRT validation, and distillation checkpoint handling receive targeted fixes that reduce failures in real-world workflows.
Cleaner integrations: Platform and other log consumers can process progress updates directly instead of guessing which log lines represent progress bars.
Release status: The package version is updated to 8.4.129.
imread_unicode for non-ASCII paths to prevent downstream imread overriding for Windows by @Laughing-q in #21070ONNXBackend in ProfileModels for CPU speed benchmark by @Laughing-q in #23924Full Changelog: v8.4.128...v8.4.129
v8.4.128 improves OpenVINO reliability and batch inference, reduces RAM use during training, strengthens export behavior, and clarifies dataset, augme
v8.4.128 improves OpenVINO reliability and batch inference, reduces RAM use during training, strengthens export behavior, and clarifies dataset, augmentation, and tracking workflows. 🚀
LATENCY performance hint. Throughput implementations remain available internally, but mode selection is forced to latency-oriented execution to avoid hangs in AsyncInferQueue, particularly for dynamic INT8 batches on CPU systems.cache='ram' 💾RANK and LOCAL_RANK values unless a real multi-process environment is detected.RANK and LOCAL_RANK variables are set outside of DDP context by @Y-T-G in #22724cache='ram' memory leak via shared image buffer by @raimbekovm in #24673Full Changelog: v8.4.127...v8.4.128
🚀 v8.4.127 makes exported YOLO models reliably load with the correct task and model family, while improving deployment stability, training recovery, a
🚀 v8.4.127 makes exported YOLO models reliably load with the correct task and model family, while improving deployment stability, training recovery, and dataset documentation.
Correct task detection for exported models across all 20 formats by @artest08
Improved OpenVINO inference reliability
Safer training resume behavior
pretrained model.last.pt.More robust result handling
Expanded CoreML export support
nms=True now works for segmentation and pose exports in addition to detection.Tracking and YOLOE fixes
Checkpoint and training reproducibility improvements
More direct Ultralytics Platform dataset access
Documentation media delivery improvements
Full Changelog: v8.4.126...v8.4.127
v8.4.126 makes restricted checkpoint loading safer and up to 40% faster in concurrent environments, while simplifying RLE loss calculations and preser
v8.4.126 makes restricted checkpoint loading safer and up to 40% faster in concurrent environments, while simplifying RLE loss calculations and preserving backward compatibility.
🔒 Thread-safe restricted checkpoint loading
⚡ Faster restricted model loading
🛡️ Improved compatibility with secure loading
🧠 Simplified RLE prior calculation
🏷️ Version update
Full Changelog: v8.4.125...v8.4.126
Ultralytics 8.4.125 makes YOLO26 model startup significantly faster—up to 46% faster in fresh Python processes—while improving dependency loading and
Ultralytics 8.4.125 makes YOLO26 model startup significantly faster—up to 46% faster in fresh Python processes—while improving dependency loading and documentation consistency. 🚀
from ultralytics import YOLO import also became faster, improving from 558 ms to 471 ms.torch.nn checkpoint classes.8.4.124 to 8.4.125.Full Changelog: v8.4.124...v8.4.125
🚀 Ultralytics v8.4.124 restores reliable dynamic-size inference for exports with embedded NMS, while improving training stability, deployment compatib
🚀 Ultralytics v8.4.124 restores reliable dynamic-size inference for exports with embedded NMS, while improving training stability, deployment compatibility, performance, and documentation.
Dynamic NMS exports restored — PR #25874
dynamic=True and nms=True once again support runtime image heights and widths.max_det limit instead of using the number of anchors from the export image size.Improved CoreML attention export
mlprogram exports now use a more compatible attention implementation to avoid GPU compilation crashes on recent Apple systems.More efficient RT-DETR training and TensorRT inference
Training reliability fixes
Prediction and result-processing improvements
predict() or track() no longer carries filters such as classes, max_det, or NMS settings into later calls.Results.save_txt(), save_crop(), and summary() are consolidated, reducing per-object synchronization overhead.Export and platform updates
max_det limit.perf_counter() for ProfileModels and time_sync() latency measurement by @raimbekovm in #25845Results.save_txt/save_crop/summary by @JESUSROYETH in #25842copy_paste candidate fraction by @raimbekovm in #25810autobatch() when no candidate batch size fits by @JESUSROYETH in #25739seed reach dataloader workers so augmentations vary between runs by @yentur in #25815Full Changelog: v8.4.123...v8.4.124
Your coding agent can read these notes before it upgrades. Set up the MCP server →