NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4303 most downloaded on PyPI
A set of easy-to-use utils that will come in handy in any Computer Vision project
Last release today
04 Oct 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 46 of 51 stable releases
Nothing withdrawn
no release was ever pulled
4 years old
111 releases · first in 2023
One column per quarter.
Nothing published for this version
…a notebook dead-link fix. No new public API, no breaking changes. The broadest-reach fix rewrites pillow_to_cv2 , the entry point every annotator exce…
supervision 0.30.6 is a patch release closing 12 library correctness bugs and landing 2 behavior refinements across datasets, annotators, key points, line-crossing counting, and the Pillow-based fallback backend, plus a new docs guide and a notebook dead-link fix. No new public API, no breaking changes. The broadest-reach fix rewrites pillow_to_cv2, the entry point every annotator except RichLabelAnnotator and several image helpers use to convert a Pillow image: it used to pass raw mode bytes straight through, so a 1-bit mask, a 16-bit depth map, an LA/PA image, or a CMYK JPEG all drew wrong or crashed. Two fixes close silent-wrong-data bugs rather than crashes: LineZone.trigger let unconfirmed tracks inflate crossing counts, and DetectionDataset.as_yolo/.as_pascal_voc silently dropped box-only objects from a mixed polygon/box COCO dataset. DetectionsSmoother no longer crashes when two tracked objects first seen on different frames carry different metadata (e.g. RF-DETR's source_image); its own docstring pipeline was broken. Every library fix ships a regression test.
sv.pillow_to_cv2 now converts every Pillow mode the way cv2.imread wouldRaw mode bytes used to pass straight through: a 1-bit mask came back as 0/1 instead of 0/255, a 16-bit depth map wrapped modulo 256 and drew as noise, an LA/PA image crashed in cvtColor on its 2-channel array, and a CMYK JPEG drew its cyan/magenta/yellow ink as red/green/blue. This function is the entry point for every annotator except RichLabelAnnotator (which stays on the Pillow path) plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image whenever handed a Pillow image, so the bug reached the whole drawing surface, not just one call site. A single-channel grayscale image also drew from a read-only buffer view before this fix. Annotating into it silently failed; it now hands back a writable copy.
from PIL import Image
depth_map = Image.open("depth.tif") # mode "I;16"
scene = sv.pillow_to_cv2(depth_map)
# clipped to the 16-bit range and keeps its high byte, instead of wrapping mod 256RGB, RGBA, grayscale, and palette images are unchanged.
sv.LineZone.trigger no longer lets unconfirmed tracks inflate countsDetections with a negative tracker_id (how trackers like ByteTrackTracker mark an unconfirmed track) all keyed under one shared id. Two distinct unconfirmed objects crossing on opposite sides in the same frame read as one track oscillating, silently inflating in_count/out_count with crossings no confirmed track ever made. Confirmed tracks (tracker_id >= 0) are counted exactly as before.
sv.DetectionDataset.as_yolo / .as_pascal_voc stop dropping box-only objectsfrom_coco gives an all-zero mask to any annotation without its own segmentation when a sibling annotation has one, and to every annotation when force_masks=True. The exporters only ever wrote the polygon traced from a mask, so a mixed polygon/box COCO dataset silently lost its box-only objects on export, and a box-only COCO dataset loaded with force_masks=True wrote empty YOLO label files or object-less Pascal VOC files. A detection with an empty or contour-less mask now exports as its bounding box instead.
sv.DetectionsSmoother no longer crashes on a second tracked objectEach smoothed track was built on the oldest frame in its window, so two objects first seen on different frames carried different frames' metadata (exactly what the RF-DETR / inference connectors attach), and Detections.merge rejected the mismatch. class_id and data (e.g. class_name) now follow the current frame instead of lagging up to length - 1 frames behind a class change; xyxy, confidence, and oriented-box corners are still averaged as before.
sv.DetectionDataset.split, sv.ClassificationDataset.split (both take split_ratio), and the internal train_test_split helper they call (takes train_ratio) never validated the ratio. A finite out-of-range value looked like a successful split: for 10 images, a ratio of -0.2 returned 8/2 via negative slicing, and 1.2 (or an accidental percentage like 80) returned 10/0 with no held-out data, no error either way.
train, test = dataset.split(split_ratio=80) # meant 80%
# now raises ValueError naming the problem, instead of silently returning
# every image as train and none as test0 and 1 keep their existing meanings. Not an API break: no signature changed, only previously-silent input now raises.
No breaking changes in this release.
No deprecations or removals landed in 0.30.6 either. All scheduled remove_in markers in the codebase target 0.31.0 or 0.32.0 and are untouched by this patch release.
Two entries change output values rather than only fixing a crash or a silently-wrong count; not API breaks, but worth checking if your code depends on the old values:
sv.KeyPoints.from_ultralytics / .from_inference / .from_detectron2: as_detections().confidence now carries the model's own detection score, not the mean of the per-keypoint confidences.sv.DetectionsSmoother: class_id and data fields on a smoothed track now follow the current frame instead of lagging behind a class change.Annotators / key points
sv.VertexEllipseHaloAnnotator now draws the whole halo of a key point whose covariance ellipse is not horizontal. A vertical or diagonal ellipse used to be clipped to a thin band because the fade box was sized as if the major axis always ran along the image x axis. Horizontal ellipses are drawn exactly as before. (#2634)sv.KeyPoints.from_ultralytics, .from_inference and .from_detectron2 now keep each object's detection score as detection_confidence instead of dropping it, which previously made sv.KeyPoints.with_nms raise ValueError and as_detections() report the mean keypoint confidence instead of the model's own score. (#2633)sv.LabelAnnotator now sizes the label background correctly for a label containing a blank line. Text drew each blank line at full height while the background measured it as zero, pushing the last line outside the box. Labels without blank lines are unchanged. (#2632)sv.IconAnnotator now draws a palette icon that carries its own alpha channel (Pillow PA mode, storable in TIFF) instead of raising ValueError. Palettes without alpha, and every read with OpenCV installed, are unchanged. (#2624)Datasets
sv.DetectionDataset.as_yolo / .as_pascal_voc now write a detection whose mask is empty or has no valid contour as its bounding box instead of silently leaving it out of the label file. (#2631)sv.DetectionDataset.from_yolo now loads a label row carrying a trailing confidence or tracker id (written by Ultralytics save_txt with save_conf=True or tracking on) instead of raising ValueError, for both box and segmentation rows. The extra token is ignored; box and polygon geometry are unchanged. An odd-length polygon row with no extra field now loses its last value rather than raising, since coordinate parity is the only signal separating the two cases. (#2619, #2626, #2635)Tracking
sv.DetectionsSmoother no longer raises ValueError: Conflicting metadata when a second tracked object, first seen on a different frame, enters the smoothing window. (#2628)sv.LineZone.trigger now ignores detections with a negative tracker_id (the value trackers such as ByteTrackTracker report for an unconfirmed track) instead of letting them silently inflate in_count/out_count. (#2623)Detection utils
sv.polygon_to_mask now accepts a list/tuple/array-like of [x, y] vertices (not just a NumPy array) and returns an all-zero mask for an empty or under-3-vertex polygon instead of crashing inside OpenCV/NumPy with an opaque error. A malformed polygon now raises ValueError naming the problem. (#2622)sv.Detections.from_sam3 now keeps SAM 3 PVS contour fragments with fewer than 3 vertices instead of dropping them. Single points and 2-point edges are rasterized directly into the mask and bounding box. (#2625)Image / IO
sv.pillow_to_cv2 now converts every Pillow mode to the 8-bit array cv2.imread would produce, reaching every annotator plus crop_image, resize_image, letterbox_image, scale_image, tint_image, grayscale_image, and plot_image. (#2614)sv.CSVSink now writes UTF-8 on every platform, preserving non-English detection labels and custom fields on Windows. (#2615)Docs
quickstart.ipynb, annotate-video-with-detections.ipynb, underestand-visitors-with-yolo-world.ipynb). (#2630)sv.DetectionDataset.split, sv.ClassificationDataset.split, and the internal train_test_split helper they call now reject a ratio outside [0, 1] (NaN and ±inf included) with a ValueError naming the problem, instead of silently mis-splitting or dropping the held-out set. (#2611)sv.tint_image and sv.grayscale_image now accept a single-channel (H, W) array or grayscale Image, as sv.letterbox_image already did. (#2614)docs/metrics/mean_average_precision.md, cross-linked from docs/how_to/benchmark_a_model.md. Detections.from_transformers gained a docstring note that it expects post-processed predictions, not raw DataLoader targets. (#2620)DetectionsSmoother conflicting-metadata crash, the KeyPoints detection-score fix, the YOLO/Pascal VOC empty-mask export fix, the VertexEllipseHaloAnnotator rotated-ellipse fix, the LabelAnnotator blank-line background fix, and the polygon_to_mask crash fixpillow_to_cv2 rewrite for every Pillow modeLineZone unconfirmed-track fixIconAnnotator fixfrom_sam3 degenerate-fragment fixFull changelog: 0.30.5...0.30.6
…drawing/loading — no new public API of note, no breaking changes. The most consequential fixes are silent, not crashes: LineZone.trigger miscounted cr…
supervision 0.30.5 is a patch release closing 11 correctness and crash bugs across line-crossing tracking, mAR@K scoring, model connectors, key points, and image drawing/loading — no new public API of note, no breaking changes. The most consequential fixes are silent, not crashes: LineZone.trigger miscounted crossings by one per flicker whenever a tracker briefly touched the far side of the line, and MeanAverageRecall scored mAR@K against the wrong predictions when a lower-ranked one fit a target more tightly than one within the top K. The remaining fixes close a hard crash in InferenceSlicer on conflicting slice metadata (plus two related Detections.__eq__ bugs), bring the OpenCV-free fallback backend to parity with OpenCV for rotated videos, CMYK images, transparent/1-bit PNGs, and default JPEG/WebP write quality, fix two drawing bugs that reproduce with OpenCV installed (16-bit draw_image, grayscale IconAnnotator icons), and fix two isolated bugs in KeyPoints.as_detections and plot_images_grid. Every fix ships a regression test.
sv.LineZone.trigger no longer counts flicker as a crossingA crossing was confirmed whenever the oldest entry of a minimum_crossing_threshold + 1 frame history differed from every later entry, which never verified the tracker had actually settled on the side it supposedly came from. With minimum_crossing_threshold=2 the side sequence A,A,A,B,A,A,A counted a crossing into A — the side the object never left — so counts drifted by one per flicker, in the wrong direction; three separated flickers gave in_count=3 instead of 0.
line_zone = sv.LineZone(
start=sv.Point(0, 0), end=sv.Point(0, 100), minimum_crossing_threshold=2
)
# a tracker that flickers to the far side for one frame and back
# no longer registers a phantom crossing; only a sustained crossing countsCrossings are now measured against the last side a tracker was confirmed on, not against the oldest history entry. Sustained crossings and minimum_crossing_threshold=1 (the default) are unchanged.
sv.metrics.MeanAverageRecall now scores mAR@K from each image's own top KThe matcher pairs predictions to targets by highest IoU, not confidence, so a prediction ranked below K could take a target away from one ranked within it. Adding a low-confidence duplicate that fit a target more tightly than the top prediction actually lowered mAR@1 — one target with a top prediction at IoU 0.71 scored mAR@1 0.5 alone but 0.0 once a second, unrelated prediction at IoU 1.0 and confidence 0.1 was added. Each detection limit now matches only its own top K predictions.
sv.InferenceSlicer no longer crashes when slices disagree on metadataRF-DETR and inference-package connectors attach a source_image array per slice; any mismatch across slices previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped with a one-time warning naming the dropped keys; source_image is reattached afterward as the full input image.
slicer = sv.InferenceSlicer(callback=callback)
detections = slicer(image) # metadata.source_image is the full input image againFixing this also closed two Detections.__eq__ bugs: NaN in a float-array metadata value now compares equal to itself, and a list-valued value against an ndarray-valued one now returns False instead of raising ValueError.
OpenCV's FFmpeg backend applies a video's display-rotation matrix to decoded frames and to the reported width/height; the OpenCV-free fallback ignored it, so a portrait phone video came back sideways with its width/height swapped.
sv.KeyPoints.as_detections drops invisible key points from the derived boxas_detections stretched each box over every key point that wasn't [0, 0] or non-finite, ignoring visible — unlike with_nms, which already respected it. A pose whose low-confidence joints were predicted off-frame no longer produces an inflated box.
No breaking changes in this release.
No deprecations or removals landed in 0.30.5 either. All scheduled remove_in markers in the codebase target 0.31.0 and are untouched by this patch release.
sv.config.SOURCE_IMAGE_METADATA_FIELD is a new public constant (added alongside the InferenceSlicer fix) naming the Detections.metadata key RF-DETR/inference-package connectors use for the source image — additive, no existing code needs to change.
Tracking / metrics
sv.LineZone.trigger no longer counts a spurious crossing in the opposite direction when a tracker flickers to the far side of the line for fewer than minimum_crossing_threshold frames. Crossings are now measured against the last side a tracker was confirmed on, and once a tracker has a confirmed side, a new one replaces it only after being held for the full threshold. Sustained crossings, minimum_crossing_threshold=1 (the default), and per-tracker isolation are unchanged. (#2600)sv.metrics.MeanAverageRecall now scores mAR@K from each image's K most confident predictions alone, instead of matching every prediction first and only then keeping the top K by confidence. mAR@1 and mAR@10 now equal the recall of the top 1 and top 10 predictions per image, as documented; mAR@100 changes only for images with more than 100 predictions. (#2604)Model connectors
sv.InferenceSlicer no longer raises when slices disagree on metadata (e.g. a source_image NumPy array attached per-slice by RF-DETR/inference-package connectors), which previously crashed the merge outright. Any metadata key that isn't identical across every slice is now dropped from the merged result, with a SupervisionWarnings warning naming the dropped keys, emitted once per slicer instance. source_image is a special case: removed from each slice before the lenient merge, then reattached afterward as the full input image. Fixing this also closed two related Detections.__eq__ bugs: metadata holding NaN in a float array now compares equal to itself instead of always reading unequal, and comparing a list-valued metadata value against an ndarray-valued one now returns False instead of raising ValueError. (#2596)OpenCV-free fallback backend
sv.VideoInfo.from_video_path, sv.get_video_frames_generator and sv.process_video now turn a rotated video upright when OpenCV is not installed, as they already do with OpenCV. The fallback now applies quarter and half turns, the same angles OpenCV applies, using the container's display-rotation matrix. Videos without a display rotation, and every read with OpenCV installed, are unchanged. (#2601)sv.IconAnnotator and sv.draw_image now draw CMYK JPEG and TIFF images in their colors when OpenCV is not installed, instead of returning the four ink channels as if they were blue/green/red/alpha. Other images, and every read with OpenCV installed, are unchanged. (#2602)sv.ImageSink and the dataset exports that encode in-memory images now write JPEG and WebP files at OpenCV's default quality (JPEG 95, lossless WebP) when OpenCV is not installed, instead of Pillow's defaults (JPEG 75, lossy WebP). The fallback's in-memory encoder also now accepts .jpe, .tif, .jp2, and .pgm, which it previously rejected even though file writes already accepted them. PNG and every write with OpenCV installed are unchanged. (#2592)sv.IconAnnotator and sv.draw_image now keep the transparency of grayscale PNGs with alpha, RGB PNGs with a transparent color, and 1-bit PNGs, when OpenCV is not installed — the fallback previously returned Pillow's pixel layout for IMREAD_UNCHANGED rather than OpenCV's, causing shape-mismatch ValueErrors or black-drawn transparent pixels. Other images, and every read with OpenCV installed, are unchanged. (#2588)Annotators / drawing
sv.draw_image now scales a 16-bit PNG down to 8 bits on load, instead of blending it into an 8-bit scene and clipping channel values above 255 — this reproduces with OpenCV installed, since IMREAD_UNCHANGED keeps a 16-bit PNG at 16 bits. Eight-bit images, and images passed as arrays, are drawn as before. (#2603)sv.IconAnnotator now draws grayscale PNG icons instead of failing with IndexError — the overlay only handled BGR and BGRA arrays, and cv2.imread(..., IMREAD_UNCHANGED) returns a 2-D array for a grayscale PNG without alpha, with or without OpenCV installed. A grayscale icon is now expanded to BGR on load, and a 16-bit icon is scaled to 8 bits, as cv2.imread does by default. Eight-bit color icons are drawn as before. (#2591)Key points / utils
sv.KeyPoints.as_detections now leaves key points marked not visible out of the box it derives for each skeleton, as sv.KeyPoints.with_nms already does. A skeleton with no visible key point is now dropped, as one with only missing key points already was. Key points without visible convert as before. (#2605)sv.plot_images_grid now plots a single image in a grid_size=(1, 1) grid, instead of failing with AttributeError — plt.subplots returns a lone Axes rather than an array for a 1x1 grid. Grids with more than one cell plot as before. (#2590)draw_image and grayscale IconAnnotator drawing fixes, plus the MeanAverageRecall/plot_images_grid/KeyPoints.as_detections fixesLineZone flicker-crossing fixInferenceSlicer conflicting-metadata fix and the two Detections.__eq__ bugs it uncoveredFull changelog: 0.30.4...0.30.5
…handling — no new public API, no breaking changes. The most consequential fixes are silent, not crashes: COCO polygon masks loaded shifted by up to a…
supervision 0.30.4 is a patch release closing 14 correctness and crash bugs across dataset loaders/exporters, model connectors, and key point/annotator/video handling — no new public API, no breaking changes. The most consequential fixes are silent, not crashes: COCO polygon masks loaded shifted by up to a pixel, EXIF-rotated photos loaded with swapped width/height across two loaders and both image backends, and non-ASCII class names were mangled or rejected on Windows across four loaders and one writer. The remaining fixes close hard crashes in the Transformers v4/v5 connectors, Ultralytics pose loading, TraceAnnotator, iterative_seek, and three more dataset-format edge cases. Every fix ships a regression test.
sv.DetectionDataset.from_coco no longer shifts mask polygons by up to a pixelCOCO polygons commonly hold sub-pixel float coordinates, but the loader cast them straight to int32, which truncates rather than rounds. Every polygon mask loaded shifted up and to the left by up to a pixel — the same polygon loaded one pixel apart depending on the format it was stored in (a square with corners at 2.6/7.6 covered pixels 2–7 from COCO but 3–8 from LabelMe, an IoU of 0.53 between the two masks). Silent, not a crash — training data was quietly misaligned.
dataset = sv.DetectionDataset.from_coco(
images_directory_path="images",
annotations_path="annotations.json",
) # polygon vertices now rounded to nearest pixel, matching from_yolo/from_labelme/from_pascal_vocA non-finite vertex now raises ValueError naming the annotation id, instead of silently propagating.
from_yolo/as_coco no longer swap width and height on EXIF-rotated photosPhotos from phones are often stored sideways with an EXIF orientation tag. cv2.imread applies that tag, but the size read used Pillow's file-header read, which doesn't. For a quarter-turned photo, from_yolo scaled normalized boxes and polygons by the swapped width/height, so they landed outside the image, and as_coco wrote the swapped dimensions. The OpenCV-free fallback backend now also applies the orientation tag, matching how OpenCV itself handles every read except IMREAD_UNCHANGED — the same file previously loaded with a different shape depending on whether opencv-python was installed.
from_coco, from_labelme, from_createml, from_yolo read JSON/YAML, and as_pascal_voc wrote XML, using the platform's default encoding — cp1252 on Windows. A class name like café loaded as café; 고양이 raised UnicodeDecodeError. All five now read/write UTF-8 explicitly.
sv.Detections.from_transformers now loads Mask2Former/MaskFormer overlap-safe binary mapspost_process_instance_segmentation(return_binary_maps=True) — the option Transformers recommends when instances can overlap — returns a (num_instances, H, W) stack of binary maps, but the v5 instance path compared it against each segment's id as if it were an id-map, producing a 4-D array mask_to_xyxy rejected outright. Each segment now indexes the stack at its own id, so overlapping instances keep their full masks.
detections = sv.Detections.from_transformers(
transformers_results=processor.post_process_instance_segmentation(
result, target_sizes=[image.size[::-1]], return_binary_maps=True
)[0],
id2label=model.config.id2label,
) # overlapping-instance results now load instead of crashingsv.DetectionDataset.from_pascal_voc no longer silently drops bmp/tif/webp imagesThe loader only listed .jpg, .jpeg, .png, so every other image was left out without a warning — even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them. A dataset exported to Pascal VOC and read back came back smaller than it went out.
No breaking changes in this release.
No deprecations or removals landed in 0.30.4 either. Eight scheduled-removal commits (ByteTrack, supervision.keypoint, create_tiles/overlay_image, keypoint validators, the legacy MeanAveragePrecision, LMM/from_lmm, and others) plus a new MetricResult ABC exist on develop but are not part of this cherry-picked patch release — they target 0.31.0 and will get their own migration guide when that release ships.
Two internal (non-underscored but unexported) helper functions changed signature as part of fixes in this release: parse_polygon_points (dataset/formats/pascal_voc.py) now returns float64 instead of int, and detections_from_xml_obj's xyxy is now list[list[float]] instead of list[list[int]]. Neither is re-exported from supervision.__init__ or supervision.dataset.__init__, so this is not a public API change — no action needed unless you import these paths directly.
Dataset loaders / exporters
sv.DetectionDataset.from_coco now rounds polygon vertices to the nearest pixel before rasterising masks, instead of truncating them to int32, as from_yolo/from_labelme/from_pascal_voc already do. Truncation shifted every mask up and to the left by up to a pixel — a polygon with corners at 2.6/7.6 covered pixels 2 to 7 from COCO but 3 to 8 from LabelMe, an IoU of 0.53 between the two masks. A non-finite vertex now raises ValueError naming the annotation id; integer vertices and RLE masks load as before. (#2587)sv.DetectionDataset.from_coco/from_labelme/from_createml/from_yolo now read JSON/YAML as UTF-8, and as_pascal_voc writes XML as UTF-8, instead of the platform default (cp1252 on Windows). A COCO category or YOLO data.yaml name outside ASCII broke on Windows: café loaded as café, 고양이 failed with UnicodeDecodeError, and as_pascal_voc either failed with UnicodeEncodeError or wrote a file from_pascal_voc rejected with ParseError: not well-formed (invalid token) — Linux/macOS, already UTF-8 by default, are unaffected. (#2585)sv.ClassificationDataset.as_folder_structure now copies source image files unchanged instead of round-tripping them through cv2.imread/cv2.imwrite, which dropped alpha channels, downcast 16-bit PNGs to 8-bit, and recompressed JPEGs on every export — matching sv.DetectionDataset's exports (as_yolo/as_pascal_voc/as_coco), which already copied; exporting into the folder the dataset was loaded from previously rewrote its own images the same lossy way. Images held in memory are still encoded with cv2.imwrite. (#2584)sv.DetectionDataset.from_labelme now finds images for LabelMe files saved on Windows when loading on Linux or macOS. imagePath is written with \, but the loader took only Path(...).name, which doesn't split on \ on POSIX — the whole value became the filename and reading failed with ValueError: Could not read image from path. The filename is now taken with either separator on every system, as LabelMe itself does when reading its own files; forward-slash paths, and every file on Windows, are unchanged. (#2581)sv.DetectionDataset.from_yolo now loads label files whose class ids are written as decimals (e.g. 1.0 0.5 0.5 0.2 0.4), which previously aborted the whole load with ValueError: invalid literal for int() — np.savetxt writes floats by default and Ultralytics tolerates them. Whole numbers load in any notation now; fractional, non-finite, or non-numeric ids still raise, naming the offending id. (#2580)sv.DetectionDataset.from_yolo/as_coco now size EXIF-oriented images the way cv2.imread loads them (swapping width/height for orientations 5 to 8), instead of reading the un-rotated file-header size via Pillow — quarter-turned photos previously scaled boxes/polygons by the swapped dimensions and produced mismatched mask shapes. The OpenCV-free fallback backend's imread/imdecode now apply EXIF orientation too, matching OpenCV's behavior for every read except IMREAD_UNCHANGED — previously the same file loaded with a different shape depending on whether opencv-python was installed. (#2577)sv.DetectionDataset.from_pascal_voc no longer fails on annotations with decimal coordinates (e.g. <xmin>48.5</xmin>), which aborted the whole load with ValueError: invalid literal for int() — Datumaro, which CVAT uses for its exports, writes VOC this way. Box coordinates are now read as floats and keep their precision; polygon vertices are rounded after the 1-index offset, as the YOLO and LabelMe loaders already do; non-finite values are still rejected. (#2568)sv.DetectionDataset.from_pascal_voc no longer skips .bmp, .tif, .tiff, and .webp images without a warning — the loader only listed .jpg/.jpeg/.png, even though from_yolo/from_folder_structure load those formats and as_pascal_voc writes annotations for them, so a dataset exported to Pascal VOC and read back came back smaller than it went out. It now accepts the same extensions as sv.ClassificationDataset.from_folder_structure. (#2569)Model connectors
sv.Detections.from_transformers now loads Transformers v5 return_binary_maps=True instance results (a (num_instances, H, W) stack), which the v5 path previously compared against each segment's id as if it were an id-map, producing a 4-D array that mask_to_xyxy rejected with ValueError: too many values to unpack (expected 3). Each segment now indexes the stack at its own id, keeping full masks for overlapping instances; segment-id-map results are unchanged. (#2576)sv.Detections.from_transformers no longer crashes on a Transformers v4 panoptic result with no segments — post_process_panoptic's empty segments_info produced a (0,) mask array instead of (0, H, W), and mask_to_xyxy raised ValueError: not enough values to unpack (expected 3, got 1). The path now builds a (0, H, W) mask stack and an integer class_id, matching the v5 paths, and yields empty Detections. (#2571)Key points / annotators / video
sv.KeyPoints.from_inference now places each key point at the slot given by its class_id (skeleton index) instead of appending in received order — Inference drops key points below keypoint_confidence and multi-skeleton models report different counts per object, so stacking as-received either raised ValueError: ... inhomogeneous shape or silently slid later key points into earlier slots, joining the wrong joints. Omitted slots now stay (0, 0) at zero confidence, already skipped by the key point annotators and as_detections; a result with every key point omitted now loads with zero key points instead of failing validation. (#2575)sv.KeyPoints.from_ultralytics no longer crashes on pose models whose key points carry no visibility score (kpt_shape=[K,2], where Results.keypoints.conf is None) — the connector called .cpu() on it unconditionally, raising AttributeError: 'NoneType' object has no attribute 'cpu' on every non-empty frame. Such results now load with keypoint_confidence=None; models that do report visibility are unaffected. (#2570)sv.TraceAnnotator.annotate no longer raises ValueError: Length of color lookup 3 does not match length of detections 2 when a custom_color_lookup is passed alongside pending (tracker_id == -1) tracks — the annotator skips those detections but still resolved colors against the full-length lookup, although the other annotators accept the same detections and lookup. The lookup is now filtered alongside the detections, so each confirmed track keeps its own color; frames without pending tracks, and calls without custom_color_lookup, are unchanged. (#2586)sv.get_video_frames_generator no longer reads start frames past end when iterative_seek=True — it counted start down to zero, then measured end from that zero instead of from start (start=2, end=5 yielded frames 2 to 6 instead of 2 to 4; start=4, end=6 yielded frames 4 to 9 instead of 4 to 6). A separate counter now keeps both seek modes aligned; start=0 and non-iterative calls are unchanged. (#2583)Full changelog: 0.30.3...0.30.4
0.30.3: Pose, VLM, and video/CSV crash and correctness fixes
supervision 0.30.3 is a bug-fix release closing crash and silent-correctness gaps across pose estimation, VLM parsing, video/CSV output, and geometry utilities. Non-finite key points — how pose estimators report an undetected joint — no longer produce duplicate poses that survive sv.KeyPoints.with_nms, or crash the key point annotators outright. sv.Detections.from_vlm now orders backwards box corners, closing a bug where such a box scored a false 0.0 IoU and both survived NMS as a duplicate and counted as a total miss in mAP. sv.TraceAnnotator and sv.CSVSink no longer crash or silently drop columns on the first frame with no detections — a case every non-ByteTrack tracker pipeline hits. sv.process_video no longer hangs forever when max_frames exceeds the video length. Continuing 0.30.2's numeric-correctness theme, sv.pad_boxes and sv.scale_boxes are fixed against integer overflow. No breaking API changes, no new public API.
sv.KeyPoints.with_nms tested key point validity with xy == 0 alone, and NaN — how pose estimators report an undetected joint — is not 0. The stale joint stayed in the NMS box, so a duplicate skeleton scored False on every IoU comparison against it and survived suppression. The key point annotators (sv.VertexAnnotator, sv.EdgeAnnotator, sv.VertexLabelAnnotator, the sv.VertexEllipse*Annotator family) had the matching crash: a single undetected joint raised ValueError: cannot convert float NaN to integer for the whole frame. Both now skip non-finite coordinates, matching sv.KeyPoints.as_detections.
keypoints = sv.KeyPoints(xy=xy, confidence=confidence)
keypoints.with_nms(
threshold=0.5
) # duplicate skeletons with a NaN joint are now suppressedsv.Detections.from_vlm no longer scores a false IoU miss on backwards box cornersA VLM that emits a corner pair backwards produced an xyxy row with x_min > x_max. Nothing downstream caught it: sv.box_iou_batch clamps intersection width at zero, so the box scored 0.0 IoU against itself — surviving NMS as a duplicate and counting as a total miss in mAP — while box_area still reported a plausible positive value. Every VLM parser now orders each box's corners before returning it.
sv.TraceAnnotator and sv.CSVSink no longer crash or silently corrupt output on an empty-detections framesv.TraceAnnotator.annotate raised ValueError: The tracker_id field is missing on the first frame with no detections, for every tracker except sv.ByteTrack. Such a frame now draws nothing and still advances the frame counter, so trace_length stays a window over elapsed frames rather than only over populated ones. sv.CSVSink had a quieter failure: an empty batch fixed the CSV header without the data/custom_data columns, and every later row was silently truncated to that schema — dropping fields like class_name for the whole file. The header is now fixed by the first batch that actually carries detections.
sv.process_video no longer hangs forever when max_frames exceeds the video lengthThe reader thread failed on the out-of-range end before enqueuing its sentinel, leaving the main loop blocked on the read queue indefinitely. max_frames is now capped at the video length, and any reader-thread error surfaces as RuntimeError("Reader thread raised: ...") instead of stalling the call.
sv.pad_boxes and sv.scale_boxesBoth computed intermediate values that could overflow or silently wrap for large integer coordinates (e.g. large int32/uint16/int64 boxes). Both now use overflow-safe arithmetic.
xyxy = np.array([[10, 20, 30, 40]], dtype=np.int64)
sv.pad_boxes(xyxy=xyxy, px=5, py=10) # int64 output, no wraparoundpad_boxes changes return dtype for integer input — see the migration guide below.
sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer crash on an extreme aspect ratio or tiny scale factorA small enough factor — or an aspect ratio too extreme for the target box — could round an output axis down to 0, and cv2.resize raised an assertion naming nothing the caller passed. Each axis now keeps at least one pixel. Two callers inherit the fix: sv.letterbox_image could not fill the resolution it was asked for, and sv.CropAnnotator with scale_factor < 1 aborted the whole frame as soon as one detection box was a few pixels across.
No breaking API changes. One fix changes return dtype for integer input:
sv.pad_boxes — integer xyxy now returns int64 (or float64 if a padded coordinate exceeds the int64 range), instead of the input's original integer dtype, which could silently overflow or wrap for small dtypes like int16/uint8.If your code assumes pad_boxes preserves the input's exact dtype (e.g. reusing the result as an int16 array), cast explicitly: sv.pad_boxes(...).astype(np.int16).
sv.scale_boxes also fixes an integer-overflow bug, but its return dtype was already float64 for integer input before this release — unaffected.
sv.Detections.from_ultralytics now assigns the placeholder class ID 0 to every mask in a masks-only result, instead of sequential IDs across masks that belong to the same image. (#2566)sv.pad_boxes now computes integer-coordinate padding without overflow or unsigned casting errors. (#2565)sv.scale_image and sv.resize_image(keep_aspect_ratio=True) no longer derive a zero-sized target, fixing a crash reached via sv.letterbox_image and sv.CropAnnotator. (#2564)sv.KeyPoints.with_nms no longer stops suppressing duplicate skeletons as soon as a key point is non-finite. (#2563)sv.tint_image no longer tints the caller's own image array in place. (#2562)sv.LineZone no longer consumes the triggering_anchors iterable during validation, so a generator or map passed in is no longer exhausted before the first trigger() call. (#2561)ValueError. (#2560)sv.PolygonZone now rejects a polygon with fewer than three vertices instead of building a zone that can never trigger; sv.Detections.from_vlm now orders each parsed box's corners. (#2554)sv.ClassificationDataset.as_folder_structure now rejects images that would overwrite the same class-relative filename before writing any files. (#2551)sv.process_video no longer hangs forever when max_frames is larger than the number of frames in the video. (#2546)sv.filter_polygons_by_area and sv.approximate_polygon now preserve local geometry for large-origin integer and float64 polygons. (#2542)sv.TraceAnnotator.annotate no longer raises on an empty-detections frame; sv.CSVSink no longer lets an empty batch fix the CSV header. (#2539)sv.scale_boxes now preserves exact integer intermediates, preventing overflow and scaled-corner rounding errors for large integer-coordinate boxes. (#2541)gh-pages write lock instead of racing each other. (#2536)KeyPoints.with_nms/key point annotators crashing on non-finite key points, tint_image image aliasing, LineZone generator exhaustion, and the scale_image/resize_image zero-target crashPolygonZone/from_vlm box-corner ordering and the TraceAnnotator/CSVSink empty-frame crashscale_boxes and dtype loss in filter_polygons_by_area/approximate_polygonfrom_ultralytics masks-only class ID sizingpad_boxesprocess_video hanging when max_frames exceeds the video lengthFull changelog: 0.30.2...0.30.3
0.30.2: Detection numeric-correctness fixes
supervision 0.30.2 fixes three silent numeric-correctness bugs in the detection utilities — integer box areas that could wrap negative on large boxes, and two coordinate converters that truncated fractional values on integer input — plus an InferenceSlicer determinism fix that restores its documented source-order result guarantee under multithreading. A set of versioned-docs reliability fixes rounds out the release. No breaking API changes, no new public API.
sv.Detections.box_area no longer overflows to a negative numberInteger-coordinate box area now computes in float64. A large int32 box (50000 x 50000) previously wrapped to a negative area.
detections = sv.Detections(xyxy=np.array([[0, 0, 50000, 50000]], dtype=np.int32))
detections.box_area # array([2.5e+09]) — was negative before the fixxcycwh_to_xyxy / denormalize_boxes stop truncating integer boxesBoth converters wrote fractional half-extent or scaled coordinates into a copy of the integer input, silently truncating toward zero — relevant when converting quantized VLM output (e.g. boxes on a 0..1000 grid).
xcycwh_to_xyxy(np.array([[10, 10, 5, 5]], dtype=np.int32))
# array([[ 7.5, 7.5, 12.5, 12.5]]) — the fractional coordinate 7.5 is no longer truncated to 7InferenceSlicer merges results in source order under multithreadingSlice results now merge in source order under thread_workers > 1, restoring the ordering guarantee its docstring documents. Row order — and, for tied confidences, which overlapping box survives with_nms/with_nmm — no longer varies between runs on identical input.
No breaking API changes, but three fixes above change return dtype for integer input:
sv.Detections.box_area / .area — integer xyxy now returns float64 (was the input's integer dtype, which could silently overflow)sv.xcycwh_to_xyxy — integer input now returns float64 (was truncated integer output)sv.denormalize_boxes — integer input now returns float64 (was truncated integer output)If your code indexes arrays with these outputs (e.g. image[y1:y2, x1:x2]), a float64 result raises TypeError: slice indices must be integers. Cast explicitly where integer indices are required: xcycwh_to_xyxy(boxes).astype(int).
sv.Detections.box_area (and sv.Detections.area for axis-aligned boxes) now computes integer-coordinate box areas in float64, preventing integer overflow for large boxes. (#2514)sv.xcycwh_to_xyxy no longer truncates coordinates for integer input arrays. (#2515)sv.denormalize_boxes no longer truncates coordinates for integer input arrays. (#2516)sv.InferenceSlicer now merges slice results in source order when thread_workers > 1, restoring the ordering guarantee its docstring documents. (#2517)latest no longer fails with error: version 'latest' already exists when latest exists as an alias of a released version. (#2512, #2513)/latest/search/ SearchAction URL when Mike removes the trailing slash from site_url; docs CI renders the custom theme under Mike version contexts to protect the URL, version banners, and star JSON-LD. Applies to future builds going forward. (#2529)develop tree; the workflow also now backs up the pre-rewrite gh-pages tip to a timestamped branch before committing over it. Applies to future builds going forward. (#2533)latest/, and backfills the outdated-version banner itself into already-published archive trees, patching the empty banner markup those pages already carry rather than rebuilding them. (#2534)box_areadenormalize_boxes and xcycwh_to_xyxyInferenceSlicer result ordering under multithreadingMIKE_DOCS_VERSION export, gh-pages backup, and canonical-backfill fixes (#2529, #2532, #2533, #2534)Full changelog: 0.30.1...0.30.2
0.30.1: Numeric-precision and stability fixes
supervision 0.30.1 is a bug-fix patch release. It corrects numeric-precision issues that only surface on specific inputs — large-coordinate oriented boxes (geospatial data, stitched frames), large integer boxes for box_iou, and rotated tracks in DetectionsSmoother — where prior versions could silently return imprecise or self-inconsistent results instead of erroring. It also fixes a duplicate-libavdevice-load crash risk on macOS when both av and opencv-python are installed, plus smaller fixes to list_files_with_extensions and the cv2-free RGBA fallback. No public API was added or removed, and no signature changed — a drop-in upgrade from 0.30.0 for virtually all users. See Migration guide below for the one narrow exception (box_iou on complex-valued coordinates) and for the precision caveats on the numeric fixes.
sv.Detections.area and sv.oriented_box_iou_batch now translate OBB coordinates to a local origin before floating-point math. Previously, large-coordinate inputs could lose enough precision that a box's IoU with itself collapsed below 1.0.
pair_origin = np.minimum(origin_i, origin_j)
offset_i = (origin_i - pair_origin).astype(np.float32, copy=False)
offset_j = (origin_j - pair_origin).astype(np.float32, copy=False)sv.box_iou no longer overflows on large integer boxesArea computation now takes coordinate differences before casting to float, avoiding int32 overflow. For realistic coordinate magnitudes, box_iou's scalar result now matches box_iou_batch.
libavdevice crash fixed on macOSimport supervision no longer loads PyAV's native libraries when the OpenCV backend is active — PyAV is now imported lazily, only where it's used, preventing a duplicate libavdevice warning (and possible crash) when both av and opencv-python are installed.
DetectionsSmoother keeps oriented-box corners consistentSmoothed OBB corners are now aligned (start index + winding) to a reference before averaging, so rotated tracks smooth correctly instead of averaging mismatched corner orderings.
sv.get_polygon_center precision fix for large-coordinate polygonsCentroid calculation now translates to the first vertex and computes in float64 before adding the origin back, preventing integer overflow and precision loss for realistic coordinate magnitudes.
No public signature changed. One item below (box_iou on complex coordinates) does make one specific previously-succeeding call now raise — narrow and deliberate, not classified as breaking since complex-valued box coordinates were never a documented/supported input. The rest only change output values for inputs that were already edge cases:
sv.box_iou on complex-valued coordinates: previously silently discarded the imaginary part and returned a real number. Now raises TypeError("box coordinates must be real-valued").sv.box_iou / box_iou_batch agreement: for realistic integer coordinate magnitudes (below 2^53), box_iou's scalar result now matches box_iou_batch. Not a universal guarantee — box_iou subtracts before casting to float, box_iou_batch still casts to float64 before subtracting, so the two can diverge at coordinates ≥ 2^53 (~9 quadrillion), far outside any real use case.rfdetr_example.py) added to the count_people_in_zone, heatmap_and_track, speed_estimation, tracking, and traffic_analysis bundled examples. (#2497)sv.box_iou now raises TypeError for complex-valued box coordinates instead of silently discarding the imaginary part. (#2485)DetectionsSmoother.update_with_detections now checks active tracker IDs via set membership instead of scanning per tracked object. No output changes. (#2496)sv.get_polygon_center now calculates polygon centroids in translated float64 coordinates, preventing integer overflow and precision loss for realistic-magnitude large-coordinate polygons. (#2491)sv.Detections.area and sv.oriented_box_iou_batch now translate oriented-box coordinates to local origins before floating-point math, preventing self-IoU collapse for large-coordinate inputs. (#2492)DetectionsSmoother now keeps oriented-box corners aligned with smoothed xyxy geometry, including rotated tracks and mixed metadata windows. (#2489)sv.box_iou now calculates overlap in float64, preventing int32 area overflow for large boxes; its scalar result now matches sv.box_iou_batch for realistic coordinate magnitudes. (#2485)sv.list_files_with_extensions no longer includes directories when listing all files without an extension filter. (#2486)sv.pillow_to_cv2 now accepts RGBA images when the cv2-free fallback backend is active, matching OpenCV by dropping alpha and returning BGR channels. (#2488)import supervision no longer loads PyAV's native libraries when the OpenCV backend is selected; PyAV is now imported lazily on first use, preventing a duplicate libavdevice warning (and possible crash) on macOS when both av and opencv-python are installed. (#2509)annotators/core.py, dataset/formats/coco.py, dataset/formats/createml.py, and Detections.from_vlm; previously wrong documented outputs corrected for several VLM examples. (#2474, #2475, #2479, #2484)Also in this release: routine dependency bumps (dependabot: astral-sh/setup-uv, wheel, pymdown-extensions x2, pypa/gh-action-pypi-publish, twine, cryptography), CI/docs-workflow maintenance, and test-only additions (geometry contract test, sklearn parity test) — none change installed package behavior. (#2028, #2470, #2472, #2473, #2480, #2481, #2482, #2483, #2499, #2501, #2506, #2507, #2508)
box_iou int32 overflow; added complex-coordinate TypeError guardDetectionsSmoother tracker-ID lookup performance improvementlist_files_with_extensions to exclude directoriesFull changelog: 0.30.0...0.30.1
…reads for InferenceSlicer, and ships five breaking changes, most notably OpenCV no longer being installed by default, JSONSink switching to native JSO…
supervision 0.30.0 makes OpenCV optional. A new private _cv2/ backend (NumPy and Pillow, with PyAV for the video path) reimplements every OpenCV call the library needs, so supervision now runs on opencv-python-headless — or no OpenCV wheel at all — instead of crashing on import. This release also adds Soft-NMS, LabelMe and CreateML dataset formats, GeoTIFF-aware windowed reads for InferenceSlicer, and ships five breaking changes, most notably OpenCV no longer being installed by default, JSONSink switching to native JSON types, and mask_non_max_merge computing exact mask overlap instead of a downscaled approximation. Python 3.9 support is dropped — 3.10 is now the minimum.
import supervision as sv
window = sv.ImageWindow("frame")
for frame in sv.get_video_frames_generator("input.mp4"):
window.show(frame)
if window.wait_key(1) == "q":
break
The largest change in this release: OpenCV stays the default backend when installed, but supervision no longer requires it — there's no opencv-python extra anymore either. sv.ImageWindow replaces cv2.imshow/cv2.waitKey for display. av>=14.2 is now a required dependency for the PyAV video path during this transition. See the OpenCV migration guide.
detections = sv.Detections.from_ultralytics(result)
softened = detections.with_soft_nms(sigma=0.5)
filtered = detections.with_soft_nms(sigma=0.5, score_threshold=0.3)
sv.Detections.with_soft_nms (plus sv.box_soft_non_max_suppression / sv.mask_soft_non_max_suppression) rescales overlapping detections' confidence instead of discarding them outright — useful in crowded scenes where hard NMS drops valid overlapping objects.
dataset = sv.DetectionDataset.from_labelme(
images_directory_path="images/",
annotations_directory_path="annotations/",
)
import rasterio
with rasterio.open("RGB.byte.tif") as raster:
slicer = sv.InferenceSlicer(callback=my_model_callback, batch_size=4)
detections = slicer(raster)
DetectionDataset.from_labelme/as_labelme and from_createml/as_createml join the existing COCO/YOLO/Pascal-VOC converters. sv.InferenceSlicer can now read an open rasterio dataset window-by-window for multi-GB aerial/drone GeoTIFFs without loading the whole image (pip install "supervision[geotiff]"), and accepts batch_size for batched-callback inference.
sv.load_image_from_urlimage = sv.load_image_from_url("https://media.roboflow.com/notebooks/examples/dog.jpeg")
Load an image straight from an HTTP(S) URL as an OpenCV array, with optional on-disk caching.
Five breaking changes. Most require no code changes beyond a type check or threshold recalibration. The two that need action from most users: the OpenCV install change below, and the Python 3.10 floor.
OpenCV is no longer installed by default. If a compatible cv2 is already importable in your environment, nothing changes for you — it's still preferred automatically. Otherwise install one wheel family yourself (pip install opencv-python or opencv-python-headless) if you need OpenCV-specific behavior, then restart the process — cv2 is detected once at import time. sv.ImageWindow replaces cv2.imshow/cv2.waitKey. Full guide: docs/how_to/opencv_migration.md.
Python 3.10+ is now required — 3.9 reached end-of-life in October 2025.
sv.JSONSink now emits native JSON types, not strings:
# before 0.30.0
row["score"] == "0.85" # str
row["is_valid"] == "True" # str
# after 0.30.0
row["score"] == 0.85 # float
row["is_valid"] is True # bool
sv.CSVSink stays textual, but its per-row custom-data slicing now matches JSONSink.
sv.mask_non_max_merge computes exact mask overlap, not a downscaled approximation, and ignores the now-deprecated mask_dimension parameter (kept for signature compatibility, removal in 0.33.0). Re-tune your overlap threshold after upgrading. Passing overlap_metric/mask_dimension positionally still works — the values are still honored — but now emits a DeprecationWarning; pass them by keyword to silence it. More than five positional arguments raises TypeError.
Detections.merge() on mixed dense + CompactMask inputs now returns a CompactMask, not a plain ndarray:
merged = sv.Detections.merge([dense_detections, compact_mask_detections])
isinstance(merged.mask, np.ndarray) # was True, now False — it's a CompactMask
Only affects code that explicitly merges a CompactMask-carrying Detections object with a dense-mask one yourself — InferenceSlicer, DetectionsSmoother, and with_nms/with_nmm always merge type-homogeneous lists internally, so they're unaffected. The all-dense merge path is also unchanged. This is a substantial performance win: ~2500× less peak memory, ~13× faster on a 1080p frame with 40 detections. If you need the old return type without touching every call site: call merged.mask = merged.mask.to_dense() right after merge(), or avoid producing CompactMask in the first place (Detections.from_inference(compact_masks=False), the default).
supervision also now requires av>=14.2 as an install-time dependency for the PyAV cv2-free video path — this doesn't change any API, so it isn't counted as breaking, but pinned/vendored environments should account for it.
Deprecation removals pushed back one release: ByteTrack, supervision.keypoint, normalized_xyxy, and supervision.dataset.utils RLE compatibility shims — originally scheduled for removal in 0.30.0 — are now scheduled for 0.31.0 instead, giving a full transition window.
sv.load_image_from_url — load an HTTP(S) image as an OpenCV array, with optional on-disk caching (#2372)_cv2 backend facade — image/geometry/drawing/text/video without OpenCV (#2430, #2431, #2432, #2433, #2435, #2438, #2439, #2440, #2441, #2443)sv.ImageWindow — tkinter+Pillow desktop window replacing cv2.imshow/cv2.waitKey (#2320)sv.box_soft_non_max_suppression, sv.mask_soft_non_max_suppression, sv.Detections.with_soft_nms (#1624)sv.VLM.GOOGLE_GEMINI_3_5 — Detections.from_vlm parses Gemini 3.5 output (#2449)get_video_frames_generator(prefetch=...) — background-thread decode into a bounded queue (#2273)PolygonZone(require_all_anchors=...) — toggle all-anchors vs. any-anchor containment (#2272)KeyPoints.merge() — combine a list of KeyPoints, mirroring Detections.merge (#2412)BaseAnnotator.requires_mask — class-level flag on all annotators (#2370)CompactMask.from_coco_rle + Detections.from_inference(compact_masks=True) (#2367)CompactMask.image_shape property (#2383)sv.mask_to_roi — exclusive mask-bound helper for slicing/crops (#2416)DetectionDataset.from_labelme/as_labelme (#2299)DetectionDataset.from_createml/as_createml (#2284)InferenceSlicer GeoTIFF support — sv.WindowedRasterDataset, pip install "supervision[geotiff]" (#2281)InferenceSlicer(batch_size=...) — batched callback contract (#1239)ConfusionMatrix.benchmark(save_directory_path=...) — adaptive TP/FP/FN validation-mosaic export (#2271)HeatMapAnnotator.reset(), TraceAnnotator.reset(), DetectionsSmoother.reset() — clear accumulated per-stream state, so a single instance can be reused across independent streams (#2418)AREA_DATA_FIELD config constant (#2428)sv.denormalize_boxes and sv.xyxyxyxy_to_xyxy now exported at the top levelsv.JSONSink emits native JSON types instead of strings; sv.CSVSink custom-data slicing now matches JSONSink (#2400)sv.mask_non_max_merge computes exact overlap, ignores mask_dimension, positional overlap_metric/mask_dimension deprecated (#2400)Detections.merge() on mixed dense + CompactMask inputs returns CompactMask (#2383)DetectionDataset/ClassificationDataset equality now compares ordered classes lists, not an unordered setsupervision now requires av>=14.2 as an install-time dependency for the cv2-free video fallback — no API change (#2438)ByteTrack, supervision.keypoint, normalized_xyxy, dataset-utils RLE compat removals moved 0.30.0 → 0.31.0count_nonzero mask pixel counts (#2361), vectorized box_iou_batch_with_jaccard (#2359), faster mask-annotation ROI blending (#2368), fewer corner circles on square label backgrounds (#2346), less compact-mask materialization in the polygon annotator (#2369)sv.Recall tracks prediction-only classes, matching Precision/F1Score (#2467, #2468)DetectionDataset.from_pascal_voc no longer raises on background images, with or without force_masks=True (#2463, #2469)import supervision no longer surfaces the deprecated ByteTrack warningsv.CSVSink/sv.JSONSink starts a fresh session — no stale rows or header (#2459)from_vlm Gemini 2.0/2.5/3.5 salvages valid entries from partially malformed JSON arrays (#2449)save_coco_annotations/as_coco read image sizes from headers, no pixel decode for labels-only export (#2442)sv.F1Score no longer emits a spurious div-by-zero RuntimeWarning (#2437)Precision/Recall/F1Score no longer miscount out-of-bucket detections (#2427, #2428, #2408)sv.box_iou_batch upcasts corners to float64, fixing int32-coordinate overflow into a wrong 0.0 IoU (#2418)from_tensorflow scales boxes by correct axes (#2360); from_inference stays aligned on partial masks (#2362) and partial tracker_id (#2353)get_anchors_coordinates is OBB-aware (#2382)CropAnnotator (#2391), HeatMapAnnotator uint8 wrap (#2393), BackgroundOverlayAnnotator negative coords (#2396); get_video_frames_generator releases capture via try/finally (#2393)ByteTrack no longer mutates input Detections; hardened edge cases (#2413)KeyPoints.as_detections accepts numpy/tuple/generator indices (#2402)hex_to_rgba rejects multiple leading # (#2421); Color(...) validates RGBA range (#2407)ColorPalette.by_idx() on empty palette raises ValueError, not ZeroDivisionError (#2407)ConfusionMatrix rejects invalid class ids, per-class recall per max-det cutoff, user ignore flags preserved; FP counted on empty-GT images (#2397)Classifications.from_timm softmaxes logits; download_assets verifies MD5 + retries once (#2414)ImageSink.save_image() raises OSError on write failure (#2416)np.cross with explicit determinant (#2386); removed defensive asserts in image annotators (#2354)KeyPoints.merge(); fixed out-of-bucket metric scoring and key_points edge casesget_anchors_coordinates OBB-aware; kept from_inference aligned on partial dataDetections doctests to runnable examplessv.load_image_from_urlInferenceSlicerInferenceSlicer supportprefetch to get_video_frames_generator, require_all_anchors to PolygonZoneshow_progress to dataset load/save (0.29.1)from_tensorflowclass_id to stay integral for VOC background imagesdraw/utils.py doctestsF1ScoreFull changelog: https://github.com/roboflow/supervision/compare/0.29.1...0.30.0
```python import supervision as sv
KeyPoints.with_nms() — NMS for pose estimationimport supervision as sv
key_points = model.predict(image) # sv.KeyPoints
key_points = key_points.with_nms(threshold=0.5) # removes duplicate skeletons
Derives axis-aligned bounding boxes from each skeleton's valid (non-zero and visible) keypoints, then applies standard box NMS. Supports class_agnostic mode and any OverlapMetric (IOU, IOS). Raises ValueError if detection_confidence is not set.
https://github.com/user-attachments/assets/ed7bd310-4868-4275-ae04-c88e7a0c2561
sv.DetectionDataset.as_pascal_voc no longer mutates bounding boxes (#2341) Previously, every export shifted every bounding box by +1 px in-place. A second call compounded the shift. Fixed by rebinding to a new array; on-disk XML output is unchanged.
sv.Precision and sv.F1Score correctly count background false positives (#2331) Predictions on images with no ground-truth objects, and predictions of classes absent from any annotation, were previously ignored. Under MICRO and MACRO averaging they are now counted as false positives. WEIGHTED averaging is unchanged. Users should re-evaluate existing metric results after upgrading.
sv.DetectionsSmoother works with confidence-free detections (#2333) The smoother no longer raises when detections have no confidence scores. Confidence is averaged over the frames that carry it; tracks without any confidence produce None.
sv.Detections.from_vlm is robust to malformed Gemini/Qwen output (#2342) Valid JSON that is not a list, or whose elements are not dicts, now degrades to empty Detections instead of raising TypeError. A malformed mask value in Gemini 2.5 responses no longer misaligns the xyxy/confidence/masks arrays.
sv.JSONSink serializes NumPy scalars in custom_data (#2334) np.int64 frame indices and other NumPy scalars in custom_data no longer raise TypeError at flush time. NumPy arrays are serialized as lists. The file handle closes even when serialization fails.
sv.approximate_polygon respects the point-count budget (#2332) The function now returns at most floor(N * (1 - percentage)) points (minimum 3). Previously it could return more points than requested. epsilon_step is now validated to be positive.
COCO export preserves all segments for multi-part masks (#2322) Previously, only the first polygon was written when a non-crowd detection had disjoint mask segments. All polygon parts are now written.
sv.HaloAnnotator is ~4× faster with CompactMask detections (#2339) HaloAnnotator now uses the same optimized CompactMask paint path as MaskAnnotator. Previously it materialized each mask full-frame; now it operates on the bounding-box crop. Annotated output is unchanged.
Mask IoU uses less peak memory (#2323) Mask IoU computation now uses matrix multiplication on flattened masks instead of an explicit (N, M, H, W) tensor. For masks larger than 4096×4096 px, computation promotes to float64 automatically. Results are numerically identical.
sv.mask_to_xyxy and sv.KeyPoints.as_detections vectorized (#2330) Both functions now use batched NumPy operations instead of per-element loops. Outputs are bit-identical.
KeyPoints.with_nms()Full Changelog: https://github.com/roboflow/supervision/compare/0.29.0...0.29.1
Full Changelog: https://github.com/roboflow/supervision/compare/0.29.0...0.29.0.post0
Full Changelog: https://github.com/roboflow/supervision/compare/0.29.0...0.29.0.post0
| Deprecated | Removal | Replacement |
Added sv.VertexEllipseAreaAnnotator, sv.VertexEllipseOutlineAnnotator, and sv.VertexEllipseHaloAnnotator for visualizing keypoint uncertainty as covariance ellipses. (#2277, #2286)
import cv2
import supervision as sv
from rfdetr import RFDETRKeypointPreview
image = cv2.imread("<SOURCE_IMAGE_PATH>")
model = RFDETRKeypointPreview()
key_points = model.predict(image)
annotator = sv.VertexEllipseAreaAnnotator(
sigma=[1.0, 2.0, 3.0],
color=[sv.Color.GREEN, sv.Color.YELLOW, sv.Color.RED],
opacity=0.4,
)
annotated = annotator.annotate(image.copy(), key_points)
https://github.com/user-attachments/assets/e01322ff-f39c-420c-bf72-81efba8b0fd3
import cv2
import supervision as sv
from rfdetr import RFDETRKeypointPreview
image = cv2.imread("<SOURCE_IMAGE_PATH>")
model = RFDETRKeypointPreview()
key_points = model.predict(image)
annotator = sv.VertexEllipseOutlineAnnotator(
sigma=[1.0, 2.0, 3.0],
color=[sv.Color.GREEN, sv.Color.YELLOW, sv.Color.RED],
thickness=2,
)
annotated = annotator.annotate(image.copy(), key_points)
https://github.com/user-attachments/assets/bce268c4-7e85-477b-b90c-9f12ef8500b2
import cv2
import supervision as sv
from rfdetr import RFDETRKeypointPreview
image = cv2.imread("<SOURCE_IMAGE_PATH>")
model = RFDETRKeypointPreview()
key_points = model.predict(image)
annotator = sv.VertexEllipseHaloAnnotator(
sigma=[1.0, 2.0, 3.0],
color=[sv.Color.GREEN, sv.Color.YELLOW, sv.Color.RED],
opacity=0.6,
)
annotated = annotator.annotate(image.copy(), key_points)
https://github.com/user-attachments/assets/b8662ea4-666a-4ac7-a2fb-11d59c68fd74
Added sv.oriented_box_non_max_suppression and sv.oriented_box_non_max_merge for performing NMS and NMM directly on oriented bounding boxes. (#2303)
Added OBB (Oriented Bounding Box) support to sv.ConfusionMatrix via MetricTarget.ORIENTED_BOUNDING_BOXES. (#2247)
Added preserve_audio parameter to sv.process_video. When enabled, the audio stream from the source video is muxed into the output using ffmpeg. (#2252)
Added is_obb parameter to sv.DetectionDataset.as_yolo for exporting oriented bounding box annotations in the YOLO OBB format (9-token lines with 4 corner coordinates). (#2302, #2289)
sv.EdgeAnnotator and sv.VertexAnnotator now respect the visible mask. Invisible keypoints and their edges are skipped during rendering. (#2286)
sv.EdgeAnnotator and sv.VertexLabelAnnotator now support per-class skeleton definitions, enabling correct rendering when multiple skeleton topologies (e.g. person + animal) coexist in one frame. (#2286)
sv.Detections.with_nms and sv.Detections.with_nmm are now OBB-aware. When data[ORIENTED_BOX_COORDINATES] is present, oriented-box IoU is used automatically instead of axis-aligned box IoU. (#2303)
<img width="1703" height="889" alt="box_nms_demo" src="https://github.com/user-attachments/assets/1845d110-040f-4ee1-9f92-a69dde6ed748" />
<img width="1703" height="889" alt="box_nmm_demo" src="https://github.com/user-attachments/assets/a7a21f62-17ef-4e20-8d78-d0cf12dae339" />
<img width="1703" height="889" alt="obb_nms_demo" src="https://github.com/user-attachments/assets/400ab62b-0202-47f3-929b-89d04160ff88" />
<img width="1703" height="889" alt="obb_nmm_demo" src="https://github.com/user-attachments/assets/18267478-2428-4052-946a-89607e97cb64" />
sv.Detections.area is now OBB-aware. When oriented box coordinates are present, the property returns the polygon area of the rotated bounding box (via the shoelace formula) instead of the axis-aligned box area. (#2306)
sv.InferenceSlicer now detects OBB outputs from callbacks and automatically falls back to sequential processing to avoid thread-safety issues when thread_workers > 1. (#2256)
Fixed sv.oriented_box_iou_batch to correctly handle non-square canvases. Previously, rasterization assumed square dimensions, leading to incorrect IoU values for tall or wide images. (#2282)
Fixed sv.process_video audio stream handling. The audio muxing path now correctly creates temp files on the same filesystem, decodes ffmpeg errors, and avoids muxing incomplete output. (#2252)
Fixed sv.Detections.from_vlm returning None for class_id on empty VLM parses. Now returns an empty int ndarray. (#2239)
Fixed sv.Detections.from_inference to preserve class_name as a string-dtype array when predictions are empty. Previously it returned an untyped empty array. (#2270)
Fixed sv.HeatMapAnnotator divide-by-zero crash when called with empty detections. (#2269)
Fixed COCO export emitting 0-indexed category_id values. Now correctly emits 1-indexed IDs as per the COCO specification. (#2276)
Fixed COCO annotation and image IDs not being chainable across dataset splits. IDs are now sequential across train/val/test. (#2267)
Fixed sv.DetectionDataset.as_yolo losing OBB rotation when exporting oriented bounding boxes. (#2289)
Fixed YOLO dataset loading to sort class names by numeric keys when the data.yaml uses integer class IDs. (#2296)
Fixed letterbox utility to support grayscale images. (#2297)
Fixed file extension filters to normalize casing (e.g. .JPG now matches .jpg). (#2298)
| Deprecated | Removal | Replacement |
|---|---|---|
KeyPoints.confidence |
0.32.0 |
KeyPoints.keypoint_confidence |
merge_inner_detection_object_pair |
0.32.0 |
None (internal use only) |
merge_inner_detections_objects |
0.32.0 |
None (internal use only) |
merge_inner_detections_objects_without_iou |
0.32.0 |
None (internal use only) |
validate_detections_fields |
0.32.0 |
None (internal use only) |
validate_vlm_parameters |
0.32.0 |
None (internal use only) |
validate_fields_both_defined_or_none |
0.32.0 |
None (internal use only) |
validate_xyxy |
0.32.0 |
None (internal use only) |
validate_mask |
0.32.0 |
None (internal use only) |
validate_class_id |
0.32.0 |
None (internal use only) |
validate_confidence |
0.32.0 |
None (internal use only) |
validate_tracker_id |
0.32.0 |
None (internal use only) |
validate_data |
0.32.0 |
None (internal use only) |
validate_xy |
0.32.0 |
None (internal use only) |
validate_key_point_confidence |
0.32.0 |
None (internal use only) |
validate_key_points_fields |
0.32.0 |
None (internal use only) |
validate_resolution |
0.32.0 |
None (internal use only) |
validate_custom_values |
0.32.0 |
None (internal use only) |
validate_input_tensors |
0.32.0 |
None (internal use only) |
validate_labels |
0.32.0 |
None (internal use only) |
@SkalskiP (Piotr Skalski), @Borda (Jirka Borovec), @kounelisagis (Agis Kounelis), @RitwijParmar (Ritwij Aryan Parmar), @Khanz9664 (Shahid Ul Islam), @satishkc7 (SATISH K C), @Ace3Z (Mahbod Tajdini), @Madhav-C, @RubenHaisma (Ruben Haisma), @adhavan18 (Tamil Adhavan), @Bortlesboat (Andrew Barnes), @Lourdhu02, @tarunbommawar27, @YousefZahran1 (Youssef Ibrahim), @JFrench-Enterprise, @Patel-Prem (Premkumar Patel)
releasing 0.29.0rc1
releasing 0.29.0rc1
releasing 0.29.0.rc0
releasing 0.29.0.rc0
Tracker implementations now live in the dedicated trackers package. sv.ByteTrack remains available in 0.28–0.29 with DeprecationWarning; removal in 0.…
sv.CompactMaskSegmentation models produce one full-resolution bitmap per instance. On a 1920×1080 image with 28 detections that is ~55 MB of mask data. Most pixels are background. sv.CompactMask stores only the tight bounding-box crop, RLE-encoded — the same 28 masks drop to ~237 KB of crops, a 240× reduction before RLE kicks in.
It's a drop-in replacement: annotators, filters, and area all work unchanged.
<img width="1185" height="697" alt="supervision-sam3" src="https://github.com/user-attachments/assets/ed6483c0-bf2e-4b1f-bbe2-a32c4e1a002b" />
import supervision as sv
# any segmentation model — RF-DETR Seg, YOLO-Seg, SAM3
detections = model.predict(image) # sv.Detections with dense masks
dense_mb = detections.mask.nbytes / 1024 / 1024
compact = sv.CompactMask.from_dense(
masks=detections.mask,
xyxy=detections.xyxy,
image_shape=image.shape[:2],
)
detections.mask = compact # swap in — API unchanged
# filter by pixel area without materialising dense masks
large = detections[compact.area > 1000]
# annotators call .to_dense() internally
annotated = sv.MaskAnnotator().annotate(image.copy(), detections)
SAM3 segments objects by free-text prompt — no class list, no bounding boxes. sv.Detections.from_sam3() parses both PCS (multi-prompt) and PVS (video) response formats into a standard sv.Detections, with class_id set to the prompt index.
<!-- IMAGE: SAM3 output on people-walking.jpg — person masks red, bag masks teal -->
import requests, base64
import supervision as sv
PROMPTS = ["person", "bag"]
with open("image.jpg", "rb") as f:
img_b64 = base64.b64encode(f.read()).decode()
response = requests.post(
f"https://api.roboflow.com/inferenceproxy/seg-preview?api_key={API_KEY}",
json={
"image": {"type": "base64", "value": img_b64},
"prompts": [{"type": "text", "text": p} for p in PROMPTS],
},
headers={"Content-Type": "application/json"},
)
sam3_result = response.json()
h, w = cv2.imread("image.jpg").shape[:2]
detections = sv.Detections.from_sam3(sam3_result=sam3_result, resolution_wh=(w, h))
# class_id == 0 → "person", class_id == 1 → "bag"
VideoInfo.fps is now floatNTSC frame rates (23.976, 29.97, 59.94) were silently truncated. fps is now the true float — cast at call sites that need an integer.
<details> <summary>Before</summary>
info = sv.VideoInfo.from_video_path("clip.mp4")
buf = collections.deque(maxlen=info.fps)
trace = sv.TraceAnnotator(trace_length=info.fps)
</details>
<details> <summary>After</summary>
info = sv.VideoInfo.from_video_path("clip.mp4")
buf = collections.deque(maxlen=int(info.fps))
trace = sv.TraceAnnotator(trace_length=int(info.fps))
</details>
sv.ByteTrack deprecated — use ByteTrackTrackerTracker implementations now live in the dedicated trackers package. sv.ByteTrack remains available in 0.28–0.29 with DeprecationWarning; removal in 0.30.0.
<details> <summary>Before</summary>
tracker = sv.ByteTrack()
detections = tracker.update_with_detections(detections)
</details>
<details> <summary>After</summary>
# pip install trackers
from trackers import ByteTrackTracker
tracker = ByteTrackTracker()
detections = tracker.update(detections)
</details>
Memory-efficient masks with sv.CompactMask. Sparse segmentation masks are now stored as a crop region plus RLE-encoded data instead of full-resolution bitmaps, cutting memory use by 10–100× for typical instance-segmentation outputs. It's a drop-in change — sv.Detections.mask, filtering, merging, and area all keep working without materialising the full array. (#2159)
SAM3 detection and PVS support in from_inference. sv.Detections.from_inference now parses SAM3 detection and point-video-segmentation outputs, both from the local inference package and from Roboflow-hosted server responses. (#2103, #2152)
Compressed COCO RLE masks in from_inference. Inference responses with rle or rle_mask fields containing a compressed counts string (as produced by pycocotools) are decoded directly into binary masks, skipping the lossy polygon round-trip. (#2178)
Standard logging module instead of print. Diagnostic output is now emitted under the supervision logger, so applications can capture, filter, or silence it through standard logging configuration. (#2154)
RGBA hex codes in sv.Color. sv.Color.from_hex accepts 8-digit hex (#ff00ff80), and Color.as_hex() round-trips alpha when not fully opaque. New top-level helpers: sv.hex_to_rgba, sv.rgba_to_hex, and sv.is_valid_hex. (#2004)
Dynamic kernel sizing in blur and pixelate annotators. BlurAnnotator(kernel_size=None) and PixelateAnnotator(pixel_size=None) (the new default) compute the kernel per detection as a fraction of the shorter bounding-box side, giving visually consistent results across object scales. (#709)
sv.ImageAssets for sample images. A counterpart to the existing video assets — downloads sample images for examples and tutorials. (#932)
Boundary warnings in InferenceSlicer. Emits a warning when callback detections fall outside tile boundaries, helping you spot coordinate-system bugs in custom callbacks early. (#2186)
sv.VideoInfo.fps is now float, not int. Frame rates like 23.976, 29.97, and 59.94 are no longer truncated. If you pass fps to APIs that require an integer (deque(maxlen=...), TraceAnnotator(trace_length=...)), wrap with int(...). (#2210)
sv.rle_to_mask returns bool, not uint8. This matches the long-declared signature. Code that does mask * 255 still works via NumPy broadcasting, but explicit casts like mask.view(np.uint8) will break. Add .astype(np.uint8) if you relied on the undocumented integer output. (#2178)
See the migration guide below for before/after snippets.
Metric arrays use float32 instead of float64. sv.MeanAveragePrecisionResult and related arrays (mAP_scores, ap_per_class, iou_thresholds, precision/recall) drop to float32, reducing memory and speeding up computation. Numerical results may differ in the last few digits. (#2169)
rle_to_mask and mask_to_rle moved. New canonical path: supervision.detection.utils.converters. The old supervision.dataset.utils import still works but is deprecated. (#2178)
normalized_xyxy argument renamed to xyxy in denormalize_boxes. sv.denormalize_boxes(normalized_xyxy=...) still works but emits a FutureWarning; switch to xyxy=. Scheduled for removal in 0.30.0.
sv.ByteTrack → ByteTrackTracker (external trackers package). Install with pip install trackers; the method renames from update_with_detections() to update(). Scheduled for removal in 0.30.0. (#2215)
supervision.keypoint → supervision.key_points. Also deprecated: the LMM enum (use VLM), from_lmm (use from_vlm), create_tiles in supervision.utils.image, ensure_cv2_image_for_processing in supervision.utils.conversion, and the keypoint validators in supervision.validators. (#2214)
PolygonZone no longer double-counts overlapping zones. When two polygons contain the same anchor, each zone now reflects its own containment instead of every zone claiming the detection. (#1991)
LineZone respects class identity across reused tracker IDs. Trackers that recycle tracker_id across classes no longer leak crossing state from one object to another. (#1868)
process_video raises immediately on callback errors. Previously the exception was swallowed and the process hung until the writer was flushed. (#2022)
DetectionDataset populates class_name. Loaded annotations now carry data["class_name"], matching what model connectors produce. (#2156)
ByteTrack preserves externally assigned tracker_id. No longer overwrites caller-assigned IDs on the first update. (#1364)
Confusion matrix double-counting fixed. evaluate_detection_batch now correctly matches multiple predictions to the same target, so FP/FN counts match expectations. (#1853)
MeanAverageRecall mAR@K is now COCO-compliant. Computed using top-K detections per image; previous values were inflated relative to pycocotools. (#2136)
Detections.is_empty() handles empty tracker_id. Returns True for zero-row detections regardless of whether tracker_id is None or an empty array. (#2209)
CSVSink and JSONSink slice custom_data per row. NumPy arrays, lists, and tuples whose length matches the detection count are now indexed per row, instead of being written whole for every detection. (#2199, #2216)
TraceAnnotator smooth mode handles stationary tracks. Deduplicates anchor points and falls back to a raw polyline when splprep cannot fit fewer than 4 unique points. (#2217)
load_coco_annotations rejects path-traversal annotations. Refuses file_name entries that escape the images directory via ../ or absolute paths. (#2218)
OBB datasets no longer blow up memory. Loading oriented-bounding-box datasets stopped allocating full-image masks per box. (#2187)
KeyPoints boolean mask indexing fixed. Uniform-count selection now works correctly when all instances share the same keypoint count. (#2188)
DetectionDataset.as_coco() preserves area and iscrowd. No longer dropped silently in the round-trip. (#2185)
force_mask=True precision and COCO empty-polygon export. Annotation conversion no longer loses precision, and COCO export tolerates empty polygons across formats. (#1746, #1086, #265)
A huge thank you to everyone who shipped this release:
rle_to_mask correctnessVideoInfo.fps as float and Detections.is_empty() fixclass_name in DetectionDatasetCSVSink NumPy slicingMeanAverageRecallPolygonZone overlap fixLineZone class-aware tracker IDsprocess_video error propagationByteTrack preserves external tracker IDsforce_masks consistencysv.Colorsv.ImageAssetsforce_mask=True precision fixCompactMask, metrics float32, deprecationsFull changelog: https://github.com/roboflow/supervision/compare/0.27.0...0.28.0
releasing 0.28.0.rc2
releasing 0.28.0.rc2
releasing 0.28.0.rc1
releasing 0.28.0.rc1
Nothing published for this version
Full Changelog: https://github.com/roboflow/supervision/compare/0.27.0...0.27.0.post2
Full Changelog: https://github.com/roboflow/supervision/compare/0.27.0...0.27.0.post2
Nothing published for this version
Nothing published for this version
Removed the deprecated overlap_ratio_wh argument from sv.InferenceSlicer. Use the pixel based overlap_wh argument to control slice overlap.
sv.filter_segments_by_distance to keep the largest connected component and any nearby components within an absolute or relative distance threshold. This helps you clean up predictions from segmentation models like SAM, SAM2, YOLO segmentation, and RF-DETR segmentation. (#2008)https://github.com/user-attachments/assets/2bdfd45d-b235-414b-91a3-6544d7c2b4ec
Added sv.edit_distance for Levenshtein distance between two strings. Supports insert, delete, substitute. (#1912)
import supervision as sv
sv.edit_distance("hello", "hello")
# 0
sv.edit_distance("hello world", "helloworld")
# 1
sv.edit_distance("YOLO", "yolo", case_sensitive=True)
# 4
Added sv.fuzzy_match_index to find the first close match in a list using edit distance. (#1912)
import supervision as sv
sv.fuzzy_match_index(["cat", "dog", "rat"], "dat", threshold=1)
# 0
sv.fuzzy_match_index(["alpha", "beta", "gamma"], "bata", threshold=1)
# 1
sv.fuzzy_match_index(["one", "two", "three"], "ten", threshold=2)
# None
Added sv.get_image_resolution_wh as a unified way to read image width and height from NumPy and PIL inputs. (#2014)
Added sv.tint_image to apply a solid color overlay to an image at a specified opacity. Works with both NumPy and PIL inputs. (#1943)
Added sv.grayscale_image to convert an image to 3-channel grayscale for compatibility with color-based drawing utilities. (#1943)
Added sv.xyxy_to_mask to convert bounding boxes into 2D boolean masks. Each mask corresponds to one bounding box. (#2006)
Added Qwen3-VL support in sv.Detections.from_vlm and legacy from_lmm mapping. Use vlm=sv.QWEN_3_VL. (#2015)
import supervision as sv
response = """```json
[
{"bbox_2d": [220, 102, 341, 206], "label": "taxi"},
{"bbox_2d": [30, 606, 171, 743], "label": "taxi"},
{"bbox_2d": [192, 451, 318, 581], "label": "taxi"},
{"bbox_2d": [358, 908, 506, 1000], "label": "taxi"},
{"bbox_2d": [735, 359, 873, 480], "label": "taxi"},
{"bbox_2d": [758, 508, 885, 617], "label": "taxi"},
{"bbox_2d": [857, 263, 988, 374], "label": "taxi"},
{"bbox_2d": [735, 243, 838, 351], "label": "taxi"},
{"bbox_2d": [303, 291, 434, 417], "label": "taxi"},
{"bbox_2d": [426, 273, 552, 382], "label": "taxi"}
]
```"""
detections = sv.Detections.from_vlm(
vlm=sv.VLM.QWEN_3_VL,
result=response,
resolution_wh=(1023, 682)
)
detections.xyxy
# array([[ 225.06 , 69.564, 348.843, 140.492],
# [ 30.69 , 413.292, 174.933, 506.726],
# [ 196.416, 307.582, 325.314, 396.242],
# [ 366.234, 619.256, 517.638, 682. ],
# [ 751.905, 244.838, 893.079, 327.36 ],
# [ 775.434, 346.456, 905.355, 420.794],
# [ 876.711, 179.366, 1010.724, 255.068],
# [ 751.905, 165.726, 857.274, 239.382],
# [ 309.969, 198.462, 443.982, 284.394],
# [ 435.798, 186.186, 564.696, 260.524]])
<img width="1023" height="682" alt="supervision-0 27 0-promo-from-qwen-3-vl" src="https://github.com/user-attachments/assets/bbef03fd-5d76-4dbf-a7a9-2fbaef758b6c" />
Added DeepSeek-VL2 support in sv.Detections.from_vlm and legacy from_lmm mapping. Use vlm=sv.VLM.DEEPSEEK_VL_2. (#1884)
Improved sv.Detections.from_vlm parsing for Qwen 2.5 VL outputs. The function now handles incomplete or truncated JSON responses. (#2015)
sv.InferenceSlicer now uses a new offset generation logic that removes redundant tiles and ensures clean border aligned slicing. This reduces the number of tiles processed, lowering inference time without hurting detection quality. (#2014)
https://github.com/user-attachments/assets/0141ff44-0269-472c-900d-610f47330d57
import supervision as sv
from PIL import Image
from rfdetr import RFDETRMedium
model = RFDETRMedium()
def callback(tile):
return model.predict(tile)
slicer = sv.InferenceSlicer(callback, slice_wh=512, overlap_wh=128)
image = Image.open("example.png")
detections = slicer(image)
sv.Detections now includes a box_aspect_ratio property for vectorized aspect ratio computation. You use it to filter detections based on box shape. (#2016)import numpy as np
import supervision as sv
xyxy = np.array([
[10, 10, 50, 50],
[60, 10, 180, 50],
[10, 60, 50, 180],
])
detections = sv.Detections(xyxy=xyxy)
ar = detections.box_aspect_ratio
# array([1.0, 3.0, 0.33333333])
detections[(ar < 2.0) & (ar > 0.5)].xyxy
# array([[10., 10., 50., 50.]])
Improved the performance of sv.box_iou_batch. Processing runs about 2x to 5x faster. (#2001)
sv.process_video now uses a threaded reader, processor, and writer pipeline. This removes I/O stalls and improves throughput while keeping the callback single threaded and safe for stateful models. (#1997)
sv.denormalize_boxes now supports batch conversion of bounding boxes. The function now accepts arrays of shape (N, 4) and returns a batch of absolute pixel coordinates.
sv.LabelAnnotator and sv.RichLabelAnnotator now accepts text_offset=(x, y) to shift the label relative to text_position. Works with smart label position and line wrapping. (#1917)
overlap_ratio_wh argument from sv.InferenceSlicer. Use the pixel based overlap_wh argument to control slice overlap. (#2014)[!TIP] Convert your old ratio based overlap to pixel based overlap. Multiply each ratio by the slice dimensions.
# before slice_wh = (640, 640) overlap_ratio_wh = (0.25, 0.25) slicer = sv.InferenceSlicer( callback=callback, slice_wh=slice_wh, overlap_ratio_wh=overlap_ratio_wh, overlap_filter=sv.OverlapFilter.NON_MAX_SUPPRESSION, ) # after overlap_wh = ( int(overlap_ratio_wh[0] * slice_wh[0]), int(overlap_ratio_wh[1] * slice_wh[1]), ) slicer = sv.InferenceSlicer( callback=callback, slice_wh=slice_wh, overlap_wh=overlap_wh, overlap_filter=sv.OverlapFilter.NON_MAX_SUPPRESSION, )
@SkalskiP (Piotr Skalski), @onuralpszr (Onuralp SEZER), @soumik12345 (Soumik Rakshit), @rcvsq, @AlexBodner (Alex Bodner), @Ashp116, @kshitijaucharmal (Kshitij Aucharmal), @ernestlwt, @AnonymDevOSS, @jackiehimel (Jackie Himel ), @dominikWin (Dominik Winecki)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fixed error in `sv.MeanAveragePrecision` where the area used for size-specific evaluation (small / medium / large) was always zero unless explicitly p
sv.MeanAveragePrecision where the area used for size-specific evaluation (small / medium / large) was always zero unless explicitly provided in sv.Detections.data. (https://github.com/roboflow/supervision/pull/1894)ID=0 bug in sv.MeanAveragePrecision where objects were getting 0.0 mAP despite perfect IoU matches due to a bug in annotation ID assignment. (https://github.com/roboflow/supervision/pull/1895)sv.MeanAveragePrecision could return negative values when certain object size categories have no data. (https://github.com/roboflow/supervision/pull/1898)match_metric support for sv.Detections.with_nms. (https://github.com/roboflow/supervision/pull/1901)border_thickness parameter usage for sv.PercentageBarAnnotator. (https://github.com/roboflow/supervision/pull/1906)@balthazur (Balthasar Huber), @onuralpszr (Onuralp SEZER), @rafaelpadilla (Rafael Padilla), @soumik12345 (Soumik Rakshit), @SkalskiP (Piotr Skalski)
Nothing published for this version
sv.LMM enum is deprecated and will be removed in supervision-0.31.0. Use sv.VLM instead.
[!WARNING]
supervision-0.26.0dropspython3.8support and upgrade all codes topython3.9syntax style.
[!TIP] Our docs page now has a fresh look that is consistent with the documentations of all Roboflow open-source projects. (#1858)
Added support for creating sv.KeyPoints objects from ViTPose and ViTPose++ inference results via sv.KeyPoints.from_transformers. (#1788)
https://github.com/user-attachments/assets/f1917032-29d8-4b88-b871-65c2e28a756e
Added support for the IOS (Intersection over Smallest) overlap metric that measures how much of the smaller object is covered by the larger one in sv.Detections.with_nms, sv.Detections.with_nmm, sv.box_iou_batch, and sv.mask_iou_batch. (#1774)
import numpy as np
import supervision as sv
boxes_true = np.array([
[100, 100, 200, 200],
[300, 300, 400, 400]
])
boxes_detection = np.array([
[150, 150, 250, 250],
[320, 320, 420, 420]
])
sv.box_iou_batch(
boxes_true=boxes_true,
boxes_detection=boxes_detection,
overlap_metric=sv.OverlapMetric.IOU
)
# array([[0.14285714, 0. ],
# [0. , 0.47058824]])
sv.box_iou_batch(
boxes_true=boxes_true,
boxes_detection=boxes_detection,
overlap_metric=sv.OverlapMetric.IOS
)
# array([[0.25, 0. ],
# [0. , 0.64]])
Added sv.box_iou that efficiently computes the Intersection over Union (IoU) between two individual bounding boxes. (#1874)
Added support for frame limitations and progress bar in sv.process_video. (#1816)
Added sv.xyxy_to_xcycarh function to convert bounding box coordinates from (x_min, y_min, x_max, y_max) into measurement space to format (center x, center y, aspect ratio, height), where the aspect ratio is width / height. (#1823)
Added sv.xyxy_to_xywh function to convert bounding box coordinates from (x_min, y_min, x_max, y_max) format to (x, y, width, height) format. (#1788)
sv.LabelAnnotator now supports the smart_position parameter to automatically keep labels within frame boundaries, and the max_line_length parameter to control text wrapping for long or multi-line labels. (#1820)
https://github.com/user-attachments/assets/361c17c7-0810-466d-907d-c752e91bc6f7
<img width="1600" height="1400" alt="Snap (25)" src="https://github.com/user-attachments/assets/7945dafc-e646-46e5-ae62-492685d1bbc0" />
sv.LabelAnnotator now supports non-string labels. (#1825)
sv.Detections.from_vlm now supports parsing bounding boxes and segmentation masks from responses generated by Google Gemini models. You can test Gemini prompting, result parsing, and visualization with Supervision using this example notebook. (#1792)
import supervision as sv
gemini_response_text = """```json
[
{"box_2d": [543, 40, 728, 200], "label": "cat", "id": 1},
{"box_2d": [653, 352, 820, 522], "label": "dog", "id": 2}
]
```"""
detections = sv.Detections.from_vlm(
sv.VLM.GOOGLE_GEMINI_2_5,
gemini_response_text,
resolution_wh=(1000, 1000),
classes=['cat', 'dog'],
)
detections.xyxy
# array([[543., 40., 728., 200.], [653., 352., 820., 522.]])
detections.data
# {'class_name': array(['cat', 'dog'], dtype='<U26')}
detections.class_id
# array([0, 1])
<img width="2200" height="1500" alt="Snap (27)" src="https://github.com/user-attachments/assets/b53c6670-49f7-49b0-99e5-90a1e8bf78f2" />
sv.Detections.from_vlm now supports parsing bounding boxes from responses generated by Moondream. (#1878)
import supervision as sv
moondream_result = {
'objects': [
{
'x_min': 0.5704046934843063,
'y_min': 0.20069346576929092,
'x_max': 0.7049859315156937,
'y_max': 0.3012596592307091
},
{
'x_min': 0.6210969910025597,
'y_min': 0.3300672620534897,
'x_max': 0.8417936339974403,
'y_max': 0.4961046129465103
}
]
}
detections = sv.Detections.from_vlm(
sv.VLM.MOONDREAM,
moondream_result,
resolution_wh=(1000, 1000),
)
detections.xyxy
# array([[1752.28, 818.82, 2165.72, 1229.14],
# [1908.01, 1346.67, 2585.99, 2024.11]])
sv.Detections.from_vlm now supports parsing bounding boxes from responses generated by Qwen-2.5 VL. You can test Qwen2.5-VL prompting, result parsing, and visualization with Supervision using this example notebook. (#1709)
import supervision as sv
qwen_2_5_vl_result = """```json
[
{"bbox_2d": [139, 768, 315, 954], "label": "cat"},
{"bbox_2d": [366, 679, 536, 849], "label": "dog"}
]
```"""
detections = sv.Detections.from_vlm(
sv.VLM.QWEN_2_5_VL,
qwen_2_5_vl_result,
input_wh=(1000, 1000),
resolution_wh=(1000, 1000),
classes=['cat', 'dog'],
)
detections.xyxy
# array([[139., 768., 315., 954.], [366., 679., 536., 849.]])
detections.class_id
# array([0, 1])
detections.data
# {'class_name': array(['cat', 'dog'], dtype='<U10')}
detections.class_id
# array([0, 1])
Significantly improved the speed of HSV color mapping in sv.HeatMapAnnotator, achieving approximately 28x faster performance on 1920x1080 frames. (#1786)
Supervision’s sv.MeanAveragePrecision is now fully aligned with pycocotools, the official COCO evaluation tool, ensuring accurate and standardized metrics. (#1834)
import supervision as sv
from supervision.metrics import MeanAveragePrecision
predictions = sv.Detections(...)
targets = sv.Detections(...)
map_metric = MeanAveragePrecision()
map_metric.update(predictions, targets).compute()
# Average Precision (AP) @[ IoU=0.50:0.95 | area= all | maxDets=100 ] = 0.464
# Average Precision (AP) @[ IoU=0.50 | area= all | maxDets=100 ] = 0.637
# Average Precision (AP) @[ IoU=0.75 | area= all | maxDets=100 ] = 0.203
# Average Precision (AP) @[ IoU=0.50:0.95 | area= small | maxDets=100 ] = 0.284
# Average Precision (AP) @[ IoU=0.50:0.95 | area=medium | maxDets=100 ] = 0.497
# Average Precision (AP) @[ IoU=0.50:0.95 | area= large | maxDets=100 ] = 0.629
[!TIP] The updated mAP implementation enabled us to build an updated version of the Computer Vision Model Leaderboard.
<img width="1492" height="1014" alt="imageedit_1_8427510007" src="https://github.com/user-attachments/assets/1f46b877-abbc-486b-b8f9-2829f43716e1" />
sv.Detections.data when detections filtering.sv.LMM enum is deprecated and will be removed in supervision-0.31.0. Use sv.VLM instead.sv.Detections.from_lmm property is deprecated and will be removed in supervision-0.31.0. Use sv.Detections.from_vlm instead.sv.DetectionDataset.images property has been removed in supervision-0.26.0. Please loop over images with for path, image, annotation in dataset:, as that does not require loading all images into memory.sv.DetectionDataset with parameter images as Dict[str, np.ndarray] is deprecated and has been removed in supervision-0.26.0. Please pass a list of paths List[str] instead.sv.BoundingBoxAnnotator is deprecated and has been removed in supervision-0.26.0. It has been renamed to sv.BoxAnnotator.@onuralpszr (Onuralp SEZER), @SkalskiP (Piotr Skalski), @SunHao-AI (Hao Sun), @rafaelpadilla Rafael Padilla, @Ashp116 (Ashp116), @capjamesg (James Gallagher), @blakeburch (Blake Burch), @hidara2000 (hidara2000), @Armaggheddon (Alessandro Brunello), @soumik12345 (Soumik Rakshit).
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Supervision 0.25.0 is here! Featuring a more robust `LineZone` crossing counter, support for tracking KeyPoints, Python 3.13 compatibility, and 3 new
Supervision 0.25.0 is here! Featuring a more robust LineZone crossing counter, support for tracking KeyPoints, Python 3.13 compatibility, and 3 new metrics: Precision, Recall and Mean Average Recall. The update also includes smart label positioning, improved Oriented Bounding Box support, and refined error handling. Thank you to all contributors - especially those who answered the call of Hacktoberfest!
LineZone: when computing line crossings, detections that jitter might be counted twice (or more!). This can now be solved with the minimum_crossing_threshold argument. If you set it to 2 or more, extra frames will be used to confirm the crossing, improving the accuracy significantly. (#1540)https://github.com/user-attachments/assets/89ca2ee6-93c9-41e6-a432-e16c4c69c695
KeyPoints. See the complete step-by-step guide in the Object Tracking Guide. (#1658)import numpy as np
import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8m-pose.pt")
tracker = sv.ByteTrack()
trace_annotator = sv.TraceAnnotator()
def callback(frame: np.ndarray, _: int) -> np.ndarray:
results = model(frame)[0]
key_points = sv.KeyPoints.from_ultralytics(results)
detections = key_points.as_detections()
detections = tracker.update_with_detections(detections)
annotated_image = trace_annotator.annotate(frame.copy(), detections)
return annotated_image
sv.process_video(
source_path="input_video.mp4",
target_path="output_video.mp4",
callback=callback
)
https://github.com/user-attachments/assets/4c3bdf54-391e-4633-9164-f15878ddfb33
<sup>See the guide for the full code used to make the video</sup>
Added is_empty method to KeyPoints to check if there are any keypoints in the object. (#1658)
Added as_detections method to KeyPoints that converts KeyPoints to Detections. (#1658)
Added a new video to supervision[assets]. (#1657)
from supervision.assets import download_assets, VideoAssets
path_to_video = download_assets(VideoAssets.SKIING)
Python 3.13. The most renowned update is the ability to run Python without Global Interpreter Lock (GIL). We expect support for this among our dependencies to be inconsistent, but if you do attempt it - let us know the results! (#1595)Mean Average Recall mAR metric, which returns a recall score, averaged over IoU thresholds, detected object classes, and limits imposed on maximum considered detections. (#1661)import supervision as sv
from supervision.metrics import MeanAverageRecall
predictions = sv.Detections(...)
targets = sv.Detections(...)
map_metric = MeanAverageRecall()
map_result = map_metric.update(predictions, targets).compute()
map_result.plot()
Precision and Recall metrics, providing a baseline for comparing model outputs to ground truth or another model (#1609)import supervision as sv
from supervision.metrics import Recall
predictions = sv.Detections(...)
targets = sv.Detections(...)
recall_metric = Recall()
recall_result = recall_metric.update(predictions, targets).compute()
recall_result.plot()
import supervision as sv
from supervision.metrics import F1_Score
predictions = sv.Detections(...)
targets = sv.Detections(...)
f1_metric = MeanAverageRecall(metric_target=sv.MetricTarget.ORIENTED_BOUNDING_BOXES)
f1_result = f1_metric.update(predictions, targets).compute()
<!-- TODO: image -->
smart_position is set for LabelAnnotator, RichLabelAnnotator or VertexLabelAnnotator, the labels will move around to avoid overlapping others. (#1625)import supervision as sv
from ultralytics import YOLO
image = cv2.imread("image.jpg")
label_annotator = sv.LabelAnnotator(smart_position=True)
model = YOLO("yolo11m.pt")
results = model(image)[0]
detections = sv.Detections.from_ultralytics(results)
annotated_frame = label_annotator.annotate(first_frame.copy(), detections)
sv.plot_image(annotated_frame)
https://github.com/user-attachments/assets/ef768db4-867d-4305-b905-80e690bb1ea7
metadata variable to Detections. It allows you to store custom data per-image, rather than per-detected-object as was possible with data variable. For example, metadata could be used to store the source video path, camera model or camera parameters. (#1589)import supervision as sv
from ultralytics import YOLO
model = YOLO("yolov8m")
result = model("image.png")[0]
detections = sv.Detections.from_ultralytics(result)
# Items in `data` must match length of detections
object_ids = [num for num in range(len(detections))]
detections.data["object_number"] = object_ids
# Items in `metadata` can be of any length.
detections.metadata["camera_model"] = "Luxonis OAK-D"
<!-- TODO: image -->
py.typed type hints metafile. It should provide a stronger signal to type annotators and IDEs that type support is available. (#1586)ByteTrack no longer requires detections to have a class_id (#1637)draw_line, draw_rectangle, draw_filled_rectangle, draw_polygon, draw_filled_polygon and PolygonZoneAnnotator now comes with a default color (#1591)ByteTrack, removing shared variables. Previously, multiple instances of ByteTrack would share some date, requiring liberal use of tracker.reset(). (#1603), (#1528)class_agnostic setting in MeanAveragePrecision would not work. (#1577) hacktoberfestByteTrack (#1603)
BaseTrack classRichLabelAnnotator, matching its contents with LabelAnnotator. (#1625)@onuralpszr (Onuralp SEZER), @kshitijaucharmal (KshitijAucharmal), @grzegorz-roboflow (Grzegorz Klimaszewski), @Kadermiyanyedi (Kader Miyanyedi), @PrakharJain1509 (Prakhar Jain), @DivyaVijay1234 (Divya Vijay), @souhhmm (Soham Kalburgi), @joaomarcoscrs (João Marcos Cardoso Ramos da Silva), @AHuzail (Ahmad Huzail Khan), @DemyCode (DemyCode), @ablazejuk (Andrey Blazejuk), @LinasKo (Linas Kondrackis)
A special thanks goes out to everyone who joined us for Hacktoberfest! We hope it was a rewarding experience and look forward to seeing you continue contributing and growing with our community. Keep building, keep innovating—your efforts make a difference! 🚀
Nothing published for this version
Nothing published for this version
Clarified documentation around the overlap_ratio_wh argument deprecation in InferenceSlicer. #1547
Supervision 0.24.0 is here! We've added many new changes, including the F1 score, enhancements to LineZone, EasyOCR support, NCNN support, and the best Cookbook to date! You can also try out our annotators directly in the browser. Check out the release notes to find out more!
Supervision is celebrating Hacktoberfest! Whether you're a newcomer to open source or a veteran contributor, we welcome you to join us in improving supervision. You can grab any issue without an assigned contributor: Hacktoberfest Issues Board. We'll be adding many more issues next week! 🎉
We recently launched the Model Leaderboard. Come check how the latest models perform! It is also open-source, so you can contribute to it as well! 🚀
import supervision as sv
from supervision.metrics import F1Score
predictions = sv.Detections(...)
targets = sv.Detections(...)
f1_metric = F1Score()
f1_result = f1_metric.update(predictions, targets).compute()
print(f1_result)
print(f1_result.f1_50)
print(f1_result.small_objects.f1_50)
InferenceSlicer for small object detection, and is one of the best cookbooks we've ever seen. Thank you @ediardo! #1483
LineZoneAnnotator, allowing the labels to align with the line, even when it's not horizontal. Also, you can now disable text background, and choose to draw labels off-center which minimizes overlaps for multiple LineZone labels. Thank you @jcruz-ferreyra! #854import supervision as sv
import cv2
image = cv2.imread("<SOURCE_IMAGE_PATH>")
line_zone = sv.LineZone(
start=sv.Point(0, 100),
end=sv.Point(50, 200)
)
line_zone_annotator = sv.LineZoneAnnotator(
text_orient_to_line=True,
display_text_box=False,
text_centered=False
)
annotated_frame = line_zone_annotator.annotate(
frame=image.copy(), line_counter=line_zone
)
sv.plot_image(frame)
https://github.com/user-attachments/assets/d7694b81-26ca-4236-bc66-af3d9e79d367
LineZone and introduced LineZoneAnnotatorMulticlass for visualizing the counts per class. This feature allows tracking of individual classes crossing a line, enhancing the flexibility of use cases like traffic monitoring or crowd analysis. #1555import supervision as sv
import cv2
image = cv2.imread("<SOURCE_IMAGE_PATH>")
line_zone = sv.LineZone(
start=sv.Point(0, 100),
end=sv.Point(50, 200)
)
line_zone_annotator = sv.LineZoneAnnotatorMulticlass()
frame = line_zone_annotator.annotate(
frame=frame, line_zones=[line_zone]
)
sv.plot_image(frame)
https://github.com/user-attachments/assets/b109f5bd-6ae7-473b-b4e8-910a869736b4
from_easyocr, allowing integration of OCR results into the supervision framework. EasyOCR is an open-source optical character recognition (OCR) library that can read text from images. Thank you @onuralpszr! #1515import supervision as sv
import easyocr
import cv2
image = cv2.imread("<SOURCE_IMAGE_PATH>")
reader = easyocr.Reader(["en"])
result = reader.readtext("<SOURCE_IMAGE_PATH>", paragraph=True)
detections = sv.Detections.from_easyocr(result)
box_annotator = sv.BoxAnnotator(color_lookup=sv.ColorLookup.INDEX)
label_annotator = sv.LabelAnnotator(color_lookup=sv.ColorLookup.INDEX)
annotated_image = image.copy()
annotated_image = box_annotator.annotate(scene=annotated_image, detections=detections)
annotated_image = label_annotator.annotate(scene=annotated_image, detections=detections)
sv.plot_image(annotated_image)
oriented_box_iou_batch function to detection.utils. This function computes Intersection over Union (IoU) for oriented or rotated bounding boxes (OBB), making it easier to evaluate detections with non-axis-aligned boxes. Thank you @patel-zeel! #1502import numpy as np
boxes_true = np.array([[[1, 0], [0, 1], [3, 4], [4, 3]]])
boxes_detection = np.array([[[1, 1], [2, 0], [4, 2], [3, 3]]])
ious = sv.oriented_box_iou_batch(boxes_true, boxes_detection)
print("IoU between true and detected boxes:", ious)
Note: the IoU is approximated as mask IoU.
Extended PolygonZoneAnnotator to allow setting opacity when drawing zones, providing enhanced visualization by filling the zone with adjustable transparency. Thank you @grzegorz-roboflow! #1527
Added from_ncnn, a connector for the NCNN. It is a powerful object detection framework from Tencent, written from ground-up in C++, with no third party dependencies. Thank you @onuralpszr! #1524
import cv2
from ncnn.model_zoo import get_model
import supervision as sv
image = cv2.imread("<SOURCE_IMAGE_PATH>")
model = get_model(
"yolov8s",
target_size=640,
prob_threshold=0.5,
nms_threshold=0.45,
num_threads=4,
use_gpu=True,
)
result = model(image)
detections = sv.Detections.from_ncnn(result)
Supervision now depends on opencv-python rather than opencv-python-headless. #1530
Fixed broken or outdated links in documentation and notebooks, improving navigation and ensuring accuracy of references. Thanks to @capjamesg for identifying these issues. #1523
Enabled and fixed Ruff rules for code formatting, including changes like avoiding unnecessary iterable allocations and using Optional for default mutable arguments. #1526
Updated the COCO 101 point Average Precision algorithm to correctly interpolate precision, providing a more precise calculation of average precision without averaging out intermediate values. #1500
Resolved miscellaneous issues highlighted when building documentation. This mostly includes whitespace adjustments and type inconsistencies. Updated documentation for clarity and fixed formatting issues. Added explicit version for mkdocstrings-python. #1549
Clarified documentation around the overlap_ratio_wh argument deprecation in InferenceSlicer. #1547
frame_resolution_wh parameter in PolygonZone has been removed due to deprecation.pip install supervision[headless] will install the base library and warn of non-existent extra.@onuralpszr (Onuralp SEZER), @joaomarcoscrs (João Marcos Cardoso Ramos da Silva), @jcruz-ferreyra (Juan Cruz), @patel-zeel (Zeel B Patel), @grzegorz-roboflow (Grzegorz Klimaszewski), @Kadermiyanyedi (Kader Miyanyedi), @ediardo (Eddie Ramirez), @CharlesCNorton, @ethanwhite (Ethan White), @josephofiowa (Joseph Nelson), @tibeoh (Thibault Itart-Longueville), @SkalskiP (Piotr Skalski), @LinasKo (Linas Kondrackis)
Thank you to Pexels for providing fantastic images and videos!
Nothing published for this version
overlap_filter_strategy in InferenceSlicer.__init__ is deprecated and will be removed in supervision-0.27.0. Use overlap_strategy instead.
BackgroundOverlayAnnotator annotates the background of your image! #1385https://github.com/user-attachments/assets/c1f3ce11-08c1-4648-9176-4e7920b91a8a
(video by Pexels)
xyxy boxes and masks. Over the next few releases, supervision will focus on adding more metrics, allowing you to evaluate your model performance. We plan to support not just boxes, masks, but oriented bounding boxes as well! #1442[!TIP] Help in implementing metrics is very welcome! Keep an eye on our issue board if you'd like to contribute!
import supervision as sv
from supervision.metrics import MeanAveragePrecision
predictions = sv.Detections(...)
targets = sv.Detections(...)
map_metric = MeanAveragePrecision()
map_result = map_metric.update(predictions, targets).compute()
print(map_result)
print(map_result.map50_95)
print(map_result.large_objects.map50_95)
map_result.plot()
Here's a very basic way to compare model results:
<details> <summary>📊 Example code</summary>
import supervision as sv
from supervision.metrics import MeanAveragePrecision
from inference import get_model
import matplotlib.pyplot as plt
# !wget https://media.roboflow.com/notebooks/examples/dog.jpeg
image = "dog.jpeg"
model_1 = get_model("yolov8n-640")
model_2 = get_model("yolov8s-640")
model_3 = get_model("yolov8m-640")
model_4 = get_model("yolov8l-640")
results_1 = model_1.infer(image)[0]
results_2 = model_2.infer(image)[0]
results_3 = model_3.infer(image)[0]
results_4 = model_4.infer(image)[0]
detections_1 = sv.Detections.from_inference(results_1)
detections_2 = sv.Detections.from_inference(results_2)
detections_3 = sv.Detections.from_inference(results_3)
detections_4 = sv.Detections.from_inference(results_4)
map_n_metric = MeanAveragePrecision().update([detections_1], [detections_4]).compute()
map_s_metric = MeanAveragePrecision().update([detections_2], [detections_4]).compute()
map_m_metric = MeanAveragePrecision().update([detections_3], [detections_4]).compute()
labels = ["YOLOv8n", "YOLOv8s", "YOLOv8m"]
map_values = [map_n_metric.map50_95, map_s_metric.map50_95, map_m_metric.map50_95]
plt.title("YOLOv8 Model Comparison")
plt.bar(labels, map_values)
ax = plt.gca()
ax.set_ylim([0, 1])
plt.show()
</details>
IconAnnotator, which allows you to place icons on your images. #930https://github.com/user-attachments/assets/ff80acf5-67f2-4c20-a3fe-b63cac07ae31
(Video by Pexels, icons by Icons8)
import supervision as sv
from inference import get_model
image = <SOURCE_IMAGE_PATH>
icon_dog = <DOG_PNG_PATH>
icon_cat = <CAT_PNG_PATH>
model = get_model(model_id="yolov8n-640")
results = model.infer(image)[0]
detections = sv.Detections.from_inference(results)
icon_paths = []
for class_name in detections.data["class_name"]:
if class_name == "dog":
icon_paths.append(icon_dog)
elif class_name == "cat":
icon_paths.append(icon_cat)
else:
icon_paths.append("")
icon_annotator = sv.IconAnnotator()
annotated_frame = icon_annotator.annotate(
scene=image.copy(),
detections=detections,
icon_path=icon_paths
)
from_sam, we've added support to from_ultralytics for loading the results if you ran it with Ultralytics. #1354import cv2
import supervision as sv
from ultralytics import SAM
image = cv2.imread("...")
model = SAM("mobile_sam.pt")
results = model(image, bboxes=[[588, 163, 643, 220]])
detections = sv.Detections.from_ultralytics(results[0])
polygon_annotator = sv.PolygonAnnotator()
mask_annotator = sv.MaskAnnotator()
annoated_image = mask_annotator.annotate(image.copy(), detections)
annoated_image = polygon_annotator.annotate(annoated_image, detections)
sv.plot_image(annoated_image, (12,12))
SAM2 with our annotators:
https://github.com/user-attachments/assets/6a98d651-2596-43e9-b485-ea6f0de4fffa
TriangleAnnotator and DotAnnotator contour color customization #1458VertexLabelAnnotator for keypoints now has text_color parameter #1409sv.Detections.from_transformers to support the transformers v5 functions. This includes the DetrImageProcessor methods post_process_object_detection, post_process_panoptic_segmentation, post_process_semantic_segmentation, and post_process_instance_segmentation. #1386InferenceSlicer now features an overlap_ratio_wh parameter, making it easier to compute slice sizes when handling overlapping slices. #1434image_with_small_objects = cv2.imread("...")
model = get_model("yolov8n-640")
def callback(image_slice: np.ndarray) -> sv.Detections:
print("image_slice.shape:", image_slice.shape)
result = model.infer(image_slice)[0]
return sv.Detections.from_inference(result)
slicer = sv.InferenceSlicer(
callback=callback,
slice_wh=(128, 128),
overlap_ratio_wh=(0.2, 0.2),
)
detections = slicer(image_with_small_objects)
plot_image now clearly states the size is in inches. #1424overlap_filter_strategy in InferenceSlicer.__init__ is deprecated and will be removed in supervision-0.27.0. Use overlap_strategy instead.overlap_ratio_wh in InferenceSlicer.__init__ is deprecated and will be removed in supervision-0.27.0. Use overlap_wh instead.track_buffer, track_thresh, and match_thresh parameters in ByteTrack are deprecated and were removed as of supervision-0.23.0. Use lost_track_buffer, track_activation_threshold, and minimum_matching_threshold instead.triggering_position parameter in sv.PolygonZone was removed as of supervision-0.23.0. Use triggering_anchors instead.@shaddu, @onuralpszr (Onuralp SEZER), @Kadermiyanyedi (Kader Miyanyedi), @xaristeidou (Christoforos Aristeidou), @Gk-rohan (Rohan Gupta), @Bhavay-2001 (Bhavay Malhotra), @arthurcerveira (Arthur Cerveira), @J4BEZ (Ju Hoon Park), @venkatram-dev, @eric220, @capjamesg (James), @yeldarby (Brad Dwyer), @SkalskiP (Piotr Skalski), @LinasKo (LinasKo)
`sv.KeyPoints.from_mediapipe` adding support for Mediapipe keypoint models (both legacy and modern), along with default visualizers for face and body
sv.KeyPoints.from_mediapipe adding support for Mediapipe keypoint models (both legacy and modern), along with default visualizers for face and body pose keypoints. (#1232, #1316)import numpy as np
import mediapipe as mp
import supervision as sv
from PIL import Image
model = mp.solutions.face_mesh.FaceMesh()
edge_annotator = sv.EdgeAnnotator(color=sv.Color.BLACK, thickness=2)
image = Image.open(<PATH_TO_IMAGE>).convert('RGB')
results = model.process(np.array(image))
key_points = sv.KeyPoints.from_mediapipe(results, resolution_wh=image.size)
annotated_image = edge_annotator.annotate(scene=image, key_points=key_points)
https://github.com/user-attachments/assets/883a6bcc-5e39-41b0-9b6d-0348b5b2fe0e
sv.KeyPoints.from_detectron2 and sv.Detections.from_detectron2 extending support for Detectron2 models. (#1310, #1300)
sv.RichLabelAnnotator allowing to draw unicode characters (e.g. from non-latin languages), as long as you provide a compatible font. (#1277)
https://github.com/user-attachments/assets/de60eeb4-1259-421b-af66-f622a15988ea
sv.DetectionsDataset and sv.ClassificationDataset allowing to load the images into memory only when necessary (lazy loading). (#1326)import roboflow
from roboflow import Roboflow
import supervision as sv
roboflow.login()
rf = Roboflow()
project = rf.workspace(<WORKSPACE_ID>).project(<PROJECT_ID>)
dataset = project.version(<PROJECT_VERSION>).download("coco")
ds_train = sv.DetectionDataset.from_coco(
images_directory_path=f"{dataset.location}/train",
annotations_path=f"{dataset.location}/train/_annotations.coco.json",
)
path, image, annotation = ds_train[0]
# loads image on demand
for path, image, annotation in ds_train:
# loads image on demand
sv.Detections.from_lmm allowing to parse Florence-2 text result into sv.Detections object. (#1296)sv.DotAnnotator and sv.TriangleAnnotator allowing to add marker outlines. (#1294)sv.ColorAnnotator and sv.CropAnnotator buggy behaviours. (#1277, #1312)This release, @onuralpszr added two new Cookbooks to our collection. Check them out to learn how to save Detections to a file and convert it back to Detections!
@onuralpszr (Onuralp SEZER), @David-rn (David Redó), @jeslinpjames (Jeslin P James), @Bhavay-2001 (Bhavay Malhotra), @hardikdava (Hardik Dava), @kirilman, @dsaha21 (Dripto Saha), @cdragos (Dragos Catarahia), @mqasim41 (Muhammad Qasim), @SkalskiP (Piotr Skalski), @LinasKo (Linas Kondrackis)
Special thanks to @rolson24 (Raif Olson) for helping the community with ByteTrack!
Nothing published for this version
The supervision-0.21.0 release is around the corner. Here is the timeline:
The supervision-0.21.0 release is around the corner. Here is the timeline:
5 Jun 2024 08:00 PM CEST (UTC +2) / 5 Jun 2024 11:00 AM PDT (UTC -7) - merge develop into main - closing list supervision-0.21.0 features6 Jun 2024 11:00 AM CEST (UTC +2) / 6 Jun 2024 02:00 AM PDT (UTC -7) - release supervision-0.21.0sv.Detections.with_nmm to perform non-maximum merging on the current set of object detections. (#500)sv.Detections.from_lmm allowing to parse Large Multimodal Model (LMM) text result into sv.Detections object. For now from_lmm supports only PaliGemma result parsing. (#1221)import supervision as sv
paligemma_result = "<loc0256><loc0256><loc0768><loc0768> cat"
detections = sv.Detections.from_lmm(
sv.LMM.PALIGEMMA,
paligemma_result,
resolution_wh=(1000, 1000),
classes=['cat', 'dog']
)
detections.xyxy
# array([[250., 250., 750., 750.]])
detections.class_id
# array([0])
sv.VertexLabelAnnotator allowing to annotate every vertex of a keypoint skeleton with custom text and color. (#1236)import supervision as sv
image = ...
key_points = sv.KeyPoints(...)
LABELS = [
"nose", "left eye", "right eye", "left ear",
"right ear", "left shoulder", "right shoulder", "left elbow",
"right elbow", "left wrist", "right wrist", "left hip",
"right hip", "left knee", "right knee", "left ankle",
"right ankle"
]
COLORS = [
"#FF6347", "#FF6347", "#FF6347", "#FF6347",
"#FF6347", "#FF1493", "#00FF00", "#FF1493",
"#00FF00", "#FF1493", "#00FF00", "#FFD700",
"#00BFFF", "#FFD700", "#00BFFF", "#FFD700",
"#00BFFF"
]
COLORS = [sv.Color.from_hex(color_hex=c) for c in COLORS]
vertex_label_annotator = sv.VertexLabelAnnotator(
color=COLORS,
text_color=sv.Color.BLACK,
border_radius=5
)
annotated_frame = vertex_label_annotator.annotate(
scene=image.copy(),
key_points=key_points,
labels=labels
)
sv.KeyPoints.from_inference and sv.KeyPoints.from_yolo_nas allowing to create sv.KeyPoints from Inference and YOLO-NAS result. (#1147 and #1138)
sv.mask_to_rle and sv.rle_to_mask allowing for easy conversion between mask and rle formats. (#1163)
sv.InferenceSlicer allowing to select overlap filtering strategy (NONE, NON_MAX_SUPPRESSION and NON_MAX_MERGE). (#1236)
sv.InferenceSlicer adding instance segmentation model support. (#1178)
import cv2
import numpy as np
import supervision as sv
from inference import get_model
model = get_model(model_id="yolov8x-seg-640")
image = cv2.imread(<SOURCE_IMAGE_PATH>)
def callback(image_slice: np.ndarray) -> sv.Detections:
results = model.infer(image_slice)[0]
return sv.Detections.from_inference(results)
slicer = sv.InferenceSlicer(callback = callback)
detections = slicer(image)
mask_annotator = sv.MaskAnnotator()
label_annotator = sv.LabelAnnotator()
annotated_image = mask_annotator.annotate(
scene=image, detections=detections)
annotated_image = label_annotator.annotate(
scene=annotated_image, detections=detections)
sv.LineZone making it 10-20 times faster, depending on the use case. (#1228)sv.DetectionDataset.from_coco and sv.DetectionDataset.as_coco adding support for run-length encoding (RLE) mask format. (#1163)@onuralpszr (Onuralp SEZER), @LinasKo (Linas Kondrackis), @rolson24 (Raif Olson), @mario-dg (Mario da Graca), @xaristeidou (Christoforos Aristeidou), @ManzarIMalik (Manzar Iqbal Malik), @tc360950 (Tomasz Cąkała), @emSko, @SkalskiP (Piotr Skalski)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
> The frame_resolution_wh parameter in sv.PolygonZone is deprecated and will be removed in supervision-0.24.0.
sv.KeyPoints to provide initial support for pose estimation and broader keypoint detection models. (#1128)
sv.EdgeAnnotator and sv.VertexAnnotator to enable rendering of results from keypoint detection models. (#1128)
import cv2
import supervision as sv
from ultralytics import YOLO
image = cv2.imread(<SOURCE_IMAGE_PATH>)
model = YOLO('yolov8l-pose')
result = model(image, verbose=False)[0]
keypoints = sv.KeyPoints.from_ultralytics(result)
edge_annotators = sv.EdgeAnnotator(color=sv.Color.GREEN, thickness=5)
annotated_image = edge_annotators.annotate(image.copy(), keypoints)
import cv2
import supervision as sv
from ultralytics import YOLO
image = cv2.imread(<SOURCE_IMAGE_PATH>)
model = YOLO('yolov8l-pose')
result = model(image, verbose=False)[0]
keypoints = sv.KeyPoints.from_ultralytics(result)
vertex_annotators = sv.VertexAnnotator(color=sv.Color.GREEN, radius=10)
annotated_image = vertex_annotators.annotate(image.copy(), keypoints)
sv.LabelAnnotator by adding an additional corner_radius argument that allows for rounding the corners of the bounding box. (#1037)
sv.PolygonZone such that the frame_resolution_wh argument is no longer required to initialize sv.PolygonZone. (#1109)
[!WARNING]
Theframe_resolution_whparameter insv.PolygonZoneis deprecated and will be removed insupervision-0.24.0.
sv.get_polygon_center to calculate a more accurate polygon centroid. (#1084)
sv.Detections.from_transformers by adding support for Transformers segmentation models and extract class names values. (#1069)
import torch
import supervision as sv
from PIL import Image
from transformers import DetrImageProcessor, DetrForSegmentation
processor = DetrImageProcessor.from_pretrained("facebook/detr-resnet-50-panoptic")
model = DetrForSegmentation.from_pretrained("facebook/detr-resnet-50-panoptic")
image = Image.open(<SOURCE_IMAGE_PATH>)
inputs = processor(images=image, return_tensors="pt")
with torch.no_grad():
outputs = model(**inputs)
width, height = image.size
target_size = torch.tensor([[height, width]])
results = processor.post_process_segmentation(
outputs=outputs, target_sizes=target_size)[0]
detections = sv.Detections.from_transformers(results, id2label=model.config.id2label)
mask_annotator = sv.MaskAnnotator()
label_annotator = sv.LabelAnnotator(text_position=sv.Position.CENTER)
annotated_image = mask_annotator.annotate(
scene=image, detections=detections)
annotated_image = label_annotator.annotate(
scene=annotated_image, detections=detections)
sv.ByteTrack.update_with_detections which was removing segmentation masks while tracking. Now, ByteTrack can be used alongside segmentation models. (#787)@onuralpszr (Onuralp SEZER), @rolson24 (Raif Olson), @xaristeidou (Christoforos Aristeidou), @jeslinpjames (Jeslin P James), @Griffin-Sullivan (Griffin Sullivan), @PawelPeczek-Roboflow (Paweł Pęczek), @pirnerjonas (Jonas Pirner), @sharingan000, @macc-n, @LinasKo (Linas Kondrackis), @SkalskiP (Piotr Skalski)
Nothing published for this version
Nothing published for this version
> The track_buffer, track_thresh, and match_thresh parameters in sv.ByterTrack are deprecated and will be removed in supervision-0.23.0. Use lost_trac…
Supervision Cookbooks - A curated open-source collection crafted by the community, offering practical examples, comprehensive guides, and walkthroughs for leveraging Supervision alongside diverse Computer Vision models. (#860)
sv.CSVSink allowing for the straightforward saving of image, video, or stream inference results in a .csv file. (#818)import supervision as sv
from ultralytics import YOLO
model = YOLO(<SOURCE_MODEL_PATH>)
csv_sink = sv.CSVSink(<RESULT_CSV_FILE_PATH>)
frames_generator = sv.get_video_frames_generator(<SOURCE_VIDEO_PATH>)
with csv_sink:
for frame in frames_generator:
result = model(frame)[0]
detections = sv.Detections.from_ultralytics(result)
csv_sink.append(detections, custom_data={<CUSTOM_LABEL>:<CUSTOM_DATA>})
https://github.com/roboflow/supervision/assets/26109316/621588f9-69a0-44fe-8aab-ab4b0ef2ea1b
sv.JSONSink allowing for the straightforward saving of image, video, or stream inference results in a .json file. (#819)import supervision as sv
from ultralytics import YOLO
model = YOLO(<SOURCE_MODEL_PATH>)
json_sink = sv.JSONSink(<RESULT_JSON_FILE_PATH>)
frames_generator = sv.get_video_frames_generator(<SOURCE_VIDEO_PATH>)
with json_sink:
for frame in frames_generator:
result = model(frame)[0]
detections = sv.Detections.from_ultralytics(result)
json_sink.append(detections, custom_data={<CUSTOM_LABEL>:<CUSTOM_DATA>})
sv.mask_iou_batch allowing to compute Intersection over Union (IoU) of two sets of masks. (#847)sv.mask_non_max_suppression allowing to perform Non-Maximum Suppression (NMS) on segmentation predictions. (#847)sv.CropAnnotator allowing users to annotate the scene with scaled-up crops of detections. (#888)import cv2
import supervision as sv
from inference import get_model
image = cv2.imread(<SOURCE_IMAGE_PATH>)
model = get_model(model_id="yolov8n-640")
result = model.infer(image)[0]
detections = sv.Detections.from_inference(result)
crop_annotator = sv.CropAnnotator()
annotated_frame = crop_annotator.annotate(
scene=image.copy(),
detections=detections
)
https://github.com/roboflow/supervision/assets/26109316/0a5b67ce-55e7-4e26-9495-a68f9ad97ec7
sv.ByteTrack.reset allowing users to clear trackers state, enabling the processing of multiple video files in sequence. (#827)sv.LineZoneAnnotator allowing to hide in/out count using display_in_count and display_out_count properties. (#802)sv.ByteTrack input arguments and docstrings updated to improve readability and ease of use. (#787)[!WARNING]
Thetrack_buffer,track_thresh, andmatch_threshparameters insv.ByterTrackare deprecated and will be removed insupervision-0.23.0. Uselost_track_buffer,track_activation_threshold, andminimum_matching_thresholdinstead.
sv.PolygonZone to now accept a list of specific box anchors that must be in zone for a detection to be counted. (#910)[!WARNING]
Thetriggering_positionparameter insv.PolygonZoneis deprecated and will be removed insupervision-0.23.0. Usetriggering_anchorsinstead.
sv.DetectionsSmoother removing tracking_id from sv.Detections. (#944)sv.DetectionDataset which, after changes introduced in supervision-0.18.0, failed to load datasets in YOLO, PASCAL VOC, and COCO formats.@onuralpszr (Onuralp SEZER), @LinasKo (Linas Kondrackis), @LeviVasconcelos (Levi Vasconcelos), @AdonaiVera (Adonai Vera), @xaristeidou (Christoforos Aristeidou), @Kadermiyanyedi (Kader Miyanyedi), @NickHerrig (Nick Herrig), @PacificDou (Shuyang Dou), @iamhatesz (Tomasz Wrona), @capjamesg (James Gallagher), @sansyo, @SkalskiP (Piotr Skalski)
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →