NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2387 most downloaded on PyPI
Fast, flexible, and advanced augmentation library for deep learning, computer vision, and medical imaging. Albumentations offers a wide range of transformations for both 2D (images, masks, bboxes, keypoints) and 3D (volumes, volumetric masks, keypoints) data, with optimized performance and seamless integration into ML workflows.
Last release 1 years ago
27 May 2025
Ships fairly regularly
a new release about every 2 weeks
Most releases are documented
notes for 48 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
87 releases · first in 2018
Want to stay in the loop? Get updates on new features, documentation changes, and tools like the UI explorer by subscribing to our mailing list: 👉 htt
Want to stay in the loop?
Get updates on new features, documentation changes, and tools like the UI explorer by subscribing to our mailing list:
👉 https://albumentations.ai/subscribe/
You can unsubscribe anytime.
area_for_downscale to RandomResizedCrop and RandomSizedCropparameter may have value:
imageimage_maskWhen enabled will use interpolation that was passed to the transform for upscale, but cv2.INTER_AREA for downscale, as for downscale INTER_AREA generates the least amount of artifacts.
Vectorized application to videos and volume in
When applied to videos Albumentations on 1 CPU core is still, slower than torchvision on GTX 4090, (Benchmark on videos). But with such pull requests, the gap. hopefully, will get smaller.
One column per quarter.
Want to stay in the loop? Get updates on new features, documentation changes, and tools like the UI explorer by subscribing to our mailing list: 👉 htt
Want to stay in the loop?
Get updates on new features, documentation changes, and tools like the UI explorer by subscribing to our mailing list:
👉 https://albumentations.ai/subscribe/
You can unsubscribe anytime.
transform(image=image, masks=[])
area_for_downscale parameterAdded to: • RandomScale • LongestMaxSize • SmallestMaxSize • Resize
area_for_downscale options:
• None – default behavior
• "image" – use cv2.INTER_AREA when image is downscaled
• "image_mask" – use cv2.INTER_AREA for both image and mask
Using cv2.INTER_AREA for downscaling helps reduce artifacts, unlike other interpolation methods.
🐛 Bug Fixes • Fixed serialization in ToFloat
✅ TL;DR • ✅ Empty mask lists now supported • ✅ area_for_downscale improves downscaling quality • 🐞 Serialization fix in ToFloat • 💌 Subscribe for updates
Help Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
<img width="1134" alt="Screenshot 2025-04-28 at 6 33 47 PM" src="https://github.com/user-attachments/assets/44252890-28ff-4576-aec3-3bea37cbc554" />
Generalization of Mosaic from Ultralitics and YOLO4, and works per image an not on "batch" => can choose what additional images to pass, could be hard or rare classes.
by @Shysto and @ternaus
Changed functionality to a more intuituve
Now it works as:
n transforms with equal probabilityRemoved to pass labels when apply to bounding boxes.
In [9]: bboxes = np.array([[0.2, 0.2, 0.4, 0.4], [0.3, 0.4, 0.7, 0.9]])
In [10]: transform = A.Compose([A.HorizontalFlip(p=1)], bbox_params={"format": "albumentations"})
In [11]: image = np.random.rand(640, 640, 3)
In [12]: transformed = transform(image=image, bboxes=bboxes)
=> we can just pass coordinates, without bounding box labels
When applied to uint images on 1 CPU core Albumentations outperforms Kornia and torchvision: Image benchmark
But when we compare:
Albumentations has a lot to improve. Benchmark on videos
=> Speedups on videos in this release:
drop_length was not used beforeexact and approximate modeHelp Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
Reverted to the earlier version of the algorithm, as it generated more naturally looking effects.
This is a direct alias to the D4 transform.
If your problem has square symmetry, meaning you can perform flips, transpose, rotation by 90 degrees but it is better to use SquareSymmetry or D4 directly as it ensures that all 8 orientations are applied with the same probability.
Help Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
Applies H&E (Hematoxylin and Eosin) stain augmentation to histopathology images.
This transform simulates different H&E staining conditions using either:
1. Predefined stain matrices (8 standard references)
2. Vahadane method for stain extraction
3. Macenko method for stain extraction
4. Custom stain matrices
<img width="1484" alt="Screenshot 2025-02-09 at 6 15 50 PM" src="https://github.com/user-attachments/assets/cdb9d823-aef3-4045-8f79-75a0b3334b1b" />
There was a lot of deprecation in the last year. It may happen that your augmentation pipeline does not behave as expected, as parameters that use in…
Extended the functionality of the strict parameter in Compose.
Now, if strict=True and you pass incorrect arguments to transforms in Compose => will get error.
if strict=False you will get only a warning
There was a lot of deprecation in the last year. It may happen that your augmentation pipeline does not behave as expected, as parameters that use in transforms are ignored, and default values are used instead.
np.arrayas labels in BboxParamsHelp Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
Added parameter max_accept_ratio to BBoxParams
max_accept_ratio (float | None): Maximum allowed aspect ratio for bounding boxes.
The aspect ratio is calculated as max(width/height, height/width), so it's always >= 1.
Boxes with aspect ratio greater than this value will be filtered out.
For example, if `max_accept_ratio=3.0`, boxes with width:height or height:width ratios
greater than 3:1 will be removed. Set to None to disable aspect ratio filtering. Default: None.
clip=True in BboxParams, was clipping not only boxes, but class labels if passed as numpy arrayPlasmaShadow, PlasmaBrightnessContrast, ChannelShuffleHelp Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
Added parameter filter_invalid_bboxes to BboxParams.
If True, filters out invalid bounding boxes (e.g., boxes with negative dimensions or boxes where x_max < x_min or y_max < y_min at the beginning of the pipeline. If clip=True, filtering is applied after clipping. Default: False.
FromFloat, when applied to images, volume, volumesfillsigma. Fixed. Also now matches behavior in PIL pretty closeall parameter renaming was moved through deprecations => you got deprecation warning for months
This is major release, meaning
only one new transform
a lot of changes.
If you have questions or proposals:
If you have complaints:
<img width="1224" alt="Screenshot 2025-01-03 at 5 58 27 PM" src="https://github.com/user-attachments/assets/71669a29-fb4d-43a9-9e77-d0746cfdcfb2" />
by @vedantdalimkar
always_apply => use p=1 to always apply and p=0 for not applying.update_params, get_params_dependent_on_targets => use get_params_dependent_on_datavar_limit, meanstd_range, mean_rangeIt is not just a renaming, var_limit and std_range sample from different distributions. Sampling from std_range matches with other libraries like torchvision.
sigmaX_limit, sigmaY_limitsigma_x_limit, sigma_y_limitpad_mode, pad_val_mask, pad_cvlborder_mode, fill_mask, fillpad_mode, pad_val_mask, pad_cvlborder_mode, fill_mask, fillpad_mode, pad_val_mask, pad_cvlborder_mode, fill_mask, fillheight, widthsizeheight, widthsizecropping_box_keycropping_bbox_keypad_mode, pad_val_mask, pad_cvlborder_mode, fill_mask, filltemplate_weightfill_valuefillmin_holes, max_holes, min_height, max_height, min_width, max_width, mask_fill_value, fill_valuenum_holes_range, hole_height_range, hole_width_range, fill, fill_maskAlso default parameters changed:
num_height_range = (8, 8) => num_height_range = (0.1, 0.2)
num_width_range = (8, 8) => num_width_range = (0.1, 0.2)
unit_size_min, unit_size_max, holes_number_x, holes_number_y, shift_x, shift_y, fill_value, mask_fill_valueunit_size_range, holes_number_xy, fill, fill_maskimage_fill_value, mask_fill_valuefill, fill_maskmask_fill_value, fill_valuefill, fill_maskvalue, mask_valuefill, fill_maskChanged default value for border_mode from cv2.BORDER_REFLECT_101 to cv2.BORDER_CONSTANT
value, mask_valuefill, fill_maskChanged default value for border_mode from cv2.BORDER_REFLECT_101 to cv2.BORDER_CONSTANT
border_mode, value, mask_valuepad_mode, pad_val, mask_pad_valcval, cval_mask, modefill, fill_mask, border_modevalue, mask_valuefill, fill_maskChanged default border_mode from cv2.BORDER_REFLECT_101 to cv2.BORDER_CONSTANT
cval, cval_mask, mode, keypoints_thresholdshift_limit, value, mask_value, border_modevalue, mask_value, border_modeChanged default probability from p=0.5 to p=1
value, mask_valuefill, fill_maskChanged default value for border_mode from cv2.BORDER_REFLECT_101 to cv2.BORDER_CONSTANT
quality_lower, quality_upperquality_rangesnow_point_lower, snow_point_uppersnow_point_rangeslant_lower, slant_upperslant_rangefog_coef_lower, fog_coef_upperfog_coef_rangeangle_lower, angle_upper, num_flare_circles_lower, num_flare_circles_uppernum_flare_circles_range, angle_rangenum_shadows_lower, num_shadows_uppernum_shadows_limitthresholdthreshold_rangeinterpolation, scale_min, scale_maxinterpolation_pair, scale_rangeby @ternaus
Help Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
xyz for ImageOnly and Dual transforms (z coordinate stays unchanged)Crop an area from image while ensuring at least one bounding box is present in the crop.
<img width="1201" alt="Screenshot 2024-12-24 at 1 46 24 PM" src="https://github.com/user-attachments/assets/21fa3971-3aa2-442e-bdfa-a076541c1096" />
by @guillaume-rochette-oxb
CenterCrop3D, CoarseDropout3D, CubicSymmetry, Pad3D, PadIfNeeded3D, RandomCrop3D (by @ternaus)eval-type-backport for python 3.10 and older. by @PerchunPakToTensorV2 by @matejpekarHelp Us Grow - If you find value in Albumentations, consider becoming a sponsor. Every contribution, no matter the size, helps us maintain and improve
images as numpy arrayNow supports numpy arrays with shape (num_images, height, width, num_channels) or (num_images, height, width) as images in Compose
(depth, height, width) or (depth, height, width, num_channels)(depth, height, width) or (depth, height, width, num_channels)(num_volumes, depth, height, width) for batch processing(num_volumes, depth, height, width) for batch processingvolume = np.random.rand(96, 256, 256) # Your 3D medical volume
mask = np.zeros((96, 256, 256)) # Your 3D segmentation mask
transformed = transform(volume=volume, mask3d=mask)
transformed_volume = transformed['volume']
transformed_mask = transformed['mask3d']
Added 3D transforms by @ternaus
transform = A.Compose([
# Crop volume to a fixed size for memory efficiency
A.RandomCrop3D(size=(64, 128, 128), p=1.0),
# Randomly remove cubic regions to simulate occlusions
A.CoarseDropout3D(
num_holes_range=(2, 6),
hole_depth_range=(0.1, 0.3),
hole_height_range=(0.1, 0.3),
hole_width_range=(0.1, 0.3),
p=0.5
),
])
volume = np.random.rand(96, 256, 256) # Your 3D medical volume
mask = np.zeros((96, 256, 256)) # Your 3D segmentation mask
transformed = transform(volume=volume, mask3d=mask)
transformed_volume = transformed['volume']
transformed_mask = transformed['mask3d']
Deprecated parameters border_mode, value, mask_value - you can specify them, but will not have any effect.
noise_distribution that allows sampling displacement fields from gaussian and from uniform distributions.border_mode, value, mask_value - you can specify them, but will not have any effect.<img width="831" alt="Screenshot 2024-12-06 at 10 34 34" src="https://github.com/user-attachments/assets/b1fd6ffc-ed35-4065-bafa-9ea679eea176">
Apply shot noise to the image by modeling photon counting as a Poisson process.
Shot noise (also known as Poisson noise) occurs in imaging due to the quantum nature of light.
When photons hit an imaging sensor, they arrive at random times following Poisson statistics.
This transform simulates this physical process in linear light space by:
1. Converting to linear space (removing gamma)
2. Treating each pixel value as an expected photon count
3. Sampling actual photon counts from a Poisson distribution
4. Converting back to display space (reapplying gamma)
The noise characteristics follow real camera behavior:
- Noise variance equals signal mean in linear space (Poisson statistics)
- Brighter regions have more absolute noise but less relative noise
- Darker regions have less absolute noise but more relative noise
- Noise is generated independently for each pixel and color channel
Addes support for bounding boxes
<img width="823" alt="Screenshot 2024-12-06 at 10 38 44" src="https://github.com/user-attachments/assets/e7fbeac0-f92b-4097-838f-d5ddaab9c68f">
Added an option to inpaint holes using inpaint_ns and inpaint_telea from OpenCV
Added an option to inpaint holes using inpaint_ns and inpaint_telea from OpenCV
Added an option to inpaint holes using inpaint_ns and inpaint_telea from OpenCV
Added an option to inpaint holes using inpaint_ns and inpaint_telea from OpenCV
Added NewTransform TimeReverse
Reverse the time axis of a spectrogram image, also known as time inversion.
Time inversion of a spectrogram is analogous to the random flip of an image,
an augmentation technique widely used in the visual domain. This can be relevant
in the context of audio classification tasks when working with spectrograms.
The technique was successfully applied in the AudioCLIP paper, which extended
CLIP to handle image, text, and audio inputs.
This transform is implemented as a subclass of HorizontalFlip since reversing
time in a spectrogram is equivalent to flipping the image horizontally.
Added NewTransform TimeMasking
Apply masking to a spectrogram in the time domain.
This transform masks random segments along the time axis of a spectrogram,
implementing the time masking technique proposed in the SpecAugment paper.
Time masking helps in training models to be robust against temporal variations
and missing information in audio signals.
This is a specialized version of XYMasking configured for time masking only.
For more advanced use cases (e.g., multiple masks, frequency masking, or custom
fill values), consider using XYMasking directly.
Apply masking to a spectrogram in the frequency domain.
This transform masks random segments along the frequency axis of a spectrogram,
implementing the frequency masking technique proposed in the SpecAugment paper.
Frequency masking helps in training models to be robust against frequency variations
and missing spectral information in audio signals.
This is a specialized version of XYMasking configured for frequency masking only.
For more advanced use cases (e.g., multiple masks, time masking, or custom
fill values), consider using XYMasking directly.
Added NewTransform FrequencyMasking
It is a specialized version of XYMasking that has the similar API as FrequencyMasking from torchaudio
<img width="1192" alt="Screenshot 2024-12-06 at 11 19 42" src="https://github.com/user-attachments/assets/60d597ac-9c3a-4324-9b30-d66c37c6dd18">
Pad the sides of an image by specified number of pixels.
Args:
padding (int, tuple[int, int] or tuple[int, int, int, int]): Padding values. Can be:
* int - pad all sides by this value
* tuple[int, int] - (pad_x, pad_y) to pad left/right by pad_x and top/bottom by pad_y
* tuple[int, int, int, int] - (left, top, right, bottom) specific padding per side
This is the generalization of the torchvision transform with the same name
<img width="1199" alt="Screenshot 2024-12-06 at 11 23 25" src="https://github.com/user-attachments/assets/8bf42b14-7c09-4bb3-8e61-2ea7b1ea16e7">
This is the generalization of the similar torchvision transform
Randomly erases rectangular regions in an image, following the Random Erasing Data Augmentation technique.
This augmentation helps improve model robustness by randomly masking out rectangular regions in the image,
simulating occlusions and encouraging the model to learn from partial information. It's particularly
effective for image classification and person re-identification tasks.
<img width="1198" alt="Screenshot 2024-12-06 at 11 26 17" src="https://github.com/user-attachments/assets/557c6dff-01a7-4fe2-a0fd-9073568bcd87">
Apply random noise to image channels using various noise distributions.
This transform generates noise using different probability distributions and applies it
to image channels. The noise can be generated in three spatial modes and supports
multiple noise distributions, each with configurable parameters.
Args:
noise_type: Type of noise distribution to use. Options:
- "uniform": Uniform distribution, good for simple random perturbations
- "gaussian": Normal distribution, models natural random processes
- "laplace": Similar to Gaussian but with heavier tails, good for outliers
- "beta": Flexible bounded distribution, can be symmetric or skewed
spatial_mode: How to generate and apply the noise. Options:
- "constant": One noise value per channel, fastest
- "per_pixel": Independent noise value for each pixel and channel, slowest
- "shared": One noise map shared across all channels, medium speed
Added 'gaussian' method for image sharpening.
<img width="1181" alt="Screenshot 2024-12-06 at 11 52 54" src="https://github.com/user-attachments/assets/b93c1863-7db1-4aac-ba18-97cefad43dad">
Apply salt and pepper noise to the input image.
Salt and pepper noise is a form of impulse noise that randomly sets pixels to either maximum value (salt)
or minimum value (pepper). The amount and proportion of salt vs pepper noise can be controlled.
<img width="1169" alt="Screenshot 2024-12-06 at 11 54 34" src="https://github.com/user-attachments/assets/b783a2ad-3757-401d-8964-29728d829dd3">
Apply plasma fractal pattern to modify image brightness and contrast.
This transform uses the Diamond-Square algorithm to generate organic-looking fractal patterns
that are then used to create spatially-varying brightness and contrast adjustments.
The result is a natural-looking, non-uniform modification of the image.
<img width="1180" alt="Screenshot 2024-12-06 at 11 56 21" src="https://github.com/user-attachments/assets/fc4e6ab9-54e0-4442-9088-cb1e62d2cb8c">
Apply plasma-based shadow effect to the image.
Creates organic-looking shadows using plasma fractal noise pattern.
The shadow intensity varies smoothly across the image, creating natural-looking
darkening effects that can simulate shadows, shading, or lighting variations.
Added angle_range and direction_range parameters.
Apply motion blur to the input image using a directional kernel.
This transform simulates motion blur effects that occur during image capture,
such as camera shake or object movement. It creates a directional blur using
a line-shaped kernel with controllable angle, direction, and position.
Args:
blur_limit (int | tuple[int, int]): Maximum kernel size for blurring.
Should be in range [3, inf).
- If int: kernel size will be randomly chosen from [3, blur_limit]
- If tuple: kernel size will be randomly chosen from [min, max]
Larger values create stronger blur effects.
Default: (3, 7)
angle_range (tuple[float, float]): Range of possible angles in degrees.
Controls the rotation of the motion blur line:
- 0°: Horizontal motion blur →
- 45°: Diagonal motion blur ↗
- 90°: Vertical motion blur ↑
- 135°: Diagonal motion blur ↖
Default: (0, 360)
direction_range (tuple[float, float]): Range for motion bias.
Controls how the blur extends from the center:
- -1.0: Blur extends only backward (←)
- 0.0: Blur extends equally in both directions (←→)
- 1.0: Blur extends only forward (→)
For example, with angle=0:
- direction=-1.0: ←•
- direction=0.0: ←•→
- direction=1.0: •→
Default: (-1.0, 1.0)
<img width="1207" alt="Screenshot 2024-12-06 at 12 00 10" src="https://github.com/user-attachments/assets/8b334b83-83ad-4397-9628-6bcad21fb7d0">
Apply Thin Plate Spline (TPS) transformation to create smooth, non-rigid deformations.
Imagine the image printed on a thin metal plate that can be bent and warped smoothly:
- Control points act like pins pushing or pulling the plate
- The plate resists sharp bending, creating smooth deformations
- The transformation maintains continuity (no tears or folds)
- Areas between control points are interpolated naturally
The transform works by:
1. Creating a regular grid of control points (like pins in the plate)
2. Randomly displacing these points (like pushing/pulling the pins)
3. Computing a smooth interpolation (like the plate bending)
4. Applying the resulting deformation to the image
Apply various illumination effects to the image.
This transform simulates different lighting conditions by applying controlled
illumination patterns. It can create effects like:
- Directional lighting (linear mode)
- Corner shadows/highlights (corner mode)
- Spotlights or local lighting (gaussian mode)
These effects can be used to:
- Simulate natural lighting variations
- Add dramatic lighting effects
- Create synthetic shadows or highlights
- Augment training data with different lighting conditions
Args:
mode (Literal["linear", "corner", "gaussian"]): Type of illumination pattern:
- 'linear': Creates a smooth gradient across the image,
simulating directional lighting like sunlight
through a window
- 'corner': Applies gradient from any corner,
simulating light source from a corner
- 'gaussian': Creates a circular spotlight effect,
simulating local light sources
Default: 'linear'
fisheye method.border_mode, value, mask_value<img width="1177" alt="Screenshot 2024-12-06 at 12 11 54" src="https://github.com/user-attachments/assets/71d85dd7-2036-4e74-b544-8b2db6277af4">
Apply random auto contrast to images.
Auto contrast enhances image contrast by stretching the intensity range
to use the full range while preserving relative intensities. For each
color channel:
1. Compute histogram
2. Find cumulative percentiles
3. Clip and scale intensities to full range
Unified naming for border_mode and filling constants.
value, fill_value, cval, pad_val => fillmask_value, cval_mask, fill_mask_value, pad_mask_value => fill_maskpad_mode, mode => border_modeby @ternaus
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Added option to pad the image if crop size is larger than the crop size
Old way
[
A.PadIfNeeded(min_height=1024, min_width=1024, p=1),
A.RandomCrop(height=1204, width=1024, p=1)
]
New way:
A.RandomCrop(height=1204, width=1024, p=1, pad_if_needed=True)
Works for:
You may also use it to pad image to a desired size.
Now random state for the pipeline does not depend on the global random state
Before
random.seed(seed)
np.random.seed(seed)
transform = A.Compose(...)
Now
transform = A.Compose(seed=seed, ...)
or
transform = A.Compose(...)
transform.set_random_seed(seed)
Now you can get exact parameters that were used in the pipeline on a given sample with
transform = A.Compose(save_applied_params=True, ...)
result = transform(image=image, bboxes=bboxes, mask=mask, keypoints=keypoints)
print(result["applied_transforms"])
Moved benchmark to a separate repo
https://github.com/albumentations-team/benchmark/
Current result for uint8 images:
| Transform | albumentations<br>1.4.20 | augly<br>1.0.0 | imgaug<br>0.4.0 | kornia<br>0.7.3 | torchvision<br>0.20.0 |
|---|---|---|---|---|---|
| HorizontalFlip | 8325 ± 955 | 4807 ± 818 | 6042 ± 788 | 390 ± 106 | 914 ± 67 |
| VerticalFlip | 20493 ± 1134 | 9153 ± 1291 | 10931 ± 1844 | 1212 ± 402 | 3198 ± 200 |
| Rotate | 1272 ± 12 | 1119 ± 41 | 1136 ± 218 | 143 ± 11 | 181 ± 11 |
| Affine | 967 ± 3 | - | 774 ± 97 | 147 ± 9 | 130 ± 12 |
| Equalize | 961 ± 4 | - | 581 ± 54 | 152 ± 19 | 479 ± 12 |
| RandomCrop80 | 118946 ± 741 | 25272 ± 1822 | 11503 ± 441 | 1510 ± 230 | 32109 ± 1241 |
| ShiftRGB | 1873 ± 252 | - | 1582 ± 65 | - | - |
| Resize | 2365 ± 153 | 611 ± 78 | 1806 ± 63 | 232 ± 24 | 195 ± 4 |
| RandomGamma | 8608 ± 220 | - | 2318 ± 269 | 108 ± 13 | - |
| Grayscale | 3050 ± 597 | 2720 ± 932 | 1681 ± 156 | 289 ± 75 | 1838 ± 130 |
| RandomPerspective | 410 ± 20 | - | 554 ± 22 | 86 ± 11 | 96 ± 5 |
| GaussianBlur | 1734 ± 204 | 242 ± 4 | 1090 ± 65 | 176 ± 18 | 79 ± 3 |
| MedianBlur | 862 ± 30 | - | 813 ± 30 | 5 ± 0 | - |
| MotionBlur | 2975 ± 52 | - | 612 ± 18 | 73 ± 2 | - |
| Posterize | 5214 ± 101 | - | 2097 ± 68 | 430 ± 49 | 3196 ± 185 |
| JpegCompression | 845 ± 61 | 778 ± 5 | 459 ± 35 | 71 ± 3 | 625 ± 17 |
| GaussianNoise | 147 ± 10 | 67 ± 2 | 206 ± 11 | 75 ± 1 | - |
| Elastic | 171 ± 15 | - | 235 ± 20 | 1 ± 0 | 2 ± 0 |
| Clahe | 423 ± 10 | - | 335 ± 43 | 94 ± 9 | - |
| CoarseDropout | 11288 ± 609 | - | 671 ± 38 | 536 ± 87 | - |
| Blur | 4816 ± 59 | 246 ± 3 | 3807 ± 325 | - | - |
| ColorJitter | 536 ± 41 | 255 ± 13 | - | 55 ± 18 | 46 ± 2 |
| Brightness | 4443 ± 84 | 1163 ± 86 | - | 472 ± 101 | 429 ± 20 |
| Contrast | 4398 ± 143 | 736 ± 79 | - | 425 ± 52 | 335 ± 35 |
| RandomResizedCrop | 2952 ± 24 | - | - | 287 ± 58 | 511 ± 10 |
| Normalize | 1016 ± 84 | - | - | 626 ± 40 | 519 ± 12 |
| PlankianJitter | 1844 ± 208 | - | - | 813 ± 211 | - |
cv2.addWeighted with wsum from simsimd packageFix in RandomSizedCrop and RandomResizedCrop
Hotfix version.
RandomOrderLove the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Added mask_interpolation to all transforms that use mask interpolation, including:
by @ternaus
cv2.LUT to stringzilla lutmask_interpolation to Compose that overrides mask interpolation value in all transforms in that Compose, now can use more accurate cv2.INTER_NEAREST_EXACT for semantic segmentation and can work with depth and heatmap estimation using cubic, area, linear, etcLove the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Added support for keypoints
Added support for keypoints and bounding boxes
Added support for keypoints and bounding boxes
Added support for keypoints and bounding boxes
Added support for bounding boxes and keypoints
Added support for keypoints
Added support for keypoints and bonding boxes
Added support for bounding boxes and keypoints
Added support for masks as numpy arrays of the shape (num_masks, height, width)
Now you can apply transforms to masks as:
masks = <numpy array with shape (num_masks, height, width)>
transform(image=image, masks=masks)
Removed MixUp as it was doing almost exactly the same as TemplateTransform
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
remove_invisible=False keeps keypoints
<img width="1224" alt="Screenshot 2024-09-30 at 15 25 53" src="https://github.com/user-attachments/assets/55cd90f4-7c9a-4409-91d3-b0bec69e99f0">by @ternaus
Added support for keypoints
<img width="1228" alt="Screenshot 2024-09-30 at 15 29 36" src="https://github.com/user-attachments/assets/7dc62f05-e8b2-4e49-b501-4327761dc4e3">
by @ternaus
Added RandomOrder Compose
Select N transforms to apply. Selected transforms will be called in random order with force_apply=True.
Transforms probabilities will be normalized to one 1, so in this case transforms probabilities works as weights.
This transform is like SomeOf, but transforms are called with random order.
It will not replay random order in ReplayCompose.
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
For visual debug wrote a tool that allows visually inspect effects of augmentations on the image.
You can find it at https://explore.albumentations.ai/
it is work in progress. It is not stable and polished yet, but if you have feedback or proposals - just write in the Discord Server mentioned above.
uint8 and float32 inputsAdded texture method to RandomSnow
<img width="1226" alt="Screenshot 2024-09-14 at 19 09 52" src="https://github.com/user-attachments/assets/19a3d0b1-a51c-46c5-b5de-810402f5e489">
Added physics_based method to RandomSunFlare
<img width="1232" alt="Screenshot 2024-09-14 at 19 10 41" src="https://github.com/user-attachments/assets/89d2c5f7-410e-429c-9636-ddb72f8db600">
Albumnetations version is tailored to a specific albucore version. Added pre-commit hook to automatically check it on every commit.Still works, but deprecated. It was a very strange transform, I cannot find use case, where you needed to use it.
For visual debug wrote a tool that allows visually inspect effects of augmentations on the image.
You can find it at https://explore.albumentations.ai/
RIght now supports only ImageOnly transforms, and not all but a subset of them.
it is work in progress. It is not stable and polished yet, but if you have feedback or proposals - just write in the Discord Server mentioned above.
Affine and ShiftScaleRotateStill works, but deprecated. It was a very strange transform, I cannot find use case, where you needed to use it.
It was equivalent to:
OneOf([Transpose, VerticalFlip, HorizontalFlip])
Most likely if you needed transform that does not create artifacts, you should look at:
HorizontalFlip (Symmetry group has 2 elements, meaning will effectively increase your dataset 2x)VerticalFlip (Symmetry group has 2 elements, meaning will effectively increase your dataset 2x)RandomRotate90 (Symmetry group has 2 elements, meaning will effectively increase your dataset 4x)D4 (Symmetry group has 8 elements, meaning will effectively increase your dataset 8x)Now you can define the number of output channels in the resulting gray image. All channels will be the same.
Extended ways one can get grayscale image. Most of them can work with any number of channels as input
weighted_average: Uses a weighted sum of RGB channels (0.299R + 0.587G + 0.114B)
Works only with 3-channel images. Provides realistic results based on human perception.from_lab: Extracts the L channel from the LAB color space.
Works only with 3-channel images. Gives perceptually uniform results.desaturation: Averages the maximum and minimum values across channels.
Works with any number of channels. Fast but may not preserve perceived brightness well.average: Simple average of all channels.
Works with any number of channels. Fast but may not give realistic results.max: Takes the maximum value across all channels.
Works with any number of channels. Tends to produce brighter results.pca: Applies Principal Component Analysis to reduce channels.
Works with any number of channels. Can preserve more information but is computationally intensive.Now uses Affine under the hood.
GridElasticDeform by @4pygmalionto_float and from_floatLove the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
GridElasticDeform transformGrid-based Elastic deformation Albumentation implementation
This class applies elastic transformations using a grid-based approach.
The granularity and intensity of the distortions can be controlled using
the dimensions of the overlaying distortion grid and the magnitude parameter.
Larger grid sizes result in finer, less severe distortions.
Args:
num_grid_xy (tuple[int, int]): Number of grid cells along the width and height.
Specified as (grid_width, grid_height). Each value must be greater than 1.
magnitude (int): Maximum pixel-wise displacement for distortion. Must be greater than 0.
interpolation (int): Interpolation method to be used for the image transformation.
Default: cv2.INTER_LINEAR
mask_interpolation (int): Interpolation method to be used for mask transformation.
Default: cv2.INTER_NEAREST
p (float): Probability of applying the transform. Default: 1.0.
Targets:
image, mask
Image types:
uint8, float32
Example:
>>> transform = GridElasticDeform(num_grid_xy=(4, 4), magnitude=10, p=1.0)
>>> result = transform(image=image, mask=mask)
>>> transformed_image, transformed_mask = result['image'], result['mask']
Note:
This transformation is particularly useful for data augmentation in medical imaging
and other domains where elastic deformations can simulate realistic variations.
by @4pygmalion
Now reflection padding correctly with bounding boxes and keypoints
by @ternaus
Simulates shadows for the image by reducing the brightness of the image in shadow regions.
Args:
shadow_roi (tuple): region of the image where shadows
will appear (x_min, y_min, x_max, y_max). All values should be in range [0, 1].
num_shadows_limit (tuple): Lower and upper limits for the possible number of shadows.
Default: (1, 2).
shadow_dimension (int): number of edges in the shadow polygons. Default: 5.
shadow_intensity_range (tuple): Range for the shadow intensity.
Should be two float values between 0 and 1. Default: (0.5, 0.5).
p (float): probability of applying the transform. Default: 0.5.
Targets:
image
Image types:
uint8, float32
Reference:
https://github.com/UjjwalSaxena/Automold--Road-Augmentation-Library
by @JonasKlotz
Affine. Now fit_output=True works correctly with bounding boxes. by @ternausColorJitter. By @maremunCoarseDropout. By @thomaoc1logger anymore. by @ternausHistorgramMatching. Before it output array of ones. Now works as expected. by @ternausUpdated mixing parameters by @ternaus in https://github.com/albumentations-team/albumentations/pull/1859
Full Changelog: https://github.com/albumentations-team/albumentations/compare/1.4.12...1.4.13
Deprecated parameter alpha_affine in ElasticTransform. To have Affine effects on your image, use the Affine transform.
Allows adding text on top of images. Works with np,unit8 and np.float32 images with any number of channels.
Additional functionalities:
images targetYou can now apply the same transform to a list of images of the same shape, not just one image.
Use cases:
import albumentations as A
transform = A.Compose([A.Affine(p=1)])
transformed = transform(images=<list of images>)
transformed_images = transformed["images"]
Note: You can apply the same transform to any number of images, masks, bounding boxes, and sets of keypoints using the additional_targets functionality notebook with examples
Contributors @ternaus, @ayasyrev
get_params_dependent_on dataRelevant for those who build custom transforms.
Old way
@property
def targets_as_params(self) -> list[str]:
return <list of targets>
def get_params_dependent_on_targets(self, params: dict[str, Any]) -> dict[str, np.ndarray]:
image = params["image"]
....
New way
def get_params_dependent_on_data(self, params: dict[str, Any], data: dict[str, Any]) -> dict[str, np.ndarray]:
image = data["image"]
Contributor @ayasyrev
shape to paramsOld way:
def get_params_dependent_on_targets(self, params: dict[str, Any]) -> dict[str, np.ndarray]:
image = params["image"]
shape = image.shape
New way:
def get_params_dependent_on_data(self, params: dict[str, Any], data: dict[str, Any]) -> dict[str, np.ndarray]:
shape = params["shape"]
Contributor @ayasyrev
Deprecated parameter alpha_affine in ElasticTransform. To have Affine effects on your image, use the Affine transform.
Contributor @ternaus
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Allows to paste set of images + corresponding masks to the image.
It is not entirely CopyAndPaste as "masks", "bounding boxes" and "keypoints" are not supported, but step in that direction.
Added balanced sampling for scale_limit
From FAQ:
The default scaling logic in RandomScale, ShiftScaleRotate, and Affine transformations is biased towards upscaling.
For example, if scale_limit = (0.5, 2), a user might expect that the image will be scaled down in half of the cases and scaled up in the other half. However, in reality, the image will be scaled up in 75% of the cases and scaled down in only 25% of the cases. This is because the default behavior samples uniformly from the interval [0.5, 2], and the interval [0.5, 1] is three times smaller than [1, 2].
To achieve balanced scaling, you can use Affine with balanced_scale=True, which ensures that the probability of scaling up and scaling down is equal.
balanced_scale_transform = A.Compose([A.Affine(scale=(0.5, 2), balanced_scale=True)])
by @ternaus
Added support for keypoints
by @ternaus
Added support for keypoints
by @ternaus
by @zakajd
by @ternaus
strict parameter to ComposeIf strict=True only targets that are expected could be passed.
If strict = False, user can pass data with extra keys. Such data would not be affected by transforms.
Request came from users that use pipelines in the form:
transform = A.Compose([....])
data = A.Compose(**data)
by @ayasyrev
Crop module was heavily refactored, all tests and checks pass, but we will see.
Old way:
GridDropout(
holes_number_x=XXX,
holes_numver_y=YYY,
unit_size_min=ZZZ,
unit_size_max=PPP
)
New way:
GridDropout(
holes_number_xy = (XXX, YYY),
unit_size_range = (ZZZ, PPP)
)
by @ternaus
Old way:
RandomSunFlare(
num_flare_circles_lower = XXX,
num_flare_circles_upper = YYY
)
new way:
RandomSunFlare(num_flare_circles_range = (XXX, YYY))
ISONoise, as it returned zeros. by @ternausAffine as during rotation image, mask, keypoints have one center point for rotation and bounding box another => we need to create two separate affine matrices. by @ternausp=number. Say for VerticalFlip(0.5) you could expect 50% chance, but 0.5 was attributed not to p but to always_apply which meant that the transform was always applied. by @ayasyrevHotfix release that addresses issues introduced in 1.4.9
Hotfix release that addresses issues introduced in 1.4.9
There were two issues in GaussNoise that this release addresses:
noise_scale_factor, which is different from the behavior before version 1.4.9. Now default value = 1, which means random noise is created for every point independentlygauss >=0. Fixed.always_apply is deprecated now. always_apply=True still works, but it will be deprecated in the future. Use p=1 instead
New transform, based on
<img width="634" alt="Screenshot 2024-06-17 at 17 53 00" src="https://github.com/albumentations-team/albumentations/assets/5481618/d042299a-3fcd-47e2-a2f8-c023646659d1">
Statements from the paper on why PlanckianJitter is superior to ColorJitter:
Realistic Color Variations: PlanckianJitter applies physically realistic illuminant variations based on Planck’s Law for black-body radiation. This leads to more natural and realistic variations in chromaticity compared to the arbitrary changes in hue, saturation, brightness, and contrast applied by ColorJitter.
Improved Representation for Color-Sensitive Tasks: The transformations in PlanckianJitter maintain the ability to discriminate image content based on color information, making it particularly beneficial for tasks where color is a crucial feature, such as classifying natural objects like birds or flowers. ColorJitter, on the other hand, can significantly alter colors, potentially degrading the quality of learned color features.
Robustness to Illumination Changes: PlanckianJitter produces models that are robust to illumination changes commonly observed in real-world images. This robustness is advantageous for applications where lighting conditions can vary widely.
Enhanced Color Sensitivity: Models trained with PlanckianJitter show a higher number of color-sensitive neurons, indicating that these models retain more color information compared to those trained with ColorJitter, which tends to induce color invariance.
by @zakajd
Added option to approximate GaussNoise.
Generation of random Noise for large images is slow.
Added scaling factor for noise generation. Value should be in the range (0, 1]. When set to 1, noise is sampled for each pixel independently. If less, noise is sampled for a smaller size and resized to fit the shape of the image. Smaller values make the transform much faster. Default: 0.5
Added integration wit HFHub. Now you can load and save augmentation pipeline to HuggingFace and reuse it in the future or share with others.
import albumentations as A
import numpy as np
transform = A.Compose([
A.RandomCrop(256, 256),
A.HorizontalFlip(),
A.RandomBrightnessContrast(),
A.RGBShift(),
A.Normalize(),
])
evaluation_transform = A.Compose([
A.PadIfNeeded(256, 256),
A.Normalize(),
])
transform.save_pretrained("qubvel-hf/albu", key="train")
# ^ this will save the transform to a directory "qubvel-hf/albu" with filename "albumentations_config_train.json"
transform.save_pretrained("qubvel-hf/albu", key="train", push_to_hub=True)
# ^ this will save the transform to a directory "qubvel-hf/albu" with filename "albumentations_config_train.json"
# + push the transform to the Hub to the repository "qubvel-hf/albu"
transform.push_to_hub("qubvel-hf/albu", key="train")
# ^ this will push the transform to the Hub to the repository "qubvel-hf/albu" (without saving it locally)
loaded_transform = A.Compose.from_pretrained("qubvel-hf/albu", key="train")
# ^ this will load the transform from local folder if exist or from the Hub repository "qubvel-hf/albu"
evaluation_transform.save_pretrained("qubvel-hf/albu", key="eval", push_to_hub=True)
# ^ this will save the transform to a directory "qubvel-hf/albu" with filename "albumentations_config_eval.json"
loaded_evaluation_transform = A.Compose.from_pretrained("qubvel-hf/albu", key="eval")
# ^ this will load the transform from the Hub repository "qubvel-hf/albu"
by @qubvel
These transforms should be faster for all types of images. But measured only for three channel uint8
always_applyFor years we had two parameters in constructors - probability and always_apply. The interplay between them is not always obvious and intuitively always_apply=True should be equivalent to p=1.
always_apply is deprecated now. always_apply=True still works, but it will be deprecated in the future. Use p=1 instead
by @ayasyrev
Updated interface for RandomFog
Old way:
RandomFog(fog_coef_lower=0.3, fog_coef_upper=1)
New way:
RandomFog(fog_coef_range=(0.3, 1))
by @ternaus
When one imports Albumentations library, there is a check that it is the latest version installed.
To disable this check you can set up environmental variable: NO_ALBUMENTATIONS_UPDATE to 1
by @lerignoux
For a set of transforms we were throwing deprecation warnings, even when modern version of the interface was used. Fixed. by @ternaus
We moved low level operations like add, multiply, normalize, etc to a separate library: https://github.com/albumentations-team/albucore
There are numerous ways to perform such operations in opencv and numpy. And there is no clear winner. Results depend on image type.
Separate library gives us confidence that we picked the fastest version that works on any image type.
by @ternaus
Various bugfixes by @ayasyrev @immortalCO
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Added to the documentation links to the UI on HuggingFace to explore hyperparameters visually.
<div style="display: flex; justify-content: space-around; align-items: center;"> <img width="730" alt="Screenshot 2024-05-28 at 16 27 09" src="https://github.com/albumentations-team/albumentations/assets/5481618/525ca812-a2ad-46cb-9fb2-b89ec3a119a3"> <img width="885" alt="Screenshot 2024-05-28 at 16 28 03" src="https://github.com/albumentations-team/albumentations/assets/5481618/ff81c193-4355-4aee-962c-77459c8a1292"> </div>
Updated interface:
Old way:
transform = A.Compose([A.RandomSnow(
snow_point_lower=0.1,
snow_point_upper=0.3,
p=0.5
)])
New way:
transform = A.Compose([A.RandomSnow(
snow_point_range=(0.1, 0.3),
p=0.5
)])
by @MarognaLorenzo
Old way
transform = A.Compose([A.RandomSnow(
slant_lower=-10,
slant_upper=10,
p=0.5
)])
New way:
transform = A.Compose([A.RandomRain(
slant_range=(-10, 10),
p=0.5
)])
by @MarognaLorenzo
Created library with core functions albucore. Moved a few helper functions there. We need this library to be sure that transforms are:
numpy and opencv. For some functions it is possible to be faster than both of them.check_for_updates. Now the pipeline does not throw an error regardless of why we cannot check for update.RandomShadow. Does not create unexpected purple color on bright white regions with shadow overlay anymore.Compose. Now Compose([]) does not throw an error, but just works as NoOp by @ayasyrevmin_max normalization. Now return 0 and not NaN on constant images. by @ternausCropAndPad. Now we can sample pad/crop values for all sides with interface like ((-0.1, -0.2), (-0.2, -0.3), (0.3, 0.4), (0.4, 0.5)) by @christian-steinmeyerLove the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Old way:
transform = A.Compose([A.ImageCompression(
quality_lower=75,
quality_upper=100,
p=0.5
)])
New way:
transform = A.Compose([A.ImageCompression(
quality_range=(75, 100),
p=0.5
)])
by @MarognaLorenzo
Old way:
transform = A.Compose([A.Downscale(
scale_min=0.25,
scale_max=1,
interpolation= {"downscale": cv2.INTER_AREA, "upscale": cv2.INTER_CUBIC},
p=0.5
)])
New way:
transform = A.Compose([A.Downscale(
scale_range=(0.25, 1),
interpolation_pair = {"downscale": cv2.INTER_AREA, "upscale": cv2.INTER_CUBIC},
p=0.5
)])
As of now both ways work and will provide the same result, but old functionality will be removed in later releases.
by @ternaus
Blur.bbox clipping, it could be not intuitive, but boxes should be clipped by height, width and not height - 1, width -1 by @ternausPadIfNeeded if value parameter is not None, but border mode is reflection, border mode is changed to cv2.BORDER_CONSTANT by @ternausIn version 1.4.5 there was a bug that went unnoticed - if you used pipeline that consisted only of ImageOnly transforms but pass bounding boxes into i
In version 1.4.5 there was a bug that went unnoticed - if you used pipeline that consisted only of ImageOnly transforms but pass bounding boxes into it, you would get an error.
If you had in such pipeline at least one non ImageOnly transform, say HorizontalFlip or Crop, everything would work as expected.
We fixed the issue and added tests to be sure that it will not happen in the future.
Love the library? You can contribute to its development by becoming a sponsor for the library. Your support is invaluable, and every contribution make
Applies one of the eight possible D4 dihedral group transformations to a square-shaped input, maintaining the square shape. These transformations correspond to the symmetries of a square, including rotations and reflections by @ternaus
The D4 group transformations include:
- e (identity): No transformation is applied.
- r90 (rotation by 90 degrees counterclockwise)
- r180 (rotation by 180 degrees)
- r270 (rotation by 270 degrees counterclockwise)
- v (reflection across the vertical midline)
- hvt (reflection across the anti-diagonal)
- h (reflection across the horizontal midline)
- t (reflection across the main diagonal)
Could be applied to:
Does not generate interpolation artifacts as there is no interpolation.
Provides the most value in tasks where data is invariant to rotations and reflections like:
Example:
<img width="831" alt="Screenshot 2024-04-16 at 19 00 05" src="https://github.com/albumentations-team/albumentations/assets/5481618/141a778e-33d5-4804-8a96-167b9bcbe621">
standard - subtract fixed mean, divide by fixed stdimage - the same as standard, but mean and std computed for each image independently.image_per_channel - the same as before, but per channelmin_max - subtract min(image)and divide by max(image) - min(image)min_max_per_channel - the same, but per channel
by @ternausNew, preferred wat is to use num_shadows_limit instead of num_shadows_lower / num_shadows_upper by @ayasyrev
Now all input parameters are validated and prepared with Pydantic. This will prevent bugs, when transforms are initialized without errors with parameters that are outside of allowed ranges. by @ternaus
Example:
Standard way uses additional_targets
transform = A.Compose(
transforms=[A.Rotate(limit=(90.0, 90.0), p=1.0)],
keypoint_params=A.KeypointParams(
angle_in_degrees=True,
check_each_transform=True,
format="xyas",
label_fields=None,
remove_invisible=False,
),
additional_targets={"keypoints2": "keypoints"},
)
Now you can also add them using add_targets:
transform = A.Compose(
transforms=[A.Rotate(limit=(90.0, 90.0), p=1.0)],
keypoint_params=A.KeypointParams(
angle_in_degrees=True,
check_each_transform=True,
format="xyas",
label_fields=None,
remove_invisible=False,
),
)
transform.add_targets({"keypoints2": "keypoints"})
by @ayasyrev
add_weighted function by @gogetronMinor improvements and bug fixes
<img width="1659" alt="Screenshot 2024-04-02 at 18 43 51" src="https://github.com/albumentations-team/albumentations/assets/5481618/e9c95aab-b2a8-4b12-9d72-86041b08f3ed">
Morphological transform that modifies the structure of the image. Dilation expands the white (foreground) regions in a binary or grayscale image, while erosion shrinks them.Do not throw deprecation warning when people do not use deprecated parameters in AdvancedBlur by @Aloqeely
<div align="center"> <a href="https://i.imgur.com/8wWkMmL.jpeg"> <img src="https://i.imgur.com/8wWkMmL.jpeg" width="30%"> </a> <a href="https://i.imgur.com/B687Opr.jpeg"> <img src="https://i.imgur.com/B687Opr.jpeg" width="30%"> </a> <a href="https://i.imgur.com/jkjwFMB.jpeg"> <img src="https://i.imgur.com/jkjwFMB.jpeg" width="30%"> </a> <p> <b>Left:</b> Original, <b>Middle:</b> Chromatic aberration (default args, mode="green_purple"), <b>Right:</b> Chromatic aberration (default args, mode="red_blue") <br>(Image is from our internal mobile mapping dataset) </p> </div>
ChromaticAbberation transform that adds chromatic distortion to the image. Wiki by @mrsmrynkmixing parameter for MixUp transform by @Dipet. For more details Tutorial on MixUpAdvancedBlur by @AloqeelyCONTRIBUTORS.md for Windows users by @AloqeelyDownScale transform by @ryoryon66PadIfNeeded serialization @ternausIf you enjoy using the library as an individual developer or during the day job as a part of the company, please consider becoming a sponsor for the l
<img width="660" alt="Screenshot 2024-03-04 at 14 52 15" src="https://github.com/albumentations-team/albumentations/assets/5481618/68e5031b-e45e-4578-abe8-1d8e33db4831">
MixUp transform: which linearly combines an input (image, mask, and class label) with another set from a predefined reference dataset. The mixing degree is controlled by a parameter λ (lambda), sampled from a Beta distribution. This method is known for improving model generalization by promoting linear behavior between classes and smoothing decision boundaries.isort, flake8, black to ruffopencv library inconsistencies issuesThe deprecated code, including 15 transforms, was removed. Dependency on the imgaug library was removed.
In this release, we mainly focused on the technical debt as its decrease allows faster iterations and bug fixes in the codebase. We added only one new transform, did not work on speeding up transforms, and other changes are minor.
But, somehow, we are cutting this dependency only in 2024.
<img width="560" alt="Screenshot 2024-02-17 at 13 09 01" src="https://github.com/albumentations-team/albumentations/assets/5481618/18aaebad-4b58-4cc6-932f-e2d8a1f352ab">
XYMasking transform: applies masking strips to an image, either horizontally (X axis) or vertically (Y axis), simulating occlusions. This transform is helpful for training models to recognize images with varied visibility conditions. It's particularly effective for spectrogram images, allowing spectral and frequency masking to improve model robustness.
As other dropout transforms CoarseDropout, MaskDropout, GridDropout it supports images, masks and keypoints as targets. (https://github.com/albumentations-team/albumentations/commit/004fabbf90794fbc21ee356e2dde6637b7fecbd4 by @ternaus )The deprecated code, including 15 transforms, was removed. Dependency on the imgaug library was removed.
(https://github.com/albumentations-team/albumentations/commit/be6a217b207b3d7ebe792caabb438d660b45f2a5 by @ternaus )
JpegCompression. Use ImageCompression instead.RandomBrightness. Use RandomBrigtnessContrast instead.RandomContrast. Use RandomBrigtnessContrast instead.Cutout. Use CoarseDropout instead.ToTensor. Use ToTensorV2 instead.IAAAdditiveGaussianNoise. Use GaussNoise instead.IAAAffine. Use Affine instead.IAAEmboss. Use Emboss instead.IAAFliplr. Use HorizontalFlip instead.IAAFlipud. Use VerticalFlip instead.IAAPerspective. Use Perspective instead.IAAPiecewiseAffine. Use PiecewiseAffine instead.IAASharpen. Use Sharpen instead.IAASuperpixels. Use Superpixels instead.eps parameter in RandomGammalambda_transformsin serialization.from_dict function.matrix=None case for Piecewise affine transform (https://github.com/albumentations-team/albumentations/commit/c70e664e060bfd7463c20674927aed217f72d437 @Dipet )Fixed deprecated imports of scipy.ndimage.gaussian_filter (#1311 by @rbu)
ToRGB transform (#1323 by @kinoooshnik)RandomGravel transform (#1365 by @onurtore)Spatter (#1305 by @Andredance)rotate_method in Affine (#1394 by @i-aki-y)CoarseDropout (#1330 by @domef)scipy.ndimage.gaussian_filter (#1311 by @rbu)RandomSunFlare (#1333 by @jasonrock-a3)ToSepia transform (#1397 by @ifeherva)skimage deprecetions (#1421 by @Dipet)Renamed method to rotate_method inside `Rotate` to keep consistency between naming parameters. (#1258 by @Dipet, thanks to @MichaelMonashev)
method to rotate_method inside Rotate to keep consistency between naming parameters. (#1258 by @Dipet, thanks to @MichaelMonashev)RandomCropFromBorders - Crops image based on indents from image borders. (#1240 by @Dipet based on #476 by @ZFTurbo)BBoxSafeRandomCrop - Crops image without loss of bboxes. Instead of RandomSizedBBoxSafeCrop this implementation do not apply resize to target size. (#579 by @SunQpark)Spatter - Simulates corruption which can occlude a lens in the form of rain or mud. (#573 by @akarsakov)Defocus - Imitates lens defocusing. (#551 by @akarsakov)ZoomBlur - Imitates lens blur on zoomig. (#551 by @akarsakov)RandomBrightnessContrast when brightness_by_max=False. (#487 by @Dipet)Perspective and Affine. (#1231 by @Dipet)min_visibility=0 or min_visibility=1. (#616 by @IlyaOvodov)Rotate when crop_border=True. (#1250 by @Dipet, thanks to @jonkoi)always_apply Compose children. (#561 by @albu)src_color, and use all three color values. (#1285 by @hoel-bagard)gamma_limit. (#1286 by @zahragolpa)Normalize in some case up to 2 times. (#563 by @Dipet)GridDistortion, ElasticTransform and OpticalDistortion now supports bbox targets. (#476, #1262 by @ZFTurbo and @Dipet)MotionBlur now supports allow_shifted flag. When it's value is False only non shifted kernels generated. (#1239 by @Dipet)GridDistortion now supports normalized flag. When it is set to True will be applied distortion inside image border. (#722 by @poke1024)Downscale. This is needed to avoid interpolation artefacts. (#584 by @nathanhubens)albumentations.augmentations.utils.py. (#1260 by @Dipet)albumentations.augmentations.blur. (#1259 by @Dipet)Fixed a deprecation warning in match_histograms. (#1121 by @BloodAxe)
A.Rotate and A.ShiftScaleRotate now support new rotation method for bounding boxes, ellipse. (#1203 by @victor1cea)A.Rotate now supports new argument crop_border. If set to True, the rotated image will be cropped as much as possible to eliminate pixel values at the edges that were not well defined after rotation. (#1214 by @bonlime)match_histograms. (#1121 by @BloodAxe)A.CropNonEmptyMaskIfExists modified the first element of masks in-place. Now, this behavior is fixed and A.CropNonEmptyMaskIfExists doesn't do in-place modification of input masks. (#1193 by @ORippler).fill_value and mask_fill_value parameters for A.GridDropout. (#1191 by @victor1cea)A.ColorJitter now correctly works with A.ReplayCompose. (#1199 by @zakajd)A.ColorJitter for np.float32 input images when contrast is set to 0 (previously, all values were set to 0.5 instead of using the average value).. (#1207 by @Dipet)A.Rotate, A.Affine and A.ShiftScaleRotate now do rotation in the same way. Fixed incorrect rotation angle for A.Affine. A.Rotate and A.ShiftScaleRotate now correctly rotate the keypoints 90 degrees and don't leave black lines around the edges of the image. (#1091 by @Dipet )`A.UnsharpMask`. This transform sharpens the input image using Unsharp Masking processing and overlays the result with the original image. (#1063 by @
A.UnsharpMask. This transform sharpens the input image using Unsharp Masking processing and overlays the result with the original image. (#1063 by @zakajd)A.RingingOvershoot. This transform creates ringing or overshoot artifacts by convolving the image with a 2D sinc filter. (#1064 by @zakajd)A.AdvancedBlur. This transform blurs the input image using a Generalized Normal filter with randomly selected parameters. It also adds multiplicative noise to generated kernel before convolution. (#1066 by @zakajd)A.PixelDropout. This transformation randomly replaces pixels with the passed value. (#1082 by @Dipet)A.RandomShadow from working with non-contiguous input. (#1117 by @i-aki-y)A.PadIfNeeded now works with an arbitrary number of channels. (#1069 by @BloodAxe)np.random use cases to prevent identical values when using multiprocessing. (#1070 by @Dipet)slant param now has an effect in A.RandomRain. (#1179 by @victor1cea)translate_percent now uses 0 as a default value in the A.Affine transform. (#1183 by @victor1cea)A.SafeRotate no longer loses blocks and keypoints. (#1109 by @Dipet)A.CropAndPad now correctly handles bboxes when keep_size=True. (#1059 by @cannon)A.RandomCrop, A.RandomSizedCrop, and A.RandomSizedBBoxSafeCrop now sample last pixel. (#1080 by @Multihuntr)A.Compose now warns the user if it receives a single augmentation instead of a sequence of augmentations. (#1055 by @Dipet)A.CoarseDropout and A.RandomGridShuffle now support keypoints. (#1084 by @BloodAxe)A.ToTensorV2 now supports the masks target. (#1097 by @alessiobonfiglio)A.PadIfNeeded now supports random padding. (#1160 by @mys007 )A.Affine now has keep_ratio flag. (#1104 by @i-aki-y)!133947365-6cba891b-4537-4d97-8b84-5ac9ce908d1d
TemplateTransform. This transform allows the blending of an input image with specified templates. (#572 by @akarsakov )PixelDistributionAdaptation. A new domain adaptation augmentation. It fits a simple transform on both the original and reference image, transforms the original image with transform trained on this image, and performs inverse transformation using transform fitted on the reference image. See the examples of this transform in the qudida repository. (#959 by @arsenyinfo)LongestMaxSize and SmallestMaxSize now can also accept a list of sizes as their max_size argument and the actual max_size value will be sampled randomly from this list. (#930 by @kmistry-wx )A.Affine now accepts mask_interpolation as a parameter. (#975 by @dskkato )A.RandomRain now alters brightness in HSV space instead of HLS space to prevent image corruption. (#990 by @ErlingLie)ValueError if bbox_params is not specified and bbox transformation is called (#1013 by @VirajBagal)CoarseDropout can now set the height and width of holes based on the fraction of original image height and width (#1014 by @VirajBagal )ElasticTransform got performance optimizations. (#1004 by @b0nce)CropNonEmptyMaskIfExists thrown an error when it was used with a keypoint even though keypoints were mentioned as a correct target. (#986 by @GalDude33 )RandomCropNearBBox when it received values with x_min <= 0 or y_min <= 0 (#993 by @Dipet )Fixed problem with incorrect shape at keypoints and bboxes processors after ToTensorV2 #963
ToTensorV2 #963Fixed YOLO format conversion problem when bbox greater than image by 1 pixel. Now YOLO bbox will be converted to Albumentations format without bbox de
Added position argument to PadIfNeeded (#933 by @yisaienkov)
Added position argument to PadIfNeeded (#933 by @yisaienkov)
Possible values: center top_left, top_right, bottom_left, bottom_right, with center being the default value.
One possible use case for this feature is object detection where you need to pad an image to square, but you want predicted bounding boxes being equal to the bounding box of the unpadded image.
Deprecated augmentation ToTensor that converts NumPy arrays to PyTorch tensors is completely removed from Albumentations. You will get a RuntimeError…
imgaug dependency is now optional, and by default, Albumentations won't install it. This change was necessary to prevent simultaneous install of both opencv-python-headless and opencv-python (you can read more about the problem in this issue). If you still need imgaug as a dependency, you can use the pip install -U albumentations[imgaug] command to install Albumentations with imgaug.ToTensor that converts NumPy arrays to PyTorch tensors is completely removed from Albumentations. You will get a RuntimeError exception if you try to use it. Please switch to ToTensorV2 in your pipelines.A.RandomToneCurve. See a notebook for examples of this augmentation (#839 by @aaroswings)SafeRotate. Safely Rotate Images Without Cropping (#888 by @deleomike)SomeOf transform that applies N augmentations from a list. Generalizing of OneOf (#889 by @henrique)By default, Albumentations doesn't require imgaug as a dependency. But if you need imgaug, you can install it along with Albumentations by running pip install -U albumentations[imgaug].
Here is a table of deprecated imgaug augmentations and respective augmentations from Albumentations that you should use instead:
| Old deprecated augmentation | New augmentation |
|---|---|
| IAACropAndPad | CropAndPad |
| IAAFliplr | HorizontalFlip |
| IAAFlipud | VerticalFlip |
| IAAEmboss | Emboss |
| IAASharpen | Sharpen |
| IAAAdditiveGaussianNoise | GaussNoise |
| IAAPerspective | Perspective |
| IAASuperpixels | Superpixels |
| IAAAffine | Affine |
| IAAPiecewiseAffine | PiecewiseAffine |
Serialization logic is updated. Previously, Albumentations used the full classpath to identify an augmentation (e.g. albumentations.augmentations.transforms.RandomCrop). With the updated logic, Albumentations will use only the class name for augmentations defined in the library (e.g., RandomCrop). For custom augmentations created by users and not distributed with Albumentations, the library will continue to use the full classpath to avoid name collisions (e.g., when a user creates a custom augmentation named RandomCrop and uses it in a pipeline).
This new logic will allow us to refactor the code without breaking serialized augmentation pipelines created using previous versions of Albumentations. This change will also reduce the size of YAML and JSON files with serialized data.
The new serialization logic is backward compatible. You can load serialized augmentation pipelines created in previous versions of Albumentations because Albumentations supports the old format.
A.ReplayCompose to work with bounding boxes and keypoints correctly. (#748)A.GlassBlur now correctly works with float32 inputs (#826)MultiplicativeNoise now correctly works with gray images with shape [h, w, 1]. (#793)albumentations.augmentations.geometric. (#784)albumentations.augmentations.crops. (#791)setup.py that detects existing installations of OpenCV now also looks for opencv-contrib-python and opencv-contrib-python-headless (#837 by @agchang-cgl)ToTensorV2 now automatically expands grayscale images with the shape [H, W] to the shape [H, W, 1]. PR #604 by @Ingwar.
[H, W] to the shape [H, W, 1]. PR #604 by @Ingwar.masks argument to the transform function. Previously this augmentation worked only with a single mask provided by the mask argument. PR #761API for `A.FDA` is changed to resemble API of `A.HistogramMatching`. Now, both transformations expect to receive a list of reference images, a functio
A.FDA is changed to resemble API of A.HistogramMatching. Now, both transformations expect to receive a list of reference images, a function to read those image, and additional augmentation parameters. (#734)A.HistogramMatching now usesread_rgb_image as a default read_fn. This function reads an image from the disk as an RGB NumPy array. Previously, the default read_fn was cv2.imread which read an image as a BGR NumPy array. (#734)A.Sequential transform that can apply augmentations in a sequence. This transform is not intended to be a replacement for A.Compose. Instead, it should be used inside A.Compose the same way A.OneOf or A.OneOrOther. For instance, you can combine A.OneOf with A.Sequential to create an augmentation pipeline containing multiple sequences of augmentations and apply one randomly chosen sequence to input data. (#735)A.ShiftScaleRotate now has two additional optional parameters: shift_limit_x and shift_limit_y. If either of those parameters (or both of them) is set A.ShiftScaleRotate will use the set values to shift images on the respective axis. (#735)A.ToTensorV2 now supports an additional argument transpose_mask (False by default). If the argument is set to True and an input mask has 3 dimensions, A.ToTensorV2 will transpose dimensions of a mask tensor in addition to transposing dimensions of an image tensor. (#735)A.FDA now correctly uses coordinates of the center of an image. (#730)A.HistogramMatching. (#734)A.load() was called to deserialize a pipeline that contained A.ToTensor or A.ToTensorV2, but those transforms were not imported in the code before the call. (#735)Albumentations now explicitly checks that all inputs to augmentations are named arguments and raise an exception otherwise. So if an augmentation rece
A.FDA transform for Fourier-based domain adaptation. (#685)A.HistogramMatching transform that applies histogram matching. (#708)A.ColorJitter transform that behaves similarly to ColorJitter from torchvision (though there are some minor differences due to different internal logic for working with HSV colorspace in Pillow, which is used in torchvision and OpenCV, which is used in Albumentations). (#705)A.PadIfNeeded now accepts additional pad_width_divisor, pad_height_divisor (None by default) to ensure image has width & height that is dividable by given values. (#700)A.CoarseDropout to masks via mask_fill_value. (#699)A.GaussianBlur now supports the sigma parameter that sets standard deviation for Gaussian kernel. (#674, #673) .A.HueSaturationValue for float dtype. (#696, #710)YOLO format. (#688)Change the ImgAug dependency version from “imgaug>=0.2.5,<0.2.7” to “imgaug>=0.4.0". Now Albumentations won’t downgrade your existing ImgAug installat
ReplayCompose is now serializable. PR #623 by IlyaOvodovPadIfNeeded). That happened because Albumentations checked which bounding boxes and keypoints lie outside the image only after applying all augmentations. Now Albumentations will check and remove keypoints and bounding boxes that lie outside the image after each augmentation. If, for some reason, you need the old behavior, pass check_each_transform=False in your KeypointParams or BboxParams. Issue #565 and PR #566.ImageCompression and GaussNoise. PR #569label_fields in BboxParams. PR #504 by IlyaOvodovNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Added YOLO format to bounding boxes
New transforms
New features
Improvements
fill_value to Cutoutfill_value for image and mask targetsBug Fixes
Documentation Updated
https://github.com/albu/albumentations/commit/2e25667f8c39eba3e6be0e85719e5156422ee9a9 Target: image
This transform mimics the noise that images will have if the ISO parameter of the camera is high. Wiki
https://github.com/albu/albumentations/commit/e365b52df6c6535a1bf06733b607915231f2f9d4 Targets: image
Solarize inverts all pixels above some threshold. It is an essential part of the work AutoAugment: Learning Augmentation Policies from Data.
https://github.com/albu/albumentations/commit/9f71038c95c4124bdaf3ee13a9823225bb8d85da Target: image
Equalizes image histogram. It is an essential part of the work AutoAugment: Learning Augmentation Policies from Data.
https://github.com/albu/albumentations/commit/ad95fa005fd5325deb73461bfb6e543fca342f45 Target: image
Reduce the number of bits for each pixel. It is an essential part of the work AutoAugment: Learning Augmentation Policies from Data.
Target: image https://github.com/albu/albumentations/commit/b6127864d45cfa5b5299578d309680baa0ce7aa3 Decrease Jpeg or WebP compression to the image.
https://github.com/albu/albumentations/commit/df831d6605140e7aa013deab6012d85af9854be3 Target: image
Decreases image quality by downscaling and upscaling back.
https://github.com/albu/albumentations/commit/4dbe41e8795c7b7d48e0cc4501efe8046e21765b Targets: image, mask, bboxes, keypoints
Crop the given Image to the random size and aspect ratio. This transform is an essential part of many image classification pipelines. Very popular for ImageNet classification.
It has the same API as RandomResizedCrop in torchvision.
https://github.com/albu/albumentations/commit/4cf6c36bc2332729d91e44f58f18f44b66db3c6f Targets: image, mask
Partition an image into tiles. Shuffle them and merge back.
Targets: image, mask, bboxes, keypoints
Crop area with a mask if the mask is non-empty, else make a random crop.
https://github.com/albu/albumentations/commit/a5026800d84c6c1998f224b86dedbf3f005ae994 Targets: image, mask
Convert image and mask to torch.Tensor
https://github.com/albu/albumentations/commit/d05db9e9aae6b7607c33c4cdce69be011c2f8802
The Yolo format of a bounding box has a format [x, y, width, height], where values normalized to the size of the image. Ex: [0.3, 0.1, 0.05, 0.07]
https://github.com/albu/albumentations/commit/9942689f9846c59006c80718ee8db38e02ee2104
Augmentations pipeline has a lot of randomnesses, which is hard to debug. We added Determentsic / Replay mode in which you can track what parameters were applied to the input and use precisely the same transform to another input if necessary.
Jupyter notebook with an example.
fill_value to the Cutout transform.https://github.com/albu/albumentations/commit/d85bab59eb8ccb0a2fec86750f94173e18e86395
fill_value for images and maskshttps://github.com/albu/albumentations/commit/2c1a1485f690b4e8ead50f5bb29d3838fbbc177d
One of the use cases is it to use mask_value, which is equal to the ignore_index of your loss. This will decrease the level of noise and may improve convergence.
https://github.com/albu/albumentations/commit/c3cc277f37b172bebf7177c779a7cf3cdf7120d3
3.2 times faster for uint8 images.
https://github.com/albu/albumentations/commit/448761df9a008384cf914343f25e3cfb7c4d7551
2 times faster for uint8 images.
https://github.com/albu/albumentations/commit/4e12c6ec3e55cf79cf242a09c5cdc813bcfc6401
2.7 times faster for uint8 images.
https://github.com/albu/albumentations/commit/ac499d0365bfb2494cb535e82591fc3460d4595a
4 times faster for uint8 images.
https://github.com/albu/albumentations/commit/c028a9557cc960da11720a0a505a19cdd4fe0b24
https://github.com/albu/albumentations/commit/30a3f3024dc34597307c466a6307e2e6d27e9d3e Not all spatial tranforms jave keypoints support yet. In this release we added Crop, CropNonEmptyMaskIfExists, LongestMaxSize, RandomCropNearBBox, Resize, SmallestMaxSize, and Transpose.
We are delighted that albumentations are helpful to the academic community. We extended documentation with a page that lists all papers and preprints that cite albumentations in their work. This page is automatically generated by parsing Google Scholar. At this moment, this number is 24.
We are delighted that albumentations help people to get top results in machine learning competitions at Kaggle and other platforms. We added a "Hall of Fame" where people can share their achievements. This page is manually created. We encourage people to add more information about their results with pull requests, following the contributing guide.
@albu @Dipet @creafz @BloodAxe @ternaus @vfdev-5 @arsenyinfo @qubvel @toshiks @Jae-Hyuck @BelBES @alekseynp @timeous @jveitchmichaelis @bfialkoff
Nothing published for this version
Nothing published for this version
Nothing published for this version
Now we can define transformations in a python dictionary, json, yaml files and they will be deserialized and used in the code.
json, yaml files and they will be deserialized and used in the code.json and yaml files.Jupyter notebook with an example
Special thanks to @creafz
Special thanks to @vfdev-5 @ternaus @BloodAxe @kirillbobyrev
fill_value parameter to CutOutSpecial thanks to @qubvel @ternaus @albu @BloodAxe
Nothing published for this version
Nothing published for this version
Nothing published for this version
Special thanks to the Evegene Khvedchenya (@BloodAxe) for the work.
Special thanks to the Evegene Khvedchenya (@BloodAxe) for the work.
The possible use case are image2image or stereo-image pipelines.
Special thanks to Alexander Buslaev (@albu) for the work.
And many others.
@BloodAxe @albu @creafz @ternaus @erikgaas @marcocaccin @libfun @DBusAI @alexobednikov @StrikerRUS @IlyaOvodov @ZFTurbo @Vcv85 @georgymironov @LinaShiryaeva @vfdev-5 @daisukelab @cdicle
Your coding agent can read these notes before it upgrades. Set up the MCP server →