NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #1610 most downloaded on PyPI
Toolbox for imbalanced dataset in machine learning
Last release 3 months ago
07 Jun 2026
Ships fairly regularly
a new release about every 4 months
Most releases are documented
notes for 27 of 40 stable releases
Nothing withdrawn
no release was ever pulled
10 years old
40 releases · first in 2016
Compatibility with scikit-learn 1.9 #1178 by Guillaume Lemaitre.
Compatibility with scikit-learn 1.8 #1158 by Guillaume Lemaitre and stratakis.
One column per quarter.
Add InstanceHardnessCV to split data and ensure that samples are distributed in folds based on their instance hardness. #1125 by Frits Hermans.
algorithm parameter in RUSBoostClassifier is now deprecated and will be removed in 0.14. #1109 by @glemaitre.
get_metadata_routing in Pipeline such that one can use a sampler with metadata routing. #1115 by @glemaitre.Pipeline now uses check_is_fitted. In 0.15, it will raise an error instead of a warning. #1109 by @glemaitre.
algorithm parameter in RUSBoostClassifier is now deprecated and will be removed in 0.14. #1109 by @glemaitre.
Compatibility with NumPy 2.0+ #1097 by Guillaume Lemaitre.
Compatibility with scikit-learn 1.5 #1074 and #1084 by Guillaume Lemaitre.
Fix the way we check for a specific Python version in the test suite. #1075 by Guillaume Lemaitre.
Do not use distutils in tests due to deprecation. #1065 by Michael R. Crusoe.
Deprecate estimator_ argument in favor of estimators_ for the classes CondensedNearestNeighbour and OneSidedSelection. estimator_ will be removed in 0…
The fitted attribute ohe_ in SMOTENC is deprecated and will be removed in version 0.13. Use categorical_encoder_ instead. #1000 by Guillaume Lemaitre.
Fix a bug in classification_report_imbalanced where the parameter target_names was not taken into account when output_dict=True. #989 by AYY7.
SMOTENC now handles mix types of data type such as bool and pd.CategoricalDtype by delegating the conversion to scikit-learn encoder. #1002 by Guillaume Lemaitre.
Handle sparse matrices in SMOTEN and raise a warning since it requires a conversion to dense matrices. #1003 by Guillaume Lemaitre.
Remove spurious warning raised when minority class get over-sampled more than the number of sample in the majority class. #1007 by Guillaume Lemaitre.
The fitted attribute ohe_ in SMOTENC is deprecated and will be removed in version 0.13. Use categorical_encoder_ instead. #1000 by Guillaume Lemaitre.
The default of the parameters sampling_strategy and replacement will change in BalancedRandomForestClassifier to follow the implementation of the original paper. This changes will take effect in version 0.13. #1006 by Guillaume Lemaitre.
SMOTENC now accepts a parameter categorical_encoder allowing to specify a OneHotEncoder with custom parameters. #1000 by Guillaume Lemaitre.
SMOTEN now accepts a parameter categorical_encoder allowing to specify a OrdinalEncoder with custom parameters. A new fitted parameter categorical_encoder_ is exposed to access the fitted encoder. #1001 by Guillaume Lemaitre.
RandomUnderSampler and RandomOverSampler (when shrinkage is not None) now accept any data types and will not attempt any data conversion. #1004 by Guillaume Lemaitre.
SMOTENC now support passing array-like of str when passing the categorical_features parameter. #1008 by :userGuillaume Lemaitre <glemaitre>.
SMOTENC now support automatic categorical inference when categorical_features is set to "auto". #1009 by :userGuillaume Lemaitre <glemaitre>.
Fix a regression in over-sampler where the string minority was rejected as an unvalid sampling strategy. #964 by Prakhyath07.
minority was rejected as an unvalid sampling strategy. #964 by Prakhyath07.The parameter n_jobs has been deprecated from the classes ADASYN, BorderlineSMOTE, SMOTE, SMOTENC, SMOTEN, and SVMSMOTE. Instead, pass a nearest neigh…
python -OO that replaces doc by None. #953 bu Guillaume Lemaitre.feature_names_in_ as well as get_feature_names_out for all samplers. #959 by Guillaume Lemaitre.n_jobs has been deprecated from the classes ADASYN, BorderlineSMOTE, SMOTE, SMOTENC, SMOTEN, and SVMSMOTE. Instead, pass a nearest neighbors estimator where n_jobs is set. #887 by Guillaume Lemaitre.base_estimator is deprecated and will be removed in version 0.12. It is impacted the following classes: BalancedBaggingClassifier, EasyEnsembleClassifier, RUSBoostClassifier. #946 by Guillaume Lemaitre.Compatibility with scikit-learn 1.1.0
Compatibility with scikit-learn 1.1.0
Compatibility with scikit-learn 1.0.2
Compatibility with scikit-learn 1.0.2
Make imbalanced-learn compatible with scikit-learn 1.0. #864 by Guillaume Lemaitre.
September 29, 2021
Make imbalanced-learn compatible with scikit-learn 1.0. #864 by Guillaume Lemaitre.
The context manager imblearn.utils.testing.warns is deprecated in 0.8 and will be removed 1.0. #815 by Guillaume Lemaitre.
February 18, 2021
imblearn.metrics.macro_averaged_mean_absolute_error returning the average across class of the MAE. This metric is used in ordinal classification. #780 by Aurélien Massiot.imblearn.metrics.pairwise.ValueDifferenceMetric to compute pairwise distances between samples containing only categorical values. #796 by Guillaume Lemaitre.imblearn.over_sampling.SMOTEN to over-sample data only containing categorical features. #802 by Guillaume Lemaitre.imblearn.ensemble.BalancedBaggingClassifier unlocking the implementation of methods based on resampled bagging. #808 by Guillaume Lemaitre.output_dict in imblearn.metrics.classification_report_imbalanced to return a dictionary instead of a string. #770 by Guillaume Lemaitre.imblearn.under_sampling.ClusterCentroids where voting="hard" could have lead to select a sample from any class instead of the targeted class. #769 by Guillaume Lemaitre.imblearn.FunctionSampler where validation was performed even with validate=False when calling fit. #790 by Guillaume Lemaitre.extras_require within the setup.py file. #816 by Guillaume Lemaitre.pydata-sphinx-theme. #801 by Guillaume Lemaitre.imblearn.utils.testing.warns is deprecated in 0.8 and will be removed 1.0. #815 by Guillaume Lemaitre.A release to bump the minimum version of scikit-learn to 0.23 with a couple of bug fixes. Check the what's new for more information.
A release to bump the minimum version of scikit-learn to 0.23 with a couple of bug fixes. Check the what's new for more information.
This is a bug-fix release to resolve some issues regarding the handling the input and the output format of the arrays.
This is a bug-fix release to resolve some issues regarding the handling the input and the output format of the arrays.
This is a bug-fix release to primarily resolve some packaging issues in version 0.6.0. It also includes minor documentation improvements and some bug
This is a bug-fix release to primarily resolve some packaging issues in version 0.6.0. It also includes minor documentation improvements and some bug fixes.
Fix a bug in imblearn.ensemble.BalancedRandomForestClassifier leading to a wrong number of samples used during fitting due max_samples and therefore a bad computation of the OOB score. 656 by Guillaume Lemaitre.
The following classes have been removed after 2 deprecation cycles: ensemble.BalanceCascade and ensemble.EasyEnsemble. 617 by Guillaume Lemaitre .
The following models might give some different sampling due to changes in scikit-learn:
imblearn.under_sampling.ClusterCentroids
imblearn.under_sampling.InstanceHardnessThreshold
The following samplers will give different results due to change linked to the random state internal usage:
imblearn.over_sampling.SMOTENC
imblearn.under_sampling.InstanceHardnessThreshold now take into account the random_state and will give deterministic results. In addition, cross_val_predict is used to take advantage of the parallelism. 599 by Shihab Shahriar Khan.
Fix a bug in imblearn.ensemble.BalancedRandomForestClassifier leading to a wrong computation of the OOB score. 656 by Guillaume Lemaitre.
Update imports from scikit-learn after that some modules have been privatize. The following import have been changed: sklearn.ensemble._base._set_random_states, sklearn.ensemble._forest._parallel_build_trees, sklearn.metrics._classification._check_targets, sklearn.metrics._classification._prf_divide, sklearn.utils.Bunch, sklearn.utils._safe_indexing, sklearn.utils._testing.assert_allclose, sklearn.utils._testing.assert_array_equal, sklearn.utils._testing.SkipTest. 617 by Guillaume Lemaitre.
Synchronize imblearn.pipeline with sklearn.pipeline. 620 by Guillaume Lemaitre.
Synchronize imblearn.ensemble.BalancedRandomForestClassifier and add parameters max_samples and ccp_alpha. 621 by Guillaume Lemaitre.
imblearn.under_sampling.RandomUnderSampling, imblearn.over_sampling.RandomOverSampling, imblearn.datasets.make_imbalance accepts Pandas DataFrame in and will output Pandas DataFrame. Similarly, it will accepts Pandas Series in and will output Pandas Series. 636 by Guillaume Lemaitre.
imblearn.FunctionSampler accepts a parameter validate allowing to check or not the input X and y. 637 by Guillaume Lemaitre.
imblearn.under_sampling.RandomUnderSampler, imblearn.over_sampling.RandomOverSampler can resample when non finite values are present in X. 643 by Guillaume Lemaitre.
All samplers will output a Pandas DataFrame if a Pandas DataFrame was given as an input. 644 by Guillaume Lemaitre.
The samples generation in imblearn.over_sampling.SMOTE, imblearn.over_sampling.BorderlineSMOTE, imblearn.over_sampling.SVMSMOTE, imblearn.over_sampling.KMeansSMOTE, imblearn.over_sampling.SMOTENC is now vectorize with giving an additional speed-up when X in sparse. 596 by Matt Eding.
The following classes have been removed after 2 deprecation cycles: ensemble.BalanceCascade and ensemble.EasyEnsemble. 617 by Guillaume Lemaitre.
The following functions have been removed after 2 deprecation cycles: utils.check_ratio. 617 by Guillaume Lemaitre.
The parameter ratio and return_indices has been removed from all samplers. 617 by Guillaume Lemaitre.
The parameters m_neighbors, out_step, kind, svm_estimator have been removed from the imblearn.over_sampling.SMOTE. 617 by Guillaume Lemaitre.
Remove support for Python 2, remove deprecation warning from scikit-learn 0.21. 576 by Guillaume Lemaitre .
Changed models ---
The following models or function might give different results even if the same data X and y are the same.
imblearn.ensemble.RUSBoostClassifier default estimator changed from sklearn.tree.DecisionTreeClassifier with full depth to a decision stump (i.e., tree with max_depth=1).
Documentation ---
Correct the definition of the ratio when using a float in sampling strategy for the over-sampling and under-sampling. 525 by Ariel Rossanigo.
Add imblearn.over_sampling.BorderlineSMOTE and imblearn.over_sampling.SVMSMOTE in the API documenation. 530 by Guillaume Lemaitre.
Enhancement --- - Add Parallelisation for SMOTEENN and SMOTETomek.
547 by Michael Hsieh.
Add imblearn.utils._show_versions. Updated the contribution guide and issue template showing how to print system and dependency information from the command line. 557 by Alexander L. Hayes.
Add imblearn.over_sampling.KMeansSMOTE which is an over-sampler clustering points before to apply SMOTE. 435 by Stephan Heijl.
Maintenance ---
Make it possible to import imblearn and access submodule. 500 by Guillaume Lemaitre.
Remove support for Python 2, remove deprecation warning from scikit-learn 0.21. 576 by Guillaume Lemaitre.
Fix wrong usage of keras.layers.BatchNormalization in porto_seguro_keras_under_sampling.py example. The batch normalization was moved before the activation function and the bias was removed from the dense layer. 531 by Guillaume Lemaitre.
Fix bug which converting to COO format sparse when stacking the matrices in imblearn.over_sampling.SMOTENC. This bug was only old scipy version. 539 by Guillaume Lemaitre.
Fix bug in imblearn.pipeline.Pipeline where None could be the final estimator. 554 by Oliver Rausch.
Fix bug in imblearn.over_sampling.SVMSMOTE and imblearn.over_sampling.BorderlineSMOTE where the default parameter of n_neighbors was not set properly. 578 by Guillaume Lemaitre.
Fix bug by changing the default depth in imblearn.ensemble.RUSBoostClassifier to get a decision stump as a weak learner as in the original paper. 545 by Christos Aridas.
Allow to import keras directly from tensorflow in the imblearn.keras. 531 by Guillaume Lemaitre.
Nothing published for this version
Fix a bug in imblearn.over_sampling.SMOTENC in which the the median of the standard deviation instead of half of the median of the standard deviation.
Version 0.4.2
Bug fixes
In addition, the return_indices argument has been deprecated and all samplers will exposed a sample_indices_ whenever this is possible.
October, 2018
Version 0.4 is the last version of imbalanced-learn to support Python 2.7 and Python 3.4. Imbalanced-learn 0.5 will require Python 3.5 or higher.
This release brings its set of new feature as well as some API changes to strengthen the foundation of imbalanced-learn.
As new feature, 2 new modules imblearn.keras and
imblearn.tensorflow have been added in which imbalanced-learn samplers
can be used to generate balanced mini-batches.
The module imblearn.ensemble has been consolidated with new classifier:
imblearn.ensemble.BalancedRandomForestClassifier,
imblearn.ensemble.EasyEnsembleClassifier,
imblearn.ensemble.RUSBoostClassifier.
Support for string has been added in
imblearn.over_sampling.RandomOverSampler and
imblearn.under_sampling.RandomUnderSampler. In addition, a new class
imblearn.over_sampling.SMOTENC allows to generate sample with data
sets containing both continuous and categorical features.
The imblearn.over_sampling.SMOTE has been simplified and break down
to 2 additional classes:
imblearn.over_sampling.SVMSMOTE and
imblearn.over_sampling.BorderlineSMOTE.
There is also some changes regarding the API:
the parameter sampling_strategy has been introduced to replace the
ratio parameter. In addition, the return_indices argument has been
deprecated and all samplers will exposed a sample_indices_ whenever this is
possible.
In addition, the return_indices argument has been deprecated and all samplers will exposed a sample_indices_ whenever this is possible.
October, 2018
Warning
Version 0.4 is the last version of imbalanced-learn to support Python 2.7 and Python 3.4. Imbalanced-learn 0.5 will require Python 3.5 or higher.
This release brings its set of new feature as well as some API changes to strengthen the foundation of imbalanced-learn.
As new feature, 2 new modules imblearn.keras and imblearn.tensorflow have been added in which imbalanced-learn samplers can be used to generate balanced mini-batches.
The module imblearn.ensemble has been consolidated with new classifier: imblearn.ensemble.BalancedRandomForestClassifier, imblearn.ensemble.EasyEnsembleClassifier, imblearn.ensemble.RUSBoostClassifier.
Support for string has been added in imblearn.over_sampling.RandomOverSampler and imblearn.under_sampling.RandomUnderSampler. In addition, a new class imblearn.over_sampling.SMOTENC allows to generate sample with data sets containing both continuous and categorical features.
The imblearn.over_sampling.SMOTE has been simplified and break down to 2 additional classes: imblearn.over_sampling.SVMSMOTE and imblearn.over_sampling.BorderlineSMOTE.
There is also some changes regarding the API: the parameter sampling_strategy has been introduced to replace the ratio parameter. In addition, the return_indices argument has been deprecated and all samplers will exposed a sample_indices_ whenever this is possible.
Bug fix in the classification report
Bug fix in the classification report
Nothing published for this version
Nothing published for this version
Deprecation of the use of min_c_ in datasets.make_imbalance. 312 by Guillaume Lemaitre_
# What's new in version 0.3.0
## Testing
Pytest is used instead of nosetests. 321 by `Joan Massich`_.
## Documentation
Added a User Guide and extended some examples. 295 by `Guillaume Lemaitre`_.
# Bug fixes
Fixed a bug in utils.check_ratio such that an error is raised when the number of samples required is negative. 312 by `Guillaume Lemaitre`_.
Fixed a bug in under_sampling.NearMiss version 3. The indices returned were wrong. 312 by `Guillaume Lemaitre`_.
Fixed bug for ensemble.BalanceCascade and combine.SMOTEENN and SMOTETomek. 295 by `Guillaume Lemaitre`_.`
Fixed bug for check_ratio to be able to pass arguments when ratio is a callable. 307 by `Guillaume Lemaitre`_.`
## New features
Turn off steps in pipeline.Pipeline using the None object. By `Christos Aridas`_.
Add a fetching function datasets.fetch_datasets in order to get some imbalanced datasets useful for benchmarking. 249 by `Guillaume Lemaitre`_.
## Enhancement
All samplers accepts sparse matrices with defaulting on CSR type. 316 by `Guillaume Lemaitre`_.
datasets.make_imbalance take a ratio similarly to other samplers. It supports multiclass. 312 by `Guillaume Lemaitre`_.
All the unit tests have been factorized and a utils.check_estimators has been derived from scikit-learn. By `Guillaume Lemaitre`_.
Script for automatic build of conda packages and uploading. 242 by `Guillaume Lemaitre`_
Remove seaborn dependence and improve the examples. 264 by `Guillaume Lemaitre`_.
adapt all classes to multi-class resampling. 290 by `Guillaume Lemaitre`_
## API changes summary
__init__ has been removed from the base.SamplerMixin to create a real mixin class. 242 by `Guillaume Lemaitre`_.
creation of a module exceptions to handle consistant raising of errors. 242 by `Guillaume Lemaitre`_.
creation of a module utils.validation to make checking of recurrent patterns. 242 by `Guillaume Lemaitre`_.
move the under-sampling methods in prototype_selection and prototype_generation submodule to make a clearer dinstinction. 277 by `Guillaume Lemaitre`_.
change ratio such that it can adapt to multiple class problems. 290 by `Guillaume Lemaitre`_.
## Deprecation
Deprecation of the use of min_c_ in datasets.make_imbalance. 312 by `Guillaume Lemaitre`_
Deprecation of the use of float in datasets.make_imbalance for the ratio parameter. 290 by `Guillaume Lemaitre`_.
deprecate the use of float as ratio in favor of dictionary, string, or callable. 290 by `Guillaume Lemaitre`_.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Release created after transferring the repository to scikit-learn-contrib.
Release created after transferring the repository to scikit-learn-contrib.
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →