NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4796 most downloaded on PyPI
Machine Learning Library Extensions
Last release 4 months ago
06 Jun 2026
Ships fairly regularly
a new release about every 6 months
Nearly every release is documented
notes for 51 of 54 stable releases
Nothing withdrawn
no release was ever pulled
12 years old
55 releases · first in 2014
MAINT: remove deprecated multi_class="multinomial" from LogisticRegression usage by @sachinn854 in https://github.com/rasbt/mlxtend/pull/1147
Full Changelog: https://github.com/rasbt/mlxtend/compare/v0.24.0...v0.25.0
One column per quarter.
Source code (zip)
Source code (tar.gz)
Removed explicit multi_class="multinomial" arguments, which are deprecated in newer versions of scikit-learn, from LogisticRegression usage in examples / notebooks / tests ( #1147 via sachinn854 )
Added multiprocessing support for apriori via the n_jobs parameter ( #1151 via mariam851 )
Fixes an edge-case bug where decision regions plots didn't have unique colors ( #1157 via mariam851 )
Fix preprocessing.standardize so a constant column is mapped to all-zeros (as the docstring promises) instead of -mean(column) ( #1058 via jbbqqf )
Reject min_support values outside the documented (0, 1] interval in apriori , fpgrowth , fpmax , and hmine . The previous check only caught <= 0 , so passing e.g. min_support=2 silently returned an empty result ( #864 via jbbqqf )
Add a top_k argument to ExhaustiveFeatureSelector.get_metric_dict() so callers can request only the highest-scoring subsets before converting the result to a DataFrame ( #610 via jbbqqf )
minmax_scaling no longer returns silent NaNs for constant columns; constant columns are now collapsed to min_val , mirroring the existing contract of standardize . ( #1167 via jbbqqf )
bias_variance_decomp now accepts pandas DataFrames and Series as input, in addition to NumPy arrays. ( #1070 via berns722 )
Clarified in the bias_variance_decomp docstring that, for the mse loss, avg_bias is the average squared bias (the Bias^2 term in Loss = Bias^2 + Variance ), and fixed a typo in the Returns section that previously read "average bias, and average bias". Behaviour unchanged. ( #1083 via jbbqqf )
Compatibility with latest scikit-learn (1.8.0) and pandas versions (2.3.3)
mlxtend/classifier/stacking_cv_classification.py and mlxtend/regressor/stacking_cv_regression.py
- Modified meta_features to ensure compatibility with scikit-learn versions 1.4 and above by dynamically selecting between fit_params and params in cross_val_predict.Source code (zip)
Source code (tar.gz)
Compatibility with latest scikit-learn (1.8.0) and pandas versions (2.3.3)
mlxtend/classifier/stacking_cv_classification.py and mlxtend/regressor/stacking_cv_regression.py
Modified meta_features to ensure compatibility with scikit-learn versions 1.4 and above by dynamically selecting between fit_params and params in cross_val_predict .
fix(test): replace np.float_ to np.float64 by @Bot-wxt1221 in https://github.com/rasbt/mlxtend/pull/1119
Full Changelog: https://github.com/rasbt/mlxtend/compare/v0.23.3...v0.23.4
Source code (zip)
Source code (tar.gz)
Improved plot_splits for time series splits by @d-kleine in https://github.com/rasbt/mlxtend/pull/1113
plot_splits for time series splits by @d-kleine in https://github.com/rasbt/mlxtend/pull/1113publish CI/CD workflow by @d-kleine in https://github.com/rasbt/mlxtend/pull/1111Full Changelog: https://github.com/rasbt/mlxtend/compare/v0.23.2...v0.23.3
Source code (zip)
Source code (tar.gz)
Files updated:
Don't include tests in built wheel by @carlsmedstad in https://github.com/rasbt/mlxtend/pull/1076
set_output method into TransactionEncoder by @it176131 in https://github.com/rasbt/mlxtend/pull/1087_calc_score for scikit-learn version compatibility by @d-kleine in https://github.com/rasbt/mlxtend/pull/1109Full Changelog: https://github.com/rasbt/mlxtend/compare/v0.23.1...v0.23.2
Source code (zip)
Source code (tar.gz)
Implement the FP-Growth and FP-Max algorithms with the possibility of missing values in the input dataset. Added a new metric Representativity for the association rules generated ( #1004 via zazass8 ). Files updated:
['mlxtend.frequent_patterns.fpcommon']
'mlxtend.frequent_patterns.fpgrowth'
'mlxtend.frequent_patterns.fpmax'
'mlxtend/feature_selection/utilities.py'
Modified _calc_score function to ensure compatibility with scikit-learn versions 1.4 and above by dynamically selecting between fit_params and params in cross_val_score .
mlxtend.feature_selection.SequentialFeatureSelector
Updated negative infinity constant to be compatible with old and new (>=2.0) numpy versions
mlxtend.frequent_patterns.association_rules
Implemented three new metrics: Jaccard, Certainty, and Kulczynski. ( #1096 )
Integrated scikit-learn's set_output method into TransactionEncoder ( #1087 via it176131 )
[ mlxtend.frequent_patterns.fpcommon ] Added the null_values parameter in valid_input_check signature to check in case the input also includes null values. Changes the returns statements and function signatures for setup_fptree and generated_itemsets respectively to return the disabled array created and to include it as a parameter. Added code in [ mlxtend.frequent_patterns.fpcommon ] and mlxtend.frequent_patterns.association_rules to implement the algorithms in case null values exist when null_values is True.
mlxtend.frequent_patterns.association_rules Added optional parameter 'return_metrics' to only return a given list of metrics, rather than every possible metric.
Add n_classes_ attribute to stacking classifiers for compatibility with scikit-learn 1.3 ( #1091 )
Updated dependency on distutils for python 3.12 and above ([#1072](https://github.com/rasbt/mlxtend/issues/1072) via [peanutsee](https://github.com/pe
Source code (zip)
Source code (tar.gz)
Address NumPy deprecations to make mlxtend compatible to NumPy 1.24
[Source code (zip)](https://github.com/rasbt/mlxtend/archive/v0.21.1.zip)
[Source code (tar.gz)](https://github.com/rasbt/mlxtend/archive/v0.22.1.tar.gz)
LinearRegression model of sklearn in the test removing the normalize parameter as it is deprecated. ([#1036](https://github.com/rasbt/mlxtend/issues/1036))pyproject.toml to support PEP 518 builds ([#1065](https://github.com/rasbt/mlxtend/issues/1065) via [jmahlik](https://github.com/jmahlik))pyproject.toml ([#1065](https://github.com/rasbt/mlxtend/issues/1065) via [jmahlik](https://github.com/jmahlik))mlxtend.image submodule with face recognition functions due to poor dlib support in modern environments.SequentialFeatureSelector and multiclass ROC AUC.When `ExhaustiveFeatureSelector` is run with n_jobs == 1, joblib is now disabled, which enables more immediate (live) feedback when the verbose mode i
ExhaustiveFeatureSelector is run with n_jobs == 1, joblib is now disabled, which enables more immediate (live) feedback when the verbose mode is enabled. (#985 via Nima Sarajpoor)EnsembleVoteClassifier (#941)mlxtend.frequent_patterns.association_rules function has a new metric - Zhang's Metric, which measures both association and dissociation. (#980)mlxtend.frequent_patterns.fpmax code improvement that avoids casting a sparse DataFrame into a dense NumPy array. (#1000 via Tim Kellogg)plot_decision_regions function now has a n_jobs parameter to parallelize the computation. (In a particular use case, on a small dataset, there was a 21x speed-up (449 seconds vs 21 seconds on local HPC instance of 36 cores). (#998 via Khalid ElHaj)mlxtend.frequent_patterns.hmine algorithm and documentation for mining frequent itemsets using the H-Mine algorithm. (#1020 via Fatih Sen)The mlxtend.evaluate.feature_importance_permutation function has a new feature_groups argument to treat user-specified feature groups as single featur
mlxtend.evaluate.feature_importance_permutation function has a new feature_groups argument to treat user-specified feature groups as single features, which is useful for one-hot encoded features. (#955)mlxtend.feature_selection.ExhaustiveFeatureSelector and SequentialFeatureSelector also gained support for feature_groups with a behavior similar to the one described above. (#957 and #965 via Nima Sarajpoor)custom_feature_names parameter was removed from the ExhaustiveFeatureSelector due to redundancy and to simplify the code base. The ExhaustiveFeatureSelector documentation illustrates how the same behavior and outcome can be achieved using pandas DataFrames. (#957)The mlxtend.evaluate.bootstrap_point632_score now supports fit_params.
mlxtend.evaluate.bootstrap_point632_score now supports fit_params. (#861)mlxtend/plotting/decision_regions.py function now has a contourf_kwargs for matplotlib to change the look of the decision boundaries if desired. (#881 via [pbloem])norm_colormap parameter to mlxtend.plotting.plot_confusion_matrix, to allow normalizing the colormap, e.g., using matplotlib.colors.LogNorm() (#895)GroupTimeSeriesSplit class for evaluation in time series tasks with support of custom groups and additional parameters in comparison with scikit-learn's TimeSeriesSplit. (#915 via Dmitry Labazkin)mlxtend.plotting.heatmap and mlxtend.plotting.plot_confusion_matrix (#872)apriori, fpmax, and fpgrowth. (#934 via NimaSarajpoor)Removes deprecated res argument from plot_decision_regions.
evaluate.accuracy_score in addition to the existing "average" option to compute the scikit-learn-style balanced accuracy. (#764)scatter_hist function to mlxtend.plotting for generating a scattered histogram. (#757 via Maitreyee Mhasaka)evaluate.permutation_test function now accepts a paired argument to specify to support paired permutation/randomization tests. (#768)StackingCVRegressor now also supports multi-dimensional targets similar to StackingRegressor via StackingCVRegressor(..., multi_output=True). (#802 via Marco Tiraboschi)StackingRegressor now requires setting StackingRegressor(..., multi_output=True) if the target is multi-dimensional; this allows for better input validation. (#802)res argument from plot_decision_regions. (#803)title_fontsize parameter to plot_learning_curves for controlling the title font size; also the plot style is now the matplotlib default. (#818)'c': 'none' instead of 'c': '' in mlxtend.plotting.plot_decision_regions's scatterplot highlights to stay compatible with Matplotlib 3.4 and newer. (#822)fontcolor_threshold parameter to the mlxtend.plotting.plot_confusion_matrix function as an additional option for determining the font color cut-off manually. (#827)frequent_patterns.association_rules now raises a ValueError if an empty frequent itemset DataFrame is passed. (#843)mlxtend.evaluate.bootstrap_point632_score function now use the whole training set for the resubstitution weighting term instead of the internal training set that is a new bootstrap sample in each round. (#844)The bias_variance_decomp function now supports optional fit_params for the estimators that are fit on bootstrap samples.
bias_variance_decomp function now supports optional fit_params for the estimators that are fit on bootstrap samples. (#748)bias_variance_decomp function now supports Keras estimators. (#725 via @hanzigs)mlxtend.classifier.OneRClassifier (One Rule Classfier) class, a simple rule-based classifier that is often used as a performance baseline or simple interpretable model. (#726create_counterfactual method for creating counterfactuals to explain model predictions. (#740)permutation_test (mlxtend.evaluate.permutation) ìs corrected to give the proportion of permutations whose statistic is at least as extreme as the one observed. (#721 via Florian Charlier)LogisticRegression for logging purposes didn't include the L2 penalty for the first weight in the weight vector (this is not the bias unit). However, since this loss function was only used for logging purposes, and the gradient remains correct, this does not have an effect on the main code. (#741)bias_variance_decomp where when the mse loss was used, downcasting to integers caused imprecise results for small numbers. (#749)Switched to using raw strings for regex in mlxtend.text to prevent deprecation warning in Python 3.8
predict_proba kwarg to bootstrap methods, to allow bootstrapping of scoring functions that take in probability values. (#700 via Adam Li)cell_values parameter to mlxtend.plotting.heatmap() to optionally suppress cell annotations by setting cell_values=False. (#703use_clones and fit_base_estimators (previously refit in EnsembleVoteClassifier) for EnsembleVoteClassifier and StackingClassifier. (#670 via Katrina Ni)mlxtend.text to prevent deprecation warning in Python 3.8 (#688)meshgrid in no_information_rate function used by the bootstrap_point632_score function for the .632+ estimate. (#688)fpmax that could lead to incorrect support values. (#692 via Steve Harenberg)The previously deprecated OnehotTransactions has been removed in favor of the TransactionEncoder.
OnehotTransactions has been removed in favor of the TransactionEncoder.SparseDataFrame support in frequent pattern mining functions in favor of pandas >=1.0's new way for working sparse data. If you used SparseDataFrame formats, please see pandas' migration guide at https://pandas.pydata.org/pandas-docs/stable/user_guide/sparse.html#migrating (#667)Source code (zip)
Source code (tar.gz)
The previously deprecated OnehotTransactions has been removed in favor of the TransactionEncoder.
Removed SparseDataFrame support in frequent pattern mining functions in favor of pandas >=1.0's new way for working sparse data. If you used SparseDataFrame formats, please see pandas' migration guide at https://pandas.pydata.org/pandas-docs/stable/user_guide/sparse.html#migrating. ( #667 )
The plot_confusion_matrix.py now also accepts a matplotlib figure and axis as input to which the confusion matrix plot can be added. ( #671 via Vahid Mirjalili )
The SequentialFeatureSelector now supports using pre-specified feature sets via the fixed_features parameter.
SequentialFeatureSelector now supports using pre-specified feature sets via the fixed_features parameter. (#578)accuracy_score function to mlxtend.evaluate for computing basic classifcation accuracy, per-class accuracy, and average per-class accuracy. (#624 via Deepan Das)StackingClassifier and StackingCVClassifiernow have a decision_function method, which serves as a preferred choice over predict_proba in calculating roc_auc and average_precision scores when the meta estimator is a linear model or support vector classifier. (#634 via Qiang Gu)apriori frequent itemset generating function when low_memory=True. Setting low_memory=False (default) is still faster for small itemsets, but low_memory=True can be much faster for large itemsets and requires less memory. Also, input validation for apriori, ̀ fpgrowthandfpmaxtakes a significant amount of time when input pandas DataFrame is large; this is now dramatically reduced when input contains boolean values (and not zeros/ones), which is the case when usingTransactionEncoder`. (#619 via Denis Barbier)apriori, ̀ fpgrowthandfpmax` runs much faster on sparse DataFrame when input pandas DataFrame contains integer values. (#621 via Denis Barbier)fpgrowth and fpmax directly work on sparse DataFrame, they were previously converted into dense Numpy arrays. (#622 via Denis Barbier)mlxtend.plotting.plot_pca_correlation_graph that caused the explaind variances not summing up to 1. Also, improves the runtime performance of the correlation computation and adds a missing function argument for the explained variances (eigenvalues) if users provide their own principal components. (#593 via Gabriel Azevedo Ferreira)fpgrowth and apriori consistent for edgecases such as min_support=0. (#573 via Steve Harenberg)fpmax returns an empty data frame now instead of raising an error if the frequent itemset set is empty. (#573 via Steve Harenberg)mlxtend.plotting.plot_confusion_matrix, where the font-color choice for medium-dark cells was not ideal and hard to read. #588 via sohrabtowfighi)svd mode of mlxtend.feature_extraction.PrincipalComponentAnalysis now also n-1 degrees of freedom instead of n d.o.f. when computing the eigenvalues to match the behavior of eigen. #595StackingCVClassifier because it causes issues if pipelines are used as input. #606Added an enhancement to the existing iris_data() such that both the UCI Repository version of the Iris dataset as well as the corrected, original vers
iris_data() such that both the UCI Repository version of the Iris dataset as well as the corrected, original
version of the dataset can be loaded, which has a slight difference in two data points (consistent with Fisher's paper; this is also the same as in R). (via #539 via janismdhanbad)groups parameter to SequentialFeatureSelector and ExhaustiveFeatureSelector fit() methods for forwarding to sklearn CV (#537 via arc12)plot_pca_correlation_graph function to the mlxtend.plotting submodule for plotting a PCA correlation graph. (#544 via Gabriel-Azevedo-Ferreira)zoom_factor parameter to the mlxten.plotting.plot_decision_region function that allows users to zoom in and out of the decision region plots. (#545)fpgrowth that implements the FP-Growth algorithm for mining frequent itemsets as a drop-in replacement for the existing apriori algorithm. (#550 via Steve Harenberg)heatmap function in mlxtend.plotting. (#552)fpmax that implements the FP-Max algorithm for mining maximal itemsets as a drop-in replacement for the fpgrowth algorithm. (#553 via Steve Harenberg)figsize parameter for the plot_decision_regions function in mlxtend.plotting. (#555 via Mirza Hasanbasic)low_memory option for the apriori frequent itemset generating function. Setting low_memory=False (default) uses a substantially optimized version of the algorithm that is 3-6x faster than the original implementation (low_memory=True). (#567 via jmayse)sklearn.externals.joblib. (#547)StackingCVClassifier and StackingCVRegressor such that first-level models are allowed to generate output of non-numeric type. (#562)iris_data() under iris.py by adding a note about differences in the iris data in R and UCI machine learning repo.'svd' mode is used in PCA, the number of eigenvalues is the same as when using 'eigen' (append 0's zeros in that case) (#565)StackingCVClassifier and StackingCVRegressor now support random_state parameter, which, together with shuffle, controls the randomness in the cv split
StackingCVClassifier and StackingCVRegressor now support random_state parameter, which, together with shuffle, controls the randomness in the cv splitting. (#523 via Qiang Gu)StackingCVClassifier and StackingCVRegressor now have a new drop_last_proba parameter. It drops the last "probability" column in the feature set since if True,
because it is redundant: p(y_c) = 1 - p(y_1) + p(y_2) + ... + p(y_{c-1}). This can be useful for meta-classifiers that are sensitive to perfectly collinear features. (#532)StackingClassifier, StackingCVClassifier and StackingRegressor, support grid search over the regressors and even a single base regressor. (#522 via Qiang Gu)StackingCVClassifier. (#522 via Qiang Gu)StackingCVRegressor. (#512 via Qiang Gu)StackingCVRegressor also enables grid search over the regressors and even a single base regressor. When there are level-mixed parameters, GridSearchCV will try to replace hyperparameters in a top-down order (see the documentation for examples details). (#515 via Qiang Gu)verbose parameter to apriori to show the current iteration number as well as the itemset size currently being sampled. (#519class_name parameter to the confusion matrix function to display class names on the axis as tick marks. (#487 via sandpiturtle)GridSearchCV, etc.) the StackingCVRegressor's meta regressor is now being accessed via 'meta_regressor__* in the parameter grid. E.g., if a RandomForestRegressor as meta- egressor was previously tuned via 'randomforestregressor__n_estimators', this has now changed to 'meta_regressor__n_estimators'. (#515 via Qiang Gu)StackingClassifier, StackingCVClassifier and StackingRegressor. (#522 via Qiang Gu)feature_selection.ColumnSelector now also supports column names of type int (in addition to str names) if the input is a pandas DataFrame. (#500 via tetrar124plot_confusion_matrix for imbalanced datasets if show_absolute=True and show_normed=True. (#504)SparseDataFrame is passed to apriori and the dataframe has integer column names that don't start with 0 due to current limitations of the SparseDataFrame implementation in pandas. (#503)mlxtend.evaluate.feature_importance_permutation now correctly accepts scoring functions with proper function signature as metric argument. #528Source code (zip)
Source code (tar.gz)
StackingCVClassifier and StackingCVRegressor now support random_state parameter, which, together with shuffle , controls the randomness in the cv splitting. ( #523 via Qiang Gu )
StackingCVClassifier and StackingCVRegressor now have a new drop_last_proba parameter. It drops the last "probability" column in the feature set since if True , because it is redundant: p(y_c) = 1 - p(y_1) + p(y_2) + ... + p(y_{c-1}). This can be useful for meta-classifiers that are sensitive to perfectly collinear features. ( #532 )
Other stacking estimators, including StackingClassifier , StackingCVClassifier and StackingRegressor , support grid search over the regressors and even a single base regressor. ( #522 via Qiang Gu )
Adds multiprocessing support to StackingCVClassifier . ( #522 via Qiang Gu )
Adds multiprocessing support to StackingCVRegressor . ( #512 via Qiang Gu )
Now, the StackingCVRegressor also enables grid search over the regressors and even a single base regressor. When there are level-mixed parameters, GridSearchCV will try to replace hyperparameters in a top-down order (see the documentation for examples details). ( #515 via Qiang Gu )
Adds a verbose parameter to apriori to show the current iteration number as well as the itemset size currently being sampled. ( #519
Adds an optional class_name parameter to the confusion matrix function to display class names on the axis as tick marks. ( #487 via sandpiturtle )
Adds a pca.e_vals_normalized_ attribute to PCA for storing the eigenvalues also in normalized form; this is commonly referred to as variance explained ratios. #545
Due to new features, restructuring, and better scikit-learn support (for GridSearchCV , etc.) the StackingCVRegressor 's meta regressor is now being accessed via 'meta_regressor__* in the parameter grid. E.g., if a RandomForestRegressor as meta- egressor was previously tuned via 'randomforestregressor__n_estimators' , this has now changed to 'meta_regressor__n_estimators' . ( #515 via Qiang Gu )
The same change mentioned above is now applied to other stacking estimators, including StackingClassifier , StackingCVClassifier and StackingRegressor . ( #522 via Qiang Gu )
Automatically performs mean centering for PCA solver 'SVD' such that using SVD is always equal to using the covariance matrix approach #545
The feature_selection.ColumnSelector now also supports column names of type int (in addition to str names) if the input is a pandas DataFrame. ( #500 via tetrar124
Fix unreadable labels in plot_confusion_matrix for imbalanced datasets if show_absolute=True and show_normed=True . ( #504 )
Raises a more informative error if a SparseDataFrame is passed to apriori and the dataframe has integer column names that don't start with 0 due to current limitations of the SparseDataFrame implementation in pandas. ( #503 )
SequentialFeatureSelector now supports DataFrame as input for all operating modes (forward/backward/floating). #506
mlxtend.evaluate.feature_importance_permutation now correctly accepts scoring functions with proper function signature as metric argument. #528
Nothing published for this version
Addressed deprecations warnings in NumPy 0.15.
scatterplotmatrix function to the plotting module. (#437)sample_weight option to StackingRegressor, StackingClassifier, StackingCVRegressor, StackingCVClassifier, EnsembleVoteClassifier. (#438)RandomHoldoutSplit class to perform a random train/valid split without rotation in SequentialFeatureSelector, scikit-learn GridSearchCV etc. (#442)PredefinedHoldoutSplit class to perform a train/valid split, based on user-specified indices, without rotation in SequentialFeatureSelector, scikit-learn GridSearchCV etc. (#443)mlxtend.image submodule for working on image processing-related tasks. (#457)extract_face_landmarks based on dlib to mlxtend.image. (#458)method='oob' option to the mlxtend.evaluate.bootstrap_point632_score method to compute the classic out-of-bag bootstrap estimate (#459)method='.632+' option to the mlxtend.evaluate.bootstrap_point632_score method to compute the .632+ bootstrap estimate that addresses the optimism bias of the .632 bootstrap (#459)mlxtend.evaluate.ftest function to perform an F-test for comparing the accuracies of two or more classification models. (#460)mlxtend.evaluate.combined_ftest_5x2cv function to perform an combined 5x2cv F-Test for comparing the performance of two models. (#461)mlxtend.evaluate.difference_proportions test for comparing two proportions (e.g., classifier accuracies) (#462)mlxtend.plotting.plot_confusion_matrix. (#428)A meaningful error message is now raised when a cross-validation generator is used with SequentialFeatureSelector.
SequentialFeatureSelector. (#377)SequentialFeatureSelector now accepts custom feature names via the fit method for more interpretable feature subset reports. (#379)SequentialFeatureSelector is now also compatible with Pandas DataFrames and uses DataFrame column-names for more interpretable feature subset reports. (#379)ColumnSelector now works with Pandas DataFrames columns. (#378 by Manuel Garrido)ExhaustiveFeatureSelector estimator in mlxtend.feature_selection now is safely stoppable mid-process by control+c. (#380)vectorspace_orthonormalization and vectorspace_dimensionality were added to mlxtend.math to use the Gram-Schmidt process to convert a set of linearly independent vectors into a set of orthonormal basis vectors, and to compute the dimensionality of a vectorspace, respectively. (#382)mlxtend.frequent_patterns.apriori now supports pandas SparseDataFrames to generate frequent itemsets. (#404 via Daniel Morales)plot_confusion_matrix function now has the ability to show normalized confusion matrix coefficients in addition to or instead of absolute confusion matrix coefficients with or without a colorbar. The text display method has been changed so that the full range of the colormap is used. The default size is also now set based on the number of classes.StackingRegressor (via use_features_in_secondary) like it is already supported in the other Stacking classes. (#418)support_only to the association_rules function, which allow constructing association rules (based on the support metric only) for cropped input DataFrames that don't contain a complete set of antecedent and consequent support values. (#421)apriori are now frozensets (#393 by William Laney and #394)apriori contains non 0, 1, True, False values. #419)clone function. (#374)refit=False in StackingRegressor and StackingCVRegressor (#384 and (#385) by selay01)StackingClassifier to work with sparse matrices when use_features_in_secondary=True (#408 by Floris Hoogenbook)StackingCVRegressor to work with sparse matrices when use_features_in_secondary=True (#416)StackingCVClassifier to work with sparse matrices when use_features_in_secondary=True (#417)A new feature_importance_permuation function to compute the feature importance in classifiers and regressors via the *permutation importance* method
feature_importance_permuation function to compute the feature importance in classifiers and regressors via the permutation importance method (#358)ExhaustiveFeatureSelector now optionally accepts **fit_params for the estimator that is used for the feature selection. (#354 by Zach Griffith)SequentialFeatureSelector now optionally accepts
**fit_params for the estimator that is used for the feature selection. (#350 by Zach Griffith)plot_decision_regions colors by a colorblind-friendly palette and adds contour lines for decision regions. (#348)NonFittedErrors if any method for inference is called prior to fitting the estimator. (#353)refit parameter of both the StackingClassifier and StackingCVClassifier to use_clones to be more explicit and less misleading. (#368)StackingCVClassifier's meta features were not stored in the original order when shuffle=True (#370)The old res parameter has been deprecated. (#309 by Guillaume Poirier-Morency)
paired_ttest_resampled)
to compare the performance of two models
(also called k-hold-out paired t-test). (#323)paired_ttest_kfold_cv)
to compare the performance of two models
(also called k-hold-out paired t-test). (#324)paired_ttest_5x2cv) proposed by Dieterrich (1998)
to compare the performance of two models. (#325)refit parameter was added to stacking classes (similar to the refit parameter in the EnsembleVoteClassifier), to support classifiers and regressors that follow the scikit-learn API but are not compatible with scikit-learn's clone function. (#325)ColumnSelector now has a drop_axis argument to use it in pipelines with CountVectorizers. (#333)predict or predict_meta_features is called prior to calling the fit method in StackingRegressor and StackingCVRegressor. (#315)plot_decision_regions function now automatically determines the optimal setting based on the feature dimensions and supports anti-aliasing. The old res parameter has been deprecated. (#309 by Guillaume Poirier-Morency)onehot transformation and the amount of candidates generated by the apriori algorithm. (#327 by Jakub Smid)OnehotTransactions class (which is typically often used in combination with the apriori function for association rule mining) is now more memory efficient as it uses boolean arrays instead of integer arrays. In addition, the OnehotTransactions class can be now be provided with sparse argument to generate sparse representations of the onehot matrix to further improve memory efficiency. (#328 by Jakub Smid)OneHotTransactions has been deprecated and replaced by the TransactionEncoder. (#332plot_decision_regions function now has three new parameters, scatter_kwargs, contourf_kwargs, and scatter_highlight_kwargs, that can be used to modify the plotting style. (#342 by James Bourbeau)EnsembleVoteClassifier when refit was set to false. (#322)plot_decision_regions function. (#337)New store_train_meta_features parameter for fit in StackingCVRegressor. if True, train meta-features are stored in self.train_meta_features_. New pred
store_train_meta_features parameter for fit in StackingCVRegressor. if True, train meta-features are stored in self.train_meta_features_.
New pred_meta_features method for StackingCVRegressor. People can get test meta-features using this method. (#294 via takashioya)store_train_meta_features attribute and pred_meta_features method for the StackingCVRegressor were also added to the StackingRegressor, StackingClassifier, and StackingCVClassifier (#299 & #300)evaluate.mcnemar_tables) for creating multiple 2x2 contigency from model predictions arrays that can be used in multiple McNemar (post-hoc) tests or Cochran's Q or F tests, etc. (#307)evaluate.cochrans_q) for performing Cochran's Q test to compare the accuracy of multiple classifiers. (#310)requirements.txt to setup.py. (#304 via Colin Carrol)Fixed an deprecation error that occured with McNemar test when using SciPy 1.0.
mlxtend.evaluate.bootstrap_point632_score to evaluate the performance of estimators using the .632 bootstrap. (#283)max_len parameter for the frequent itemset generation via the apriori function to allow for early stopping. (#270)SequentialFeatureSelector or now in sorted order. (#262)SequentialFeatureSelector now runs the continuation of the floating inclusion/exclusion as described in Novovicova & Kittler (1994).
Note that this didn't cause any difference in performance on any of the test scenarios but could lead to better performance in certain edge cases.
(#262)utils.Counter now accepts a name variable to help distinguish between multiple counters, time precision can be set with the 'precision' kwarg and the new attribute end_time holds the time the last iteration completed. (#278 via Mathew Savage)Added evaluate.permutation_test, a permutation test for hypothesis testing (or A/B testing) to test if two samples come from the same distribution. Or
evaluate.permutation_test, a permutation test for hypothesis testing (or A/B testing) to test if two samples come from the same distribution. Or in other words, a procedure to test the null hypothesis that that two groups are not significantly different (e.g., a treatment and a control group). (#250)'leverage' and 'conviction as evaluation metrics to the frequent_patterns.association_rules function. (#246 & #247)loadings_ attribute to PrincipalComponentAnalysis to compute the factor loadings of the features on the principal components. (#251)make_multiplexer_dataset function that creates a dataset generated by a n-bit Boolean multiplexer for evaluating supervised learning algorithms. (#263)BootstrapOutOfBag class, an implementation of the out-of-bag bootstrap to evaluate supervised learning algorithms. (#265)StackingClassifier, StackingCVClassifier, StackingRegressor, StackingCVRegressor, and EnsembleVoteClassifier can now be tuned using scikit-learn's GridSearchCV (#254 via James Bourbeau)'support' column returned by frequent_patterns.association_rules was changed to compute the support of "antecedant union consequent", and new antecedant support' and 'consequent support' column were added to avoid ambiguity. (#245)OnehotTransactions to be cloned via scikit-learn's clone function, which is required by e.g., scikit-learn's FeatureUnion or GridSearchCV (via Iaroslav Shcherbatyi). (#249)self._init_time parameter in _IterativeModel subclasses. (#256)plot_ecdf when run on Python 2.7. (264)PrincipalComponentAnalysis are no being scaled so that the eigenvalues via solver='eigen' and solver='svd' now store eigenvalues that have the same magnitudes. (#251)Added a mlxtend.evaluate.bootstrap that implements the ordinary nonparametric bootstrap to bootstrap a single statistic (for example, the mean. median
mlxtend.evaluate.bootstrap that implements the ordinary nonparametric bootstrap to bootstrap a single statistic (for example, the mean. median, R^2 of a regression fit, and so forth) #232SequentialFeatureSelecor's k_features now accepts a string argument "best" or "parsimonious" for more "automated" feature selection. For instance, if "best" is provided, the feature selector will return the feature subset with the best cross-validation performance. If "parsimonious" is provided as an argument, the smallest feature subset that is within one standard error of the cross-validation performance will be selected. #238SequentialFeatureSelector now uses np.nanmean over normal mean to support scorers that may return np.nan #211 (via mrkaiser)skip_if_stuck parameter was removed from SequentialFeatureSelector in favor of a more efficient implementation comparing the conditional inclusion/exclusion results (in the floating versions) to the performances of previously sampled feature sets that were cached #237ExhaustiveFeatureSelector was modified to consume substantially less memory #195 (via Adam Erickson)SequentialFeatureSelector selected a feature subset larger than then specified via the k_features tuple max-value #213New mlxtend.plotting.ecdf function for plotting empirical cumulative distribution functions (#196).
StackingCVRegressor for stacking regressors with out-of-fold predictions to prevent overfitting (#201via Eike Dehling).plot_decision_regions now supports plotting decision regions for more than 2 training features #189, via James Bourbeau).mlxtend.feature_selection.SequentialFeatureSelector and mlxtend.feature_selection.ExhaustiveFeatureSelector is now performed over different feature subsets instead of the different cross-validation folds to better utilize machines with multiple processors if the number of features is large (#193, via @whalebot-helmsman).DataFrames or Python lists of lists are fed into the StackingCVClassifer as a fit arguments (198).n_folds parameter of the StackingCVClassifier was changed to cv and can now accept any kind of cross validation technique that is available from scikit-learn. For example, StackingCVClassifier(..., cv=StratifiedKFold(n_splits=3)) or StackingCVClassifier(..., cv=GroupKFold(n_splits=3)) (#203, via Konstantinos Paliouras).SequentialFeatureSelector now correctly accepts a None argument for the scoring parameter to infer the default scoring metric from scikit-learn classifiers and regressors (#171).plot_decision_regions function now supports pre-existing axes objects generated via matplotlib's plt.subplots. (#184, see example)math.num_combinations and math.num_permutations numerically stable for large numbers of combinations and permutations (#200).An association_rules function is implemented that allows to generate rules based on a list of frequent itemsets (via Joshua Goerner).
association_rules function is implemented that allows to generate rules based on a list of frequent itemsets (via Joshua Goerner).edgecolor to plots via plotting.plot_decision_regions to make markers more distinguishable from the background in matplotlib>=2.0.association submodule was renamed to frequent_patterns.DataFrame index of apriori results are now unique and ordered.Source code (zip)
Source code (tar.gz)
Adds a black edgecolor to plots via plotting.plot_decision_regions to make markers more distinguishable from the background in matplotlib>=2.0 .
The association submodule was renamed to frequent_patterns .
The DataFrame index of apriori results are now unique and ordered.
Fixed typos in autompg and wine datasets (via James Bourbeau ).
The CHANGELOG for the current development version is available at https://github.com/rasbt/mlxtend/blob/master/docs/sources/CHANGELOG.md.
The CHANGELOG for the current development version is available at https://github.com/rasbt/mlxtend/blob/master/docs/sources/CHANGELOG.md.
EnsembleVoteClassifier has a new refit attribute that prevents refitting classifiers if refit=False to save computational time.lift_score function in evaluate to compute lift score (via Batuhan Bardak).StackingClassifier and StackingRegressor support multivariate targets if the underlying models do (via kernc).StackingClassifier has a new use_features_in_secondary attribute like StackingCVClassifier.SequentialFeatureSelector to 0EnsembleVoteClassifier now raises a NotFittedError if the estimator wasn't fit before calling predict. (via Anton Loss)k_features in SequentialFeatureSelectorSequentialFeautureSelector as sets to prevent the iterator from getting stuck if the k_idx are different permutations of the same combination (via Zac Wellmer).SequentialFeatureSelector if there are similarly-well performing subsets in the floating variants (via Zac Wellmer).The StackingClassifier has a new parameter average_probas that is set to True by default to maintain the current behavior. A deprecation warning was a…
ExhaustiveFeatureSelector estimator in mlxtend.feature_selection for evaluating all feature combinations in a specified rangeStackingClassifier has a new parameter average_probas that is set to True by default to maintain the current behavior. A deprecation warning was added though, and it will default to False in future releases (0.6.0); average_probas=False will result in stacking of the level-1 predicted probabilities rather than averaging these.StackingCVClassifier estimator in 'mlxtend.classifier' for implementing a stacking ensemble that uses cross-validation techniques for training the meta-estimator to avoid overfitting (Reiichiro Nakano)OnehotTransactions encoder class added to the preprocessing submodule for transforming transaction data into a one-hot encoded arraySequentialFeatureSelector estimator in mlxtend.feature_selection now is safely stoppable mid-process by control+c, and deprecated print_progress in favor of a more tunable verbose parameter (Will McGinnis)apriori function in association to extract frequent itemsets from transaction data for association rule miningcheckerboard_plot function in plotting to plot checkerboard tables / heat mapsmcnemar_table and mcnemar functions in evaluate to compute 2x2 contingency tables and McNemar's testmlxtend.plotting for compatibility reasons with continuous integration services and to make the installation of matplotlib optional for users of mlxtend's core functionalityscikit-learn 0.18 using the new model_selection module while maintaining backwards compatibility to scikit-learn 0.17.mlxtend.plotting.plot_decision_regions now draws decision regions correctly if more than 4 class labels are presentAttributeError in plot_decision_regions when the X_higlight argument is a 1D array (chkoar)Nothing published for this version
Added preprocessing.CopyTransformer, a mock class that returns copies of imput arrays via transform and fit_transform
preprocessing.CopyTransformer, a mock class that returns copies of
imput arrays via transform and fit_transformfeature_selection.SequentialFeatureSelector now supports the selection of k_features using a tuple to specify a "min-max" k_features rangePrincipalComponentAnalysisAttributeError with "not fitted" message in SequentialFeatureSelector if transform or get_metric_dict are called prior to fitTfMultiLayerPerceptron's hidden layer(s) if the activations are ReLUs in order to avoid dead neuronsclone_estimator parameter to the SequentialFeatureSelector that defaults to True, avoiding the modification of the original estimator objectsevaluate.plot_decision_regions functionDenseTransformer now doesn't raise and error if the input array is not sparseBaseEstimator as parent class for feature_selection.ColumnSelectorSequentialFeatureSelector's k_features parameter and the scoring metric was more negative than -1 (e.g., as in scikit-learn's MSE scoring function) via wahutchAttributeError issue when verbose > 1 in StackingClassifierclassifier.SoftmaxRegression where the mean values of the offsets were used to update the bias units rather than their sum_layer_mapping functions that caused a swap between the random number generation seed when initializing weights and biasesNew TensorFlow estimator for Linear Regression (`tf_regressor.TfLinearRegression`)
tf_regressor.TfLinearRegression)cluster.Kmeans)tf_cluster.Kmeans)init_weights parameter of the fit methods was globally renamed to init_paramsdropout to the tf_classifier.TfMultiLayerPerceptron classifier for regularizationdecay parameter to the tf_classifier.TfMultiLayerPerceptron classifier for adaptive learning via an exponential decay of the learning rate etaNeuralNetMLP by more streamlined MultiLayerPerceptron (classifier.MultiLayerPerceptron); now also with softmax in the output layer and categorical cross-entropy loss.init_params parameter for fit functions to continue training where the algorithm left off (if supported)Nothing published for this version
The mlxtend.preprocessing.standardize function now optionally returns the parameters, which are estimated from the array, for re-use. A further improv
mlxtend.preprocessing.standardize function now optionally returns the parameters, which are estimated from the array, for re-use. A further improvement makes the standardize function smarter in order to avoid zero-division errorsclassifier.NeuralNetMLPevaluate.scoringevaluate.confusion_matrix) and plot (evaluate.plot_confusion_matrix) confusion matricesevaluate.plot_decision_regions function such as hiding plot axesclassifier.EnsembleClassfier to classifier.EnsembleVoteClassifierPerceptron, Adaline, LinearRegression, and LogisticRegressionlearning parameter of mlxtend.classifier.Adaline to solver and added "normal equation" as closed-form solution solvermlxtend.evaluate.plot_learning_curvesmlxtend.evaluate.plot_decision_regions in 1 dimensional evaluationsloadlocal_mnist to mlxtend.data for streaming MNIST from a local byte files into numpy arraysNeuralNetMLP parameters: random_weights, shuffle_init, shuffle_epochSequentialFeatureSelector class with parameters to enable floating selection and toggle between forward and backward selection.SFS features such as the generation of pandas DataFrame results tables and plotting functions (with confidence intervals, standard deviation, and standard error bars)SFShousing datasetmlxtend.plotting to mlxtend.general_plotting in order to distinguish general plotting function from specialized utility function such as evaluate.plot_decision_regionsSource code (zip)
Source code (tar.gz)
Added a progress bar tracker to classifier.NeuralNetMLP
Added a function to score predicted vs. target class labels evaluate.scoring
Added confusion matrix functions to create ( evaluate.confusion_matrix ) and plot ( evaluate.plot_confusion_matrix ) confusion matrices
New style parameter and improved axis scaling in mlxtend.evaluate.plot_learning_curves
Added loadlocal_mnist to mlxtend.data for streaming MNIST from a local byte files into numpy arrays
New NeuralNetMLP parameters: random_weights , shuffle_init , shuffle_epoch
New SFS features such as the generation of pandas DataFrame results tables and plotting functions (with confidence intervals, standard deviation, and standard error bars)
Added support for regression estimators in SFS
Added Boston housing dataset
New shuffle parameter for classifier.NeuralNetMLP
The mlxtend.preprocessing.standardize function now optionally returns the parameters, which are estimated from the array, for re-use. A further improvement makes the standardize function smarter in order to avoid zero-division errors
Cosmetic improvements to the evaluate.plot_decision_regions function such as hiding plot axes
Renaming of classifier.EnsembleClassfier to classifier.EnsembleVoteClassifier
Improved random weight initialization in Perceptron , Adaline , LinearRegression , and LogisticRegression
Changed learning parameter of mlxtend.classifier.Adaline to solver and added "normal equation" as closed-form solution solver
Hide y-axis labels in mlxtend.evaluate.plot_decision_regions in 1 dimensional evaluations
Sequential Feature Selection algorithms were unified into a single SequentialFeatureSelector class with parameters to enable floating selection and toggle between forward and backward selection.
Stratified sampling of MNIST (now 500x random samples from each of the 10 digit categories)
Renaming mlxtend.plotting to mlxtend.general_plotting in order to distinguish general plotting function from specialized utility function such as evaluate.plot_decision_regions
Sequential Feature Selection algorithms: SFS, SFFS, SBS, and SFBS
Source code (zip)
Source code (tar.gz)
mlxtend.sklearn.EnsembleClassifier -> mlxtend.classifier.EnsembleClassifier
API changes:
mlxtend.sklearn.EnsembleClassifier -> mlxtend.classifier.EnsembleClassifier
mlxtend.sklearn.ColumnSelector -> mlxtend.feature_selection.ColumnSelector
mlxtend.sklearn.DenseTransformer -> mlxtend.preprocessing.DenseTransformer
mlxtend.pandas.standardizing -> mlxtend.preprocessing.standardizing
mlxtend.pandas.minmax_scaling -> mlxtend.preprocessing.minmax_scaling
mlxtend.matplotlib -> mlxtend.plotting
Added momentum learning parameter (alpha coefficient) to mlxtend.classifier.NeuralNetMLP .
Added adaptive learning rate (decrease constant) to mlxtend.classifier.NeuralNetMLP .
mlxtend.pandas.minmax_scaling became mlxtend.preprocessing.minmax_scaling and also supports NumPy arrays now
mlxtend.pandas.standardizing became mlxtend.preprocessing.standardizing and now supports both NumPy arrays and pandas DataFrames; also, now ddof parameters to set the degrees of freedom when calculating the standard deviation
Added multilayer perceptron (feedforward artificial neural network) classifier as mlxtend.classifier.NeuralNetMLP .
Added multilayer perceptron (feedforward artificial neural network) classifier as mlxtend.classifier.NeuralNetMLP .
Added 5000 labeled trainingsamples from the MNIST handwritten digits dataset to mlxtend.data
Added ordinary least square regression using different solvers (gradient and stochastic gradient descent, and the closed form solution (normal equatio
Added ordinary least square regression using different solvers (gradient and stochastic gradient descent, and the closed form solution (normal equation)
Added option for random weight initialization to logistic regression classifier and updated l2 regularization
Added wine dataset to mlxtend.data
Added invert_axes parameter mlxtend.matplotlib.enrichtment_plot to optionally plot the "Count" on the x-axis
New verbose parameter for mlxtend.sklearn.EnsembleClassifier by Alejandro C. Bahnsen
Added mlxtend.pandas.standardizing to standardize columns in a Pandas DataFrame
Added parameters linestyles and markers to mlxtend.matplotlib.enrichment_plot
mlxtend.regression.lin_regplot automatically adds np.newaxis and works w. python lists
Added tokenizers: mlxtend.text.extract_emoticons and mlxtend.text.extract_words_and_emoticons
Added Sequential Backward Selection (mlxtend.sklearn.SBS)
Added Sequential Backward Selection (mlxtend.sklearn.SBS)
Added X_highlight parameter to mlxtend.evaluate.plot_decision_regions for highlighting test data points.
Added mlxtend.regression.lin_regplot to plot the fitted line from linear regression.
Added mlxtend.matplotlib.stacked_barplot to conveniently produce stacked barplots using pandas DataFrame s.
Added mlxtend.matplotlib.enrichment_plot
Added scoring to mlxtend.evaluate.learning_curves (by user pfsq)
Added scoring to mlxtend.evaluate.learning_curves (by user pfsq)
Fixed setup.py bug caused by the missing README.html file
matplotlib.category_scatter for pandas DataFrames and Numpy arrays
Gradient descent and stochastic gradient descent perceptron was changed to Adaline (Adaptive Linear Neuron)
Added Logistic regression
Gradient descent and stochastic gradient descent perceptron was changed to Adaline (Adaptive Linear Neuron)
Perceptron and Adaline for {0, 1} classes
Added mlxtend.preprocessing.shuffle_arrays_unison function to shuffle one or more NumPy arrays.
Added shuffle and random seed parameter to stochastic gradient descent classifier.
Added rstrip parameter to mlxtend.file_io.find_filegroups to allow trimming of base names.
Added ignore_substring parameter to mlxtend.file_io.find_filegroups and find_files .
Replaced .rstrip in mlxtend.file_io.find_filegroups with more robust regex.
Gridsearch support for mlxtend.sklearn.EnsembleClassifier
Improved robustness of EnsembleClassifier.
Improved robustness of EnsembleClassifier.
Extended plot_decision_regions() functionality for plotting 1D decision boundaries.
Function matplotlib.plot_decision_regions was reorganized to evaluate.plot_decision_regions .
evaluate.plot_learning_curves() function added.
Added Rosenblatt, gradient descent, and stochastic gradient descent perceptrons.
Added mlxtend.pandas.minmax_scaling - a function to rescale pandas DataFrame columns.
Added mlxtend.pandas.minmax_scaling - a function to rescale pandas DataFrame columns.
Slight update to the EnsembleClassifier interface (additional voting parameter)
Fixed EnsembleClassifier to return correct class labels if class labels are not integers from 0 to n.
Added new matplotlib function to plot decision regions of classifiers.
Improved mlxtend.text.generalize_duplcheck to remove duplicates and prevent endless looping issue.
Improved mlxtend.text.generalize_duplcheck to remove duplicates and prevent endless looping issue.
Added recursive search parameter to mlxtend.file_io.find_files.
Added check_ext parameter mlxtend.file_io.find_files to search based on file extensions.
Default parameter to ignore invisible files for mlxtend.file_io.find.
Added transform and fit_transform to the EnsembleClassifier .
Added mlxtend.file_io.find_filegroups function.
Implemented scikit-learn EnsembleClassifier (majority voting rule) class.
Improvements to mlxtend.text.generalize_names to handle certain Dutch last name prefixes (van, van der, de, etc.).
Improvements to mlxtend.text.generalize_names to handle certain Dutch last name prefixes (van, van der, de, etc.).
Added mlxtend.text.generalize_name_duplcheck function to apply mlxtend.text.generalize_names function to a pandas DataFrame without creating duplicates.
Added text utilities with name generalization function.
Added text utilities with name generalization function.
Added and file_io utilities.
Added combinations and permutations estimators.
Added DenseTransformer for pipelines and grid search.
mean_centering function is now a Class that creates MeanCenterer objects that can be used to fit data via the fit method, and center data at the colum
Added preprocessing module and mean_centering function.
Added matplotlib utilities and remove_borders function.
Simplified code for ColumnSelector.
From here you can search these documents. Enter your search terms below.
Keys Action
? Open this help
n Next page
p Previous page
s Search
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →