NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3157 most downloaded on PyPI
A package for encoding categorical variables for machine learning
Last release 26 days ago
08 Sep 2026
Release timing varies
gaps range from 2 weeks to 11 months
Nearly every release is documented
notes for 37 of 40 stable releases
Nothing withdrawn
no release was ever pulled
11 years old
40 releases · first in 2016
Fix/526 row order upstream by @wdm0006 in #527
NestedCVWrapper.fit_transform now returns out-of-fold rows in the
input's original order (issue#526). Previously the per-fold encodings were
concatenated in cross-validation iteration order, so with a shuffling
splitter (including the wrapper's default StratifiedKFold) positional
consumers (NumPy arrays, .iloc, scikit-learn estimators) received rows
misaligned with y even though the index labels were correct.One column per quarter.
Remove deprecated copy keyword from astype calls in benchmarking loaders by @xyf5432 in #501
Full Changelog: 2.10.0...v2.11.0
OneHotEncoder's handle_missing='value' docstring wrongly
claimed missing values are zero-filled; corrected it to describe the actual (and, per
test_missing_values, intended library-wide) behavior of treating a missing value seen
during fit as its own category, same as MultiHotEncoder's existing value/ignore
split. No behavior change: the existing ignore option already zero-fills missing values.handle_missing='ignore' to RankHotEncoder, matching the option already
available on OneHotEncoder and MultiHotEncoder (issue#400's RankHotEncoder gap).
Zero-fills a missing value at both fit and transform time, without adding an extra
thermometer column for it; the default value behavior (own category, unchanged) still
matches test_missing_values' library-wide invariant.CountTargetEncoder, a supervised count-based
encoder that stores per-class category counts and encodes categories as
smoothing-adjusted log-odds against the global target prior (binary targets
emit one WOE-style column, multiclass targets one column per class).
Classification-only for now; a binning-based regression variant is planned
as a follow-up.MultiHotEncoder, an unsupervised encoder for delimiter-separated
multi-value cells (e.g. 'mathematics|physics'): every distinct item seen at
fit gets one binary column, with configurable delimiter and the usual
handle_unknown/handle_missing support (issue#161).composite_cols on supervised encoders (Target, MEstimate, WOE,
JamesStein, CatBoost, GLMM, LeaveOneOut, Quantile): joint ("composite") target
encoding of column groups passed as tuples, learning statistics over
combinations rather than per column (issue#429). A new keep_components
flag controls whether component columns are kept alongside the output.min_group_size/min_group_name rare-category lumping
from CountEncoder to every encoder derived from BaseEncoder; small
groups are merged before the encoder's own statistics are computed (issue#279).handle_unknown/handle_missing now accept callables and integers
across the supervised and ordinal encoders: a callable receives the
unknown/missing value and returns the encoding to use, an integer assigns a
dedicated code (issues #283, #344). The ad-hoc -1/-2 codes are now
shared sentinel constants and mapping finalization is centralized in one helper.OneHotEncoder.transform's dummy expansion now assembles the output
with a single allocation instead of a per-column accumulator: measured peak
memory on a 100k×10×20 frame dropped 626 → 337 MB (−46%) (issue#362). Repro
script: benchmarks/repro-362.py.transform
(e.g. the numpy array emitted by the previous step of a scikit-learn pipeline):
the fitted column names are re-attached positionally and object/category dtypes
are restored, so the output matches transforming the equivalent DataFrame
(GH #406). OrdinalEncoder.inverse_transform re-attaches the fitted output
names for arraylike input, matching the BaseN/OneHot precedent.cols now explains in the
error that the names cannot be recovered positionally and how to proceed.transform now accepts DataFrames carrying extra pass-through columns
beyond the encoded ones (GH #355 scenario, GH #367). Arraylike inputs keep the
strict dimension check, and missing encoded columns raise a clear error naming
them instead of a KeyError.GrayEncoder.inverse_transform now recovers the original categories.
Previously it inherited the positional base-N decoder from BaseNEncoder,
which cannot invert a Gray code (consecutive code words differ by a single
bit and are not positional), so it returned the wrong categories.CatBoostEncoder and RankHotEncoder
are now handled consistently instead of producing degenerate encodings.QuantileEncoder collapses singleton (unique-value) levels to the prior
instead of emitting degenerate quantile statistics (issue#327).fit now raise an
informative NotFittedError naming the encoder instead of an opaque
AttributeError (issue#232 follow-up).patsy. The four contrast-coding
encoders (Polynomial, Helmert, BackwardDifference, Sum) now build their
contrast matrices in-tree.statsmodels is now an optional extra. GLMMEncoder is
loaded lazily and raises a clear ImportError if statsmodels is not
installed. Install with pip install category_encoders[glmm] to use
it. Removes both statsmodels and patsy from the default install.Fix CountEncoder with normalize=True incorrectly dropping columns by @wdm0006 in #469
Full Changelog: 2.9.0...2.10.0
.zenodo.json and CITATION.cff files, and pointed the
README DOI badge at the Zenodo concept DOI (always resolves to the latest
version).index_start parameter on OrdinalEncoder for
zero-indexed labels.cols="all" encodes every column regardless of
dtype (useful inside sklearn ColumnTransformer).sigma parameter on TargetEncoder/LOO/CatBoost/WOE/GLMM/
JamesStein/MEstimate) is multiplicative (X * N(1, sigma)), not
additive.n_components
docstring (it is the number of output hash buckets, not bits).normalize=True and
drop_invariant=True no longer drops all output columns.Series[0] access that raised KeyError on recent pandas; use
.iloc[0] instead. Also fixed the bundled examples script
(removed the LogisticRegression(multi_class=...) argument removed in
scikit-learn 1.9).NotFittedError during
fit/fit_transform when scikit-learn's transform_output='pandas'
is configured globally or via a ColumnTransformer.Release version 2.9.0 with bug fixes and version upgrades
Release version 2.9.0 with bug fixes and version upgrades
Release with changes according to changelog
Release with changes according to changelog
pd.Categorical targets.Release 2.8.0 with changes according to changelog
Release 2.8.0 with changes according to changelog
Release 2.7.0 with changes according to changelog
Release 2.7.0 with changes according to changelog
feature_names_in_ and feature_names_out_ to np.ndarray instead of lists.Release 2.6.4 with some fixes and improvements according to changelog
Release 2.6.4 with some fixes and improvements according to changelog
process_creation_method parameterrelease 2.6.3 with some bugfixes according to changelog
release 2.6.3 with some bugfixes according to changelog
Release 2.6.2 containing minor bug fixes according to changelog
Release 2.6.2 containing minor bug fixes according to changelog
importlib instead of pkg_resourcesrelease 2.6.1 with some bugfixes according to changelog
release 2.6.1 with some bugfixes according to changelog
feature_names_in_ attributeget_feature_names_out function has the correct signaturecategory_mapping property (issue #256)Release 2.6.0 with new features and fixes to be found in changelog
Release 2.6.0 with new features and fixes to be found in changelog
feature_names_out_* fix pypi sdist
changes according to changelog
changes according to changelog
Release 2.5.0 with changes according to changelog
Release 2.5.0 with changes according to changelog
release 2.4.1 with minor fixes
release 2.4.1 with minor fixes
Releasing changes described in changelog.md
Releasing changes described in changelog.md
Largely a bugfix release after a period of no maintenance. May still be issues with GLMEncoder.
Largely a bugfix release after a period of no maintenance. May still be issues with GLMEncoder.
Added generalized linear mixed model encoder
Added experimental support for multithreading in hashing encoder
Added James-Stein, CatBoost and m-estimate encoders
Added Weight of Evidence encoder
Critical bugfix in hashing encoder
Bugfixes related to missing value imputation
A copy of the v1.2.6 release for zenodo
A copy of the v1.2.6 release for zenodo
Onehot transform returns same columns always
Added more sophisticated missing value or unknown category handling to ordinal
Full support for numpy arrays as input, not just dataframes.
All encoders handle missing values and are tested for their handling
Better handling for missing values in hashing encoder
Hash type in hashing encoder now defaults to md5 using hashlib, but can be set to any valid hashlib hash
Added optional parameter to return a numpy array rather than a dataframe from all transformers.
Immediately return if cols is empty.
Optionally pass drop_invariant to any encoder to consistently drop columns with 0 variance from the output (based on training set data in fit())
Changed setup.py to not explicitly force reinstalls of other packages
* Bugfixes
Nothing published for this version
Nothing published for this version
Nothing published for this version
First real usable release, includes sklearn compatible encoders.
Your coding agent can read these notes before it upgrades. Set up the MCP server →