NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #913 most downloaded on PyPI
A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner
Last release 3 years ago
no release in 18 months
Release timing varies
gaps range from 3 weeks to 10 months
Most releases are documented
notes for 44 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
86 releases · first in 2018
Removed deprecated loffset parameter
loffset parameter
Add secondary fallback for series applies
One column per quarter.
Enable indexing after a groupby, e.g. df.swifter.groupby(by)[key].apply(func)
df.swifter.groupby(by)[key].apply(func)rayEnable users to pass in df.index as the by parameter for the df.swifter.groupby(by).apply(func) command
df.index as the by parameter for the df.swifter.groupby(by).apply(func) commandEnable users to df.swifter.groupby.apply, which requires a new package (ray) that now available as an extra_requires.
df.swifter.groupby.apply, which requires a new package (ray) that now available as an extra_requires.pip install -U swifter[groupby]Nothing published for this version
Include log10 performance plot
Include log10 performance plot
Update swifter apply examples
Update swifter apply examples
force_parallel which immediately forces swifter to jump to using dask apply. This enables a simple interface for parallel processing, but disables swifter's algorithm to determine the fastest apply solution possible.Enable users to leverage set_defaults functionality so they don't have to keep invoking individual settings on a per swifter invocation basis
set_defaults functionality so they don't have to keep invoking individual settings on a per swifter invocation basisMerge pull request #179 from jmcarpenter2/jmc/sample-index-robustness…
Merge pull request #179 from jmcarpenter2/jmc/sample-index-robustness…
Resolve installation issue by removing import from setup.py
Resolve installation issues by removing modin dependency, and modin apply route for axis=1 string applies
Nothing published for this version
Nothing published for this version
Nothing published for this version
Sample applies now suppress logging in addition to stdout and stderr
offset and origin for pandas df.resampleFix warnings introduced in 1.0.5 to appropriate warn when using an apply function
Added warnings/errors for swifter methods which do not exist when using modin dataframes
* Broken release
Fixed bug with string, axis=1 applies for pandas dataframes that prevented swifter from leveraging modin for parallelization when returning a series i
* Remove pickle5 hard dependency
Reduce resources consumed by swifter by only importing modin/ray when necessary.
swifter.register_modin() function, which gives access to modin.DataFrame.swifter.apply(...), but is only required if modin is imported after swifter. If you import modin before swifter, this is not necessary.Two major enhancements are included in this release, both involving the use of modin in swifter. Special thanks to Devin Petersohn for the collaborati
Two major enhancements are included in this release, both involving the use of modin in swifter. Special thanks to Devin Petersohn for the collaboration.
df.swifter.apply(...), but still attempts to vectorize the operation which can lead to a performance boost.Example:
import modin.pandas as pd
df = pd.DataFrame(...)
df.swifter.apply(...)
(1) Remove Numba hard dependency, but still handle TypingErrors when numba is installed (2) Only call tqdm's progress_apply on transformations (e.g. R
(1) Remove Numba hard dependency, but still handle TypingErrors when numba is installed
(2) Only call tqdm's progress_apply on transformations (e.g. Resampler, Rolling) when tqdm has an implementation for that object.
Swifter performance consistency improved in two ways:
Swifter performance consistency improved in two ways:
(1) The validation check of the vectorized form of swifter was always failing, because of assumption of dataframe type, when usually a vectorized function form results in array type.
(2) When using a dataframe with duplicated indices, swifter was silently failing to leverage dask. Now swifter raises a warning when the dataframe has duplicated indices, with a recommendation to call df.reset_index(drop=True).
Nothing published for this version
Nothing published for this version
Following pandas release v1.0.0, removing deprecated keyword args "broadcast" and "reduce"
Following pandas release v1.0.0, removing deprecated keyword args "broadcast" and "reduce"
Added new applymap method for pandas dataframes. df.swifter.applymap(...)
Added new applymap method for pandas dataframes. df.swifter.applymap(...)
Fixed issue causing errors when using swifter on empty dataframes. Now swifter will perform a pandas apply on empty dataframes.
Fixed issue causing errors when using swifter on empty dataframes. Now swifter will perform a pandas apply on empty dataframes.
Added support for resample objects in syntax that refects pandas. df.swifter.resample(...).apply(...)
Added support for resample objects in syntax that refects pandas. df.swifter.resample(...).apply(...)
Context Manager to suppress print statements during the sample/test applies. Now if a print is part of the function that is applied, the only print th
Context Manager to suppress print statements during the sample/test applies. Now if a print is part of the function that is applied, the only print that will occur is the final apply's print.
Made several code simplifications, thanks to @ianozvsald's suggestions. One of these code changes avoids the issue where assertions are ignored accord
Made several code simplifications, thanks to @ianozvsald's suggestions. One of these code changes avoids the issue where assertions are ignored according to python -O flag, which would effectively break swifter.
Require tqdm>=4.33.0, which resolves a bug with the progress bar that stems from pandas itself.
Require tqdm>=4.33.0, which resolves a bug with the progress bar that stems from pandas itself.
Fix known security vulnerability in parso <= 0.4.0 by requiring parso > 0.4.0
Fix known security vulnerability in parso <= 0.4.0 by requiring parso > 0.4.0
Change import from tqdm.auto instead of tqdm.autenook. Less warnings will show when importing swifter.
Change import from tqdm.auto instead of tqdm.autenook. Less warnings will show when importing swifter.
df.swifter.progress_bar(desc=" ") now allows for a custom description.
df.swifter.progress_bar(desc="<Your description>") now allows for a custom description.
Nothing published for this version
Very minor bug fixes for edge cases, e.g. KeyError for applying on a dataframe with a dictionary as a nested element
Very minor bug fixes for edge cases, e.g. KeyError for applying on a dataframe with a dictionary as a nested element
Fixed bugs with rolling apply. Added unit test coverage.
Fixed bugs with rolling apply. Added unit test coverage.
Fixed a bug that prevented result_type kwarg from being passed to the dask apply function. Now you can use this functionality and it will rely on dask
Fixed a bug that prevented result_type kwarg from being passed to the dask apply function. Now you can use this functionality and it will rely on dask rather than pandas.
Additionally adjusted the speed estimation for data sets < 25000 rows so that it doesn't spend a lot of time estimating how long to run an apply for on the first 1000 rows when the data set is tiny. We want to asymptote to near-pandas performance even on tiny data sets.
Uses tqdm.autonotebook to dynamically switch between beautiful notebook progress bar and CLI version of the progress bar
Uses tqdm.autonotebook to dynamically switch between beautiful notebook progress bar and CLI version of the progress bar
Minor ipywidgets requirements update
Minor ipywidgets requirements update
Allowed user to override scheduler default to multithreading if desired.
Allowed user to override scheduler default to multithreading if desired.
Add an option allow_dask_on_strings to DataFrameAccessor. This is a non-recommended option if you are doing string processing. It is intended for usin
Add an option allow_dask_on_strings to DataFrameAccessor. This is a non-recommended option if you are doing string processing. It is intended for using the string as a lookup for the rest of the dataframe processing. This override is also included in SeriesAccessor, but there I am not aware of a use-case that it makes sense to use this.
Nothing published for this version
Swifter now defaults to axis=0, with a NotImplementedError for when trying to use dask on large datasets, because dask hasn't implemented axis=0 appli
Swifter now defaults to axis=0, with a NotImplementedError for when trying to use dask on large datasets, because dask hasn't implemented axis=0 applies yet.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Added documentation and code styling thanks to @msampathkumar. Also included override options for dask_threshold and npartitions parameters.
Added documentation and code styling thanks to @msampathkumar. Also included override options for dask_threshold and npartitions parameters.
Nothing published for this version
Added support for rolling objects in syntax that reflects pandas. df.swifter.rolling(..).apply(...)
Added support for rolling objects in syntax that reflects pandas. df.swifter.rolling(..).apply(...)
Nothing published for this version
Nothing published for this version
Fixed a bug that would call a vectorized function when in fact the vectorization was wrong to apply. We have to ensure that output data shape is align
Fixed a bug that would call a vectorized function when in fact the vectorization was wrong to apply. We have to ensure that output data shape is aligned regardless of apply type.
Nothing published for this version
Nothing published for this version
Added TQDM support (to disable do df.swifter.progress_bar(False).apply(...), removed groupby_apply (because it's too slow), and tweaked some under-the
Added TQDM support (to disable do df.swifter.progress_bar(False).apply(...), removed groupby_apply (because it's too slow), and tweaked some under-the-hood _dask_apply functionality. Specific functionality changes for pd.Series.swifter.apply include falling back to dask apply if dask map_partitions fails. Specific functionality changes for pd.DataFrame.swifter.apply include falling back to pandas apply if dask apply fails.
Made a change so that swifter uses pandas apply when input is series/dataframe of dtype string. This is a temporary solution to slow dask apply proces
Made a change so that swifter uses pandas apply when input is series/dataframe of dtype string. This is a temporary solution to slow dask apply processing of strings.
Your coding agent can read these notes before it upgrades. Set up the MCP server →