NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3798 most downloaded on PyPI
Modin: Make your pandas code run faster by changing one line of code.
Last release 1 years ago
02 Oct 2025
Ships unpredictably
gaps range from 2 weeks to 8 months
Most releases are documented
notes for 51 of the last 60 stable releases
1 version withdrawn
withdrawn after publishing
8 years old
110 releases · first in 2018
This release includes a bug fix and a test fix.
This release includes a bug fix and a test fix.
@sfc-gh-jkew @sfc-gh-mvashishtha
This release includes bugfixes for Series.json, DataFrame.rename, and eval, and performance improvements for joins with AutoSwitchBackend enabled.
This release includes bugfixes for Series.json, DataFrame.rename, and eval, and performance improvements for joins with AutoSwitchBackend enabled.
Series.to_json (#7673)axis=None case for DataFrame.rename (#7674)@sfc-gh-dpetersohn @sfc-gh-joshi @sfc-gh-mvashishtha
One column per quarter.
This release includes a bug fix, a performance improvement for query() and eval(), and changes to the testing suite.
This release includes a bug fix, a performance improvement for query() and eval(), and changes to the testing suite.
AutoSwitchBackend for DataFrame.T/`Series.T~ (#7654)@sfc-gh-dpetersohn @sfc-gh-joshi @sfc-gh-mvashishtha
This release includes various bug fixes and improvements, and adds support for pandas 2.3.
This release includes various bug fixes and improvements, and adds support for pandas 2.3.
@sfc-gh-joshi @sfc-gh-mvashishtha @sfc-gh-vrpatel
This release includes a bug fix.
This release includes a bug fix.
This release includes various bug fixes and improvements.
This release includes various bug fixes and improvements.
value_counts() on Pandas backend (#7585)@sfc-gh-vrpatel @sfc-gh-joshi @sfc-gh-mvashishtha @sfc-gh-jkew
This patch release includes some bug fixes.
This patch release includes some bug fixes.
value_counts() on Pandas backend (#7585)@sfc-gh-vrpatel @sfc-gh-mvashishtha
This patch releases fixes a regression introduced in Modin 0.33.0.
This patch releases fixes a regression introduced in Modin 0.33.0.
@sfc-gh-mvashishtha
This release introduces a set of features for switching Modin execution between multiple backends (e.g. Ray and local Pandas) manually or automaticall
This release introduces a set of features for switching Modin execution between multiple backends (e.g. Ray and local Pandas) manually or automatically. It also includes several bug fixes.
.astype(pandas.CategoricalDtype(…)) (#7487)@CRiddler @YarShev @anmyachev @data-makerman @devin-petersohn @emmanuel-ferdman @mpeleshenko @noloerino @sfc-gh-dpetersohn @sfc-gh-jkew @sfc-gh-joshi @sfc-gh-mvashishtha
This release introduces support for Polars API, a new query compiler for small data, more functions that can use dynamic partitioning, as well as seve
This release introduces support for Polars API, a new query compiler for small data, more functions that can use dynamic partitioning, as well as several bug fixes.
BasePandasDataset (#7353)df.update (#7330)NoAttributeError on DataFrame.copy (#7358)motoserver/moto service, pin to 5.0.13 (#7374)__imul__ performing addition instead of multiplication (#7380)broadcast_apply (#7338)@MortalHappiness @Retribution98 @YarShev @ZhipengXue97 @anmyachev @arunjose696 @devin-petersohn @likawind @sfc-gh-joshi @sfc-gh-mvashishtha
REFACTOR-#7273: Remove deprecated functions from utils.py, accessor.py and io.py
First release compatible with NumPy 2.0.
enable_logging decorator preserve type hints (#7279)versioneer install --vendor (#7311)C engine instead of pyarrow for getting metadata in read_csv (#7322)synchronize_labels for combine function (#7300)instance_type attribute of axis partitions (#7268)_modin_frame methods from _query_compiler (#7297)reload_modin feature (#7280)MinRowPartitionSize and MinColumnPartitionSize (#7284)@Jayson729 @Retribution98 @YarShev @anmyachev @arunjose696 @kurtmckee @sfc-gh-dpetersohn @vsreekanti
Key Features and Updates Since 0.30.0
This release pins numpy<2.
@anmyachev
FIX-#6967: Remove read_pickle_distributed/to_pickle_distributed functions as deprecated
This release introduces support for DataFrame API standard, a distributed implementation for right merge/join, more efficient implementation of internal operators, which gives a performance boost to almost all distributed Modin functions, improved compatibility with pandas on pyarrow backend, type hints for pandas API to improve UX.
read_pickle_distributed/to_pickle_distributed functions as deprecated (#7258)idxmax and idxmin can work with string columns (#7193)enable_api_only mode in modin logging (#7194)df.melt handle duplicate value_vars correctly (#7208)dataframe-api-compat>=0.2.7 (#7220)use_legacy_dataset=False for ParquetDataset (#7222)modin.pandas.api.extensions overwrites re-export of pandas.api submodules (#7225)default_to_pandas error messages (#7269)cached_property and use it (#7239)doc_checker.py works with functools.cached_property (#7241)pyarrow>=10.0.1 as pandas 2.2.* does (#7247)_validate_dtypes_sum_prod_mean works correctly with datetime types (#7237)modin_frame.combine() for merge and join only when necessary (#7228)merge (#7229)modin/core/dataframe/algebra/ (#7243)extract_dtype internal function in more places (#7261)MODIN_CPUS number (#7198)from_* functions (#7256)Map operator (#7136)TreeReduce and GroupByReduce operators (#7245)from_map feature to create dataframe (#7215)Fold operator more flexible (#7257)__arrow_array__ for Series (#7200)ray-core instead of ray-default (#6955)master branch to main (#7188)base.py (#7253)merge/join (#7226)@Retribution98 @YarShev @anmyachev @arunjose696 @noloerino @sfc-gh-jkew
Key Features and Updates Since 0.29.0
This release pins numpy<2.
@anmyachev @sfc-gh-dpetersohn
REFACTOR-#7105: Deprecate cfg.RangePartitioningGroupby
This release introduces modin.pandas.testing and modin.pandas.arrays modules, faster implementation (range-partitioning) for
pivot_table, unique, drop_duplicates, nunique, df.resample functions, new functions to interact with Dask: to/from_dask,
distributed implementation for Series.case_when, optimization for astype function with scalar dtype.
Series.unique() with pyarrow dtype returns ArrowExtensionArray (#7042)pandas_dtype instead of np.dtype for some more places in Modin code (#6794)astype query compiler (#7152)astype function (#7052)shift function (#7055)iloc/loc functions (#7057)insert function (#7059)pivot when index or columns are of Index type (#7061)aggregate function (#7063)MaterializationHook with the materialized object on serialization (#7075)rank raises No axis named None... exception (#7089)Series.corr with method != pearson (#7158)quantile function works with numeric_only=True (#7160)MinPartitionSize configuration variable in remote context (#7177)shape_hint="column" for some more operations with Series (#7069)shape_hint for dropna (#7124)apply_full_axis with keep_partitioning=True (#7131)apply_full_axis with keep_partitioning=False (#7133)gen_data internal function (#7046)cfg.RangePartitioningGroupby (#7161)from/to_ray_dataset to from/to_ray (#7107)aws_example.yaml file (#7110)eval_general doesn't expect exceptions by default (#6954)test_groupby.py (#7065)test_io.py (#7067)test_default.py (#7074)test_map_metadata.py (#7077)test_series.py (#7083)test_indexing.py (#7085)test_reduce.py (#7087)raising_exceptions argument of eval_general testing function (#7095)--signoff option (#7145)AxisPartition and BlockPartition classes (#7079)modin.pandas.testing module (#7045)Series.case_when in a distributed way (#6972)_deploy_ray_func remote function. (#7005)to/from_dask functions (#7022).pivot_table() (#7048)modin.pandas.arrays module (#7071)modin_layer names to classes that inherit ClassLogger (#7099).unique() and .drop_duplicates() (#7091)nunique() (#7101)enable_api_only mode in modin logging (#7114)@remote_function decorator with cache (#7112)df.resample() (#7140)BaseQueryCompiler, BasePandasDataset, DataFrame or Series type hints at a high level (#7147)Series (#7154)DataFrame (#7179)modin.pandas.[functions] (#7181)@AndreyPavlenko @Retribution98 @YarShev @anmyachev @arunjose696 @dchigarev @sfc-gh-mvashishtha
Key Features and Updates Since 0.28.2
This release pins numpy<2.
@anmyachev @sfc-gh-dpetersohn
This release reverts the pandas requirement from 2.2.1 to >=2.2,<2.3
This release reverts the pandas requirement from 2.2.1 to >=2.2,<2.3
@sfc-gh-mvashishtha
Nothing published for this version
TEST-#6932: Don't use deprecated pandas._testing.makeStringIndex
This release introduces modin.pandas.api.extensions module, faster implementations for merge and
groupby.rolling(by default) functions, and new functions to work with Ray Dataset: to/from_ray_dataset.
It also includes some other new features, performance optimizations and bug fixes.
merge when right operand is an empty dataframe (#6941)read_parquet when dataset is created with to_parquet and index=False (#6937)isort formatting for scripts from tutorials (#6945)needs: [lint-black-isort, ...] (#6947)groupby when Modin dataframe has several column partitions (#6951)render_as_string to get sqlalchemy engine url (#6953)test_all_urls_exist (#6975)._propagate_index_objs() (#6977)._copartition() for identical indices on binary operations (#6980)read_pickle_distributed/to_pickle_distributed to read_pickle_glob/to_pickle_glob (#6957)modin.pandas.DataFrame._to_pandas a public method (#6940)DataFrame.to_pickle_distributed in favour of DataFrame.modin.to_pickle_distributed (#6959)eval_general utility (#7003)check_exception_type argument of eval_general function (#7009)to_pandas and to_ray_dataset into modin namespace (#7014)to_hdf and hist signatures to pandas (#7018)pandas._testing.makeStringIndex (#6933)test_series.py (#6995)test_io.py (#6997)log_level in logging module (#6992)read_sql by getting connection url (#6956)include_groups=False parameter in groupby.apply() (#6938)groupby().rolling() by default (#6943).merge() using range-partitioning implementation (#6966)to/from_ray_dataset functions (#6971)MetaList.__getitem__() (#7006)@AndreyPavlenko @Retribution98 @YarShev @anmyachev @arunjose696 @dchigarev @sfc-gh-dpetersohn @tochigiv
Key Features and Updates Since 0.27.0
This release pins numpy<2.
@anmyachev @dchigarev @sfc-gh-dpetersohn
This release updates pandas to 2.2, introduces lazy execution mode on Ray, new functions that support glob syntax and speeds up several more groupby c
This release updates pandas to 2.2, introduces lazy execution mode on Ray, new functions that support glob syntax and speeds up several more groupby cases. It also includes some other new features, performance optimizations and many bug fixes.
tolist function in DtypesDescriptor._merge_dtypes (#6844)read_parquet works with integer columns for pyarrow engine (#6874)query_compiler.merge (#6880)astype works correctly with int32 and float32 dtypes (#6884).from_pandas() (#6912)pydantic dependency (#6917)JoinNode instead of MaskNode for non-range row_position (#6926)iloc where beneficial (#6878)DaskThreadsPerWorker to 1 (#6923)missmatch to mismatch in ErrorMessage.missmatch_with_pandas method (#6901)PyarrowOnRay execution in favour of pyarrow-backed pandas dataframes (#6848)SocksProxy, DoLogRpyc, DoTraceRpyc outdated classes (#6834)OrderedDict in favor of builtin dict (#6853)_get_dimensions and change arguments (#6859)__all__ in modin.config.__init__.py (#6886)create_test_series (#6910)tmp_path fixture (#6709)to_csv tests on Unidist more stable (for test-all-unidist CI job) (#6851)to_csv tests (#6847)gs remote protocol since we rely on fsspec (#6882)black>=24.1.0 (#6887)pytest 8.0.0 (#6894)read_json_glob and to_json_glob (#6873)read_parquet_glob and to_parquet_glob (#6854)read_xml_glob, to_xml_glob (#6930).length() and .width() being called in a loop (#6842)to_pandas call in merge and join functions (#6850)2.2.* (#6907)@AndreyPavlenko @YarShev @anmyachev @arunjose696 @dchigarev @leshikus @vedant
This release includes a fix for concat function.
This release includes a fix for concat function.
tolist function in DtypesDescriptor._merge_dtypes (#6844)to_csv tests on Unidist more stable (for test-all-unidist CI job) (#6851)to_csv tests (#6847)@leshikus @anmyachev
This release introduces a new, faster implementation for groupby.apply, as well as many performance fixes related to improving asynchronous execution,
This release introduces a new, faster implementation for groupby.apply, as well as many performance fixes related to improving asynchronous execution, a new namespace for accessing experimental functions (for example, DataFrame.modin.to_pickle_distributed), a fix for a long-standing problem with the use of Modin objects inside UDFs for apply and many other fixes.
Note: to get Modin on MPI through unidist (as of unidist 0.5.0) fully working by installing with pip it is required to have a working MPI implementation installed beforehand.
apply (#6673)@lazy_metadata_decorator for PandasDataFrame.finalize (#6720)astype op (#6692)set_index_name(None) (#6698)unidist <= 0.4.1 (#6746).insert() (#6757)to_numpy use **kwargs after #6704 (#6769)ValueError: assignment destination is read-only for cumsum (#6772)_to_pandas return mutable pandas objects (#6775)loc to get similar behavior to pandas (#6798)Series.__getitem__ (#6780)pandas.api.types.pandas_dtype to convert to valid numpy and pandas only dtypes (#6788)DataFrame.join (#6787)ModinIndex objects (#6800)NotImplementedError to a user on a set_columns() with dupl labels (#6823)ModinIndex._lengths_id on empty partitions filtering (#6825)copy=True parameter for concat calls inside to_pandas (#4778)broadcast_apply_full_axis (#6760)reset_index for left merge (#6665)copy=False for internal usage of set_axis (#6667)copy() call for Series.reset_index (#6670)Series.tolist function (#6672)sync_labels=False for rank function (#6689)lazy_map_partitions() for dtypes conversion (#6695)get_axis internal function instead of axes property (#6700)to_numpy (#6699)_groupby_shuffle internal function (#6707)_shape_hint in query_complier.copy function (#6713)qc._shape_hint = column in columnarize function (#6715)_filter_empties (#6717)_get_axis_lengths function instead of _axes_lengths property (#6719)keep_partitioning=True, for duplicated implementation (#6722)_shape_hint = "column" in DataFrame.squeeze (#6724)result.name = None in groupby code (#6726)reset_index() (#6751).__setitem__() (#6758).concat(axis=0) (#6759)execution_wrapper instead of directly addressing DaskWrapper (#6740)modin.pandas.io module (#6806)modin.experimental folder (#6813)--extra-test-parameters option (#6730)to_csv tests on Unidist more stable (#6776)int type (#6796)DataFrame.__rdivmod__/__divmod__ (#6785)modin.pandas.error module (#6802)groupby.apply() by default (#6804)@AndreyPavlenko @JignyasAnand @RehanSD @YarShev @anmyachev @devin-petersohn @dchigarev @mvashishtha @seydar
Key Features and Updates Since 0.25.0
Hotfix for Unidist.
unidist<=0.4.1 (f0c1c03)@anmyachev
REFACTOR-#6622: Don't use deprecated random_integers func
This release introduces modin.utils.execute function to improve benchmarking experience, includes new version of HDK 0.9.
It also includes performance optimizations for sort_values, value_counts, 2D setitem and several others, as well as many bug fixes.
ray.get() inside of the kernel executing call queues (#6633)FutureWarnings in rolling unless necessary (#6586)Series.groupby.agg (#6613)join to avoid distributing a dict object warning (#6612)DataFrame.agg() (#6606).sort_values() (#6608)FutureWarnings for first/last/bool (#6625)groupby.diff() for dates (#6631)groupby.apply in case of experimental groupby (#6649)read_csv(): treat object dtype as string (#6636)skiprows parameter usage for read_excel (#6638)modin.numpy.array.sum on HDK (#6643)modin/experimental/sql/hdk/query.py part of modin package (#6646)Series.between works correctly (#6656)navigation_with_keys=True to fix docs build (#6681)from_pandas() for numerical data in Ray (#6640)sort_values by reducing the number of partitions (#6589)MODIN_CPUS instead of os.cpu_count() for the fragment size calculation (#6615)LazyProxyCategoricalDtype materialization on merge (#6630)dot operation (#6644)value_counts(): Eliminate redundant sorting. (#6654)random_integers func (#6623)pytest to print warnings in tests output (#6621)execute to trigger lazy computations and wait for them to complete (#6648)materialize parameter for partition.ip func (#6650)pyhdk version to 0.9 (#6676)@AndreyPavlenko @Egor-Krivov @Garra1980 @YarShev @anmyachev @dchigarev
Nothing published for this version
Key Features and Updates Since 0.24.1
Hotfix for Unidist.
Note: broken pip wheel, use https://github.com/modin-project/modin/releases/tag/0.24.1.post1 instead
@anmyachev @dchigarev
Key Features and Updates Since 0.24.0
Hotfix for sort_values.
DataFrame.agg() (#6606).sort_values() (#6608)@AndreyPavlenko @dchigarev
REFACTOR-#6576: Don't use deprecated is_int64_dtype and is_period_dtype function
This release upgrades the pandas version to 2.1, updates the minimum supported python version up to 3.9, introduces ModinDataLoader to improve interaction with PyTorch, fixes several issues with interchange protocol that solved known compatibility issues with Plotly, Seaborn and Altair, includes new version of HDK 0.8. It also includes some other new features, and many bug fixes.
ray>2.6.0 (#6425)read_csv (#5507)read_excel: defaults to pandas for unsupported types of 'io' (#6462)query and eval (#6488)Column.null_count to return a built-in int instead of NumPy scalar (#6526)unwrap_partitions for virtual partitions when axis=None (#6560)__getattribute__ for experimental mode (#6529)temp_df.dtype == 'category' (#6360)Series.str.find/index/rfind/rindex (#6426)copy on empty DataFrame/Series objects (#6371)__array__ method always returns array of vanilla numpy (#6300)BenchmarkMode.put(True) (#6365)groupby.size() in reshuffling groupby (#6370)sum operation (#6421)astype calls for modin.array.sum op (#6395)DataFrame.mean() result (#6520)__setitem__ op when using not hashable key (#6547)__factory to None in case of any problems during initialization (#6397)diff (#6403)disable_logging to __getattr__ (#6406)read_feather with pyarrow<11.0 (#6415)flake8==6.1.0 (#6428)pymssql==2.2.8 from environments (#6430)~ in paths in IO functions correctly (#6448)sum|mean|median groupby aggregations (#6444)fastparquet>=2023.1.0 (#6458)groupby.apply() for UDFs that change the output's shape (#6506)is_bool_dtype() for categorical (#6480)__array_ufunc__ (#6486)test_sort_cols_str from test_dataframe.py crashed on HDK 0.7.0 and python 3.9 (#6515)botocore as an optional dependency (#6521)read_excel so that it doesn't use rich_text param for old openpyxl (#6534)s3fs<2023.9.0 (#6536)s3fs<2023.9.0 (#6544)read_parquet (#6545)ValueError: buffer source array is read-only for iloc (#6538)dfsql module (#6550)FutureWarnings in groupby unless necessary (#6595)read_csv with iterator=True (#6554).read_parquet() (#6559)MODIN_OMNISCI_* env vars in favor of MODIN_HDK_* (#6562)map function via applymap (#6566)FutureWarnings in bfill/backfill/ffill/pad unless necessary (#6599)sort_values shouldn't affect source dataframe/series (#6603)concat operation (#6381)_repartition (#6376)numpy.array operations in internals of iloc/loc operation (#6393)__getitem__ when the number of rows to be taken > 90% (#6423).dropna() using map-reduce pattern (#6472)reindex (#6438)qc.to_datetime() (#6525)query() (#6584).from_pandas() (#6591)BasePandasDataset.apply (#6451)isort (#6551)Patcher internal class (#6471)__invert__ (#6490)contextlib.nullcontext instead of custom one (#6570)is_int64_dtype and is_period_dtype function (#6577)time_groupby_agg_nunique ASV bench (#6564)psycopg2-binary for testing and developing purpose (#6573)df.eval with scalar and groupby.transofm call in the expr (#6546)repr to force materialization (#6461)numexpr<2.8.5 (#6474)boto3 from environments to speedup creation (#6496)read_parquet supported parameters (#6420)dataframe.insert function (#6400)DataLoader interplay. (#6140).modin folder (#6390)to_parquet (#6404)read_parquet (#6442)nlargest/nsmallest groupby aggregation (#6485)datetime64 to int64 cast (#6501)enable_multifrag_execution_result=1 HDK launch parameter (#6503)@AndreyPavlenko @RehanSD @YarShev @anmyachev @dchigarev @mvashishtha @vnlitvinov @abykovsk @zmbc @noloerino @rentruewang
The main purpose of this release is to port as many fixes as possible to the latest version, which supports Python 3.8.
The main purpose of this release is to port as many fixes as possible to the latest version, which supports Python 3.8.
unidist<=0.4.1read_excel: defaults to pandas for unsupported types of io (#6462)ray.get() inside of the kernel executing call queues (#6633)Column.null_count to return a built-in int instead of NumPy scalar (#6526)unwrap_partitions for virtual partitions when axis=None (#6560)__getattribute__ for experimental mode (#6529)groupby.apply() for UDFs that change the output's shape (#6506)is_bool_dtype() for categorical (#6480)reshuffling in case of a string key (#6510)test_sort_cols_str from test_dataframe.py crashed on HDK 0.7.0 and python 3.9 (#6515)test_dataframe.py is crashed if Calcite is disabled (#6517)botocore as an optional dependency (#6521)read_excel so that it doesn't use rich_text param for old openpyxl (#6534)s3fs<2023.9.0 (#6536)s3fs<2023.9.0 (#6544)ValueError: buffer source array is read-only for iloc (#6538)read_csv with iterator=True (#6554)apply (#6673)Series.groupby.agg (#6613)sort_values shouldn't affect source dataframe/series (#6603)join to avoid distributing a dict object warning (#6612).sort_values() (#6608)groupby.apply in case of experimental groupby (#6649)read_csv: treat object dtype as string (#6636)skiprows parameter usage for read_excel (#6638)modin.numpy.array.sum on HDK (#6643)modin/experimental/sql/hdk/query.py part of modin package (#6646)Series.between works correctly (#6656)navigation_with_keys=True to fix docs build (#6681)@AndreyPavlenko @Egor-Krivov @Garra1980 @RehanSD @anmyachev @dchigarev @vnlitvinov
This release contains fixes that improve Modin's performance for both the NumPy and pandas APIs, as well as removes the Modin In the Cloud experimenta
This release contains fixes that improve Modin's performance for both the NumPy and pandas APIs, as well as removes the Modin In the Cloud experimental feature. This release also includes upgrades to Modin's testing suite that significantly speed up CI.
Series.str.find/index/rfind/rindex (#6426)diff (#6403)disable_logging to __getattr__ (#6406)@AndreyPavlenko @RehanSD @YarShev @anmyachev @dchigarev @mvashishtha @vnlitvinov
REFACTOR-#6329: deprecate cloud feature
Modin 0.23.0
This release upgrades the pandas version to 2.0. It also includes '.corr' speed-up, new features, and bug fixes.
con parameter for to_sql (#5940)read_json in case of rows having different columns (#5946)read_excel and unpin openpyxl (#6247)Series.equals/DataFrame.equals with NA entries (#6270)wait method for Dask/Ray/Unidist wrappers (#6049)groupby.rolling API (#6292)@AndreyPavlenko @YarShev @alexbaden @anmyachev @dchigarev @kurapov-peter @mvashishtha @vnlitvinov
This release includes support for pandas 2.0, '.corr' speed-up, new features and bug fixes.
This release includes support for pandas 2.0, '.corr' speed-up, new features and bug fixes.
Note: this is a release candidate. If everything goes well, we'll release Modin 0.23.0 in two weeks.
read_json in case of rows having different columns (#5946)read_excel and unpin openpyxl (#6247)wait method for Dask/Ray/Unidist wrappers (#6049)@AndreyPavlenko @YarShev @anmyachev @dchigarev @mvashishtha @vnlitvinov
Patch release with main point of pinning pydantic<2 to resolve Ray issues, plus a few bugfixes.
Patch release with main point of pinning pydantic<2 to resolve Ray issues, plus a few bugfixes.
@AndreyPavlenko @anmyachev
This release includes several bug fixes.
This release includes several bug fixes.
to_dict (https://github.com/modin-project/modin/pull/6260)astype("category") causing read-only buffer error (https://github.com/modin-project/modin/pull/6267)@mvashishtha
This release includes a bug fix.
This release includes a bug fix.
@mvashishtha
This release includes support for pyhdk=0.6, a few performance enhancements, new features and bug fixes.
This release includes support for pyhdk=0.6, a few performance enhancements, new features and bug fixes.
@mvashishtha @AndreyPavlenko @anmyachev @dchigarev @jkew @YarShev
This release includes many bug fixes, performance enhancements, and new features.
Modin 0.21.0
This release includes many bug fixes, performance enhancements, and new features.
dict_apply_builder use keyword argument internal_indices (#5945)AttributeError: 'list' object has no attribute '_query_compiler' in join op (#5939)build-docs CI job regardless of the files being changed (#5998)modin.pandas module (#6023)shift (#6168)tz_convert and tz_localize to QC if possible (#6137)truncate verifies that before <= after (#6134)_to_datetime attribute in pd.to_datetime (#6133)diff (#6167)pivot when values=None (#6166)numeric_only default to True (#6162)copy argument in tz_convert and tz_localize (#6182)runtime_env for a single-node case (#6028)inplace kwarg from query compiler clip arguments (#5954)to_pickle_distributed (#5950)modin/experimental/... folder (#6011)to_* functions (#5953)pd.cut to QC layer (#6136)groupby_ohlc implementation to QC layer (#6132)cancel-in-progress only for PRs (#5917)read_orc, read_spss, json_normalize, read_xml, read_gbq (#5983)modin/test/test_partition_api.py on unidist and dask (#6003)tmp_path fixture instead of ensure_clean_dir as pandas 2.0.0 does (#6008)Series.str accessor for pandas equivalence (#6033)Series.str through CachedAccessor (#6043)@AndreyPavlenko @RehanSD @YarShev @anmyachev @arunjose696 @dchigarev @devin-petersohn @helmeleegy @jkew @labanyamukhopadhyay @mdatre @mvashishtha @noloerino @pyrito @vnlitvinov @naren-ponder
This release includes some fixes.
Modin 0.20.1
This release includes some fixes.
dict_apply_builder use keyword argument internal_indices (https://github.com/modin-project/modin/pull/5945)AttributeError: 'list' object has no attribute '_query_compiler' in join op (https://github.com/modin-project/modin/pull/5939)build-docs CI job regardless of the files being changed (https://github.com/modin-project/modin/pull/5998)modin.pandas module (https://github.com/modin-project/modin/pull/6023)@AndreyPavlenko @anmyachev @dchigarev
REFACTOR-#5417: fix FutureWarning: the mangle_dupe_cols keyword is deprecated
Modin 0.20.0
This release adds parallel implementations for some functions on Dask that were previously implemented for other engines. It also includes support for pyhdk 0.5, many bug fixes and some performance enhancements.
where func (#5883)FactoryDispatcher.get_factory also initializes the engine (#4228)apply (#5915)astype used with copy=False parameter (#5918)SeriesGroupBy, DataFrameGroupBy objects (#5866)convert_dtypes as a full-axis operation instead of using map approach (#5885)pathlib.Path to str for read_parquet (#5860)Series.dt.day_of_week/day_of_year/isocalendar/asfreq methods (#5848)test_map_metadata test on the HDK engine and add to ci (#5929)test_window test on the HDK engine and add to ci (#5935)concat op (#5975)wrapper.materialize instead of wait_partitions; use AWS env vars in pytest_sessionstart function (#5981)self._identity in partitions only for "debug" logging level (#5679)_launch_tasks function (#5678)read_csv function lazy; introduce ModinIndex (#5677)read_csv, read_fwf, read_table, read_custom_text functions be executed fully asynchronous; introduce ModinDtypes (#5713)partition.get into base class (#5408)mangle_dupe_cols keyword is deprecated (#5407)upload-coverage action fail if there is no .coverage file (#5921)pragma: no cover for functions that used in apply_full_axis (#5920)codecov notifications until all reports have been sent (#5782)Import.gif as it's too large (#5958)to_parquet parallel implementation for Dask (#5876)to_sql parallel implementation for Dask (#5879)read_fwf parallel implementation for Dask (#5899)@MSHADroo @AndreyPavlenko @RehanSD @YarShev @anmyachev @dchigarev @mvashishtha @noloerino @pyrito @vnlitvinov
Nothing published for this version
Nothing published for this version
FIX-#5488: Remove usage of deprecated numpy types
Modin 0.19.0
This release introduces Modin's new, experimental NumPy API. It also features many bug fixes, improvements to documentation, and performance optimizations, including faster initialization with NumPy arrays.
expr.py (#5757)RecursionError for __int__ and __float__ (#5502)Series.values (#5469)skipfooter!=0 (#5522)dialect!=None (#5512)read_excel when usecols and index_cols parameters are provided (#5508)__repr__ of Modin categorical Series (#5516)__repr__ when display.max_rows=None (#5504)ParquetFileToRead a named tuple (#5352)DataFrameGroupBy.take() (#5474)Series.values when Series.dtype==ExtensionDtype (#5493)inplace parameter for set_axis function (#5579)PyArrowDataset.files work for 3.0.0 <= pyarrow < 8.0.0 (#5592)keep_partitioning=False (#5622)GroupBy.skew implementation via MapReduce pattern (#5318)to_pandas function (#5544)_repartition (#5543)drop_duplicates via new duplicated (#5587)pivot_table (#5546)columnarize function (#5548)reset_index function (#5547)duplicated in case there is only one column partition (#5640).str.* methods (#5658)isin (#5683)read_callback from dispatchers into parsers (#5689).loc without converting a Series to np.array (#5693)df.__setitem__ (#5708)Series.cat.codes (#5706).map() (#5704).isin() (#5707)default_to_pandas into base query_compiler class (#5479)__constructor__ in DataFrame and Series classes (#5485)FutureWarning: the mangle_dupe_cols keyword is deprecated for read_excel (#5415)modin.core.execution.dask module (#5418)df.iloc[:, i] = newvals (#5468)FutureWarning for DataFrameGroupBy.backfill (#5472)RayWrapper.put implementation (#5686)UnidistWrapper.put implementation (#5688)columns parameter for get_dtypes function (#5717)Post Run conda-incubator/setup-miniconda@v2 step on Windows (#5662)apply_full_axis with broadcast_apply_full_axis (#5637)@AndreyPavlenko @Egor-Krivov @RehanSD @YarShev @anmyachev @arunjose696 @dchigarev @devin-petersohn @mvashishtha @noloerino @vnlitvinov @Billy2551 @Retribution98 @shalearkane
FIX-#5488: Remove usage of deprecated numpy types
Modin 0.18.1
This release includes pandas 1.5.3 support and a bunch of bug fixes.
RecursionError for __int__ and __float__ (#5502)Series.values (#5469)skipfooter!=0 (#5522)dialect!=None (#5512)__repr__ of Modin categorical Series (#5516)ParquetFileToRead a named tuple (#5352)DataFrameGroupBy.take() (#5474)Series.values when Series.dtype==ExtensionDtype (#5493)@AndreyPavlenko @YarShev @anmyachev @dchigarev @vnlitvinov @Retribution98
FIX-https://github.com/modin-project/modin/issues/5319: Do not use deprecated '.iteritems()'
This release includes support for MPI backend using Unidist, improvements to the shuffling mechanism, SQL query execution on the HDK backend (currently pyhdk==0.3), support for pandas 1.5.2 and external query compilers. It also includes many bug fixes and some performance enhancements.
read_parquet to detect column partitioning in non-local filesystems (https://github.com/modin-project/modin/pull/5192)df.info failure with default columns (https://github.com/modin-project/modin/pull/5251)df_categories_equals typo (https://github.com/modin-project/modin/pull/5250)set_index case with multiindex (https://github.com/modin-project/modin/pull/5190)ray==2.1.0 (https://github.com/modin-project/modin/pull/5283)modin-test bucket (https://github.com/modin-project/modin/pull/5257)execute function (https://github.com/modin-project/modin/pull/5278)read_csv_glob with non-empty parse_dates dict (https://github.com/modin-project/modin/pull/5339)get_indices internal function (https://github.com/modin-project/modin/pull/5355)ray>=1.13.0 (https://github.com/modin-project/modin/pull/5390)get on all partitions at once in to_pandas (https://github.com/modin-project/modin/pull/4776)Variable defined multiple times error found by CodeQL (https://github.com/modin-project/modin/pull/5300)BaseIO._read (https://github.com/modin-project/modin/pull/5329)PQ_INDEX_REGEX as class variable (https://github.com/modin-project/modin/pull/5333)_validate as classmethod (https://github.com/modin-project/modin/pull/5331)add_to_apply_calls impl in base class (https://github.com/modin-project/modin/pull/5354)pandas.util.cache_readonly for __constructors__ (https://github.com/modin-project/modin/pull/5368)Index.dtype instead of isinstance(obj, Int64Index) (https://github.com/modin-project/modin/pull/5406)cache_readonly to avoid errors in doc_checker.py (https://github.com/modin-project/modin/pull/5365)reindex method (https://github.com/modin-project/modin/pull/4434)str.extract when expand==True (https://github.com/modin-project/modin/pull/5243)rebalance_partitions for Unidist (https://github.com/modin-project/modin/pull/5385)@AndreyPavlenko @Billy2551 @Garra1980 @RehanSD @YarShev @anmyachev @arunjose696 @dchigarev @devin-petersohn @lgtm-migrator @mvashishtha @noloerino @pyrito @trgiangdo @vnlitvinov @Retribution98
This release includes pandas 1.5.2 support and a bunch of bug fixes.
This release includes pandas 1.5.2 support and a bunch of bug fixes.
read_parquet to detect column partitioning in non-local filesystems (#5192)set_index case with multiindex (#5190)modin-test bucket (#5257)@AndreyPavlenko @Billy2551 @RehanSD @YarShev @anmyachev @dchigarev @mvashishtha @noloerino
FIX-#5097: Stop using deprecated mangle_dup_cols.
This release includes support for pyhdk 0.2. It also includes many bug fixes and some performance enhancements.
fillna when Modin series object is an argument (#4674)df.get() (#5035)PandasQueryCompiler.groupby_mean with timestamp in by (#5140)query_compiler.dt_prop_map (#5133)to_parquet (#5161)get_dummies to respect passed columns to be encoded (#5185)getitem_bool when the key is Series with empty partition (#5189)_compute_axis_labels_and_lengths for computing _row_lengths/_column_widths (#5030)set_axis function (#5093)join and merge ops from pandas (#5021).__setitem__ (#5142)@AndreyPavlenko @Billy2551 @RehanSD @YarShev @anmyachev @dchigarev @devin-petersohn @ienkovich @mvashishtha @noloerino @pyrito @rosdyana @shalearkane @suhailrehman @vnlitvinov
This release includes pandas 1.5.1 support and two bug fixes.
This release includes pandas 1.5.1 support and two bug fixes.
@AndreyPavlenko @mvashishtha @YarShev
This release features a bug fix, as well as fixes for deprecation warnings introduced by pandas 1.5.
This release features a bug fix, as well as fixes for deprecation warnings introduced by pandas 1.5.
df.get() (#5035)@mvashishtha @pyrito @anmyachev @vnlitvinov
This release includes support for pandas 1.5, support for the latest version of dask, and backwards compatibility with python 3.6 and pandas 1.1. Addi
This release includes support for pandas 1.5, support for the latest version of dask, and backwards compatibility with python 3.6 and pandas 1.1. Additionally, it includes many performance enhancements, bug fixes, and documentation improvements.
np.bool -> np.bool_ (#4571)read_csv in case skiprows=<0, []> (#4544)groupby + agg in case when multicolumn can arise (#4642)storage_options usage for read_csv and read_csv_glob (#4644)df.describe() (#4651)iloc/loc assignment when dataframe is empty (#4677)by in df.groupby() (#4667)read_csv that started defaulting to pandas again in case of reading from a buffer and when a buffer has a non-zero starting position (#4681)loc shouldn't drop levels for full-key lookups (#4608)read_* functions from s3 storages (#4659)frame.index or frame.columns (#4721)from_dataframe (#4737)storage_options in read_parquet (#4764)fsspec for handling s3/http-like paths instead of s3fs (#4710)read_parquet (#4783)read_parquet (#4837)base_lengths should be computed from base_frame instead of self in copartition (#4915)dtypes computation in dataframe.filter (#4928)radd for Series and DataFrame (#4908)_take_2d_positional that loses indexes due to filtering empty dataframes (#4951)__getitem_bool for single row dataframes (#4845)frac to None in _sample when n=0 (#4984)_default_to_pandas in df.attrs (#4995)execute function in ASV utils failed if len(partitions) == 0 (#5044)to_timedelta to return Series instead of TimedeltaIndex (#5028)groupby.mean for narrow data (#4591)df.copy call from from_pandas since it is not needed for Ray and Dask (#4781)__setitem__ when no new column names are assigning (#4455)drop operation (#4694)concat operation (#4728)Series objects with shared .index (#4689)ser.cat.categories, ser.cat.ordered, and ser.__array_priority__ (#4704)read_parquet over row groups (#4700)lengths and widths in put method of Dask partition like Ray do (#4780)PandasOnRayDataframePartition._length_cache and PandasOnRayDataframePartition._width_cache (#4754)compute_sliced_len.remote when row_labels/col_labels == slice(None) (#4863)Series.cat.codes, Series.dt.tz, and Series.dt.to_pytimedelta (#4833)dtypes for binary operations that can only return bool type and the right operand is not a Modin object (#4852)copy should not trigger any previous computations (#4843)dtypes in concat also for ROW_WISE case when possible (#4850)dtype when using Series.dt accessor (#4930)lengths in rebalance_partitions when possible (#4893)_propagate_index_objs (#4888)PandasDataframeAxisPartition.deploy_axis_func should be serialized only once (#4861)PandasDataframeAxisPartition.drain should be serialized only once (#4891)__getattribute__ and __getitem__ (4911)query method (#4887)iloc function that used in partition.mask should be serialized only once (#4901)take_2d_labels_or_positional unless they are needed (#4921)apply in virtual partition' drain_call_queue if call_queue is empty (#4975)reset_index shouldn't trigger index materialization if possible (#5018)width/length methods instead of _compute_axis_labels_and_lengths if index is already known (#4964)concatenate (#4953)TimeConcat benchmark (#5067)merge op with categorical data (#5084)TimeConcat benchmark with new parameter ignore_index (#5065)_build_treereduce_func call from _compute_dtypes (#4775)split_result_of_axis_func_pandas (#4831)PandasOnRayDataframePartitionManager (#4895)PandasOnDaskDataframe (#3781)modin/core/execution/dask/common/__init__.py with modin/core/execution/ray/common/__init__.py (#4979)default2pandas/dataframe.py and default2pandas/any.py (#4950)RayTask to RayWrapper in accordance with Dask (#4977)finalize method instead of list comprehension + drain_call_queue (#5006)jenkins stuff (#5002)width/length (#4971)call method in favor of register due to duplication (4943)_row_lengths and _column_widths public (#5025)RayWrapper.materialize instead of ray.get (#5010)c323f7fe385011ed849300155de07645.db file (#5082)black/flake8 for omnisci ci-notebooks (#4609)ensure_clean() in place of io_tests_data (#4881)test-compat-win (#5007)read_ function defaults to pandas (#4647)read_parquet (#4807)read_csv and read_csv_glob (#4898)infer_types dataframe algebra operator (#4871)@mvashishtha @NickCrews @prutskov @vnlitvinov @pyrito @suhailrehman @RehanSD @helmeleegy @anmyachev @d33bs @noloerino @devin-petersohn @YarShev @naren-ponder @jbrockmendel @ienkovich @Garra1980 @Billy2551
This release adds support for pandas 1.4.4 and includes a bunch of bugfixes.
This release adds support for pandas 1.4.4 and includes a bunch of bugfixes.
groupby + agg in case when multicolumn can arise (#4642)df.describe() (#4651)by in df.groupby() (#4667)iloc/loc assignment when dataframe is empty (#4677)read_* functions from s3 storages (#4659)frame.index or frame.columns (#4721)read_csv that started defaulting to pandas again in case of reading from a buffer and when a buffer has a non-zero starting position (#4681)fsspec for handling s3/http-like paths instead of s3fs (#4710)storage_options usage for read_csv and read_csv_glob (#4644)@helmeleegy @YarShev @anmyachev @pyrito @prutskov @jbrockmendel @mvashishtha @RehanSD @vnlitvinov
This release adds support for pandas 1.4.3, pins protobuf < 4.0.0 to ensure compatibility with ray < 1.13, and includes a bugfix for modifying columns
This release adds support for pandas 1.4.3, pins protobuf < 4.0.0 to ensure compatibility with
ray < 1.13, and includes a bugfix for modifying columns via attribute access.
@mvashishtha @pyrito @RehanSD
This release pins Ray < 1.13.0 to avoid deserialization race condition.
This release pins Ray < 1.13.0 to avoid deserialization race condition.
@mvashishtha
FIX-https://github.com/modin-project/modin/issues/4385: Get rid of use-deprecated option in pip
This release includes updated support for pandas 1.4.2, new Batch and Logging APIs, and a plethora of bug fixes and documentation improvements.
use-deprecated option in pip (https://github.com/modin-project/modin/pull/4386)insert function with pandas in case of numpy array with several columns (https://github.com/modin-project/modin/pull/4408)read_csv_glob with usecols parameter (https://github.com/modin-project/modin/pull/4405)reindex function that doesn't preserve initial index metadata (https://github.com/modin-project/modin/pull/4442)loc in case when need reindex item (https://github.com/modin-project/modin/pull/4457)drop_duplicates no longer removes items based on index values (https://github.com/modin-project/modin/pull/4468)deploy function (https://github.com/modin-project/modin/pull/4285)warns_that_defaulting_to_pandas (https://github.com/modin-project/modin/pull/4423)eval_insert utility that doesn't actually check results of insert function (https://github.com/modin-project/modin/pull/4410)pathlib from deps (https://github.com/modin-project/modin/pull/4384)redis to Modin dependencies (https://github.com/modin-project/modin/pull/4396)@YarShev @Garra1980 @prutskov @alexander3774 @amyskov @wangxiaoying @jeffreykennethli @mvashishtha @anmyachev @dchigarev @devin-petersohn @jrsacher @orcahmlee @naren-ponder @RehanSD
This release contains a few key bugfixes and pandas version update.
This release contains a few key bugfixes and pandas version update.
@Garra1980, @devin-petersohn, @dchigarev, @jeffreykennethli, @mvashishtha, @YarShev, @anmyachev
This release contains significant upgrades to Developer API, as well as to Modin's documentation, some refactor codebase and performance enhancements,
This release contains significant upgrades to Developer API, as well as to Modin's documentation, some refactor codebase and performance enhancements, and multiple bugfixes.
OptionError (https://github.com/modin-project/modin/pull/4109)skipif instead of skip for compatibility with pytest 7.0 (https://github.com/modin-project/modin/pull/4163)PandasDataFrame.from_labels (https://github.com/modin-project/modin/pull/4209)read_csv_glob (https://github.com/modin-project/modin/pull/4074)wait method for PandasOnRayDataframeColumnPartition class (https://github.com/modin-project/modin/pull/4231)PandasDataframePartition hierarchy (https://github.com/modin-project/modin/pull/3991)dask_client global variable in modin\pandas\__init__.py (https://github.com/modin-project/modin/pull/4230)broadcast_apply method (https://github.com/modin-project/modin/pull/3996)get_indices function (https://github.com/modin-project/modin/pull/3995)to_pandas, to_numpy functions in QueryCompiler hierarchy (https://github.com/modin-project/modin/pull/4332)modin/examples/tutorial/ directory (https://github.com/modin-project/modin/pull/4214)__init__ method of PandasOnDaskDataframePartition class (https://github.com/modin-project/modin/pull/4207)cluster directory to cloud in examples (https://github.com/modin-project/modin/pull/4212)DaskWrapper class (https://github.com/modin-project/modin/pull/3854)read_custom_text function that can read custom line-by-line text files (https://github.com/modin-project/modin/pull/3441)test-internals CI job (https://github.com/modin-project/modin/pull/4198)BasePandasDataSet docstrings warnings (https://github.com/modin-project/modin/pull/4333)black formatting, fix pydocstyle check and readthedocs build (https://github.com/modin-project/modin/pull/4114)pip in readthedocs deps list (https://github.com/modin-project/modin/pull/4170)Dask<2022.2.0 as a temporary fix of CI (https://github.com/modin-project/modin/pull/4218)@prutskov, @amyskov, @paulovn, @anmyachev, @YarShev, @RehanSD, @devin-petersohn, @dchigarev, @Garra1980, @mvashishtha, @naren-ponder, @jeffreykennethli, @dorisjlee, @Rubtsowa
This release contains a few key bugfixes and pandas version update.
This release contains a few key bugfixes and pandas version update.
@mvashishtha @anmyachev @prutskov @devin-petersohn @naren-ponder @YarShev @Garra1980
This release contains documentation polishing and small user experience improvements.
This release contains documentation polishing and small user experience improvements.
@RehanSD, @YarShev, @dchigarev, @prutskov, @Garra1980
This release contains a few key bugfixes and updates to the documentation.
This release contains a few key bugfixes and updates to the documentation.
OptionError (#4109)PandasDataframe in docs (#4108)@prutskov, @paulovn, @YarShev, @RehanSD, @devin-petersohn, @mvashishtha
This release contains significant upgrades to Modin's documentation, support for pandas 1.4, new algebra and partitioning layer APIs, and some bugfixe
This release contains significant upgrades to Modin's documentation, support for pandas 1.4, new algebra and partitioning layer APIs, and some bugfixes.
by (a04d7b7)sort=False with categorical keys (c67a7c5)REDIS_PASSWORD with Ray's DEFAULT_REDIS_PASSWORD (f79cb85)by and columns to aggregate overlap (d42c070)read_csv when callables are provided for skip_rows parameter (7c84758)ray.init when running Ray in local mode (02a23d4)groupby.indices returns positional indices (e9c06f2)df.__getitem__ respects step attribute of slice (7e85c5d)apply result type inference (ac17ca1)df.to_csv propagates metadata (e.g. index) (154697b)pyarrow requirement in environment files (b55b08d)__getitem__ flow for .loc/.iloc (0947ee8)dtypes on transpose (cd8db0c)FactoryDispatcher in Modin experimental pandas IO (2cfabaf)storage_options argument to read_csv_glob (7c33afe)dropna argument for groupby.indices and groupby.groups (144a613)Series.values to default to to_numpy() (67228ef)modin.pandas.show_versions and python -m modin --versions (efe717f)getArrowTable() (6882ec2)init when only OmniSci is present (8c8a6a3)append with default arguments (67013f9)read_csv_glob and ensure warning raised if wildcard not in filepath_or_buffer (be10ba9)modin.core.dataframe.base and modin.core.dataframe.pandas (cf1e541)pandas.read_json (0315823)ModinDataframe (4b70725)@anmyachev, @prutskov, @Rubtsowa, @vnlitvinov, @dchigarev, @YarShev, @amyskov, @mvashishtha, @dorisjlee, @devin-petersohn, @jeffreykennethli, @RehanSD, @novichkovg, @Lozovskii-Aleksandr, @naren-ponder, @ahallermed, @fexolm, @adityagp, @susmitpy, @ienkovich
Your coding agent can read these notes before it upgrades. Set up the MCP server →