NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #807 most downloaded on PyPI
Parallel PyData with Task Scheduling
Last release 1 months ago
24 Aug 2026
Ships fairly regularly
a new release about every 5 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
12 years old
239 releases · first in 2015
This is a hotfix release that fixes a P2P shuffling bug introduced in the 2023.9.0 release (see dask#10493 ).
Released on September 6, 2023
Note
This is a hotfix release that fixes a P2P shuffling bug introduced in the 2023.9.0 release (see dask#10493 ).
Stricter data type for dask keys ( dask#10485 ) crusaderky
Special handling for None in DASK_ environment variables ( dask#10487 ) crusaderky
Fix _partitions dtype in meta for DataFrame.set_index and DataFrame.sort_values ( dask#10493 ) Hendrik Makait
Handle cached_property decorators in derived_from ( dask#10490 ) Lawrence Mitchell
Bump actions/checkout from 3.6.0 to 4.0.0 ( dask#10492 )
Simplify some tests that import distributed ( dask#10484 ) crusaderky
Released on September 6, 2023
Note
This is a hotfix release that fixes a P2P shuffling bug introduced in the 2023.9.0 release (see 10493).
Stricter data type for dask keys (10485) `crusaderky`_
Special handling for None in DASK_ environment variables (10487) `crusaderky`_
Fix _partitions dtype in meta for DataFrame.set_index and DataFrame.sort_values (10493) `Hendrik Makait`_
Handle cached_property decorators in derived_from (10490) `Lawrence Mitchell`_
Bump actions/checkout from 3.6.0 to 4.0.0 (10492)
Simplify some tests that import distributed (10484) `crusaderky`_
One column per quarter.
- Remove support for np.int64 in keys ( dask#10483 ) crusaderky
Released on September 1, 2023
Remove support for np.int64 in keys ( dask#10483 ) crusaderky
Fix _partitions dtype in meta for shuffling ( dask#10462 ) Hendrik Makait
Don’t use exception hooks to shorten tracebacks ( dask#10456 ) crusaderky
Add p2p shuffle option to DataFrame docs ( dask#10477 ) Patrick Hoefler
Skip failing tests for pandas=2.1.0 ( dask#10488 ) Patrick Hoefler
Update tests for pandas=2.1.0 ( dask#10439 ) Patrick Hoefler
Enable pytest-timeout ( dask#10482 ) crusaderky
Bump actions/checkout from 3.5.3 to 3.6.0 ( dask#10470 )
Released on September 1, 2023
Remove support for np.int64 in keys (10483) `crusaderky`_
Fix _partitions dtype in meta for shuffling (10462) `Hendrik Makait`_
Don't use exception hooks to shorten tracebacks (10456) `crusaderky`_
Add p2p shuffle option to DataFrame docs (10477) `Patrick Hoefler`_
Skip failing tests for pandas=2.1.0 (10488) `Patrick Hoefler`_
Update tests for pandas=2.1.0 (10439) `Patrick Hoefler`_
Enable pytest-timeout (10482) `crusaderky`_
Bump actions/checkout from 3.5.3 to 3.6.0 (10470)
- Adding support for cgroup v2 to cpu_count ( dask#10419 ) Johan Olsson
Released on August 18, 2023
Adding support for cgroup v2 to cpu_count ( dask#10419 ) Johan Olsson
Support multi-column groupby with sort=True and split_out>1 ( dask#10425 ) Richard (Rick) Zamora
Add DataFrame.enforce_runtime_divisions method ( dask#10404 ) Richard (Rick) Zamora
Enable file mode="x" with a single_file=True for Dask DataFrame to_csv ( dask#10443 ) Genevieve Buckley
Fix ValueError when running to_csv in append mode with single_file as True ( dask#10441 ) Ben
Add default types_mapper to from_pyarrow_table_dispatch for pandas ( dask#10446 ) Richard (Rick) Zamora
Released on August 18, 2023
Adding support for cgroup v2 to cpu_count (10419) `Johan Olsson`_
Support multi-column groupby with sort=True and split_out>1 (10425) `Richard (Rick) Zamora`_
Add DataFrame.enforce_runtime_divisions method (10404) `Richard (Rick) Zamora`_
Enable file mode="x" with a single_file=True for Dask DataFrame to_csv (10443) `Genevieve Buckley`_
Fix ValueError when running to_csv in append mode with single_file as True (10441) `Ben`_
Add default types_mapper to from_pyarrow_table_dispatch for pandas (10446) `Richard (Rick) Zamora`_
- Fix for make_timeseries performance regression ( dask#10428 ) Irina Truong
Released on August 4, 2023
Fix for make_timeseries performance regression ( dask#10428 ) Irina Truong
Add distributed.print to debugging docs ( dask#10435 ) James Bourbeau
Documenting compatibility of NumPy functions with Dask functions ( dask#9941 ) Chiara Marmo
Use SPDX in license metadata ( dask#10437 ) John A Kirkham
Require dask[array] in dask[dataframe] ( dask#10357 ) John A Kirkham
Update gpuCI RAPIDS_VER to 23.10 ( dask#10427 )
Simplify compatibility code ( dask#10426 ) Hendrik Makait
Fix compatibility variable naming ( dask#10424 ) Hendrik Makait
Fix a few errors with upstream pandas and pyarrow ( dask#10412 ) Irina Truong
Released on August 4, 2023
Fix for make_timeseries performance regression (10428) `Irina Truong`_
Add distributed.print to debugging docs (10435) `James Bourbeau`_
Documenting compatibility of NumPy functions with Dask functions (9941) `Chiara Marmo`_
Use SPDX in license metadata (10437) `John A Kirkham`_
Require dask[array] in dask[dataframe] (10357) `John A Kirkham`_
Update gpuCI RAPIDS_VER to 23.10 (10427)
Simplify compatibility code (10426) `Hendrik Makait`_
Fix compatibility variable naming (10424) `Hendrik Makait`_
Fix a few errors with upstream pandas and pyarrow (10412) `Irina Truong`_
This release updates Dask DataFrame to automatically convert text data using object data types to string[pyarrow] if pandas>=2 and pyarrow>=12 are ins
Released on July 20, 2023
Note
This release updates Dask DataFrame to automatically convert text data using object data types to string[pyarrow] if pandas>=2 and pyarrow>=12 are installed.
This should result in significantly reduced memory consumption and increased computation performance in many workflows that deal with text data.
You can disable this change by setting the dataframe.convert-string configuration value to False with
dask . config . set ({ "dataframe.convert-string" : False })
Convert to pyarrow strings if proper dependencies are installed ( dask#10400 ) James Bourbeau
Avoid repartition before shuffle for p2p ( dask#10421 ) Patrick Hoefler
API to generate random Dask DataFrames ( dask#10392 ) Irina Truong
Speed up dask.bag.Bag.random_sample ( dask#10356 ) crusaderky
Raise helpful ValueError for invalid time units ( dask#10408 ) Nat Tabris
Make repartition a no-op when divisions match (divisions provided as a list) ( dask#10395 ) Nicolas Grandemange
Use dataframe.convert-string in read_parquet token ( dask#10411 ) James Bourbeau
Category dtype is lost when concatenating MultiIndex ( dask#10407 ) Irina Truong
Fix FutureWarning: The provided callable... ( dask#10405 ) Irina Truong
Enable non-categorical hive-partition columns in read_parquet ( dask#10353 ) Richard (Rick) Zamora
concat ignoring DataFrame withouth columns ( dask#10359 ) Patrick Hoefler
Released on July 20, 2023
Note
This release updates Dask DataFrame to automatically convert text data using object data types to string[pyarrow] if pandas>=2 and pyarrow>=12 are installed.
This should result in significantly reduced memory consumption and increased computation performance in many workflows that deal with text data.
You can disable this change by setting the dataframe.convert-string configuration value to False with
dask.config.set({"dataframe.convert-string": False})
Convert to pyarrow strings if proper dependencies are installed (10400) `James Bourbeau`_
Avoid repartition before shuffle for p2p (10421) `Patrick Hoefler`_
API to generate random Dask DataFrames (10392) `Irina Truong`_
Speed up dask.bag.Bag.random_sample (10356) `crusaderky`_
Raise helpful ValueError for invalid time units (10408) `Nat Tabris`_
Make repartition a no-op when divisions match (divisions provided as a list) (10395) `Nicolas Grandemange`_
Use dataframe.convert-string in read_parquet token (10411) `James Bourbeau`_
Category dtype is lost when concatenating MultiIndex (10407) `Irina Truong`_
Fix FutureWarning: The provided callable... (10405) `Irina Truong`_
Enable non-categorical hive-partition columns in read_parquet (10353) `Richard (Rick) Zamora`_
concat ignoring DataFrame withouth columns (10359) `Patrick Hoefler`_
- Fix test_first_and_last to accommodate deprecated last ( dask#10373 ) James Bourbeau
Released on July 7, 2023
Catch exceptions when attempting to load CLI entry points ( dask#10380 ) Jacob Tomlinson
Fix typo in _clean_ipython_traceback ( dask#10385 ) Alexander Clausen
Ensure that df is immutable after from_pandas ( dask#10383 ) Patrick Hoefler
Warn consistently for inplace in Series.rename ( dask#10313 ) Patrick Hoefler
Add clarification about output shape and reshaping in rechunk documentation ( dask#10377 ) Swayam Patil
Simplify astype implementation ( dask#10393 ) Patrick Hoefler
Fix test_first_and_last to accommodate deprecated last ( dask#10373 ) James Bourbeau
Add level to create_merge_tree ( dask#10391 ) Patrick Hoefler
Do not derive from scipy.stats.chisquare docstring ( dask#10382 ) Doug Davis
Released on July 7, 2023
Catch exceptions when attempting to load CLI entry points (10380) `Jacob Tomlinson`_
Fix typo in _clean_ipython_traceback (10385) `Alexander Clausen`_
Ensure that df is immutable after from_pandas (10383) `Patrick Hoefler`_
Warn consistently for inplace in Series.rename (10313) `Patrick Hoefler`_
Add clarification about output shape and reshaping in rechunk documentation (10377) `Swayam Patil`_
Simplify astype implementation (10393) `Patrick Hoefler`_
Fix test_first_and_last to accommodate deprecated last (10373) `James Bourbeau`_
Add level to create_merge_tree (10391) `Patrick Hoefler`_
Do not derive from scipy.stats.chisquare docstring (10382) `Doug Davis`_
- Deprecate DataFrame.fillna / Series.fillna with method ( dask#10349 ) Irina Truong
Released on June 26, 2023
Remove no longer supported clip_lower and clip_upper ( dask#10371 ) Patrick Hoefler
Support DataFrame.set_index(..., sort=False) ( dask#10342 ) Miles
Cleanup remote tracebacks ( dask#10354 ) Irina Truong
Add dispatching mechanisms for pyarrow.Table conversion ( dask#10312 ) Richard (Rick) Zamora
Choose P2P even if fusion is enabled ( dask#10344 ) Hendrik Makait
Validate that rechunking is possible earlier in graph generation ( dask#10336 ) Hendrik Makait
Fix issue with header passed to read_csv ( dask#10355 ) GALI PREM SAGAR
Respect dropna and observed in GroupBy.var and GroupBy.std ( dask#10350 ) Patrick Hoefler
Fix H5FD_lock error when writing to hdf with distributed client ( dask#10309 ) Irina Truong
Fix for total_mem_usage of bag.map() ( dask#10341 ) Irina Truong
Deprecate DataFrame.fillna / Series.fillna with method ( dask#10349 ) Irina Truong
Deprecate DataFrame.first and Series.first ( dask#10352 ) Irina Truong
Deprecate numpy.compat ( dask#10370 ) Irina Truong
Fix annotations and spans leaking between threads ( dask#10367 ) Irina Truong
Use general kwargs in pyarrow_table_dispatch functions ( dask#10364 ) Richard (Rick) Zamora
Remove unnecessary try / except in isna ( dask#10363 ) Patrick Hoefler
mypy support for numpy 1.25 ( dask#10362 ) crusaderky
Bump actions/checkout from 3.5.2 to 3.5.3 ( dask#10348 )
Restore numba in upstream build ( dask#10330 ) James Bourbeau
Update nightly wheel index for pandas / numpy / scipy ( dask#10346 ) Matthew Roeschke
Add rechunk config values to yaml ( dask#10343 ) Hendrik Makait
Released on June 26, 2023
Remove no longer supported clip_lower and clip_upper (10371) `Patrick Hoefler`_
Support DataFrame.set_index(..., sort=False) (10342) `Miles`_
Cleanup remote tracebacks (10354) `Irina Truong`_
Add dispatching mechanisms for pyarrow.Table conversion (10312) `Richard (Rick) Zamora`_
Choose P2P even if fusion is enabled (10344) `Hendrik Makait`_
Validate that rechunking is possible earlier in graph generation (10336) `Hendrik Makait`_
Fix issue with header passed to read_csv (10355) `GALI PREM SAGAR`_
Respect dropna and observed in GroupBy.var and GroupBy.std (10350) `Patrick Hoefler`_
Fix H5FD_lock error when writing to hdf with distributed client (10309) `Irina Truong`_
Fix for total_mem_usage of bag.map() (10341) `Irina Truong`_
Deprecate DataFrame.fillna/Series.fillna with method (10349) `Irina Truong`_
Deprecate DataFrame.first and Series.first (10352) `Irina Truong`_
Deprecate numpy.compat (10370) `Irina Truong`_
Fix annotations and spans leaking between threads (10367) `Irina Truong`_
Use general kwargs in pyarrow_table_dispatch functions (10364) `Richard (Rick) Zamora`_
Remove unnecessary try/except in isna (10363) `Patrick Hoefler`_
mypy support for numpy 1.25 (10362) `crusaderky`_
Bump actions/checkout from 3.5.2 to 3.5.3 (10348)
Restore numba in upstream build (10330) `James Bourbeau`_
Update nightly wheel index for pandas/numpy/scipy (10346) `Matthew Roeschke`_
Add rechunk config values to yaml (10343) `Hendrik Makait`_
- Add missing not in predicate support to read_parquet ( dask#10320 ) Richard (Rick) Zamora
Released on June 9, 2023
Add missing not in predicate support to read_parquet ( dask#10320 ) Richard (Rick) Zamora
Fix for incorrect value_counts ( dask#10323 ) Irina Truong
Update empty describe top and freq values ( dask#10319 ) James Bourbeau
Fix hetzner typo ( dask#10332 ) Sarah Charlotte Johnson
Test with numba and sparse on Python 3.11 ( dask#10329 ) Thomas Grainger
Remove numpy.find_common_type warning ignore ( dask#10311 ) James Bourbeau
Update gpuCI RAPIDS_VER to 23.08 ( dask#10310 )
Released on June 9, 2023
Add missing not in predicate support to read_parquet (10320) `Richard (Rick) Zamora`_
Fix for incorrect value_counts (10323) `Irina Truong`_
Update empty describe top and freq values (10319) `James Bourbeau`_
Fix hetzner typo (10332) `Sarah Charlotte Johnson`_
Test with numba and sparse on Python 3.11 (10329) `Thomas Grainger`_
Remove numpy.find_common_type warning ignore (10311) `James Bourbeau`_
Update gpuCI RAPIDS_VER to 23.08 (10310)
- Avoid pandas Series.__getitem__ deprecation in tests ( dask#10308 ) James Bourbeau
Released on May 26, 2023
Note
This release drops support for Python 3.8. As of this release Dask supports Python 3.9, 3.10, and 3.11. See this community issue for more details.
Drop Python 3.8 support ( dask#10295 ) Thomas Grainger
Change Dask Bag partitioning scheme to improve cluster saturation ( dask#10294 ) Jacob Tomlinson
Generalize dd.to_datetime for GPU-backed collections, introduce get_meta_library utility ( dask#9881 ) Charles Blackmon-Luca
Add na_action to DataFrame.map ( dask#10305 ) Patrick Hoefler
Raise TypeError in DataFrame.nsmallest and DataFrame.nlargest when columns is not given ( dask#10301 ) Patrick Hoefler
Improve sizeof for pd.MultiIndex ( dask#10230 ) Patrick Hoefler
Support duplicated columns in a bunch of DataFrame methods ( dask#10261 ) Patrick Hoefler
Add numeric_only support to DataFrame.idxmin and DataFrame.idxmax ( dask#10253 ) Patrick Hoefler
Implement numeric_only support for DataFrame.quantile ( dask#10259 ) Patrick Hoefler
Add support for numeric_only=False in DataFrame.std ( dask#10251 ) Patrick Hoefler
Implement numeric_only=False for GroupBy.cumprod and GroupBy.cumsum ( dask#10262 ) Patrick Hoefler
Implement numeric_only for skew and kurtosis ( dask#10258 ) Patrick Hoefler
mask and where should accept a callable ( dask#10289 ) Irina Truong
Fix conversion from Categorical to pa.dictionary in read_parquet ( dask#10285 ) Patrick Hoefler
Spurious config on nested annotations ( dask#10318 ) crusaderky
Fix rechunking behavior for dimensions with known and unknown chunk sizes ( dask#10157 ) Hendrik Makait
Enable drop to support mismatched partitions ( dask#10300 ) James Bourbeau
Fix divisions construction for to_timestamp ( dask#10304 ) Patrick Hoefler
pandas ExtensionDtype raising in Series reduction operations ( dask#10149 ) Patrick Hoefler
Fix regression in da.random interface ( dask#10247 ) Eray Aslan
da.coarsen doesn’t trim an empty chunk in meta ( dask#10281 ) Irina Truong
Fix dtype inference for engine="pyarrow" in read_csv ( dask#10280 ) Patrick Hoefler
Add meta_from_array to API docs ( dask#10306 ) Ruth Comer
Update Coiled links ( dask#10296 ) Sarah Charlotte Johnson
Add docs for demo day ( dask#10288 ) Matthew Rocklin
Explicitly install anaconda-client from conda-forge when uploading conda nightlies ( dask#10316 ) Charles Blackmon-Luca
Configure isort to add from future import annotations ( dask#10314 ) Thomas Grainger
Avoid pandas Series.getitem deprecation in tests ( dask#10308 ) James Bourbeau
Ignore numpy.find_common_type warning from pandas ( dask#10307 ) James Bourbeau
Add test to check that DataFrame.setitem does not modify df inplace ( dask#10223 ) Patrick Hoefler
Clean up default value of dropna in value_counts ( dask#10299 ) Patrick Hoefler
Add pytest-cov to test extra ( dask#10271 ) James Bourbeau
Released on May 26, 2023
Note
This release drops support for Python 3.8. As of this release Dask supports Python 3.9, 3.10, and 3.11. See this community issue for more details.
Drop Python 3.8 support (10295) `Thomas Grainger`_
Change Dask Bag partitioning scheme to improve cluster saturation (10294) `Jacob Tomlinson`_
Generalize dd.to_datetime for GPU-backed collections, introduce get_meta_library utility (9881) `Charles Blackmon-Luca`_
Add na_action to DataFrame.map (10305) `Patrick Hoefler`_
Raise TypeError in DataFrame.nsmallest and DataFrame.nlargest when columns is not given (10301) `Patrick Hoefler`_
Improve sizeof for pd.MultiIndex (10230) `Patrick Hoefler`_
Support duplicated columns in a bunch of DataFrame methods (10261) `Patrick Hoefler`_
Add numeric_only support to DataFrame.idxmin and DataFrame.idxmax (10253) `Patrick Hoefler`_
Implement numeric_only support for DataFrame.quantile (10259) `Patrick Hoefler`_
Add support for numeric_only=False in DataFrame.std (10251) `Patrick Hoefler`_
Implement numeric_only=False for GroupBy.cumprod and GroupBy.cumsum (10262) `Patrick Hoefler`_
Implement numeric_only for skew and kurtosis (10258) `Patrick Hoefler`_
mask and where should accept a callable (10289) `Irina Truong`_
Fix conversion from Categorical to pa.dictionary in read_parquet (10285) `Patrick Hoefler`_
Spurious config on nested annotations (10318) `crusaderky`_
Fix rechunking behavior for dimensions with known and unknown chunk sizes (10157) `Hendrik Makait`_
Enable drop to support mismatched partitions (10300) `James Bourbeau`_
Fix divisions construction for to_timestamp (10304) `Patrick Hoefler`_
pandas ExtensionDtype raising in Series reduction operations (10149) `Patrick Hoefler`_
Fix regression in da.random interface (10247) `Eray Aslan`_
da.coarsen doesn't trim an empty chunk in meta (10281) `Irina Truong`_
Fix dtype inference for engine="pyarrow" in read_csv (10280) `Patrick Hoefler`_
Add meta_from_array to API docs (10306) `Ruth Comer`_
Update Coiled links (10296) `Sarah Charlotte Johnson`_
Add docs for demo day (10288) `Matthew Rocklin`_
Explicitly install anaconda-client from conda-forge when uploading conda nightlies (10316) `Charles Blackmon-Luca`_
Configure isort to add from __future__ import annotations (10314) `Thomas Grainger`_
Avoid pandas Series.__getitem__ deprecation in tests (10308) `James Bourbeau`_
Ignore numpy.find_common_type warning from pandas (10307) `James Bourbeau`_
Add test to check that DataFrame.__setitem__ does not modify df inplace (10223) `Patrick Hoefler`_
Clean up default value of dropna in value_counts (10299) `Patrick Hoefler`_
Add pytest-cov to test extra (10271) `James Bourbeau`_
- Adjust for DataFrame.applymap deprecation and all NA concat behaviour change ( dask#10245 ) Patrick Hoefler
Released on May 12, 2023
Implement numeric_only=False for GroupBy.corr and GroupBy.cov ( dask#10264 ) Patrick Hoefler
Add support for numeric_only=False in DataFrame.var ( dask#10250 ) Patrick Hoefler
Add numeric_only support to DataFrame.mode ( dask#10257 ) Patrick Hoefler
Add DataFrame.map to dask.DataFrame API ( dask#10246 ) Patrick Hoefler
Adjust for DataFrame.applymap deprecation and all NA concat behaviour change ( dask#10245 ) Patrick Hoefler
Enable numeric_only=False for DataFrame.count ( dask#10234 ) Patrick Hoefler
Disallow array input in mask/where ( dask#10163 ) Irina Truong
Support numeric_only=True in GroupBy.corr and GroupBy.cov ( dask#10227 ) Patrick Hoefler
Add numeric_only support to GroupBy.median ( dask#10236 ) Patrick Hoefler
Support mimesis=9 in dask.datasets ( dask#10241 ) James Bourbeau
Add numeric_only support to min , max and prod ( dask#10219 ) Patrick Hoefler
Add numeric_only=True support for GroupBy.cumsum and GroupBy.cumprod ( dask#10224 ) Patrick Hoefler
Add helper to unpack numeric_only keyword ( dask#10228 ) Patrick Hoefler
Fix clone + from_array failure ( dask#10211 ) crusaderky
Fix dataframe reductions for ea dtypes ( dask#10150 ) Patrick Hoefler
Avoid scalar conversion deprecation warning in numpy=1.25 ( dask#10248 ) James Bourbeau
Make sure transform output has the same index as input ( dask#10184 ) Irina Truong
Fix corr and cov on a single-row partition ( dask#9756 ) Irina Truong
Fix test_groupby_numeric_only_supported and test_groupby_aggregate_categorical_observed upstream errors ( dask#10243 ) Irina Truong
Clean up futures docs ( dask#10266 ) Matthew Rocklin
Add Index API reference ( dask#10263 ) hotpotato
Warn when meta is passed to apply ( dask#10256 ) Patrick Hoefler
Remove imageio version restriction in CI ( dask#10260 ) Patrick Hoefler
Remove unused DataFrame variance methods ( dask#10252 ) Patrick Hoefler
Un- xfail test_categories with pyarrow strings and pyarrow>=12 ( dask#10244 ) Irina Truong
Bump gpuCI PYTHON_VER 3.8->3.9 ( dask#10233 ) Charles Blackmon-Luca
Released on May 12, 2023
Implement numeric_only=False for GroupBy.corr and GroupBy.cov (10264) `Patrick Hoefler`_
Add support for numeric_only=False in DataFrame.var (10250) `Patrick Hoefler`_
Add numeric_only support to DataFrame.mode (10257) `Patrick Hoefler`_
Add DataFrame.map to dask.DataFrame API (10246) `Patrick Hoefler`_
Adjust for DataFrame.applymap deprecation and all NA concat behaviour change (10245) `Patrick Hoefler`_
Enable numeric_only=False for DataFrame.count (10234) `Patrick Hoefler`_
Disallow array input in mask/where (10163) `Irina Truong`_
Support numeric_only=True in GroupBy.corr and GroupBy.cov (10227) `Patrick Hoefler`_
Add numeric_only support to GroupBy.median (10236) `Patrick Hoefler`_
Support mimesis=9 in dask.datasets (10241) `James Bourbeau`_
Add numeric_only support to min, max and prod (10219) `Patrick Hoefler`_
Add numeric_only=True support for GroupBy.cumsum and GroupBy.cumprod (10224) `Patrick Hoefler`_
Add helper to unpack numeric_only keyword (10228) `Patrick Hoefler`_
Fix clone + from_array failure (10211) `crusaderky`_
Fix dataframe reductions for ea dtypes (10150) `Patrick Hoefler`_
Avoid scalar conversion deprecation warning in numpy=1.25 (10248) `James Bourbeau`_
Make sure transform output has the same index as input (10184) `Irina Truong`_
Fix corr and cov on a single-row partition (9756) `Irina Truong`_
Fix test_groupby_numeric_only_supported and test_groupby_aggregate_categorical_observed upstream errors (10243) `Irina Truong`_
Clean up futures docs (10266) `Matthew Rocklin`_
Add Index API reference (10263) `hotpotato`_
Warn when meta is passed to apply (10256) `Patrick Hoefler`_
Remove imageio version restriction in CI (10260) `Patrick Hoefler`_
Remove unused DataFrame variance methods (10252) `Patrick Hoefler`_
Un-xfail test_categories with pyarrow strings and pyarrow>=12 (10244) `Irina Truong`_
Bump gpuCI PYTHON_VER 3.8->3.9 (10233) `Charles Blackmon-Luca`_
- Avoid deprecated is_categorical_dtype from pandas ( dask#10180 ) Patrick Hoefler
Released on April 28, 2023
Implement numeric_only support for DataFrame.sum ( dask#10194 ) Patrick Hoefler
Add support for numeric_only=True in GroupBy operations ( dask#10222 ) Patrick Hoefler
Avoid deep copy in DataFrame.setitem for pandas 1.4 and up ( dask#10221 ) Patrick Hoefler
Avoid calling Series.apply with _meta_nonempty ( dask#10212 ) Patrick Hoefler
Unpin sqlalchemy and fix compatibility issues ( dask#10140 ) Patrick Hoefler
Partially revert default client discovery ( dask#10225 ) Florian Jetter
Support arrow dtypes in Index meta creation ( dask#10170 ) Patrick Hoefler
Repartitioning raises with extension dtype when truncating floats ( dask#10169 ) Patrick Hoefler
Adjust empty Index from fastparquet to object dtype ( dask#10179 ) Patrick Hoefler
Update Kubernetes docs ( dask#10232 ) Jacob Tomlinson
Add DataFrame.reduction to API docs ( dask#10229 ) James Bourbeau
Add DataFrame.persist to docs and fix links ( dask#10231 ) Patrick Hoefler
Add documentation for GroupBy.transform ( dask#10185 ) Irina Truong
Fix formatting in random number generation docs ( dask#10189 ) Eray Aslan
Pin imageio to <2.28 ( dask#10216 ) Patrick Hoefler
Add note about importlib_metadata backport ( dask#10207 ) James Bourbeau
Add xarray back to Python 3.11 CI builds ( dask#10200 ) James Bourbeau
Add mindeps build with all optional dependencies ( dask#10161 ) Charles Blackmon-Luca
Provide proper like value for array_safe in percentiles_summary ( dask#10156 ) Charles Blackmon-Luca
Avoid re-opening hdf file multiple times in read_hdf ( dask#10205 ) Thomas Grainger
Add merge tests on nullable columns ( dask#10071 ) Charles Blackmon-Luca
Fix coverage configuration ( dask#10203 ) Thomas Grainger
Remove is_period_dtype and is_sparse_dtype ( dask#10197 ) Patrick Hoefler
Bump actions/checkout from 3.5.0 to 3.5.2 ( dask#10201 )
Avoid deprecated is_categorical_dtype from pandas ( dask#10180 ) Patrick Hoefler
Adjust for deprecated is_interval_dtype and is_datetime64tz_dtype ( dask#10188 ) Patrick Hoefler
Released on April 28, 2023
Implement numeric_only support for DataFrame.sum (10194) `Patrick Hoefler`_
Add support for numeric_only=True in GroupBy operations (10222) `Patrick Hoefler`_
Avoid deep copy in DataFrame.__setitem__ for pandas 1.4 and up (10221) `Patrick Hoefler`_
Avoid calling Series.apply with _meta_nonempty (10212) `Patrick Hoefler`_
Unpin sqlalchemy and fix compatibility issues (10140) `Patrick Hoefler`_
Partially revert default client discovery (10225) `Florian Jetter`_
Support arrow dtypes in Index meta creation (10170) `Patrick Hoefler`_
Repartitioning raises with extension dtype when truncating floats (10169) `Patrick Hoefler`_
Adjust empty Index from fastparquet to object dtype (10179) `Patrick Hoefler`_
Update Kubernetes docs (10232) `Jacob Tomlinson`_
Add DataFrame.reduction to API docs (10229) `James Bourbeau`_
Add DataFrame.persist to docs and fix links (10231) `Patrick Hoefler`_
Add documentation for GroupBy.transform (10185) `Irina Truong`_
Fix formatting in random number generation docs (10189) `Eray Aslan`_
Pin imageio to <2.28 (10216) `Patrick Hoefler`_
Add note about importlib_metadata backport (10207) `James Bourbeau`_
Add xarray back to Python 3.11 CI builds (10200) `James Bourbeau`_
Add mindeps build with all optional dependencies (10161) `Charles Blackmon-Luca`_
Provide proper like value for array_safe in percentiles_summary (10156) `Charles Blackmon-Luca`_
Avoid re-opening hdf file multiple times in read_hdf (10205) `Thomas Grainger`_
Add merge tests on nullable columns (10071) `Charles Blackmon-Luca`_
Fix coverage configuration (10203) `Thomas Grainger`_
Remove is_period_dtype and is_sparse_dtype (10197) `Patrick Hoefler`_
Bump actions/checkout from 3.5.0 to 3.5.2 (10201)
Avoid deprecated is_categorical_dtype from pandas (10180) `Patrick Hoefler`_
Adjust for deprecated is_interval_dtype and is_datetime64tz_dtype (10188) `Patrick Hoefler`_
- Avoid deprecated GroupBy.dtypes ( dask#10111 ) Irina Truong
Released on April 14, 2023
Override old default values in update_defaults ( dask#10159 ) Gabe Joseph
Add a CLI command to list and get a value from dask config ( dask#9936 ) Irina Truong
Handle string-based engine argument to read_json ( dask#9947 ) Richard (Rick) Zamora
Avoid deprecated GroupBy.dtypes ( dask#10111 ) Irina Truong
Revert grouper -related changes ( dask#10182 ) Irina Truong
GroupBy.cov raising for non-numeric grouping column ( dask#10171 ) Patrick Hoefler
Updates for Index supporting numpy numeric dtypes ( dask#10154 ) Irina Truong
Preserve dtype for partitioning columns when read with pyarrow ( dask#10115 ) Patrick Hoefler
Fix annotations for to_hdf ( dask#10123 ) Hendrik Makait
Handle None column name when checking if columns are all numeric ( dask#10128 ) Lawrence Mitchell
Fix valid_divisions when passed a tuple ( dask#10126 ) Brian Phillips
Maintain annotations in DataFrame.categorize ( dask#10120 ) Hendrik Makait
Fix handling of missing min/max parquet statistics during filtering ( dask#10042 ) Richard (Rick) Zamora
Deprecate use_nullable_dtypes= and add dtype_backend= ( dask#10076 ) Irina Truong
Deprecate convert_dtype in Series.apply ( dask#10133 ) Irina Truong
Document Generator based random number generation ( dask#10134 ) Eray Aslan
Update dataframe.convert_string to dataframe.convert-string ( dask#10191 ) Irina Truong
Add python-cityhash to CI environments ( dask#10190 ) Charles Blackmon-Luca
Temporarily pin scikit-image to fix Windows CI ( dask#10186 ) Patrick Hoefler
Handle pandas deprecation warnings for to_pydatetime and apply ( dask#10168 ) Patrick Hoefler
Drop bokeh<3 restriction ( dask#10177 ) James Bourbeau
Fix failing tests under copy-on-write ( dask#10173 ) Patrick Hoefler
Allow pyarrow CI to fail ( dask#10176 ) James Bourbeau
Switch to Generator for random number generation in dask.array ( dask#10003 ) Eray Aslan
Bump peter-evans/create-pull-request from 4 to 5 ( dask#10166 )
Fix flaky modf operation in test_arithmetic ( dask#10162 ) Irina Truong
Temporarily remove xarray from CI with pandas 2.0 ( dask#10153 ) James Bourbeau
Fix update_graph counting logic in test_default_scheduler_on_worker ( dask#10145 ) James Bourbeau
Fix documentation build with pandas 2.0 ( dask#10138 ) James Bourbeau
Remove dask/gpu from gpuCI update reviewers ( dask#10135 ) Charles Blackmon-Luca
Update gpuCI RAPIDS_VER to 23.06 ( dask#10129 )
Bump actions/stale from 6 to 8 ( dask#10121 )
Use declarative setuptools ( dask#10102 ) Thomas Grainger
Relax assert_eq checks on Scalar -like objects ( dask#10125 ) Matthew Rocklin
Upgrade readthedocs config to ubuntu 22.04 and Python 3.11 ( dask#10124 ) Thomas Grainger
Bump actions/checkout from 3.4.0 to 3.5.0 ( dask#10122 )
Fix test_null_partition_pyarrow in pyarrow CI build ( dask#10116 ) Irina Truong
Drop distributed pack ( dask#9988 ) Florian Jetter
Make dask.compatibility private ( dask#10114 ) Jacob Tomlinson
Released on April 14, 2023
Override old default values in update_defaults (10159) `Gabe Joseph`_
Add a CLI command to list and get a value from dask config (9936) `Irina Truong`_
Handle string-based engine argument to read_json (9947) `Richard (Rick) Zamora`_
Avoid deprecated GroupBy.dtypes (10111) `Irina Truong`_
Revert grouper-related changes (10182) `Irina Truong`_
GroupBy.cov raising for non-numeric grouping column (10171) `Patrick Hoefler`_
Updates for Index supporting numpy numeric dtypes (10154) `Irina Truong`_
Preserve dtype for partitioning columns when read with pyarrow (10115) `Patrick Hoefler`_
Fix annotations for to_hdf (10123) `Hendrik Makait`_
Handle None column name when checking if columns are all numeric (10128) `Lawrence Mitchell`_
Fix valid_divisions when passed a tuple (10126) `Brian Phillips`_
Maintain annotations in DataFrame.categorize (10120) `Hendrik Makait`_
Fix handling of missing min/max parquet statistics during filtering (10042) `Richard (Rick) Zamora`_
Deprecate use_nullable_dtypes= and add dtype_backend= (10076) `Irina Truong`_
Deprecate convert_dtype in Series.apply (10133) `Irina Truong`_
Document Generator based random number generation (10134) `Eray Aslan`_
Update dataframe.convert_string to dataframe.convert-string (10191) `Irina Truong`_
Add python-cityhash to CI environments (10190) `Charles Blackmon-Luca`_
Temporarily pin scikit-image to fix Windows CI (10186) `Patrick Hoefler`_
Handle pandas deprecation warnings for to_pydatetime and apply (10168) `Patrick Hoefler`_
Drop bokeh<3 restriction (10177) `James Bourbeau`_
Fix failing tests under copy-on-write (10173) `Patrick Hoefler`_
Allow pyarrow CI to fail (10176) `James Bourbeau`_
Switch to Generator for random number generation in dask.array (10003) `Eray Aslan`_
Bump peter-evans/create-pull-request from 4 to 5 (10166)
Fix flaky modf operation in test_arithmetic (10162) `Irina Truong`_
Temporarily remove xarray from CI with pandas 2.0 (10153) `James Bourbeau`_
Fix update_graph counting logic in test_default_scheduler_on_worker (10145) `James Bourbeau`_
Fix documentation build with pandas 2.0 (10138) `James Bourbeau`_
Remove dask/gpu from gpuCI update reviewers (10135) `Charles Blackmon-Luca`_
Update gpuCI RAPIDS_VER to 23.06 (10129)
Bump actions/stale from 6 to 8 (10121)
Use declarative setuptools (10102) `Thomas Grainger`_
Relax assert_eq checks on Scalar-like objects (10125) `Matthew Rocklin`_
Upgrade readthedocs config to ubuntu 22.04 and Python 3.11 (10124) `Thomas Grainger`_
Bump actions/checkout from 3.4.0 to 3.5.0 (10122)
Fix test_null_partition_pyarrow in pyarrow CI build (10116) `Irina Truong`_
Drop distributed pack (9988) `Florian Jetter`_
Make dask.compatibility private (10114) `Jacob Tomlinson`_
- Deprecate observed=False for groupby with categoricals ( dask#10095 ) Irina Truong
Released on March 24, 2023
Deprecate observed=False for groupby with categoricals ( dask#10095 ) Irina Truong
Deprecate axis= for some groupby operations ( dask#10094 ) James Bourbeau
The axis keyword in DataFrame.rolling/Series.rolling is deprecated ( dask#10110 ) Irina Truong
DataFrame._data deprecation in pandas ( dask#10081 ) Irina Truong
Use importlib_metadata backport to avoid CLI UserWarning ( dask#10070 ) Thomas Grainger
Port option parsing logic from dask.dataframe.read_parquet to to_parquet ( dask#9981 ) Anton Loukianov
Avoid using dd.shuffle in groupby-apply ( dask#10043 ) Richard (Rick) Zamora
Enable null hive partitions with pyarrow parquet engine ( dask#10007 ) Richard (Rick) Zamora
Support unknown shapes in *_like functions ( dask#10064 ) Doug Davis
Add to_backend methods to API docs ( dask#10093 ) Lawrence Mitchell
Remove broken gpuCI link in developer docs ( dask#10065 ) Charles Blackmon-Luca
Configure readthedocs sphinx warnings as errors ( dask#10104 ) Thomas Grainger
Un- xfail test_division_or_partition with pyarrow strings active ( dask#10108 ) Irina Truong
Un- xfail test_different_columns_are_allowed with pyarrow strings active ( dask#10109 ) Irina Truong
Restore Entrypoints compatibility ( dask#10113 ) Jacob Tomlinson
Un- xfail test_to_dataframe_optimize_graph with pyarrow strings active ( dask#10087 ) Irina Truong
Only run test_development_guidelines_matches_ci on editable install ( dask#10106 ) Charles Blackmon-Luca
Un- xfail test_dataframe_cull_key_dependencies_materialized with pyarrow strings active ( dask#10088 ) Irina Truong
Install mimesis in CI environments ( dask#10105 ) Charles Blackmon-Luca
Fix for no module named ipykernel ( dask#10101 ) Irina Truong
Fix docs builds by installing ipykernel ( dask#10103 ) Thomas Grainger
Allow pyarrow build to continue on failures ( dask#10097 ) James Bourbeau
Bump actions/checkout from 3.3.0 to 3.4.0 ( dask#10096 )
Fix test_set_index_on_empty with pyarrow strings active ( dask#10054 ) Irina Truong
Un- xfail pyarrow pickling tests ( dask#10082 ) James Bourbeau
CI environment file cleanup ( dask#10078 ) James Bourbeau
Un- xfail more pyarrow tests ( dask#10066 ) Irina Truong
Temporarily skip pyarrow_compat tests with p`andas 2.0 ( dask#10063 ) James Bourbeau
Fix test_melt with pyarrow strings active ( dask#10052 ) Irina Truong
Fix test_str_accessor with pyarrow strings active ( dask#10048 ) James Bourbeau
Fix test_better_errors_object_reductions with pyarrow strings active ( dask#10051 ) James Bourbeau
Fix test_loc_with_non_boolean_series with pyarrow strings active ( dask#10046 ) James Bourbeau
Fix test_values with pyarrow strings active ( dask#10050 ) James Bourbeau
Temporarily xfail test_upstream_packages_installed ( dask#10047 ) James Bourbeau
Released on March 24, 2023
Deprecate observed=False for groupby with categoricals (10095) `Irina Truong`_
Deprecate axis= for some groupby operations (10094) `James Bourbeau`_
The axis keyword in DataFrame.rolling/Series.rolling is deprecated (10110) `Irina Truong`_
DataFrame._data deprecation in pandas (10081) `Irina Truong`_
Use importlib_metadata backport to avoid CLI UserWarning (10070) `Thomas Grainger`_
Port option parsing logic from dask.dataframe.read_parquet to to_parquet (9981) `Anton Loukianov`_
Avoid using dd.shuffle in groupby-apply (10043) `Richard (Rick) Zamora`_
Enable null hive partitions with pyarrow parquet engine (10007) `Richard (Rick) Zamora`_
Support unknown shapes in *_like functions (10064) `Doug Davis`_
Add to_backend methods to API docs (10093) `Lawrence Mitchell`_
Remove broken gpuCI link in developer docs (10065) `Charles Blackmon-Luca`_
Configure readthedocs sphinx warnings as errors (10104) `Thomas Grainger`_
Un-xfail test_division_or_partition with pyarrow strings active (10108) `Irina Truong`_
Un-xfail test_different_columns_are_allowed with pyarrow strings active (10109) `Irina Truong`_
Restore Entrypoints compatibility (10113) `Jacob Tomlinson`_
Un-xfail test_to_dataframe_optimize_graph with pyarrow strings active (10087) `Irina Truong`_
Only run test_development_guidelines_matches_ci on editable install (10106) `Charles Blackmon-Luca`_
Un-xfail test_dataframe_cull_key_dependencies_materialized with pyarrow strings active (10088) `Irina Truong`_
Install mimesis in CI environments (10105) `Charles Blackmon-Luca`_
Fix for no module named ipykernel (10101) `Irina Truong`_
Fix docs builds by installing ipykernel (10103) `Thomas Grainger`_
Allow pyarrow build to continue on failures (10097) `James Bourbeau`_
Bump actions/checkout from 3.3.0 to 3.4.0 (10096)
Fix test_set_index_on_empty with pyarrow strings active (10054) `Irina Truong`_
Un-xfail pyarrow pickling tests (10082) `James Bourbeau`_
CI environment file cleanup (10078) `James Bourbeau`_
Un-xfail more pyarrow tests (10066) `Irina Truong`_
Temporarily skip pyarrow_compat tests with p`andas 2.0 (10063) `James Bourbeau`_
Fix test_melt with pyarrow strings active (10052) `Irina Truong`_
Fix test_str_accessor with pyarrow strings active (10048) `James Bourbeau`_
Fix test_better_errors_object_reductions with pyarrow strings active (10051) `James Bourbeau`_
Fix test_loc_with_non_boolean_series with pyarrow strings active (10046) `James Bourbeau`_
Fix test_values with pyarrow strings active (10050) `James Bourbeau`_
Temporarily xfail test_upstream_packages_installed (10047) `James Bourbeau`_
- Support pyarrow strings in MultiIndex ( dask#10040 ) Irina Truong
Released on March 10, 2023
Support pyarrow strings in MultiIndex ( dask#10040 ) Irina Truong
Improved support for pyarrow strings ( dask#10000 ) Irina Truong
Fix flaky RuntimeWarning during array reductions ( dask#10030 ) James Bourbeau
Extend complete extras ( dask#10023 ) James Bourbeau
Raise an error with dataframe.convert-string=True and pandas<2.0 ( dask#10033 ) Irina Truong
Rename shuffle/rechunk config option/kwarg to method ( dask#10013 ) James Bourbeau
Add initial support for converting pandas extension dtypes to arrays ( dask#10018 ) James Bourbeau
Remove randomgen support ( dask#9987 ) Eray Aslan
Skip rechunk when rechunking to the same chunks with unknown sizes ( dask#10027 ) Hendrik Makait
Custom utility to convert parquet filters to pyarrow expression ( dask#9885 ) Richard (Rick) Zamora
Consider numpy scalars and 0d arrays as scalars when padding ( dask#9653 ) Justus Magin
Fix parquet overwrite behavior after an adaptive read_parquet operation ( dask#10002 ) Richard (Rick) Zamora
Add and update docs for Data Transfer section ( dask#10022 ) Miles
Remove stale hive-partitioning code from pyarrow parquet engine ( dask#10039 ) Richard (Rick) Zamora
Increase minimum supported pyarrow to 7.0 ( dask#10024 ) James Bourbeau
Revert “Prepare drop packunpack ( dask#9994 ) ( dask#10037 ) Florian Jetter
Have codecov wait for more builds before reporting ( dask#10031 ) James Bourbeau
Prepare drop packunpack ( dask#9994 ) Florian Jetter
Add CI job with pyarrow strings turned on ( dask#10017 ) James Bourbeau
Fix test_groupby_dropna_with_agg for pandas 2.0 ( dask#10001 ) Irina Truong
Fix test_pickle_roundtrip for pandas 2.0 ( dask#10011 ) James Bourbeau
Released on March 10, 2023
Support pyarrow strings in MultiIndex (10040) `Irina Truong`_
Improved support for pyarrow strings (10000) `Irina Truong`_
Fix flaky RuntimeWarning during array reductions (10030) `James Bourbeau`_
Extend complete extras (10023) `James Bourbeau`_
Raise an error with dataframe.convert-string=True and pandas<2.0 (10033) `Irina Truong`_
Rename shuffle/rechunk config option/kwarg to method (10013) `James Bourbeau`_
Add initial support for converting pandas extension dtypes to arrays (10018) `James Bourbeau`_
Remove randomgen support (9987) `Eray Aslan`_
Skip rechunk when rechunking to the same chunks with unknown sizes (10027) `Hendrik Makait`_
Custom utility to convert parquet filters to pyarrow expression (9885) `Richard (Rick) Zamora`_
Consider numpy scalars and 0d arrays as scalars when padding (9653) `Justus Magin`_
Fix parquet overwrite behavior after an adaptive read_parquet operation (10002) `Richard (Rick) Zamora`_
Add and update docs for Data Transfer section (10022) `Miles`_
Remove stale hive-partitioning code from pyarrow parquet engine (10039) `Richard (Rick) Zamora`_
Increase minimum supported pyarrow to 7.0 (10024) `James Bourbeau`_
Revert "Prepare drop packunpack (9994) (10037) `Florian Jetter`_
Have codecov wait for more builds before reporting (10031) `James Bourbeau`_
Prepare drop packunpack (9994) `Florian Jetter`_
Add CI job with pyarrow strings turned on (10017) `James Bourbeau`_
Fix test_groupby_dropna_with_agg for pandas 2.0 (10001) `Irina Truong`_
Fix test_pickle_roundtrip for pandas 2.0 (10011) `James Bourbeau`_
- Bag must not pick p2p as shuffle default ( dask#10005 ) Florian Jetter
Released on March 1, 2023
Bag must not pick p2p as shuffle default ( dask#10005 ) Florian Jetter
Minor follow-up to P2P by default ( dask#10008 ) James Bourbeau
Add minimum version to optional jinja2 dependency ( dask#9999 ) Charles Blackmon-Luca
Released on March 1, 2023
Bag must not pick p2p as shuffle default (10005) `Florian Jetter`_
Minor follow-up to P2P by default (10008) `James Bourbeau`_
Add minimum version to optional jinja2 dependency (9999) `Charles Blackmon-Luca`_
This release changes the default DataFrame shuffle algorithm to p2p to improve stability and performance. Learn more here and please provide any feedb
Released on February 24, 2023
Note
This release changes the default DataFrame shuffle algorithm to p2p to improve stability and performance. Learn more here and please provide any feedback on this discussion .
If you encounter issues with this new algorithm, please see the documentation for more information, and how to switch back to the old mode.
Enable P2P shuffling by default ( dask#9991 ) Florian Jetter
P2P rechunking ( dask#9939 ) Hendrik Makait
Efficient dataframe.convert-string support for read_parquet ( dask#9979 ) Irina Truong
Allow p2p shuffle kwarg for DataFrame merges ( dask#9900 ) Florian Jetter
Change split_row_groups default to “infer” ( dask#9637 ) Richard (Rick) Zamora
Add option for converting string data to use pyarrow strings ( dask#9926 ) James Bourbeau
Add support for multi-column sort_values ( dask#8263 ) Charles Blackmon-Luca
Generator based random-number generation indask.array ( dask#9038 ) Eray Aslan
Support numeric_only for simple groupby aggregations for pandas 2.0 compatibility ( dask#9889 ) Irina Truong
Fix profilers plot not being aligned to context manager enter time ( dask#9739 ) David Hoese
Relax dask.dataframe assert_eq type checks ( dask#9989 ) Matthew Rocklin
Restore describe compatibility for pandas 2.0 ( dask#9982 ) James Bourbeau
Improving deploying Dask docs ( dask#9912 ) Sarah Charlotte Johnson
More docs for DataFrame.partitions ( dask#9976 ) Tom Augspurger
Update docs with more information on default Delayed scheduler ( dask#9903 ) Guillaume Eynard-Bontemps
Deployment Considerations documentation ( dask#9933 ) Gabe Joseph
Temporarily rerun flaky tests ( dask#9983 ) James Bourbeau
Update parsing of FULL_RAPIDS_VER/FULL_UCX_PY_VER ( dask#9990 ) Charles Blackmon-Luca
Increase minimum supported versions to pandas=1.3 and numpy=1.21 ( dask#9950 ) James Bourbeau
Fix std to work with numeric_only for pandas 2.0 ( dask#9960 ) Irina Truong
Temporarily xfail test_roundtrip_partitioned_pyarrow_dataset ( dask#9977 ) James Bourbeau
Fix copy on write failure in test_idxmaxmin ( dask#9944 ) Patrick Hoefler
Bump pre-commit versions ( dask#9955 ) crusaderky
Fix test_groupby_unaligned_index for pandas 2.0 ( dask#9963 ) Irina Truong
Un- xfail test_set_index_overlap_2 for pandas 2.0 ( dask#9959 ) James Bourbeau
Fix test_merge_by_index_patterns for pandas 2.0 ( dask#9930 ) Irina Truong
Bump jacobtomlinson/gha-find-replace from 2 to 3 ( dask#9953 ) James Bourbeau
Fix test_rolling_agg_aggregate for pandas 2.0 compatibility ( dask#9948 ) Irina Truong
Bump black to 23.1.0 ( dask#9956 ) crusaderky
Run GPU tests on python 3.8 & 3.10 ( dask#9940 ) Charles Blackmon-Luca
Fix test_to_timestamp for pandas 2.0 ( dask#9932 ) Irina Truong
Fix an error with groupby value_counts for pandas 2.0 compatibility ( dask#9928 ) Irina Truong
Config converter: replace all dashes with underscores ( dask#9945 ) Jacob Tomlinson
CI: use nightly wheel to install pyarrow in upstream test build ( dask#9873 ) Joris Van den Bossche
Released on February 24, 2023
Note
This release changes the default DataFrame shuffle algorithm to p2p to improve stability and performance. Learn more here and please provide any feedback on this discussion.
If you encounter issues with this new algorithm, please see the documentation for more information, and how to switch back to the old mode.
Enable P2P shuffling by default (9991) `Florian Jetter`_
P2P rechunking (9939) `Hendrik Makait`_
Efficient dataframe.convert-string support for read_parquet (9979) `Irina Truong`_
Allow p2p shuffle kwarg for DataFrame merges (9900) `Florian Jetter`_
Change split_row_groups default to "infer" (9637) `Richard (Rick) Zamora`_
Add option for converting string data to use pyarrow strings (9926) `James Bourbeau`_
Add support for multi-column sort_values (8263) `Charles Blackmon-Luca`_
Generator based random-number generation in``dask.array`` (9038) `Eray Aslan`_
Support numeric_only for simple groupby aggregations for pandas 2.0 compatibility (9889) `Irina Truong`_
Fix profilers plot not being aligned to context manager enter time (9739) `David Hoese`_
Relax dask.dataframe assert_eq type checks (9989) `Matthew Rocklin`_
Restore describe compatibility for pandas 2.0 (9982) `James Bourbeau`_
Improving deploying Dask docs (9912) `Sarah Charlotte Johnson`_
More docs for DataFrame.partitions (9976) `Tom Augspurger`_
Update docs with more information on default Delayed scheduler (9903) `Guillaume Eynard-Bontemps`_
Deployment Considerations documentation (9933) `Gabe Joseph`_
Temporarily rerun flaky tests (9983) `James Bourbeau`_
Update parsing of FULL_RAPIDS_VER/FULL_UCX_PY_VER (9990) `Charles Blackmon-Luca`_
Increase minimum supported versions to pandas=1.3 and numpy=1.21 (9950) `James Bourbeau`_
Fix std to work with numeric_only for pandas 2.0 (9960) `Irina Truong`_
Temporarily xfail test_roundtrip_partitioned_pyarrow_dataset (9977) `James Bourbeau`_
Fix copy on write failure in test_idxmaxmin (9944) `Patrick Hoefler`_
Bump pre-commit versions (9955) `crusaderky`_
Fix test_groupby_unaligned_index for pandas 2.0 (9963) `Irina Truong`_
Un-xfail test_set_index_overlap_2 for pandas 2.0 (9959) `James Bourbeau`_
Fix test_merge_by_index_patterns for pandas 2.0 (9930) `Irina Truong`_
Bump jacobtomlinson/gha-find-replace from 2 to 3 (9953) `James Bourbeau`_
Fix test_rolling_agg_aggregate for pandas 2.0 compatibility (9948) `Irina Truong`_
Bump black to 23.1.0 (9956) `crusaderky`_
Run GPU tests on python 3.8 & 3.10 (9940) `Charles Blackmon-Luca`_
Fix test_to_timestamp for pandas 2.0 (9932) `Irina Truong`_
Fix an error with groupby value_counts for pandas 2.0 compatibility (9928) `Irina Truong`_
Config converter: replace all dashes with underscores (9945) `Jacob Tomlinson`_
CI: use nightly wheel to install pyarrow in upstream test build (9873) `Joris Van den Bossche`_
- Update numeric_only default in quantile for pandas 2.0 ( dask#9854 ) Irina Truong
Released on February 10, 2023
Update numeric_only default in quantile for pandas 2.0 ( dask#9854 ) Irina Truong
Make repartition a no-op when divisions match ( dask#9924 ) James Bourbeau
Update datetime_is_numeric behavior in describe for pandas 2.0 ( dask#9868 ) Irina Truong
Update value_counts to return correct name in pandas 2.0 ( dask#9919 ) Irina Truong
Support new axis=None behavior in pandas 2.0 for certain reductions ( dask#9867 ) James Bourbeau
Filter out all-nan RuntimeWarning at the chunk level for nanmin and nanmax ( dask#9916 ) Julia Signell
Fix numeric meta_nonempty index creation for pandas 2.0 ( dask#9908 ) James Bourbeau
Fix DataFrame.info() tests for pandas 2.0 ( dask#9909 ) James Bourbeau
Fix GroupBy.value_counts handling for multiple groupby columns ( dask#9905 ) Charles Blackmon-Luca
Fix some outdated information/typos in development guide ( dask#9893 ) Patrick Hoefler
Add note about keep=False in drop_duplicates docstring ( dask#9887 ) Jayesh Manani
Add meta details to dask Array ( dask#9886 ) Jayesh Manani
Clarify task stream showing more rows than threads ( dask#9906 ) Gabe Joseph
Fix test_numeric_column_names for pandas 2.0 ( dask#9937 ) Irina Truong
Fix dask/dataframe/tests/test_utils_dataframe.py tests for pandas 2.0 ( dask#9788 ) James Bourbeau
Replace index.is_numeric with is_any_real_numeric_dtype for pandas 2.0 compatibility ( dask#9918 ) Irina Truong
Avoid pd.core import in dask utils ( dask#9907 ) Matthew Roeschke
Use label for upstream build on pull requests ( dask#9910 ) James Bourbeau
Broaden exception catching for sqlalchemy.exc.RemovedIn20Warning ( dask#9904 ) James Bourbeau
Temporarily restrict sqlalchemy < 2 in CI ( dask#9897 ) James Bourbeau
Update isort version to 5.12.0 ( dask#9895 ) Lawrence Mitchell
Remove unused skiprows variable in read_csv ( dask#9892 ) Patrick Hoefler
Released on February 10, 2023
Update numeric_only default in quantile for pandas 2.0 (9854) `Irina Truong`_
Make repartition a no-op when divisions match (9924) `James Bourbeau`_
Update datetime_is_numeric behavior in describe for pandas 2.0 (9868) `Irina Truong`_
Update value_counts to return correct name in pandas 2.0 (9919) `Irina Truong`_
Support new axis=None behavior in pandas 2.0 for certain reductions (9867) `James Bourbeau`_
Filter out all-nan RuntimeWarning at the chunk level for nanmin and nanmax (9916) `Julia Signell`_
Fix numeric meta_nonempty index creation for pandas 2.0 (9908) `James Bourbeau`_
Fix DataFrame.info() tests for pandas 2.0 (9909) `James Bourbeau`_
Fix GroupBy.value_counts handling for multiple groupby columns (9905) `Charles Blackmon-Luca`_
Fix some outdated information/typos in development guide (9893) `Patrick Hoefler`_
Add note about keep=False in drop_duplicates docstring (9887) `Jayesh Manani`_
Add meta details to dask Array (9886) `Jayesh Manani`_
Clarify task stream showing more rows than threads (9906) `Gabe Joseph`_
Fix test_numeric_column_names for pandas 2.0 (9937) `Irina Truong`_
Fix dask/dataframe/tests/test_utils_dataframe.py tests for pandas 2.0 (9788) `James Bourbeau`_
Replace index.is_numeric with is_any_real_numeric_dtype for pandas 2.0 compatibility (9918) `Irina Truong`_
Avoid pd.core import in dask utils (9907) `Matthew Roeschke`_
Use label for upstream build on pull requests (9910) `James Bourbeau`_
Broaden exception catching for sqlalchemy.exc.RemovedIn20Warning (9904) `James Bourbeau`_
Temporarily restrict sqlalchemy < 2 in CI (9897) `James Bourbeau`_
Update isort version to 5.12.0 (9895) `Lawrence Mitchell`_
Remove unused skiprows variable in read_csv (9892) `Patrick Hoefler`_
- Add to_backend method to Array and _Frame ( dask#9758 ) Richard (Rick) Zamora
Released on January 27, 2023
Add to_backend method to Array and _Frame ( dask#9758 ) Richard (Rick) Zamora
Small fix for timestamp index divisions in pandas 2.0 ( dask#9872 ) Irina Truong
Add numeric_only to DataFrame.cov and DataFrame.corr ( dask#9787 ) James Bourbeau
Fixes related to group_keys default change in pandas 2.0 ( dask#9855 ) Irina Truong
infer_datetime_format compatibility for pandas 2.0 ( dask#9783 ) James Bourbeau
Fix serialization bug in BroadcastJoinLayer ( dask#9871 ) Richard (Rick) Zamora
Satisfy broadcast argument in DataFrame.merge ( dask#9852 ) Richard (Rick) Zamora
Fix pyarrow parquet columns statistics computation ( dask#9772 ) aywandji
Fix “duplicate explicit target name” docs warning ( dask#9863 ) Chiara Marmo
Fix code formatting issue in “Defining a new collection backend” docs ( dask#9864 ) Chiara Marmo
Update dashboard documentation for memory plot ( dask#9768 ) Jayesh Manani
Add docs section about no-worker tasks ( dask#9839 ) Florian Jetter
Additional updates for detecting a distributed scheduler ( dask#9890 ) James Bourbeau
Update gpuCI RAPIDS_VER to 23.04 ( dask#9876 )
Reverse precedence between collection and distributed default ( dask#9869 ) Florian Jetter
Update xarray-contrib/issue-from-pytest-log to version 1.2.6 ( dask#9865 ) James Bourbeau
Dont require dask config shuffle default ( dask#9826 ) Florian Jetter
Un- xfail datetime64 Parquet roundtripping tests for new fastparquet ( dask#9811 ) James Bourbeau
Add option to manually run upstream CI build ( dask#9853 ) James Bourbeau
Use custom timeout in CI builds ( dask#9844 ) James Bourbeau
Remove kwargs from make_blockwise_graph ( dask#9838 ) Florian Jetter
Ignore warnings on persist call in test_setitem_extended_API_2d_mask ( dask#9843 ) Charles Blackmon-Luca
Fix running S3 tests locally ( dask#9833 ) James Bourbeau
Released on January 27, 2023
Add to_backend method to Array and _Frame (9758) `Richard (Rick) Zamora`_
Small fix for timestamp index divisions in pandas 2.0 (9872) `Irina Truong`_
Add numeric_only to DataFrame.cov and DataFrame.corr (9787) `James Bourbeau`_
Fixes related to group_keys default change in pandas 2.0 (9855) `Irina Truong`_
infer_datetime_format compatibility for pandas 2.0 (9783) `James Bourbeau`_
Fix serialization bug in BroadcastJoinLayer (9871) `Richard (Rick) Zamora`_
Satisfy broadcast argument in DataFrame.merge (9852) `Richard (Rick) Zamora`_
Fix pyarrow parquet columns statistics computation (9772) `aywandji`_
Fix "duplicate explicit target name" docs warning (9863) `Chiara Marmo`_
Fix code formatting issue in "Defining a new collection backend" docs (9864) `Chiara Marmo`_
Update dashboard documentation for memory plot (9768) `Jayesh Manani`_
Add docs section about no-worker tasks (9839) `Florian Jetter`_
Additional updates for detecting a distributed scheduler (9890) `James Bourbeau`_
Update gpuCI RAPIDS_VER to 23.04 (9876)
Reverse precedence between collection and distributed default (9869) `Florian Jetter`_
Update xarray-contrib/issue-from-pytest-log to version 1.2.6 (9865) `James Bourbeau`_
Dont require dask config shuffle default (9826) `Florian Jetter`_
Un-xfail datetime64 Parquet roundtripping tests for new fastparquet (9811) `James Bourbeau`_
Add option to manually run upstream CI build (9853) `James Bourbeau`_
Use custom timeout in CI builds (9844) `James Bourbeau`_
Remove kwargs from make_blockwise_graph (9838) `Florian Jetter`_
Ignore warnings on persist call in test_setitem_extended_API_2d_mask (9843) `Charles Blackmon-Luca`_
Fix running S3 tests locally (9833) `James Bourbeau`_
- Use distributed default clients even if no config is set ( dask#9808 ) Florian Jetter
Released on January 13, 2023
Use distributed default clients even if no config is set ( dask#9808 ) Florian Jetter
Implement ma.where and ma.nonzero ( dask#9760 ) Erik Holmgren
Update zarr store creation functions ( dask#9790 ) Ryan Abernathey
iteritems compatibility for pandas 2.0 ( dask#9785 ) James Bourbeau
Accurate sizeof for pandas string[python] dtype ( dask#9781 ) crusaderky
Deflate sizeof() of duplicate references to pandas object types ( dask#9776 ) crusaderky
GroupBy.getitem compatibility for pandas 2.0 ( dask#9779 ) James Bourbeau
append compatibility for pandas 2.0 ( dask#9750 ) James Bourbeau
get_dummies compatibility for pandas 2.0 ( dask#9752 ) James Bourbeau
is_monotonic compatibility for pandas 2.0 ( dask#9751 ) James Bourbeau
numpy=1.24 compatability ( dask#9777 ) James Bourbeau
Remove duplicated encoding kwarg in docstring for to_json ( dask#9796 ) Sultan Orazbayev
Mention SubprocessCluster in LocalCluster documentation ( dask#9784 ) Hendrik Makait
Move Prometheus docs to dask/distributed ( dask#9761 ) crusaderky
Temporarily ignore RuntimeWarning in test_setitem_extended_API_2d_mask ( dask#9828 ) James Bourbeau
Fix flaky test_threaded.py::test_interrupt ( dask#9827 ) Hendrik Makait
Update xarray-contrib/issue-from-pytest-log in upstream report ( dask#9822 ) James Bourbeau
pip install dask on gpuCI builds ( dask#9816 ) Charles Blackmon-Luca
Bump actions/checkout from 3.2.0 to 3.3.0 ( dask#9815 )
Resolve sqlalchemy import failures in mindeps testing ( dask#9809 ) Charles Blackmon-Luca
Ignore sqlalchemy.exc.RemovedIn20Warning ( dask#9801 ) Thomas Grainger
xfail datetime64 Parquet roundtripping tests for pandas 2.0 ( dask#9786 ) James Bourbeau
Remove sqlachemy 1.3 compatibility ( dask#9695 ) McToel
Reduce size of expected DoK sparse matrix ( dask#9775 ) Elliott Sales de Andrade
Remove executable flag from dask/dataframe/io/orc/utils.py ( dask#9774 ) Elliott Sales de Andrade
Released on January 13, 2023
Use distributed default clients even if no config is set (9808) `Florian Jetter`_
Implement ma.where and ma.nonzero (9760) `Erik Holmgren`_
Update zarr store creation functions (9790) `Ryan Abernathey`_
iteritems compatibility for pandas 2.0 (9785) `James Bourbeau`_
Accurate sizeof for pandas string[python] dtype (9781) `crusaderky`_
Deflate sizeof() of duplicate references to pandas object types (9776) `crusaderky`_
GroupBy.__getitem__ compatibility for pandas 2.0 (9779) `James Bourbeau`_
append compatibility for pandas 2.0 (9750) `James Bourbeau`_
get_dummies compatibility for pandas 2.0 (9752) `James Bourbeau`_
is_monotonic compatibility for pandas 2.0 (9751) `James Bourbeau`_
numpy=1.24 compatability (9777) `James Bourbeau`_
Remove duplicated encoding kwarg in docstring for to_json (9796) `Sultan Orazbayev`_
Mention SubprocessCluster in LocalCluster documentation (9784) `Hendrik Makait`_
Move Prometheus docs to dask/distributed (9761) `crusaderky`_
Temporarily ignore RuntimeWarning in test_setitem_extended_API_2d_mask (9828) `James Bourbeau`_
Fix flaky test_threaded.py::test_interrupt (9827) `Hendrik Makait`_
Update xarray-contrib/issue-from-pytest-log in upstream report (9822) `James Bourbeau`_
pip install dask on gpuCI builds (9816) `Charles Blackmon-Luca`_
Bump actions/checkout from 3.2.0 to 3.3.0 (9815)
Resolve sqlalchemy import failures in mindeps testing (9809) `Charles Blackmon-Luca`_
Ignore sqlalchemy.exc.RemovedIn20Warning (9801) `Thomas Grainger`_
xfail datetime64 Parquet roundtripping tests for pandas 2.0 (9786) `James Bourbeau`_
Remove sqlachemy 1.3 compatibility (9695) `McToel`_
Reduce size of expected DoK sparse matrix (9775) `Elliott Sales de Andrade`_
Remove executable flag from dask/dataframe/io/orc/utils.py (9774) `Elliott Sales de Andrade`_
- Avoid np.bool8 deprecation warning ( dask#9737 ) James Bourbeau
Released on December 16, 2022
Support dtype_backend="pandas|pyarrow" configuration ( dask#9719 ) James Bourbeau
Support cupy.ndarray to cudf.DataFrame dispatching in dask.dataframe ( dask#9579 ) Richard (Rick) Zamora
Make filesystem-backend configurable in read_parquet ( dask#9699 ) Richard (Rick) Zamora
Serialize all pyarrow extension arrays efficiently ( dask#9740 ) James Bourbeau
Fix bug when repartitioning with tz -aware datetime index ( dask#9741 ) James Bourbeau
Partial functions in aggs may have arguments ( dask#9724 ) Irina Truong
Add support for simple operation with pyarrow -backed extension dtypes ( dask#9717 ) James Bourbeau
Rename columns correctly in case of SeriesGroupby ( dask#9716 ) Lawrence Mitchell
Fix url link typo in collection backend doc ( dask#9748 ) Shawn
Update Prometheus docs ( dask#9696 ) Hendrik Makait
Add zarr to Python 3.11 CI environment ( dask#9771 ) James Bourbeau
Add support for Python 3.11 ( dask#9708 ) Thomas Grainger
Bump actions/checkout from 3.1.0 to 3.2.0 ( dask#9753 )
Avoid np.bool8 deprecation warning ( dask#9737 ) James Bourbeau
Make sure dev packages aren’t overwritten in upstream CI build ( dask#9731 ) James Bourbeau
Avoid adding data.h5 and mydask.html files during tests ( dask#9726 ) Thomas Grainger
Released on December 16, 2022
Support dtype_backend="pandas|pyarrow" configuration (9719) `James Bourbeau`_
Support cupy.ndarray to cudf.DataFrame dispatching in dask.dataframe (9579) `Richard (Rick) Zamora`_
Make filesystem-backend configurable in read_parquet (9699) `Richard (Rick) Zamora`_
Serialize all pyarrow extension arrays efficiently (9740) `James Bourbeau`_
Fix bug when repartitioning with tz-aware datetime index (9741) `James Bourbeau`_
Partial functions in aggs may have arguments (9724) `Irina Truong`_
Add support for simple operation with pyarrow-backed extension dtypes (9717) `James Bourbeau`_
Rename columns correctly in case of SeriesGroupby (9716) `Lawrence Mitchell`_
Fix url link typo in collection backend doc (9748) `Shawn`_
Update Prometheus docs (9696) `Hendrik Makait`_
Add zarr to Python 3.11 CI environment (9771) `James Bourbeau`_
Add support for Python 3.11 (9708) `Thomas Grainger`_
Bump actions/checkout from 3.1.0 to 3.2.0 (9753)
Avoid np.bool8 deprecation warning (9737) `James Bourbeau`_
Make sure dev packages aren't overwritten in upstream CI build (9731) `James Bourbeau`_
Avoid adding data.h5 and mydask.html files during tests (9726) `Thomas Grainger`_
- Remove statistics-based set_index logic from read_parquet ( dask#9661 ) Richard (Rick) Zamora
Released on December 2, 2022
Remove statistics-based set_index logic from read_parquet ( dask#9661 ) Richard (Rick) Zamora
Add support for use_nullable_dtypes to dd.read_parquet ( dask#9617 ) Ian Rose
Fix map_overlap in order to accept pandas arguments ( dask#9571 ) Fabien Aulaire
Fix pandas 1.5+ FutureWarning in .str.split(..., expand=True) ( dask#9704 ) Jacob Hayes
Enable column projection for groupby slicing ( dask#9667 ) Richard (Rick) Zamora
Support duplicate column cum-functions ( dask#9685 ) Ben
Improve error message for failed backend dispatch call ( dask#9677 ) Richard (Rick) Zamora
Revise meta creation in arrow parquet engine ( dask#9672 ) Richard (Rick) Zamora
Fix da.fft.fft for array-like inputs ( dask#9688 ) James Bourbeau
Fix groupby -aggregation when grouping on an index by name ( dask#9646 ) Richard (Rick) Zamora
Avoid PytestReturnNotNoneWarning in test_inheriting_class ( dask#9707 ) Thomas Grainger
Fix flaky test_dataframe_aggregations_multilevel ( dask#9701 ) Richard (Rick) Zamora
Bump mypy version ( dask#9697 ) crusaderky
Disable dashboard in test_map_partitions_df_input ( dask#9687 ) James Bourbeau
Use latest xarray-contrib/issue-from-pytest-log in upstream build ( dask#9682 ) James Bourbeau
xfail ttest_1samp for upstream scipy ( dask#9670 ) James Bourbeau
Update gpuCI RAPIDS_VER to 23.02 ( dask#9678 )
Released on December 2, 2022
Remove statistics-based set_index logic from read_parquet (9661) `Richard (Rick) Zamora`_
Add support for use_nullable_dtypes to dd.read_parquet (9617) `Ian Rose`_
Fix map_overlap in order to accept pandas arguments (9571) `Fabien Aulaire`_
Fix pandas 1.5+ FutureWarning in .str.split(..., expand=True) (9704) `Jacob Hayes`_
Enable column projection for groupby slicing (9667) `Richard (Rick) Zamora`_
Support duplicate column cum-functions (9685) `Ben`_
Improve error message for failed backend dispatch call (9677) `Richard (Rick) Zamora`_
Revise meta creation in arrow parquet engine (9672) `Richard (Rick) Zamora`_
Fix da.fft.fft for array-like inputs (9688) `James Bourbeau`_
Fix groupby -aggregation when grouping on an index by name (9646) `Richard (Rick) Zamora`_
Avoid PytestReturnNotNoneWarning in test_inheriting_class (9707) `Thomas Grainger`_
Fix flaky test_dataframe_aggregations_multilevel (9701) `Richard (Rick) Zamora`_
Bump mypy version (9697) `crusaderky`_
Disable dashboard in test_map_partitions_df_input (9687) `James Bourbeau`_
Use latest xarray-contrib/issue-from-pytest-log in upstream build (9682) `James Bourbeau`_
xfail ttest_1samp for upstream scipy (9670) `James Bourbeau`_
Update gpuCI RAPIDS_VER to 23.02 (9678)
- Restrict bokeh=3 support ( dask#9673 ) Gabe Joseph
Released on November 18, 2022
Restrict bokeh=3 support ( dask#9673 ) Gabe Joseph
Updates for fastparquet evolution ( dask#9650 ) Martin Durant
Update ga-yaml-parser step in gpuCI updating workflow ( dask#9675 ) Charles Blackmon-Luca
Revert importlib.metadata workaround ( dask#9658 ) James Bourbeau
Fix mindeps-distributed CI build to handle numpy / pandas not being installed ( dask#9668 ) James Bourbeau
Released on November 18, 2022
Restrict bokeh=3 support (9673) `Gabe Joseph`_
Updates for fastparquet evolution (9650) `Martin Durant`_
Update ga-yaml-parser step in gpuCI updating workflow (9675) `Charles Blackmon-Luca`_
Revert importlib.metadata workaround (9658) `James Bourbeau`_
Fix mindeps-distributed CI build to handle numpy/pandas not being installed (9668) `James Bourbeau`_
- Generalize from_dict implementation to allow usage from other backends ( dask#9628 ) GALI PREM SAGAR
Released on November 15, 2022
Generalize from_dict implementation to allow usage from other backends ( dask#9628 ) GALI PREM SAGAR
Avoid pandas constructors in dask.dataframe.core ( dask#9570 ) Richard (Rick) Zamora
Fix sort_values with Timestamp data ( dask#9642 ) James Bourbeau
Generalize array checking and remove pd.Index call in _get_partitions ( dask#9634 ) Benjamin Zaitlen
Fix read_csv behavior for header=0 and names ( dask#9614 ) Richard (Rick) Zamora
Update dashboard docs for queuing ( dask#9660 ) Gabe Joseph
Remove import dask as d from docstrings ( dask#9644 ) Matthew Rocklin
Fix link to partitions docs in read_parquet docstring ( dask#9636 ) qheuristics
Add API doc links to array/bag/dataframe sections ( dask#9630 ) Matthew Rocklin
Use conda-incubator/setup-miniconda@v2.2.0 ( dask#9662 ) John A Kirkham
Allow bokeh=3 ( dask#9659 ) James Bourbeau
Run upstream build with Python 3.10 ( dask#9655 ) James Bourbeau
Pin pyyaml version in mindeps testing ( dask#9640 ) Charles Blackmon-Luca
Add pre-commit to catch breakpoint() ( dask#9638 ) James Bourbeau
Bump xarray-contrib/issue-from-pytest-log from 1.1 to 1.2 ( dask#9635 )
Remove blosc references ( dask#9625 ) Naty Clementi
Upgrade mypy and drop unused comments ( dask#9616 ) Hendrik Makait
Harden test_repartition_npartitions ( dask#9585 ) Richard (Rick) Zamora
Released on November 15, 2022
Generalize from_dict implementation to allow usage from other backends (9628) `GALI PREM SAGAR`_
Avoid pandas constructors in dask.dataframe.core (9570) `Richard (Rick) Zamora`_
Fix sort_values with Timestamp data (9642) `James Bourbeau`_
Generalize array checking and remove pd.Index call in _get_partitions (9634) `Benjamin Zaitlen`_
Fix read_csv behavior for header=0 and names (9614) `Richard (Rick) Zamora`_
Update dashboard docs for queuing (9660) `Gabe Joseph`_
Remove import dask as d from docstrings (9644) `Matthew Rocklin`_
Fix link to partitions docs in read_parquet docstring (9636) `qheuristics`_
Add API doc links to array/bag/dataframe sections (9630) `Matthew Rocklin`_
Use conda-incubator/setup-miniconda@v2.2.0 (9662) `John A Kirkham`_
Allow bokeh=3 (9659) `James Bourbeau`_
Run upstream build with Python 3.10 (9655) `James Bourbeau`_
Pin pyyaml version in mindeps testing (9640) `Charles Blackmon-Luca`_
Add pre-commit to catch breakpoint() (9638) `James Bourbeau`_
Bump xarray-contrib/issue-from-pytest-log from 1.1 to 1.2 (9635)
Remove blosc references (9625) `Naty Clementi`_
Upgrade mypy and drop unused comments (9616) `Hendrik Makait`_
Harden test_repartition_npartitions (9585) `Richard (Rick) Zamora`_
This was a hotfix and has no changes in this repository. The necessary fix was in dask/distributed, but we decided to bump this version number for con
Released on October 31, 2022
This was a hotfix and has no changes in this repository. The necessary fix was in dask/distributed, but we decided to bump this version number for consistency.
- Enable named aggregation syntax ( dask#9563 ) ChrisJar
Released on October 28, 2022
Enable named aggregation syntax ( dask#9563 ) ChrisJar
Add extension dtype support to set_index ( dask#9566 ) James Bourbeau
Redesigning the array HTML repr for clarity ( dask#9519 ) Shingo OKAWA
Fix merge with emtpy left DataFrame ( dask#9578 ) Ian Rose
Add note about limiting thread oversubscription by default ( dask#9592 ) James Bourbeau
Use sphinx-click for dask CLI ( dask#9589 ) James Bourbeau
Fix Semaphore API docs ( dask#9584 ) James Bourbeau
Render meta description in map_overlap docstring ( dask#9568 ) James Bourbeau
Require Click 7.0+ in Dask ( dask#9595 ) John A Kirkham
Temporarily restrict bokeh<3 ( dask#9607 ) James Bourbeau
Resolve importlib -related failures in upstream CI ( dask#9604 ) Charles Blackmon-Luca
Improve upstream CI report ( dask#9603 ) James Bourbeau
Fix upstream CI report ( dask#9602 ) James Bourbeau
Remove setuptools host dep, add CLI entrypoint ( dask#9600 ) Charles Blackmon-Luca
More Backend dispatch class type annotations ( dask#9573 ) Ian Rose
Released on October 28, 2022
Enable named aggregation syntax (9563) `ChrisJar`_
Add extension dtype support to set_index (9566) `James Bourbeau`_
Redesigning the array HTML repr for clarity (9519) `Shingo OKAWA`_
Fix merge with emtpy left DataFrame (9578) `Ian Rose`_
Add note about limiting thread oversubscription by default (9592) `James Bourbeau`_
Use sphinx-click for dask CLI (9589) `James Bourbeau`_
Fix Semaphore API docs (9584) `James Bourbeau`_
Render meta description in map_overlap docstring (9568) `James Bourbeau`_
Require Click 7.0+ in Dask (9595) `John A Kirkham`_
Temporarily restrict bokeh<3 (9607) `James Bourbeau`_
Resolve importlib-related failures in upstream CI (9604) `Charles Blackmon-Luca`_
Improve upstream CI report (9603) `James Bourbeau`_
Fix upstream CI report (9602) `James Bourbeau`_
Remove setuptools host dep, add CLI entrypoint (9600) `Charles Blackmon-Luca`_
More Backend dispatch class type annotations (9573) `Ian Rose`_
- Backend library dispatching for IO in Dask-Array and Dask-DataFrame ( dask#9475 ) Richard (Rick) Zamora
Released on October 14, 2022
Backend library dispatching for IO in Dask-Array and Dask-DataFrame ( dask#9475 ) Richard (Rick) Zamora
Add new CLI that is extensible ( dask#9283 ) Doug Davis
Groupby median ( dask#9516 ) Ian Rose
Fix array copy not being a no-op ( dask#9555 ) David Hoese
Add support for string timedelta in map_overlap ( dask#9559 ) Nicolas Grandemange
Shuffle-based groupby for single functions ( dask#9504 ) Ian Rose
Make datetime.datetime tokenize idempotantly ( dask#9532 ) Martin Durant
Support tokenizing datetime.time ( dask#9528 ) Tim Paine
Avoid race condition in lazy dispatch registration ( dask#9545 ) James Bourbeau
Do not allow setitem to np.nan for int dtype ( dask#9531 ) Doug Davis
Stable demo column projection ( dask#9538 ) Ian Rose
Ensure pickle -able binops in delayed ( dask#9540 ) Ian Rose
Fix project CSV columns when selecting ( dask#9534 ) Martin Durant
Update Parquet best practice ( dask#9537 ) Matthew Rocklin
Restrict tiledb-py version to avoid CI failures ( dask#9569 ) James Bourbeau
Bump actions/github-script from 3 to 6 ( dask#9564 )
Bump actions/stale from 4 to 6 ( dask#9551 )
Bump peter-evans/create-pull-request from 3 to 4 ( dask#9550 )
Bump actions/checkout from 2 to 3.1.0 ( dask#9552 )
Bump codecov/codecov-action from 1 to 3 ( dask#9549 )
Bump the-coding-turtle/ga-yaml-parser from 0.1.1 to 0.1.2 ( dask#9553 )
Move dependabot configuration file ( dask#9547 ) James Bourbeau
Add dependabot for GitHub actions ( dask#9542 ) James Bourbeau
Run mypy on Windows and Linux ( dask#9530 ) crusaderky
Update gpuCI RAPIDS_VER to 22.12 ( dask#9524 )
Released on October 14, 2022
Backend library dispatching for IO in Dask-Array and Dask-DataFrame (9475) `Richard (Rick) Zamora`_
Add new CLI that is extensible (9283) `Doug Davis`_
Groupby median (9516) `Ian Rose`_
Fix array copy not being a no-op (9555) `David Hoese`_
Add support for string timedelta in map_overlap (9559) `Nicolas Grandemange`_
Shuffle-based groupby for single functions (9504) `Ian Rose`_
Make datetime.datetime tokenize idempotantly (9532) `Martin Durant`_
Support tokenizing datetime.time (9528) `Tim Paine`_
Avoid race condition in lazy dispatch registration (9545) `James Bourbeau`_
Do not allow setitem to np.nan for int dtype (9531) `Doug Davis`_
Stable demo column projection (9538) `Ian Rose`_
Ensure pickle-able binops in delayed (9540) `Ian Rose`_
Fix project CSV columns when selecting (9534) `Martin Durant`_
Update Parquet best practice (9537) `Matthew Rocklin`_
Restrict tiledb-py version to avoid CI failures (9569) `James Bourbeau`_
Bump actions/github-script from 3 to 6 (9564)
Bump actions/stale from 4 to 6 (9551)
Bump peter-evans/create-pull-request from 3 to 4 (9550)
Bump actions/checkout from 2 to 3.1.0 (9552)
Bump codecov/codecov-action from 1 to 3 (9549)
Bump the-coding-turtle/ga-yaml-parser from 0.1.1 to 0.1.2 (9553)
Move dependabot configuration file (9547) `James Bourbeau`_
Add dependabot for GitHub actions (9542) `James Bourbeau`_
Run mypy on Windows and Linux (9530) `crusaderky`_
Update gpuCI RAPIDS_VER to 22.12 (9524)
- Remove factorization logic from array auto chunking ( dask#9507 ) James Bourbeau
Released on September 30, 2022
Remove factorization logic from array auto chunking ( dask#9507 ) James Bourbeau
Add docs on running Dask in a standalone Python script ( dask#9513 ) James Bourbeau
Clarify custom-graph multiprocessing example ( dask#9511 ) nouman
Groupby sort upstream compatibility ( dask#9486 ) Ian Rose
Released on September 30, 2022
Remove factorization logic from array auto chunking (9507) `James Bourbeau`_
Add docs on running Dask in a standalone Python script (9513) `James Bourbeau`_
Clarify custom-graph multiprocessing example (9511) `nouman`_
Groupby sort upstream compatibility (9486) `Ian Rose`_
- Add DataFrame and Series median methods ( dask#9483 ) James Bourbeau
Released on September 16, 2022
Add DataFrame and Series median methods ( dask#9483 ) James Bourbeau
Shuffle groupby default ( dask#9453 ) Ian Rose
Filter by list ( dask#9419 ) Greg Hayes
Added distributed.utils.key_split functionality to dask.utils.key_split ( dask#9464 ) Luke Conibear
Fix overlap so that set_index doesn’t drop rows ( dask#9423 ) Julia Signell
Fix assigning pandas Series to column when ddf.columns.min() raises ( dask#9485 ) Erik Welch
Fix metadata comparison stack_partitions ( dask#9481 ) James Bourbeau
Provide default for split_out ( dask#9493 ) Lawrence Mitchell
Allow split_out to be None , which then defaults to 1 in groupby().aggregate() ( dask#9491 ) Ian Rose
Fixing enforce_metadata documentation, not checking for dtypes ( dask#9474 ) Nicolas Grandemange
Fix it's –> its typo ( dask#9484 ) Nat Tabris
Workaround for parquet writing failure using some datetime series but not others ( dask#9500 ) Ian Rose
Filter out numeric_only warnings from pandas ( dask#9496 ) James Bourbeau
Avoid set_index(..., inplace=True) where not necessary ( dask#9472 ) James Bourbeau
Avoid passing groupby key list of length one ( dask#9495 ) James Bourbeau
Update test_groupby_dropna_cudf based on cudf support for group_keys ( dask#9482 ) James Bourbeau
Remove dd.from_bcolz ( dask#9479 ) James Bourbeau
Added flake8-bugbear to pre-commit hooks ( dask#9457 ) Luke Conibear
Bind loop variables in function definitions ( B023 ) ( dask#9461 ) Luke Conibear
Added assert for comparisons ( B015 ) ( dask#9459 ) Luke Conibear
Set top-level default shell in CI workflows ( dask#9469 ) James Bourbeau
Removed unused loop control variables ( B007 ) ( dask#9458 ) Luke Conibear
Replaced getattr calls for constant attributes ( B009 ) ( dask#9460 ) Luke Conibear
Pin libprotobuf to allow nightly pyarrow in the upstream CI build ( dask#9465 ) Joris Van den Bossche
Replaced mutable data structures for default arguments ( B006 ) ( dask#9462 ) Luke Conibear
Changed flake8 mirror and updated version ( dask#9456 ) Luke Conibear
Released on September 16, 2022
Add DataFrame and Series median methods (9483) `James Bourbeau`_
Shuffle groupby default (9453) `Ian Rose`_
Filter by list (9419) `Greg Hayes`_
Added distributed.utils.key_split functionality to dask.utils.key_split (9464) `Luke Conibear`_
Fix overlap so that set_index doesn't drop rows (9423) `Julia Signell`_
Fix assigning pandas Series to column when ddf.columns.min() raises (9485) `Erik Welch`_
Fix metadata comparison stack_partitions (9481) `James Bourbeau`_
Provide default for split_out (9493) `Lawrence Mitchell`_
Allow split_out to be None, which then defaults to 1 in groupby().aggregate() (9491) `Ian Rose`_
Fixing enforce_metadata documentation, not checking for dtypes (9474) `Nicolas Grandemange`_
Fix it's --> its typo (9484) `Nat Tabris`_
Workaround for parquet writing failure using some datetime series but not others (9500) `Ian Rose`_
Filter out numeric_only warnings from pandas (9496) `James Bourbeau`_
Avoid set_index(..., inplace=True) where not necessary (9472) `James Bourbeau`_
Avoid passing groupby key list of length one (9495) `James Bourbeau`_
Update test_groupby_dropna_cudf based on cudf support for group_keys (9482) `James Bourbeau`_
Remove dd.from_bcolz (9479) `James Bourbeau`_
Added flake8-bugbear to pre-commit hooks (9457) `Luke Conibear`_
Bind loop variables in function definitions (B023) (9461) `Luke Conibear`_
Added assert for comparisons (B015) (9459) `Luke Conibear`_
Set top-level default shell in CI workflows (9469) `James Bourbeau`_
Removed unused loop control variables (B007) (9458) `Luke Conibear`_
Replaced getattr calls for constant attributes (B009) (9460) `Luke Conibear`_
Pin libprotobuf to allow nightly pyarrow in the upstream CI build (9465) `Joris Van den Bossche`_
Replaced mutable data structures for default arguments (B006) (9462) `Luke Conibear`_
Changed flake8 mirror and updated version (9456) `Luke Conibear`_
- Enable automatic column projection for groupby aggregations ( dask#9442 ) Richard (Rick) Zamora
Released on September 2, 2022
Enable automatic column projection for groupby aggregations ( dask#9442 ) Richard (Rick) Zamora
Accept superclasses in NEP-13/17 dispatching ( dask#6710 ) Gabe Joseph
Rename by columns internally for cumulative operations on the same by columns ( dask#9430 ) Pavithra Eswaramoorthy
Fix get_group with categoricals ( dask#9436 ) Pavithra Eswaramoorthy
Fix caching-related MaterializedLayer.cull performance regression ( dask#9413 ) Richard (Rick) Zamora
Add maintainer documentation page ( dask#9309 ) James Bourbeau
Revert skipped fastparquet test ( dask#9439 ) Pavithra Eswaramoorthy
tmpfile does not end files with period on empty extension ( dask#9429 ) Hendrik Makait
Skip failing fastparquet test with latest release ( dask#9432 ) James Bourbeau
Released on September 2, 2022
Enable automatic column projection for groupby aggregations (9442) `Richard (Rick) Zamora`_
Accept superclasses in NEP-13/17 dispatching (6710) `Gabe Joseph`_
Rename by columns internally for cumulative operations on the same by columns (9430) `Pavithra Eswaramoorthy`_
Fix get_group with categoricals (9436) `Pavithra Eswaramoorthy`_
Fix caching-related MaterializedLayer.cull performance regression (9413) `Richard (Rick) Zamora`_
Add maintainer documentation page (9309) `James Bourbeau`_
Revert skipped fastparquet test (9439) `Pavithra Eswaramoorthy`_
tmpfile does not end files with period on empty extension (9429) `Hendrik Makait`_
Skip failing fastparquet test with latest release (9432) `James Bourbeau`_
- Implement ma.*_like functions ( dask#9378 ) Ruth Comer
Released on August 19, 2022
Implement ma.*_like functions ( dask#9378 ) Ruth Comer
Fuse compatible annotations ( dask#9402 ) Ian Rose
Shuffle-based groupby aggregation for high-cardinality groups ( dask#9302 ) Richard (Rick) Zamora
Unpack namedtuple ( dask#9361 ) Hendrik Makait
Fix SeriesGroupBy cumulative functions with axis=1 ( dask#9377 ) Pavithra Eswaramoorthy
Sparse array reductions ( dask#9342 ) Ian Rose
Fix make_meta while using categorical column with index ( dask#9348 ) Pavithra Eswaramoorthy
Don’t allow incompatible keywords in DataFrame.dropna ( dask#9366 ) Naty Clementi
Make set_index handle entirely empty dataframes ( dask#8896 ) Julia Signell
Improve dataclass handling in unpack_collections ( dask#9345 ) Hendrik Makait
Fix bag sampling when there are some smaller partitions ( dask#9349 ) Ian Rose
Add support for empty partitions to da.min / da.max functions ( dask#9268 ) geraninam
Clarify that bind() etc. regenerate the keys ( dask#9385 ) crusaderky
Consolidate dashboard diagnostics documentation ( dask#9357 ) Sarah Charlotte Johnson
Remove outdated meta information Pavithra Eswaramoorthy
Use entry_points utility in sizeof ( dask#9390 ) James Bourbeau
Add entry_points compatibility utility ( dask#9388 ) Jacob Tomlinson
Upload environment file artifact for each CI build ( dask#9372 ) James Bourbeau
Remove werkzeug pin in CI ( dask#9371 ) James Bourbeau
Fix type annotations for dd.from_pandas and dd.from_delayed ( dask#9362 ) Jordan Yap
Released on August 19, 2022
Implement ma.*_like functions (9378) `Ruth Comer`_
Fuse compatible annotations (9402) `Ian Rose`_
Shuffle-based groupby aggregation for high-cardinality groups (9302) `Richard (Rick) Zamora`_
Unpack namedtuple (9361) `Hendrik Makait`_
Fix SeriesGroupBy cumulative functions with axis=1 (9377) `Pavithra Eswaramoorthy`_
Sparse array reductions (9342) `Ian Rose`_
Fix make_meta while using categorical column with index (9348) `Pavithra Eswaramoorthy`_
Don't allow incompatible keywords in DataFrame.dropna (9366) `Naty Clementi`_
Make set_index handle entirely empty dataframes (8896) `Julia Signell`_
Improve dataclass handling in unpack_collections (9345) `Hendrik Makait`_
Fix bag sampling when there are some smaller partitions (9349) `Ian Rose`_
Add support for empty partitions to da.min/da.max functions (9268) `geraninam`_
Clarify that bind() etc. regenerate the keys (9385) `crusaderky`_
Consolidate dashboard diagnostics documentation (9357) `Sarah Charlotte Johnson`_
Remove outdated meta information `Pavithra Eswaramoorthy`_
Use entry_points utility in sizeof (9390) `James Bourbeau`_
Add entry_points compatibility utility (9388) `Jacob Tomlinson`_
Upload environment file artifact for each CI build (9372) `James Bourbeau`_
Remove werkzeug pin in CI (9371) `James Bourbeau`_
Fix type annotations for dd.from_pandas and dd.from_delayed (9362) `Jordan Yap`_
- Ensure make_meta doesn’t hold ref to data ( dask#9354 ) Jim Crist-Harif
Released on August 5, 2022
Ensure make_meta doesn’t hold ref to data ( dask#9354 ) Jim Crist-Harif
Revise divisions logic in from_pandas ( dask#9221 ) Richard (Rick) Zamora
Warn if user sets index with existing index ( dask#9341 ) Julia Signell
Add keepdims keyword for da.average ( dask#9332 ) Ruth Comer
Change repr methods to avoid Layer materialization ( dask#9289 ) Richard (Rick) Zamora
Make sure order kwarg will not crash the astype method ( dask#9317 ) Genevieve Buckley
Fix bug for cumsum on cupy chunked dask arrays ( dask#9320 ) Genevieve Buckley
Match input and output structure in _sample_reduce ( dask#9272 ) Pavithra Eswaramoorthy
Include meta in array serialization ( dask#9240 ) Frédéric BRIOL
Fix Index.memory_usage ( dask#9290 ) James Bourbeau
Fix division calculation in dask.dataframe.io.from_dask_array ( dask#9282 ) Jordan Yap
Fow to use kwargs with custom task graphs ( dask#9322 ) Genevieve Buckley
Add note to da.from_array about how the order is not preserved ( dask#9346 ) Julia Signell
Add I/O info for async functions ( dask#9326 ) Logan Norman
Tidy up docs snippet for futures IO functions ( dask#9340 ) Julia Signell
Use consistent variable names for pandas df and Dask ddf in dataframe-groupby.rst ( dask#9304 ) ivojuroro
Switch js-yaml for yaml.js in config converter ( dask#9306 ) Jacob Tomlinson
Update da.linalg.solve for SciPy 1.9.0 compatibility ( dask#9350 ) Pavithra Eswaramoorthy
Update test_getitem_avoids_large_chunks_missing ( dask#9347 ) Pavithra Eswaramoorthy
Fix docs title formatting for “Extend sizeof ” Doug Davis
Import loop_in_thread fixture in tests ( dask#9337 ) James Bourbeau
Temporarily xfail test_solve_sym_pos ( dask#9336 ) Pavithra Eswaramoorthy
Fix small typo in 10 minutes to Dask page ( dask#9329 ) Shaghayegh
Temporarily pin werkzeug in CI to avoid test suite hanging ( dask#9325 ) James Bourbeau
Add tests for cupy.angle() ( dask#9312 ) Peter Andreas Entschev
Update gpuCI RAPIDS_VER to 22.10 ( dask#9314 )
Add pandas[test] to test extra ( dask#9110 ) Ben Beasley
Add bokeh and scipy to upstream CI build ( dask#9265 ) James Bourbeau
Released on August 5, 2022
Ensure make_meta doesn't hold ref to data (9354) `Jim Crist-Harif`_
Revise divisions logic in from_pandas (9221) `Richard (Rick) Zamora`_
Warn if user sets index with existing index (9341) `Julia Signell`_
Add keepdims keyword for da.average (9332) `Ruth Comer`_
Change repr methods to avoid Layer materialization (9289) `Richard (Rick) Zamora`_
Make sure order kwarg will not crash the astype method (9317) `Genevieve Buckley`_
Fix bug for cumsum on cupy chunked dask arrays (9320) `Genevieve Buckley`_
Match input and output structure in _sample_reduce (9272) `Pavithra Eswaramoorthy`_
Include meta in array serialization (9240) `Frédéric BRIOL`_
Fix Index.memory_usage (9290) `James Bourbeau`_
Fix division calculation in dask.dataframe.io.from_dask_array (9282) `Jordan Yap`_
Fow to use kwargs with custom task graphs (9322) `Genevieve Buckley`_
Add note to da.from_array about how the order is not preserved (9346) `Julia Signell`_
Add I/O info for async functions (9326) `Logan Norman`_
Tidy up docs snippet for futures IO functions (9340) `Julia Signell`_
Use consistent variable names for pandas df and Dask ddf in dataframe-groupby.rst (9304) `ivojuroro`_
Switch js-yaml for yaml.js in config converter (9306) `Jacob Tomlinson`_
Update da.linalg.solve for SciPy 1.9.0 compatibility (9350) `Pavithra Eswaramoorthy`_
Update test_getitem_avoids_large_chunks_missing (9347) `Pavithra Eswaramoorthy`_
Fix docs title formatting for "Extend sizeof" `Doug Davis`_
Import loop_in_thread fixture in tests (9337) `James Bourbeau`_
Temporarily xfail test_solve_sym_pos (9336) `Pavithra Eswaramoorthy`_
Fix small typo in 10 minutes to Dask page (9329) `Shaghayegh`_
Temporarily pin werkzeug in CI to avoid test suite hanging (9325) `James Bourbeau`_
Add tests for cupy.angle() (9312) `Peter Andreas Entschev`_
Update gpuCI RAPIDS_VER to 22.10 (9314)
Add pandas[test] to test extra (9110) `Ben Beasley`_
Add bokeh and scipy to upstream CI build (9265) `James Bourbeau`_
- Return Dask array if all axes are squeezed ( dask#9250 ) Pavithra Eswaramoorthy
Released on July 22, 2022
Return Dask array if all axes are squeezed ( dask#9250 ) Pavithra Eswaramoorthy
Make cycle reported by toposort shorter ( dask#9068 ) Erik Welch
Unknown chunk slicing - raise informative error ( dask#9285 ) Naty Clementi
Fix bug in HighLevelGraph.cull ( dask#9267 ) Richard (Rick) Zamora
Sort categories ( dask#9264 ) Pavithra Eswaramoorthy
Use max (instead of sum ) for calculating warnsize ( dask#9235 ) Pavithra Eswaramoorthy
Fix bug when filtering on partitioned column with pyarrow ( dask#9252 ) Richard (Rick) Zamora
Updated repartition documentation to add note about partition_size ( dask#9288 ) Dylan Stewart
Don’t include docs in Array methods, just refer to module docs ( dask#9244 ) Julia Signell
Remove outdated reference to scheduler and worker dashboards ( dask#9278 ) Pavithra Eswaramoorthy
Fix a few typos ( dask#9270 ) Tim Gates
Adds an custom aggregate example using numpy methods ( dask#9260 ) geraninam
Add type annotations to dd.from_pandas and dd.from_delayed ( dask#9237 ) Michael Milton
Update calculate_divisions docstring ( dask#9275 ) Tom Augspurger
Update test_plot_multiple for upcoming bokeh release ( dask#9261 ) James Bourbeau
Add typing to common array properties ( dask#9255 ) Illviljan
Released on July 22, 2022
Return Dask array if all axes are squeezed (9250) `Pavithra Eswaramoorthy`_
Make cycle reported by toposort shorter (9068) `Erik Welch`_
Unknown chunk slicing - raise informative error (9285) `Naty Clementi`_
Fix bug in HighLevelGraph.cull (9267) `Richard (Rick) Zamora`_
Sort categories (9264) `Pavithra Eswaramoorthy`_
Use max (instead of sum) for calculating warnsize (9235) `Pavithra Eswaramoorthy`_
Fix bug when filtering on partitioned column with pyarrow (9252) `Richard (Rick) Zamora`_
Updated repartition documentation to add note about partition_size (9288) `Dylan Stewart`_
Don't include docs in Array methods, just refer to module docs (9244) `Julia Signell`_
Remove outdated reference to scheduler and worker dashboards (9278) `Pavithra Eswaramoorthy`_
Fix a few typos (9270) `Tim Gates`_
Adds an custom aggregate example using numpy methods (9260) `geraninam`_
Add type annotations to dd.from_pandas and dd.from_delayed (9237) `Michael Milton`_
Update calculate_divisions docstring (9275) `Tom Augspurger`_
Update test_plot_multiple for upcoming bokeh release (9261) `James Bourbeau`_
Add typing to common array properties (9255) `Illviljan`_
- Support pathlib.PurePath in normalize_token ( dask#9229 ) Angus Hollands
Released on July 8, 2022
Support pathlib.PurePath in normalize_token ( dask#9229 ) Angus Hollands
Add AttributeNotImplementedError for properties so IPython glob search works ( dask#9231 ) Erik Welch
map_overlap : multiple dataframe handling ( dask#9145 ) Fabien Aulaire
Read entrypoints in dask.sizeof ( dask#7688 ) Angus Hollands
Fix TypeError: 'Serialize' object is not subscriptable when writing parquet dataset with Client(processes=False) ( dask#9015 ) Lucas Miguel Ponce
Correct dtypes when concat with an empty dataframe ( dask#9193 ) Pavithra Eswaramoorthy
Highlight note about persist ( dask#9234 ) Pavithra Eswaramoorthy
Update release-procedure to include more detail and helpful commands ( dask#9215 ) Julia Signell
Better SEO for Futures and Dask vs. Spark pages ( dask#9217 ) Sarah Charlotte Johnson
Use math.prod instead of np.prod on lists, tuples, and iters ( dask#9232 ) crusaderky
Only import IPython if type checking ( dask#9230 ) Florian Jetter
Tougher mypy checks ( dask#9206 ) crusaderky
Released on July 8, 2022
Support pathlib.PurePath in normalize_token (9229) `Angus Hollands`_
Add AttributeNotImplementedError for properties so IPython glob search works (9231) `Erik Welch`_
map_overlap: multiple dataframe handling (9145) `Fabien Aulaire`_
Read entrypoints in dask.sizeof (7688) `Angus Hollands`_
Fix TypeError: 'Serialize' object is not subscriptable when writing parquet dataset with Client(processes=False) (9015) `Lucas Miguel Ponce`_
Correct dtypes when concat with an empty dataframe (9193) `Pavithra Eswaramoorthy`_
Highlight note about persist (9234) `Pavithra Eswaramoorthy`_
Update release-procedure to include more detail and helpful commands (9215) `Julia Signell`_
Better SEO for Futures and Dask vs. Spark pages (9217) `Sarah Charlotte Johnson`_
Use math.prod instead of np.prod on lists, tuples, and iters (9232) `crusaderky`_
Only import IPython if type checking (9230) `Florian Jetter`_
Tougher mypy checks (9206) `crusaderky`_
- Deprecate extra format_time utility ( dask#9184 ) James Bourbeau
Released on June 24, 2022
Dask in pyodide ( dask#9053 ) Ian Rose
Create dask.utils.show_versions ( dask#9144 ) Sultan Orazbayev
Better error message for unsupported numpy operations on dask.dataframe objects. ( dask#9201 ) Julia Signell
Add allow_rechunk kwarg to dask.array.overlap function ( dask#7776 ) Genevieve Buckley
Add minutes and hours to dask.utils.format_time ( dask#9116 ) Matthew Rocklin
More retries when writing parquet to remote filesystem ( dask#9175 ) Ian Rose
Timedelta deterministic hashing ( dask#9213 ) Fabien Aulaire
Enum deterministic hashing ( dask#9212 ) Fabien Aulaire
shuffle_group() : avoid converting to arrays ( dask#9157 ) Mads R. B. Kristensen
Deprecate extra format_time utility ( dask#9184 ) James Bourbeau
Better SEO for 10 Minutes to Dask ( dask#9182 ) Sarah Charlotte Johnson
Better SEO for Delayed and Best Practices ( dask#9194 ) Sarah Charlotte Johnson
Include known inconsistency in DataFrame str.split accessor docstring ( dask#9177 ) Richard Pelgrim
Add inconsistencies keyword to derived_from ( dask#9192 ) Richard Pelgrim
Add missing append in delayed best practices example ( dask#9202 ) Ben
Fix indentation in Best Practices ( dask#9196 ) Sarah Charlotte Johnson
Add link to Genevieve Buckley ’s blog on chunk sizes ( dask#9199 ) Pavithra Eswaramoorthy
Update to_csv docstring ( dask#9094 ) Sarah Charlotte Johnson
Update versioneer: change from using SafeConfigParser to ConfigParser ( dask#9205 ) Thomas A Caswell
Remove ipython hack in CI( dask#9200 ) crusaderky
Released on June 24, 2022
Dask in pyodide (9053) `Ian Rose`_
Create dask.utils.show_versions (9144) `Sultan Orazbayev`_
Better error message for unsupported numpy operations on dask.dataframe objects. (9201) `Julia Signell`_
Add allow_rechunk kwarg to dask.array.overlap function (7776) `Genevieve Buckley`_
Add minutes and hours to dask.utils.format_time (9116) `Matthew Rocklin`_
More retries when writing parquet to remote filesystem (9175) `Ian Rose`_
Timedelta deterministic hashing (9213) `Fabien Aulaire`_
Enum deterministic hashing (9212) `Fabien Aulaire`_
shuffle_group(): avoid converting to arrays (9157) `Mads R. B. Kristensen`_
Deprecate extra format_time utility (9184) `James Bourbeau`_
Better SEO for 10 Minutes to Dask (9182) `Sarah Charlotte Johnson`_
Better SEO for Delayed and Best Practices (9194) `Sarah Charlotte Johnson`_
Include known inconsistency in DataFrame str.split accessor docstring (9177) `Richard Pelgrim`_
Add inconsistencies keyword to derived_from (9192) `Richard Pelgrim`_
Add missing append in delayed best practices example (9202) `Ben`_
Fix indentation in Best Practices (9196) `Sarah Charlotte Johnson`_
Add link to `Genevieve Buckley`_'s blog on chunk sizes (9199) `Pavithra Eswaramoorthy`_
Update to_csv docstring (9094) `Sarah Charlotte Johnson`_
Update versioneer: change from using SafeConfigParser to ConfigParser (9205) `Thomas A Caswell`_
Remove ipython hack in CI(9200) `crusaderky`_
- Add feature to show names of layer dependencies in HLG JupyterLab repr ( dask#9081 ) Angelos Omirolis
Released on June 10, 2022
Add feature to show names of layer dependencies in HLG JupyterLab repr ( dask#9081 ) Angelos Omirolis
Add arrow schema extraction dispatch ( dask#9169 ) GALI PREM SAGAR
Add sort_results argument to assert_eq ( dask#9130 ) Pavithra Eswaramoorthy
Add weeks to parse_timedelta ( dask#9168 ) Matthew Rocklin
Warn that cloudpickle is not always deterministic ( dask#9148 ) Pavithra Eswaramoorthy
Switch parquet default engine ( dask#9140 ) Jim Crist-Harif
Use deterministic hashing with _iLocIndexer / _LocIndexer ( dask#9108 ) Fabien Aulaire
Enfore consistent schema in to_parquet pyarrow ( dask#9131 ) Jim Crist-Harif
Fix pyarrow.StringArray pickle ( dask#9170 ) Jim Crist-Harif
Fix parallel metadata collection in pyarrow engine ( dask#9165 ) Richard (Rick) Zamora
Improve pyarrow partitioning logic ( dask#9147 ) James Bourbeau
pyarrow 8.0 partitioning fix ( dask#9143 ) James Bourbeau
Better SEO for Installing Dask and Dask DataFrame Best Practices ( dask#9178 ) Sarah Charlotte Johnson
Update logos page in docs ( dask#9167 ) Sarah Charlotte Johnson
Add example using pandas Series to map_partition doctring ( dask#9161 ) Alex-JG3
Update docs theme for rebranding ( dask#9160 ) Sarah Charlotte Johnson
Better SEO for docs on Dask DataFrames ( dask#9128 ) Sarah Charlotte Johnson
Remove ensure_file from recommended practice for downstream libraries ( dask#9171 ) Matthew Rocklin
Test round-tripping DataFrame parquet I/O including pyspark ( dask#9156 ) Ian Rose
Try disabling HDF5 locking ( dask#9154 ) Ian Rose
Link best practices to DataFrame-parquet ( dask#9150 ) Tom Augspurger
Fix typo in map_partitions func parameter description ( dask#9149 ) Christopher Akiki
Un- xfail test_groupby_grouper_dispatch ( dask#9139 ) GALI PREM SAGAR
Temporarily import cleanup fixture from distributed ( dask#9138 ) James Bourbeau
Simplify partitioning logic in pyarrow parquet engine ( dask#9041 ) Richard (Rick) Zamora
Released on June 10, 2022
Add feature to show names of layer dependencies in HLG JupyterLab repr (9081) `Angelos Omirolis`_
Add arrow schema extraction dispatch (9169) `GALI PREM SAGAR`_
Add sort_results argument to assert_eq (9130) `Pavithra Eswaramoorthy`_
Add weeks to parse_timedelta (9168) `Matthew Rocklin`_
Warn that cloudpickle is not always deterministic (9148) `Pavithra Eswaramoorthy`_
Switch parquet default engine (9140) `Jim Crist-Harif`_
Use deterministic hashing with _iLocIndexer / _LocIndexer (9108) `Fabien Aulaire`_
Enfore consistent schema in to_parquet pyarrow (9131) `Jim Crist-Harif`_
Fix pyarrow.StringArray pickle (9170) `Jim Crist-Harif`_
Fix parallel metadata collection in pyarrow engine (9165) `Richard (Rick) Zamora`_
Improve pyarrow partitioning logic (9147) `James Bourbeau`_
pyarrow 8.0 partitioning fix (9143) `James Bourbeau`_
Better SEO for Installing Dask and Dask DataFrame Best Practices (9178) `Sarah Charlotte Johnson`_
Update logos page in docs (9167) `Sarah Charlotte Johnson`_
Add example using pandas Series to map_partition doctring (9161) `Alex-JG3`_
Update docs theme for rebranding (9160) `Sarah Charlotte Johnson`_
Better SEO for docs on Dask DataFrames (9128) `Sarah Charlotte Johnson`_
Remove ensure_file from recommended practice for downstream libraries (9171) `Matthew Rocklin`_
Test round-tripping DataFrame parquet I/O including pyspark (9156) `Ian Rose`_
Try disabling HDF5 locking (9154) `Ian Rose`_
Link best practices to DataFrame-parquet (9150) `Tom Augspurger`_
Fix typo in map_partitions func parameter description (9149) `Christopher Akiki`_
Un-xfail test_groupby_grouper_dispatch (9139) `GALI PREM SAGAR`_
Temporarily import cleanup fixture from distributed (9138) `James Bourbeau`_
Simplify partitioning logic in pyarrow parquet engine (9041) `Richard (Rick) Zamora`_
- Add a dispatch for non-pandas Grouper objects and use it in GroupBy ( dask#9074 ) brandon-b-miller
Released on May 26, 2022
Add a dispatch for non-pandas Grouper objects and use it in GroupBy ( dask#9074 ) brandon-b-miller
Error if read_parquet & to_parquet files intersect ( dask#9124 ) Jim Crist-Harif
Visualize task graphs using ipycytoscape ( dask#9091 ) Ian Rose
Fix various typos ( dask#9126 ) Ryan Russell
Fix flaky test_filter_nonpartition_columns ( dask#9127 ) Pavithra Eswaramoorthy
Update gpuCI RAPIDS_VER to 22.08 ( dask#9120 )
Include conftest.py` in sdists ( dask#9115 ) Ben Beasley
Released on May 26, 2022
Add a dispatch for non-pandas Grouper objects and use it in GroupBy (9074) `brandon-b-miller`_
Error if read_parquet & to_parquet files intersect (9124) `Jim Crist-Harif`_
Visualize task graphs using ipycytoscape (9091) `Ian Rose`_
Fix various typos (9126) `Ryan Russell`_
Fix flaky test_filter_nonpartition_columns (9127) `Pavithra Eswaramoorthy`_
Update gpuCI RAPIDS_VER to 22.08 (9120)
Include conftest.py` in sdists (9115) `Ben Beasley`_
- Add pre-deprecation warnings for read_parquet kwargs chunksize and aggregate_files ( dask#9052 ) Richard (Rick) Zamora
Released on May 24, 2022
Add DataFrame.from_dict classmethod ( dask#9017 ) Matthew Powers
Add from_map function to Dask DataFrame ( dask#8911 ) Richard (Rick) Zamora
Improve to_parquet error for appended divisions overlap ( dask#9102 ) Jim Crist-Harif
Enabled user-defined process-initializer functions ( dask#9087 ) ParticularMiner
Mention align_dataframes=False option in map_partitions error ( dask#9075 ) Gabe Joseph
Add kwarg enforce_ndim to dask.array.map_blocks() ( dask#8865 ) ParticularMiner
Implement Series.GroupBy.fillna / DataFrame.GroupBy.fillna methods ( dask#8869 ) Pavithra Eswaramoorthy
Allow fillna with Dask DataFrame ( dask#8950 ) Pavithra Eswaramoorthy
Update error message for assignment with 1-d dask array ( dask#9036 ) Pavithra Eswaramoorthy
Collection Protocol ( dask#8674 ) Doug Davis
Patch around pandas ArrowStringArray pickling ( dask#9024 ) Jim Crist-Harif
Band-aid for compute_as_if_collection ( dask#8998 ) Ian Rose
Add p2p shuffle option ( dask#8836 ) Matthew Rocklin
Fixup column projection with no columns ( dask#9106 ) Jim Crist-Harif
Blockwise cull NumPy dtype ( dask#9100 ) Ian Rose
Fix column-projection bug in from_map ( dask#9078 ) Richard (Rick) Zamora
Prevent nulls in index for non-numeric dtypes ( dask#8963 ) Jorge López
Fix is_monotonic methods for more than 8 partitions ( dask#9019 ) Julia Signell
Handle enumerate and generator inputs to from_map ( dask#9066 ) Richard (Rick) Zamora
Revert is_dask_collection ; back to previous implementation ( dask#9062 ) Doug Davis
Fix Blockwise.clone does not handle iterable literal arguments correctly ( dask#8979 ) JSKenyon
Array setitem hardmask ( dask#9027 ) David Hassell
Fix overlapping divisions error on append ( dask#8997 ) Ian Rose
Add pre-deprecation warnings for read_parquet kwargs chunksize and aggregate_files ( dask#9052 ) Richard (Rick) Zamora
Document map_partitions handling of args vs kwargs , usage of partition_info ( dask#9084 ) Charles Blackmon-Luca
Update custom collection documentation (leverage new collection protocol) ( dask#9097 ) Doug Davis
Better SEO for docs on creating and storing Dask DataFrames ( dask#9098 ) Sarah Charlotte Johnson
Clarify chunking in imread docstring ( dask#9082 ) Genevieve Buckley
Rearrange docs TOC ( dask#9001 ) Matthew Rocklin
Corrected map_blocks() docstring for kwarg enforce_ndim ( dask#9071 ) ParticularMiner
Update DataFrame SQL docs references to other libraries ( dask#9077 ) Charles Blackmon-Luca
Update page on creating and storing Dask DataFrames ( dask#9025 ) Sarah Charlotte Johnson
Include NUMPY_LICENSE.txt in license files ( dask#9113 ) Ben Beasley
Increase retries when installing nightly pandas ( dask#9103 ) James Bourbeau
Force nightly pyarrow in the upstream build ( dask#9095 ) Joris Van den Bossche
Improve object handling & testing of ensure_unicode ( dask#9059 ) John A Kirkham
Force nightly pyarrow in the upstream build ( dask#8993 ) Joris Van den Bossche
Additional check on is_dask_collection ( dask#9054 ) Doug Davis
Update ensure_bytes ( dask#9050 ) John A Kirkham
Add end of file pre-commit hook ( dask#9045 ) James Bourbeau
Add codespell pre-commit hook ( dask#9040 ) James Bourbeau
Remove the HDFS tests ( dask#9039 ) Jim Crist-Harif
Fix flaky test_reductions_2D ( dask#9037 ) Jim Crist-Harif
Prevent codecov from notifying of failure too soon ( dask#9031 ) Jim Crist-Harif
Only test on Python 3.9 on macos ( dask#9029 ) Jim Crist-Harif
Update to_timedelta default unit ( dask#9010 ) Pavithra Eswaramoorthy
Released on May 24, 2022
Add DataFrame.from_dict classmethod (9017) `Matthew Powers`_
Add from_map function to Dask DataFrame (8911) `Richard (Rick) Zamora`_
Improve to_parquet error for appended divisions overlap (9102) `Jim Crist-Harif`_
Enabled user-defined process-initializer functions (9087) `ParticularMiner`_
Mention align_dataframes=False option in map_partitions error (9075) `Gabe Joseph`_
Add kwarg enforce_ndim to dask.array.map_blocks() (8865) `ParticularMiner`_
Implement Series.GroupBy.fillna / DataFrame.GroupBy.fillna methods (8869) `Pavithra Eswaramoorthy`_
Allow fillna with Dask DataFrame (8950) `Pavithra Eswaramoorthy`_
Update error message for assignment with 1-d dask array (9036) `Pavithra Eswaramoorthy`_
Collection Protocol (8674) `Doug Davis`_
Patch around pandas ArrowStringArray pickling (9024) `Jim Crist-Harif`_
Band-aid for compute_as_if_collection (8998) `Ian Rose`_
Add p2p shuffle option (8836) `Matthew Rocklin`_
Fixup column projection with no columns (9106) `Jim Crist-Harif`_
Blockwise cull NumPy dtype (9100) `Ian Rose`_
Fix column-projection bug in from_map (9078) `Richard (Rick) Zamora`_
Prevent nulls in index for non-numeric dtypes (8963) `Jorge López`_
Fix is_monotonic methods for more than 8 partitions (9019) `Julia Signell`_
Handle enumerate and generator inputs to from_map (9066) `Richard (Rick) Zamora`_
Revert is_dask_collection; back to previous implementation (9062) `Doug Davis`_
Fix Blockwise.clone does not handle iterable literal arguments correctly (8979) `JSKenyon`_
Array setitem hardmask (9027) `David Hassell`_
Fix overlapping divisions error on append (8997) `Ian Rose`_
Add pre-deprecation warnings for read_parquet kwargs chunksize and aggregate_files (9052) `Richard (Rick) Zamora`_
Document map_partitions handling of args vs kwargs, usage of partition_info (9084) `Charles Blackmon-Luca`_
Update custom collection documentation (leverage new collection protocol) (9097) `Doug Davis`_
Better SEO for docs on creating and storing Dask DataFrames (9098) `Sarah Charlotte Johnson`_
Clarify chunking in imread docstring (9082) `Genevieve Buckley`_
Rearrange docs TOC (9001) `Matthew Rocklin`_
Corrected map_blocks() docstring for kwarg enforce_ndim (9071) `ParticularMiner`_
Update DataFrame SQL docs references to other libraries (9077) `Charles Blackmon-Luca`_
Update page on creating and storing Dask DataFrames (9025) `Sarah Charlotte Johnson`_
Include NUMPY_LICENSE.txt in license files (9113) `Ben Beasley`_
Increase retries when installing nightly pandas (9103) `James Bourbeau`_
Force nightly pyarrow in the upstream build (9095) `Joris Van den Bossche`_
Improve object handling & testing of ensure_unicode (9059) `John A Kirkham`_
Force nightly pyarrow in the upstream build (8993) `Joris Van den Bossche`_
Additional check on is_dask_collection (9054) `Doug Davis`_
Update ensure_bytes (9050) `John A Kirkham`_
Add end of file pre-commit hook (9045) `James Bourbeau`_
Add codespell pre-commit hook (9040) `James Bourbeau`_
Remove the HDFS tests (9039) `Jim Crist-Harif`_
Fix flaky test_reductions_2D (9037) `Jim Crist-Harif`_
Prevent codecov from notifying of failure too soon (9031) `Jim Crist-Harif`_
Only test on Python 3.9 on macos (9029) `Jim Crist-Harif`_
Update to_timedelta default unit (9010) `Pavithra Eswaramoorthy`_
This is a bugfix release for this issue .
Released on May 2, 2022
This is a bugfix release for this issue .
Add highlights section to 2022.04.2 release notes ( dask#9012 ) James Bourbeau
Released on May 2, 2022
This is a bugfix release for this issue.
Add highlights section to 2022.04.2 release notes (9012) `James Bourbeau`_
This release includes several deprecations/breaking API changes to dask.dataframe.read_parquet and dask.dataframe.to_parquet :
Released on April 29, 2022
This release includes several deprecations/breaking API changes to dask.dataframe.read_parquet and dask.dataframe.to_parquet :
to_parquet no longer writes _metadata files by default. If you want to write a _metadata file, you can pass in write_metadata_file=True .
read_parquet now defaults to split_row_groups=False , which results in one Dask dataframe partition per parquet file when reading in a parquet dataset. If you’re working with large parquet files you may need to set split_row_groups=True to reduce your partition size.
read_parquet no longer calculates divisions by default. If you require read_parquet to return dataframes with known divisions, please set calculate_divisions=True .
read_parquet has deprecated the gather_statistics keyword argument. Please use the calculate_divisions keyword argument instead.
read_parquet has deprecated the require_extensions keyword argument. Please use the parquet_file_extension keyword argument instead.
Add removeprefix and removesuffix as StringMethods ( dask#8912 ) Jorge López
Call fs.invalidate_cache in to_parquet ( dask#8994 ) Jim Crist-Harif
Change to_parquet default to write_metadata_file=None ( dask#8988 ) Jim Crist-Harif
Let arg reductions pass keepdims ( dask#8926 ) Julia Signell
Change split_row_groups default to False in read_parquet ( dask#8981 ) Richard (Rick) Zamora
Improve NotImplementedError message for da.reshape ( dask#8987 ) Jim Crist-Harif
Simplify to_parquet compute path ( dask#8982 ) Jim Crist-Harif
Raise an error if you try to use vindex with a Dask object ( dask#8945 ) Julia Signell
Avoid pre_buffer=True when a precache method is specified ( dask#8957 ) Richard (Rick) Zamora
from_dask_array uses blockwise instead of merging graphs ( dask#8889 ) Bryan Weber
Use pre_buffer=True for “pyarrow” Parquet engine ( dask#8952 ) Richard (Rick) Zamora
Handle dtype=None correctly in da.full ( dask#8954 ) Tom White
Fix dask-sql bug caused by blockwise fusion ( dask#8989 ) Richard (Rick) Zamora
to_parquet errors for non-string column names ( dask#8990 ) Jim Crist-Harif
Make sure da.roll works even if shape is 0 ( dask#8925 ) Julia Signell
Fix recursion error issue with set_index ( dask#8967 ) Paul Hobson
Stringify BlockwiseDepDict mapping values when produces_keys=True ( dask#8972 ) Richard (Rick) Zamora
Use DataFram`eIOLayer in DataFrame.from_delayed ( dask#8852 ) Richard (Rick) Zamora
Check that values for the in predicate in read_parquet are correct ( dask#8846 ) Bryan Weber
Fix bug for reduction of zero dimensional arrays ( dask#8930 ) Tom White
Specify dtype when deciding division using np.linspace in read_sql_query ( dask#8940 ) Cheun Hong
Deprecate gather_statistics from read_parquet ( dask#8992 ) Richard (Rick) Zamora
Change require_extension to top-level parquet_file_extension read_parquet kwarg ( dask#8935 ) Richard (Rick) Zamora
Update write_metadata_file discussion in documentation ( dask#8995 ) Richard (Rick) Zamora
Update DataFrame.merge docstring ( dask#8966 ) Pavithra Eswaramoorthy
Added description for parameter align_arrays in array.blockwise() ( dask#8977 ) ParticularMiner
ecommend not to use map_block(drop_axis=...) on chunked axes of an array ( dask#8921 ) ParticularMiner
Add copy button to code snippets in docs ( dask#8956 ) James Bourbeau
Pandas 1.5.0 compatibility ( dask#8961 ) Ian Rose
Add pytest-timeout to distributed envs on CI ( dask#8986 ) Julia Signell
Improve read_parquet docstring formatting ( dask#8971 ) Bryan Weber
Remove pytest.warns(None) ( dask#8924 ) Pavithra Eswaramoorthy
Document Python 3.10 as supported ( dask#8976 ) Eray Aslan
parse_timedelta option to enforce explicit unit ( dask#8969 ) crusaderky
mypy compatibility ( dask#8854 ) Paul Hobson
Add a docs page for Dask & Parquet ( dask#8899 ) Jim Crist-Harif
Adds configuration to ignore revs in blame ( dask#8933 ) Bryan Weber
Released on April 29, 2022
This release includes several deprecations/breaking API changes to dask.dataframe.read_parquet and dask.dataframe.to_parquet:
to_parquet no longer writes _metadata files by default. If you want to write a _metadata file, you can pass in write_metadata_file=True.
read_parquet now defaults to split_row_groups=False, which results in one Dask dataframe partition per parquet file when reading in a parquet dataset. If you're working with large parquet files you may need to set split_row_groups=True to reduce your partition size.
read_parquet no longer calculates divisions by default. If you require read_parquet to return dataframes with known divisions, please set calculate_divisions=True.
read_parquet has deprecated the gather_statistics keyword argument. Please use the calculate_divisions keyword argument instead.
read_parquet has deprecated the require_extensions keyword argument. Please use the parquet_file_extension keyword argument instead.
Add removeprefix and removesuffix as StringMethods (8912) `Jorge López`_
Call fs.invalidate_cache in to_parquet (8994) `Jim Crist-Harif`_
Change to_parquet default to write_metadata_file=None (8988) `Jim Crist-Harif`_
Let arg reductions pass keepdims (8926) `Julia Signell`_
Change split_row_groups default to False in read_parquet (8981) `Richard (Rick) Zamora`_
Improve NotImplementedError message for da.reshape (8987) `Jim Crist-Harif`_
Simplify to_parquet compute path (8982) `Jim Crist-Harif`_
Raise an error if you try to use vindex with a Dask object (8945) `Julia Signell`_
Avoid pre_buffer=True when a precache method is specified (8957) `Richard (Rick) Zamora`_
from_dask_array uses blockwise instead of merging graphs (8889) `Bryan Weber`_
Use pre_buffer=True for "pyarrow" Parquet engine (8952) `Richard (Rick) Zamora`_
Handle dtype=None correctly in da.full (8954) `Tom White`_
Fix dask-sql bug caused by blockwise fusion (8989) `Richard (Rick) Zamora`_
to_parquet errors for non-string column names (8990) `Jim Crist-Harif`_
Make sure da.roll works even if shape is 0 (8925) `Julia Signell`_
Fix recursion error issue with set_index (8967) `Paul Hobson`_
Stringify BlockwiseDepDict mapping values when produces_keys=True (8972) `Richard (Rick) Zamora`_
Use DataFram`eIOLayer in DataFrame.from_delayed (8852) `Richard (Rick) Zamora`_
Check that values for the in predicate in read_parquet are correct (8846) `Bryan Weber`_
Fix bug for reduction of zero dimensional arrays (8930) `Tom White`_
Specify dtype when deciding division using np.linspace in read_sql_query (8940) `Cheun Hong`_
Deprecate gather_statistics from read_parquet (8992) `Richard (Rick) Zamora`_
Change require_extension to top-level parquet_file_extension read_parquet kwarg (8935) `Richard (Rick) Zamora`_
Update write_metadata_file discussion in documentation (8995) `Richard (Rick) Zamora`_
Update DataFrame.merge docstring (8966) `Pavithra Eswaramoorthy`_
Added description for parameter align_arrays in array.blockwise() (8977) `ParticularMiner`_
ecommend not to use map_block(drop_axis=...) on chunked axes of an array (8921) `ParticularMiner`_
Add copy button to code snippets in docs (8956) `James Bourbeau`_
Pandas 1.5.0 compatibility (8961) `Ian Rose`_
Add pytest-timeout to distributed envs on CI (8986) `Julia Signell`_
Improve read_parquet docstring formatting (8971) `Bryan Weber`_
Remove pytest.warns(None) (8924) `Pavithra Eswaramoorthy`_
Document Python 3.10 as supported (8976) `Eray Aslan`_
parse_timedelta option to enforce explicit unit (8969) `crusaderky`_
mypy compatibility (8854) `Paul Hobson`_
Add a docs page for Dask & Parquet (8899) `Jim Crist-Harif`_
Adds configuration to ignore revs in blame (8933) `Bryan Weber`_
- Remove unused (deprecated) code from ArrowDatasetEngine ( dask#8885 ) Richard (Rick) Zamora
Released on April 15, 2022
Add missing NumPy ufuncs: abs , left_shift , right_shift , positive . ( dask#8920 ) Tom White
Avoid collecting parquet metadata in pyarrow when write_metadata_file=False ( dask#8906 ) Richard (Rick) Zamora
Better error for failed wildcard path in dd.read_csv() (fixes #8878) ( dask#8908 ) Roger Filmyer
Return da.Array rather than dd.Series for non-ufunc elementwise functions on dd.Series ( dask#8558 ) Julia Signell
Let get_dummies use meta computation in map_partitions ( dask#8898 ) Julia Signell
Masked scalars input to da.from_array ( dask#8895 ) David Hassell
Raise ValueError in merge_asof for duplicate kwargs ( dask#8861 ) Bryan Weber
Make is_monotonic work when some partitions are empty ( dask#8897 ) Julia Signell
Fix custom getter in da.from_array when inline_array=False ( dask#8903 ) Ian Rose
Correctly handle dict-specification for rechunk. ( dask#8859 ) Richard
Fix merge_asof : drop index column if left_on == right_on ( dask#8874 ) Gil Forsyth
Warn users that engine='auto' will change in future ( dask#8907 ) Jim Crist-Harif
Remove pyarrow-legacy engine from parquet API ( dask#8835 ) Richard (Rick) Zamora
Add note on missing parameter out for dask.array.dot ( dask#8913 ) Francesco Andreuzzi
Update DataFrame.query docstring ( dask#8890 ) Pavithra Eswaramoorthy
Don’t test da.prod on large integer data ( dask#8893 ) Jim Crist-Harif
Add network marks to tests that fail without an internet connection ( dask#8881 ) Paul Hobson
Fix gpuCI GHA version ( dask#8891 ) Charles Blackmon-Luca
xfail / skip some flaky distributed tests ( dask#8887 ) Jim Crist-Harif
Remove unused (deprecated) code from ArrowDatasetEngine ( dask#8885 ) Richard (Rick) Zamora
Add mild typing to common utils functions, part 2 ( dask#8867 ) crusaderky
Documentation of Limitation of sample() ( dask#8858 ) Nadiem Sissouno
Released on April 15, 2022
Add missing NumPy ufuncs: abs, left_shift, right_shift, positive. (8920) `Tom White`_
Avoid collecting parquet metadata in pyarrow when write_metadata_file=False (8906) `Richard (Rick) Zamora`_
Better error for failed wildcard path in dd.read_csv() (fixes #8878) (8908) `Roger Filmyer`_
Return da.Array rather than dd.Series for non-ufunc elementwise functions on dd.Series (8558) `Julia Signell`_
Let get_dummies use meta computation in map_partitions (8898) `Julia Signell`_
Masked scalars input to da.from_array (8895) `David Hassell`_
Raise ValueError in merge_asof for duplicate kwargs (8861) `Bryan Weber`_
Make is_monotonic work when some partitions are empty (8897) `Julia Signell`_
Fix custom getter in da.from_array when inline_array=False (8903) `Ian Rose`_
Correctly handle dict-specification for rechunk. (8859) `Richard`_
Fix merge_asof: drop index column if left_on == right_on (8874) `Gil Forsyth`_
Warn users that engine='auto' will change in future (8907) `Jim Crist-Harif`_
Remove pyarrow-legacy engine from parquet API (8835) `Richard (Rick) Zamora`_
Add note on missing parameter out for dask.array.dot (8913) `Francesco Andreuzzi`_
Update DataFrame.query docstring (8890) `Pavithra Eswaramoorthy`_
Don't test da.prod on large integer data (8893) `Jim Crist-Harif`_
Add network marks to tests that fail without an internet connection (8881) `Paul Hobson`_
Fix gpuCI GHA version (8891) `Charles Blackmon-Luca`_
xfail/skip some flaky distributed tests (8887) `Jim Crist-Harif`_
Remove unused (deprecated) code from ArrowDatasetEngine (8885) `Richard (Rick) Zamora`_
Add mild typing to common utils functions, part 2 (8867) `crusaderky`_
Documentation of Limitation of sample() (8858) `Nadiem Sissouno`_
This is the first release with support for Python 3.10
Released on April 1, 2022
Note
This is the first release with support for Python 3.10
Add Python 3.10 support ( dask#8566 ) James Bourbeau
Add check on dtype.itemsize in order to produce a useful error ( dask#8860 ) Davide Gavio
Add mild typing to common utils functions ( dask#8848 ) Matthew Rocklin
Add sanity checks to divisions setter ( dask#8806 ) Jim Crist-Harif
Use Blockwise and map_partitions for more tasks ( dask#8831 ) Bryan Weber
Fix dataframe.merge_asof to preserve right_on column ( dask#8857 ) Sarah Charlotte Johnson
Fix “Buffer dtype mismatch” for pandas >= 1.3 on 32bit ( dask#8851 ) Ben Greiner
Fix slicing fusion by altering SubgraphCallable getter ( dask#8827 ) Ian Rose
Remove support for PyPy ( dask#8863 ) James Bourbeau
Drop setuptools at runtime ( dask#8855 ) crusaderky
Remove dataframe.tseries.resample.getnanos ( dask#8834 ) Sarah Charlotte Johnson
Organize diagnostic and performance docs ( dask#8871 ) Naty Clementi
Add image to explain drop_axis option of map_blocks ( dask#8868 ) ParticularMiner
Update gpuCI RAPIDS_VER to 22.06 ( dask#8828 )
Restore test_parquet in http ( dask#8850 ) Bryan Weber
Simplify gpuCI updating workflow ( dask#8849 ) Charles Blackmon-Luca
Released on April 1, 2022
Note
This is the first release with support for Python 3.10
Add Python 3.10 support (8566) `James Bourbeau`_
Add check on dtype.itemsize in order to produce a useful error (8860) `Davide Gavio`_
Add mild typing to common utils functions (8848) `Matthew Rocklin`_
Add sanity checks to divisions setter (8806) `Jim Crist-Harif`_
Use Blockwise and map_partitions for more tasks (8831) `Bryan Weber`_
Fix dataframe.merge_asof to preserve right_on column (8857) `Sarah Charlotte Johnson`_
Fix "Buffer dtype mismatch" for pandas >= 1.3 on 32bit (8851) `Ben Greiner`_
Fix slicing fusion by altering SubgraphCallable getter (8827) `Ian Rose`_
Remove support for PyPy (8863) `James Bourbeau`_
Drop setuptools at runtime (8855) `crusaderky`_
Remove dataframe.tseries.resample.getnanos (8834) `Sarah Charlotte Johnson`_
Organize diagnostic and performance docs (8871) `Naty Clementi`_
Add image to explain drop_axis option of map_blocks (8868) `ParticularMiner`_
Update gpuCI RAPIDS_VER to 22.06 (8828)
Restore test_parquet in http (8850) `Bryan Weber`_
Simplify gpuCI updating workflow (8849) `Charles Blackmon-Luca`_
- Deprecate bcolz support ( dask#8754 ) Pavithra Eswaramoorthy
Released on March 18, 2022
Bag: add implementation for reservoir sampling ( dask#7636 ) Daniel Mesejo-León
Add ma.count to Dask array ( dask#8785 ) David Hassell
Change to_parquet default to compression="snappy" ( dask#8814 ) Jim Crist-Harif
Add weights parameter to dask.array.reduction ( dask#8805 ) David Hassell
Add ddf.compute_current_divisions to get divisions on a sorted index or column ( dask#8517 ) Julia Signell
Pass name and doc through on DelayedLeaf ( dask#8820 ) Leo Gao
Raise exception for not implemented merge how option ( dask#8818 ) Naty Clementi
Move Bag.map_partitions to Blockwise ( dask#8646 ) Richard (Rick) Zamora
Improve error messages for malformed config files ( dask#8801 ) Jim Crist-Harif
Revise column-projection optimization to capture common dask-sql patterns ( dask#8692 ) Richard (Rick) Zamora
Useful error for empty divisions ( dask#8789 ) Pavithra Eswaramoorthy
Scipy 1.8.0 compat: copy private classes into dask/array/stats.py ( dask#8694 ) Julia Signell
Raise warning when using multiple types of schedulers where one is distributed ( dask#8700 ) Pedro Silva
Fix bug in applying != filter in read_parquet ( dask#8824 ) Richard (Rick) Zamora
Fix set_index when directly passed a dask Index ( dask#8680 ) Paul Hobson
Quick fix for unbounded memory usage in tensordot ( dask#7980 ) Genevieve Buckley
If hdf file is empty, don’t fail on meta creation ( dask#8809 ) Julia Signell
Update clone_key("x") to retain prefix ( dask#8792 ) crusaderky
Fix “physical” column bug in pyarrow-based read_parquet ( dask#8775 ) Richard (Rick) Zamora
Fix groupby.shift bug caused by unsorted partitions after shuffle ( dask#8782 ) kori73
Fix serialization bug ( dask#8786 ) Richard (Rick) Zamora
Bump diagnostics bokeh dependency to 2.4.2 ( dask#8791 ) Charles Blackmon-Luca
Deprecate bcolz support ( dask#8754 ) Pavithra Eswaramoorthy
Finish making map_overlap default boundary kwarg 'none' ( dask#8743 ) Genevieve Buckley
Custom collection example docs fix ( dask#8807 ) Doug Davis
Add Series.str , Series.dt , and Series.cat accessors to docs ( dask#8757 ) Sarah Charlotte Johnson
Fix docstring for ddf.compute_current_divisions ( dask#8793 ) Julia Signell
Dashboard docs on /status page ( dask#8648 ) Naty Clementi
Clarify divisions kwarg in repartition docstring ( dask#8781 ) Sarah Charlotte Johnson
Update Docker images to use ghcr.io ( dask#8774 ) Jacob Tomlinson
Reduce gpuci pytest parallelism ( dask#8826 ) GALI PREM SAGAR
absolufy-imports - No relative imports - PEP8 ( dask#8796 ) Julia Signell
Tidy up assert_eq calls in array tests ( dask#8812 ) Julia Signell
Avoid pytest.warns(None) ( dask#8718 ) LSturtew
Fix test_describe_empty to work without global -Werror ( dask#8291 ) Michał Górny
Temporarily xfail graphviz tests on windows ( dask#8794 ) Jim Crist-Harif
Use packaging.parse for md5 compatibility ( dask#8763 ) James Bourbeau
Make tokenize work in a FIPS 140-2 environment ( dask#8762 ) Jim Crist-Harif
Label issues and PRs on open with ‘needs triage’ ( dask#8761 ) Julia Signell
Add some extra test coverage ( dask#8302 ) lrjball
Specify action version and change from pull_request_target to pull_request ( dask#8767 ) Julia Signell
Make scheduler kwarg pass though to sub functions in da.assert_eq ( dask#8755 ) Julia Signell
Released on March 18, 2022
Bag: add implementation for reservoir sampling (7636) `Daniel Mesejo-León`_
Add ma.count to Dask array (8785) `David Hassell`_
Change to_parquet default to compression="snappy" (8814) `Jim Crist-Harif`_
Add weights parameter to dask.array.reduction (8805) `David Hassell`_
Add ddf.compute_current_divisions to get divisions on a sorted index or column (8517) `Julia Signell`_
Pass __name__ and __doc__ through on DelayedLeaf (8820) `Leo Gao`_
Raise exception for not implemented merge how option (8818) `Naty Clementi`_
Move Bag.map_partitions to Blockwise (8646) `Richard (Rick) Zamora`_
Improve error messages for malformed config files (8801) `Jim Crist-Harif`_
Revise column-projection optimization to capture common dask-sql patterns (8692) `Richard (Rick) Zamora`_
Useful error for empty divisions (8789) `Pavithra Eswaramoorthy`_
Scipy 1.8.0 compat: copy private classes into dask/array/stats.py (8694) `Julia Signell`_
Raise warning when using multiple types of schedulers where one is distributed (8700) `Pedro Silva`_
Fix bug in applying != filter in read_parquet (8824) `Richard (Rick) Zamora`_
Fix set_index when directly passed a dask Index (8680) `Paul Hobson`_
Quick fix for unbounded memory usage in tensordot (7980) `Genevieve Buckley`_
If hdf file is empty, don't fail on meta creation (8809) `Julia Signell`_
Update clone_key("x") to retain prefix (8792) `crusaderky`_
Fix "physical" column bug in pyarrow-based read_parquet (8775) `Richard (Rick) Zamora`_
Fix groupby.shift bug caused by unsorted partitions after shuffle (8782) `kori73`_
Fix serialization bug (8786) `Richard (Rick) Zamora`_
Bump diagnostics bokeh dependency to 2.4.2 (8791) `Charles Blackmon-Luca`_
Deprecate bcolz support (8754) `Pavithra Eswaramoorthy`_
Finish making map_overlap default boundary kwarg 'none' (8743) `Genevieve Buckley`_
Custom collection example docs fix (8807) `Doug Davis`_
Add Series.str, Series.dt, and Series.cat accessors to docs (8757) `Sarah Charlotte Johnson`_
Fix docstring for ddf.compute_current_divisions (8793) `Julia Signell`_
Dashboard docs on /status page (8648) `Naty Clementi`_
Clarify divisions kwarg in repartition docstring (8781) `Sarah Charlotte Johnson`_
Update Docker images to use ghcr.io (8774) `Jacob Tomlinson`_
Reduce gpuci pytest parallelism (8826) `GALI PREM SAGAR`_
absolufy-imports - No relative imports - PEP8 (8796) `Julia Signell`_
Tidy up assert_eq calls in array tests (8812) `Julia Signell`_
Avoid pytest.warns(None) (8718) `LSturtew`_
Fix test_describe_empty to work without global -Werror (8291) `Michał Górny`_
Temporarily xfail graphviz tests on windows (8794) `Jim Crist-Harif`_
Use packaging.parse for md5 compatibility (8763) `James Bourbeau`_
Make tokenize work in a FIPS 140-2 environment (8762) `Jim Crist-Harif`_
Label issues and PRs on open with 'needs triage' (8761) `Julia Signell`_
Add some extra test coverage (8302) `lrjball`_
Specify action version and change from pull_request_target to pull_request (8767) `Julia Signell`_
Make scheduler kwarg pass though to sub functions in da.assert_eq (8755) `Julia Signell`_
- Deprecate iteritems ( dask#8660 ) James Bourbeau
Released on February 25, 2022
Add aggregate functions first and last to dask.dataframe.pivot_table ( dask#8649 ) Knut Nordanger
Add std() support for datetime64 dtype for pandas-like objects ( dask#8523 ) Ben Glossner
Add materialized task counts to HighLevelGraph and Layer html reprs ( dask#8589 ) kori73
Do not allow iterating a DataFrameGroupBy ( dask#8696 ) Bryan Weber
Fix missing newline after info() call on empty DataFrame ( dask#8727 ) Naty Clementi
Add groupby.compute as a not implemented method ( dask#8734 ) Dranaxel
Improve multi dataframe join performance ( dask#8740 ) Holden Karau
Include bool type for Index ( dask#8732 ) Naty Clementi
Allow ArrowDatasetEngine subclass to override pandas->arrow conversion also for partitioned write ( dask#8741 ) Joris Van den Bossche
Increase performance of k-diagonal extraction in da.diag() and da.diagonal() ( dask#8689 ) ParticularMiner
Change linspace creation to match numpy when num equal to 0 ( dask#8676 ) Peter
Tokenize dataclasses ( dask#8557 ) Gabe Joseph
Update tokenize to treat dict and kwargs differently ( dask#8655 ) James Bourbeau
Fix bug in dask.array.roll() for roll-shifts that match the size of the input array ( dask#8723 ) ParticularMiner
Fix for normalize_function dataclass methods ( dask#8527 ) Sarah Charlotte Johnson
Fix rechunking with zero-size-chunks ( dask#8703 ) ParticularMiner
Move creation of sqlalchemy connection for picklability ( dask#8745 ) Julia Signell
Drop Python 3.7 ( dask#8572 ) James Bourbeau
Deprecate iteritems ( dask#8660 ) James Bourbeau
Deprecate dataframe.tseries.resample.getnanos ( dask#8752 ) Sarah Charlotte Johnson
Add deprecation warning for pyarrow-legacy engine ( dask#8758 ) Richard (Rick) Zamora
Update link typos in changelog ( dask#8717 ) James Bourbeau
Clarify dask.visualize docstring ( dask#8710 ) Dranaxel
Update Docker example to use current best practices ( dask#8731 ) Jacob Tomlinson
Update docs to include distributed.Client.preload ( dask#8679 ) Bryan Weber
Document monthly social meeting ( dask#8595 ) Thomas Grainger
Add docs for Gen2 access with RBAC/ACL i.e. security principal ( dask#8748 ) Martin Thøgersen
Use Dask configuration extension from dask-sphinx-theme ( dask#8751 ) Benjamin Zaitlen
Unpin coverage in CI ( dask#8690 ) James Bourbeau
Add manual trigger for running test suite ( dask#8716 ) James Bourbeau
Xfail scheduler_HLG_unpack_import ; flaky test ( dask#8724 ) Mike McCarty
Temporarily remove scipy upstream CI build ( dask#8725 ) James Bourbeau
Bump pre-release version to be greater than stable releases ( dask#8728 ) Charles Blackmon-Luca
Move custom sort function logic to internal sort_values ( dask#8571 ) Charles Blackmon-Luca
Pin cloudpickle and scipy in docs requirements ( dask#8737 ) Julia Signell
Make the labeler not delete labels, and look for the docs at the right spot ( dask#8746 ) Julia Signell
Fix docs build warnings ( dask#8432 ) Kristopher Overholt
Update test status badge ( dask#8747 ) James Bourbeau
Fix parquet test_pandas_timestamp_overflow_pyarrow test ( dask#8733 ) Joris Van den Bossche
Only run PR builds on changes to relevant files ( dask#8756 ) Charles Blackmon-Luca
Released on February 25, 2022
Add aggregate functions first and last to dask.dataframe.pivot_table (8649) `Knut Nordanger`_
Add std() support for datetime64 dtype for pandas-like objects (8523) `Ben Glossner`_
Add materialized task counts to HighLevelGraph and Layer html reprs (8589) `kori73`_
Do not allow iterating a DataFrameGroupBy (8696) `Bryan Weber`_
Fix missing newline after info() call on empty DataFrame (8727) `Naty Clementi`_
Add groupby.compute as a not implemented method (8734) `Dranaxel`_
Improve multi dataframe join performance (8740) `Holden Karau`_
Include bool type for Index (8732) `Naty Clementi`_
Allow ArrowDatasetEngine subclass to override pandas->arrow conversion also for partitioned write (8741) `Joris Van den Bossche`_
Increase performance of k-diagonal extraction in da.diag() and da.diagonal() (8689) `ParticularMiner`_
Change linspace creation to match numpy when num equal to 0 (8676) `Peter`_
Tokenize dataclasses (8557) `Gabe Joseph`_
Update tokenize to treat dict and kwargs differently (8655) `James Bourbeau`_
Fix bug in dask.array.roll() for roll-shifts that match the size of the input array (8723) `ParticularMiner`_
Fix for normalize_function dataclass methods (8527) `Sarah Charlotte Johnson`_
Fix rechunking with zero-size-chunks (8703) `ParticularMiner`_
Move creation of sqlalchemy connection for picklability (8745) `Julia Signell`_
Drop Python 3.7 (8572) `James Bourbeau`_
Deprecate iteritems (8660) `James Bourbeau`_
Deprecate dataframe.tseries.resample.getnanos (8752) `Sarah Charlotte Johnson`_
Add deprecation warning for pyarrow-legacy engine (8758) `Richard (Rick) Zamora`_
Update link typos in changelog (8717) `James Bourbeau`_
Clarify dask.visualize docstring (8710) `Dranaxel`_
Update Docker example to use current best practices (8731) `Jacob Tomlinson`_
Update docs to include distributed.Client.preload (8679) `Bryan Weber`_
Document monthly social meeting (8595) `Thomas Grainger`_
Add docs for Gen2 access with RBAC/ACL i.e. security principal (8748) `Martin Thøgersen`_
Use Dask configuration extension from dask-sphinx-theme (8751) `Benjamin Zaitlen`_
Unpin coverage in CI (8690) `James Bourbeau`_
Add manual trigger for running test suite (8716) `James Bourbeau`_
Xfail scheduler_HLG_unpack_import; flaky test (8724) `Mike McCarty`_
Temporarily remove scipy upstream CI build (8725) `James Bourbeau`_
Bump pre-release version to be greater than stable releases (8728) `Charles Blackmon-Luca`_
Move custom sort function logic to internal sort_values (8571) `Charles Blackmon-Luca`_
Pin cloudpickle and scipy in docs requirements (8737) `Julia Signell`_
Make the labeler not delete labels, and look for the docs at the right spot (8746) `Julia Signell`_
Fix docs build warnings (8432) `Kristopher Overholt`_
Update test status badge (8747) `James Bourbeau`_
Fix parquet test_pandas_timestamp_overflow_pyarrow test (8733) `Joris Van den Bossche`_
Only run PR builds on changes to relevant files (8756) `Charles Blackmon-Luca`_
- Deprecate is_monotonic ( dask#8653 ) James Bourbeau
Released on February 11, 2022
Note
This is the last release with support for Python 3.7
Add region to to_zarr when using existing array ( dask#8590 ) Chris Roat
Add engine_kwargs support to dask.dataframe.to_sql ( dask#8609 ) Amir Kadivar
Add include_path_column arg to read_json ( dask#8603 ) Bryan Weber
Add expand_dims to Dask array ( dask#8687 ) Tom White
Add scheduler option to assert_eq utilities ( dask#8610 ) Xinrong Meng
Fix eye inconsistency with NumPy for dtype=None ( dask#8685 ) Tom White
Fix concatenate inconsistency with NumPy for axis=None ( dask#8686 ) Tom White
Type annotations, part 1 ( dask#8295 ) crusaderky
Really allow any iterable to be passed as a meta ( dask#8629 ) Julia Signell
Use map_partitions (Blockwise) in to_parquet ( dask#8487 ) Richard (Rick) Zamora
Result of reducing an array should not depend on its chunk-structure ( dask#8637 ) ParticularMiner
Pass place-holder metadata to map_partitions in ACA code path ( dask#8643 ) Richard (Rick) Zamora
Deprecate is_monotonic ( dask#8653 ) James Bourbeau
Remove some deprecations ( dask#8605 ) James Bourbeau
Add Domino Data Lab to Hosted / managed Dask clusters ( dask#8675 ) Ray Bell
Fix inter-linking and remove deprecated function ( dask#8715 ) Julia Signell
Fix imbalanced backticks. ( dask#8693 ) Matthias Bussonnier
Add documentation for high level graph visualization ( dask#8483 ) Genevieve Buckley
Update documentation of ProgressBar out parameter ( dask#8604 ) Pedro Silva
Improve documentation of dask.config.set ( dask#8705 ) crusaderky
Revert mention to mypy among type checkers ( dask#8699 ) crusaderky
Update warning handling in get_dummies tests ( dask#8651 ) James Bourbeau
Add a github changelog template ( dask#8714 ) Julia Signell
Update year in LICENSE.txt ( dask#8665 ) David Hoese
Update pre-commit version ( dask#8691 ) James Bourbeau
Include scipy in upstream CI build ( dask#8681 ) James Bourbeau
Temporarily pin scipy < 1.8.0 in CI ( dask#8683 ) James Bourbeau
Pin scipy to less than 1.8.0 in GPU CI ( dask#8698 ) Julia Signell
Avoid pytest.warns(None) in test_multi.py ( dask#8678 ) James Bourbeau
Update GHA concurrent job cancellation ( dask#8652 ) James Bourbeau
Make test__get_paths robust to site.PREFIXES being set ( dask#8644 ) James Bourbeau
Bump gpuCI PYTHON_VER to 3.9 ( dask#8642 ) Charles Blackmon-Luca
Released on February 11, 2022
Note
This is the last release with support for Python 3.7
Add region to to_zarr when using existing array (8590) `Chris Roat`_
Add engine_kwargs support to dask.dataframe.to_sql (8609) `Amir Kadivar`_
Add include_path_column arg to read_json (8603) `Bryan Weber`_
Add expand_dims to Dask array (8687) `Tom White`_
Add scheduler option to assert_eq utilities (8610) `Xinrong Meng`_
Fix eye inconsistency with NumPy for dtype=None (8685) `Tom White`_
Fix concatenate inconsistency with NumPy for axis=None (8686) `Tom White`_
Type annotations, part 1 (8295) `crusaderky`_
Really allow any iterable to be passed as a meta (8629) `Julia Signell`_
Use map_partitions (Blockwise) in to_parquet (8487) `Richard (Rick) Zamora`_
Result of reducing an array should not depend on its chunk-structure (8637) `ParticularMiner`_
Pass place-holder metadata to map_partitions in ACA code path (8643) `Richard (Rick) Zamora`_
Deprecate is_monotonic (8653) `James Bourbeau`_
Remove some deprecations (8605) `James Bourbeau`_
Add Domino Data Lab to Hosted / managed Dask clusters (8675) `Ray Bell`_
Fix inter-linking and remove deprecated function (8715) `Julia Signell`_
Fix imbalanced backticks. (8693) `Matthias Bussonnier`_
Add documentation for high level graph visualization (8483) `Genevieve Buckley`_
Update documentation of ProgressBar out parameter (8604) `Pedro Silva`_
Improve documentation of dask.config.set (8705) `crusaderky`_
Revert mention to mypy among type checkers (8699) `crusaderky`_
Update warning handling in get_dummies tests (8651) `James Bourbeau`_
Add a github changelog template (8714) `Julia Signell`_
Update year in LICENSE.txt (8665) `David Hoese`_
Update pre-commit version (8691) `James Bourbeau`_
Include scipy in upstream CI build (8681) `James Bourbeau`_
Temporarily pin scipy < 1.8.0 in CI (8683) `James Bourbeau`_
Pin scipy to less than 1.8.0 in GPU CI (8698) `Julia Signell`_
Avoid pytest.warns(None) in test_multi.py (8678) `James Bourbeau`_
Update GHA concurrent job cancellation (8652) `James Bourbeau`_
Make test__get_paths robust to site.PREFIXES being set (8644) `James Bourbeau`_
Bump gpuCI PYTHON_VER to 3.9 (8642) `Charles Blackmon-Luca`_
- Pandas compat: Deprecate append when pandas >= 1.4.0 ( dask#8617 ) Julia Signell
Released on January 28, 2022
Add dask.dataframe.series.view() ( dask#8533 ) Pavithra Eswaramoorthy
Update tz for fastparquet + pandas 1.4.0 ( dask#8626 ) Martin Durant
Cleaning up misc tests for pandas compat ( dask#8623 ) Julia Signell
Moving to SQLAlchemy >= 1.4 ( dask#8158 ) McToel
Pandas compat: Filter sparse warnings ( dask#8621 ) Julia Signell
Fail if meta is not a pandas object ( dask#8563 ) Julia Signell
Use fsspec.parquet module for better remote-storage read_parquet performance ( dask#8339 ) Richard (Rick) Zamora
Move DataFrame ACA aggregations to HLG ( dask#8468 ) Richard (Rick) Zamora
Add optional information about originating function call in DataFrameIOLayer ( dask#8453 ) Richard (Rick) Zamora
Blockwise array creation redux ( dask#7417 ) Ian Rose
Refactor config default search path retrieval ( dask#8573 ) James Bourbeau
Add optimize_graph flag to Bag.to_dataframe function ( dask#8486 ) Maxim Lippeveld
Make sure that delayed output operations still return lists of paths ( dask#8498 ) Julia Signell
Pandas compat: Fix to_frame name to not pass None ( dask#8554 ) Julia Signell
Pandas compat: Fix axis=None warning ( dask#8555 ) Julia Signell
Expand Dask YAML config search directories ( dask#8531 ) abergou
Fix groupby.cumsum with series grouped by index ( dask#8588 ) Julia Signell
Fix derived_from for pandas methods ( dask#8612 ) Thomas J. Fan
Enforce boolean ascending for sort_values ( dask#8440 ) Charles Blackmon-Luca
Fix parsing of setitem indices ( dask#8601 ) David Hassell
Avoid divide by zero in slicing ( dask#8597 ) Doug Davis
Downgrade meta error in ( dask#8563 ) to warning ( dask#8628 ) Julia Signell
Pandas compat: Deprecate append when pandas >= 1.4.0 ( dask#8617 ) Julia Signell
Replace outdated columns argument with meta in DataFrame constructor ( dask#8614 ) kori73
Refactor deploying docs ( dask#8602 ) Jacob Tomlinson
Pin coverage in CI ( dask#8631 ) James Bourbeau
Move cached_cumsum imports to be from dask.utils ( dask#8606 ) James Bourbeau
Update gpuCI RAPIDS_VER to 22.04 ( dask#8600 )
Update cocstring for from_delayed function ( dask#8576 ) Kirito1397
Handle plot_width / plot_height deprecations ( dask#8544 ) Bryan Van de Ven
Remove unnecessary pyyaml importorskip ( dask#8562 ) James Bourbeau
Specify scheduler in DataFrame assert_eq ( dask#8559 ) Gabe Joseph
Released on January 28, 2022
Add dask.dataframe.series.view() (8533) `Pavithra Eswaramoorthy`_
Update tz for fastparquet + pandas 1.4.0 (8626) `Martin Durant`_
Cleaning up misc tests for pandas compat (8623) `Julia Signell`_
Moving to SQLAlchemy >= 1.4 (8158) `McToel`_
Pandas compat: Filter sparse warnings (8621) `Julia Signell`_
Fail if meta is not a pandas object (8563) `Julia Signell`_
Use fsspec.parquet module for better remote-storage read_parquet performance (8339) `Richard (Rick) Zamora`_
Move DataFrame ACA aggregations to HLG (8468) `Richard (Rick) Zamora`_
Add optional information about originating function call in DataFrameIOLayer (8453) `Richard (Rick) Zamora`_
Blockwise array creation redux (7417) `Ian Rose`_
Refactor config default search path retrieval (8573) `James Bourbeau`_
Add optimize_graph flag to Bag.to_dataframe function (8486) `Maxim Lippeveld`_
Make sure that delayed output operations still return lists of paths (8498) `Julia Signell`_
Pandas compat: Fix to_frame name to not pass None (8554) `Julia Signell`_
Pandas compat: Fix axis=None warning (8555) `Julia Signell`_
Expand Dask YAML config search directories (8531) `abergou`_
Fix groupby.cumsum with series grouped by index (8588) `Julia Signell`_
Fix derived_from for pandas methods (8612) `Thomas J. Fan`_
Enforce boolean ascending for sort_values (8440) `Charles Blackmon-Luca`_
Fix parsing of __setitem__ indices (8601) `David Hassell`_
Avoid divide by zero in slicing (8597) `Doug Davis`_
Downgrade meta error in (8563) to warning (8628) `Julia Signell`_
Pandas compat: Deprecate append when pandas >= 1.4.0 (8617) `Julia Signell`_
Replace outdated columns argument with meta in DataFrame constructor (8614) `kori73`_
Refactor deploying docs (8602) `Jacob Tomlinson`_
Pin coverage in CI (8631) `James Bourbeau`_
Move cached_cumsum imports to be from dask.utils (8606) `James Bourbeau`_
Update gpuCI RAPIDS_VER to 22.04 (8600)
Update cocstring for from_delayed function (8576) `Kirito1397`_
Handle plot_width / plot_height deprecations (8544) `Bryan Van de Ven`_
Remove unnecessary pyyaml importorskip (8562) `James Bourbeau`_
Specify scheduler in DataFrame assert_eq (8559) `Gabe Joseph`_
- Add groupby.shift method ( dask#8522 ) kori73
Released on January 14, 2022
Add groupby.shift method ( dask#8522 ) kori73
Add DataFrame.nunique ( dask#8479 ) Sarah Charlotte Johnson
Add da.ndim to match np.ndim ( dask#8502 ) Julia Signell
Only show percentile interpolation= keyword warning if NumPy version >= 1.22 ( dask#8564 ) Julia Signell
Raise PerformanceWarning when limit and "array.slicing.split-large-chunks" are None ( dask#8511 ) Julia Signell
Define normalize_seq function at import time ( dask#8521 ) Illviljan
Ensure that divisions are alway tuples ( dask#8393 ) Charles Blackmon-Luca
Allow a callable scheduler for bag.groupby ( dask#8492 ) Julia Signell
Save Zarr arrays with dask-on-ray scheduler ( dask#8472 ) TnTo
Make byte blocks more even in read_bytes ( dask#8459 ) Martin Durant
Improved the efficiency of matmul() by completely removing concatenation ( dask#8423 ) ParticularMiner
Limit max chunk size when reshaping dask arrays ( dask#8124 ) Genevieve Buckley
Changes for fastparquet superthrift ( dask#8470 ) Martin Durant
Fix boolean indices in array assignment ( dask#8538 ) David Hassell
Detect default dtype on array-likes ( dask#8501 ) aeisenbarth
Fix optimize_blockwise bug for duplicate dependency names ( dask#8542 ) Richard (Rick) Zamora
Update warnings for DataFrame.GroupBy.apply and transform ( dask#8507 ) Sarah Charlotte Johnson
Track HLG layer name in Delayed ( dask#8452 ) Gabe Joseph
Fix single item nanmin and nanmax reductions ( dask#8484 ) Julia Signell
Make read_csv with comment kwarg work even if there is a comment in the header ( dask#8433 ) Julia Signell
Replace interpolation with method and method with internal_method ( dask#8525 ) Julia Signell
Remove daily stock demo utility ( dask#8477 ) James Bourbeau
Add a join example in docs that be run with copy/paste ( dask#8520 ) kori73
Mention dashboard link in config ( dask#8510 ) Ray Bell
Fix changelog section hyperlinks ( dask#8534 ) Aneesh Nema
Hyphenate “single-machine scheduler” for consistency ( dask#8519 ) Deepyaman Datta
Normalize whitespace in doctests in slicing.py ( dask#8512 ) Maren Westermann
Best practices storage line typo ( dask#8529 ) Michael Delgado
Update figures ( dask#8401 ) Sarah Charlotte Johnson
Remove pyarrow -only reference from split_row_groups in read_parquet docstring ( dask#8490 ) Naty Clementi
Remove obsolete LocalFileSystem tests that fail for fsspec>=2022.1.0 ( dask#8565 ) Richard (Rick) Zamora
Tweak: “RuntimeWarning: invalid value encountered in reciprocal” ( dask#8561 ) crusaderky
Fix skipna=None for DataFrame.sem ( dask#8556 ) Julia Signell
Fix PANDAS_GT_140 ( dask#8552 ) Julia Signell
Collections with HLG must always implement dask_layers ( dask#8548 ) crusaderky
Work around race condition in import llvmlite ( dask#8550 ) crusaderky
Set a minimum version for pyyaml ( dask#8545 ) Gaurav Sheni
Adding nodefaults to environments to fix tiledb + mac issue ( dask#8505 ) Julia Signell
Set ceiling for setuptools ( dask#8509 ) Julia Signell
Add workflow / recipe to generate Dask nightlies ( dask#8469 ) Charles Blackmon-Luca
Bump gpuCI CUDA_VER to 11.5 ( dask#8489 ) Charles Blackmon-Luca
Released on January 14, 2022
Add groupby.shift method (8522) `kori73`_
Add DataFrame.nunique (8479) `Sarah Charlotte Johnson`_
Add da.ndim to match np.ndim (8502) `Julia Signell`_
Only show percentile interpolation= keyword warning if NumPy version >= 1.22 (8564) `Julia Signell`_
Raise PerformanceWarning when limit and "array.slicing.split-large-chunks" are None (8511) `Julia Signell`_
Define normalize_seq function at import time (8521) `Illviljan`_
Ensure that divisions are alway tuples (8393) `Charles Blackmon-Luca`_
Allow a callable scheduler for bag.groupby (8492) `Julia Signell`_
Save Zarr arrays with dask-on-ray scheduler (8472) `TnTo`_
Make byte blocks more even in read_bytes (8459) `Martin Durant`_
Improved the efficiency of matmul() by completely removing concatenation (8423) `ParticularMiner`_
Limit max chunk size when reshaping dask arrays (8124) `Genevieve Buckley`_
Changes for fastparquet superthrift (8470) `Martin Durant`_
Fix boolean indices in array assignment (8538) `David Hassell`_
Detect default dtype on array-likes (8501) `aeisenbarth`_
Fix optimize_blockwise bug for duplicate dependency names (8542) `Richard (Rick) Zamora`_
Update warnings for DataFrame.GroupBy.apply and transform (8507) `Sarah Charlotte Johnson`_
Track HLG layer name in Delayed (8452) `Gabe Joseph`_
Fix single item nanmin and nanmax reductions (8484) `Julia Signell`_
Make read_csv with comment kwarg work even if there is a comment in the header (8433) `Julia Signell`_
Replace interpolation with method and method with internal_method (8525) `Julia Signell`_
Remove daily stock demo utility (8477) `James Bourbeau`_
Add a join example in docs that be run with copy/paste (8520) `kori73`_
Mention dashboard link in config (8510) `Ray Bell`_
Fix changelog section hyperlinks (8534) `Aneesh Nema`_
Hyphenate "single-machine scheduler" for consistency (8519) `Deepyaman Datta`_
Normalize whitespace in doctests in slicing.py (8512) `Maren Westermann`_
Best practices storage line typo (8529) `Michael Delgado`_
Update figures (8401) `Sarah Charlotte Johnson`_
Remove pyarrow-only reference from split_row_groups in read_parquet docstring (8490) `Naty Clementi`_
Remove obsolete LocalFileSystem tests that fail for fsspec>=2022.1.0 (8565) `Richard (Rick) Zamora`_
Tweak: "RuntimeWarning: invalid value encountered in reciprocal" (8561) `crusaderky`_
Fix skipna=None for DataFrame.sem (8556) `Julia Signell`_
Fix PANDAS_GT_140 (8552) `Julia Signell`_
Collections with HLG must always implement __dask_layers__ (8548) `crusaderky`_
Work around race condition in import llvmlite (8550) `crusaderky`_
Set a minimum version for pyyaml (8545) `Gaurav Sheni`_
Adding nodefaults to environments to fix tiledb + mac issue (8505) `Julia Signell`_
Set ceiling for setuptools (8509) `Julia Signell`_
Add workflow / recipe to generate Dask nightlies (8469) `Charles Blackmon-Luca`_
Bump gpuCI CUDA_VER to 11.5 (8489) `Charles Blackmon-Luca`_
- Deprecate token keyword argument to map_blocks ( dask#8464 ) James Bourbeau
Released on December 10, 2021
Add Series and Index is_monotonic* methods ( dask#8304 ) Daniel Mesejo-León
Blockwise map_partitions with partition_info ( dask#8310 ) Gabe Joseph
Better error message for length of array with unknown chunk sizes ( dask#8436 ) Doug Davis
Use by instead of index internally on the Groupby class ( dask#8441 ) Julia Signell
Allow custom sort functions for sort_values ( dask#8345 ) Charles Blackmon-Luca
Add warning to read_parquet when statistics and partitions are misaligned ( dask#8416 ) Richard (Rick) Zamora
Support where argument in ufuncs ( dask#8253 ) mihir
Make visualize more consistent with compute ( dask#8328 ) JSKenyon
Fix map_blocks not using own arguments in name generation ( dask#8462 ) David Hoese
Fix for index error with reading empty parquet file ( dask#8410 ) Sarah Charlotte Johnson
Fix nullable-dtype error when writing partitioned parquet data ( dask#8400 ) Richard (Rick) Zamora
Fix CSV header bug ( dask#8413 ) Richard (Rick) Zamora
Fix empty chunk causes exception in nanmin / nanmax ( dask#8375 ) Boaz Mohar
Deprecate token keyword argument to map_blocks ( dask#8464 ) James Bourbeau
Deprecation warning for default value of boundary kwarg in map_overlap ( dask#8397 ) Genevieve Buckley
Clarify block_info documentation ( dask#8425 ) Genevieve Buckley
Output from alt text sprint ( dask#8456 ) Sarah Charlotte Johnson
Update talks and presentations ( dask#8370 ) Naty Clementi
Update Anaconda link in “Paid support” section of docs ( dask#8427 ) Martin Durant
Fixed broken dask-gateway link in ecosystem.rst ( dask#8424 ) ofirr
Fix CuPy doctest error ( dask#8412 ) Genevieve Buckley
Bump Bokeh min version to 2.1.1 ( dask#8431 ) Bryan Van de Ven
Fix following fsspec=2021.11.1 release ( dask#8428 ) Martin Durant
Add dask/ml.py to pytest exclude list ( dask#8414 ) Genevieve Buckley
Update gpuCI RAPIDS_VER to 22.02 ( dask#8394 )
Unpin graphviz and improve package management in environment-3.7 ( dask#8411 ) Julia Signell
Released on December 10, 2021
Add Series and Index is_monotonic* methods (8304) `Daniel Mesejo-León`_
Blockwise map_partitions with partition_info (8310) `Gabe Joseph`_
Better error message for length of array with unknown chunk sizes (8436) `Doug Davis`_
Use by instead of index internally on the Groupby class (8441) `Julia Signell`_
Allow custom sort functions for sort_values (8345) `Charles Blackmon-Luca`_
Add warning to read_parquet when statistics and partitions are misaligned (8416) `Richard (Rick) Zamora`_
Support where argument in ufuncs (8253) `mihir`_
Make visualize more consistent with compute (8328) `JSKenyon`_
Fix map_blocks not using own arguments in name generation (8462) `David Hoese`_
Fix for index error with reading empty parquet file (8410) `Sarah Charlotte Johnson`_
Fix nullable-dtype error when writing partitioned parquet data (8400) `Richard (Rick) Zamora`_
Fix CSV header bug (8413) `Richard (Rick) Zamora`_
Fix empty chunk causes exception in nanmin/nanmax (8375) `Boaz Mohar`_
Deprecate token keyword argument to map_blocks (8464) `James Bourbeau`_
Deprecation warning for default value of boundary kwarg in map_overlap (8397) `Genevieve Buckley`_
Clarify block_info documentation (8425) `Genevieve Buckley`_
Output from alt text sprint (8456) `Sarah Charlotte Johnson`_
Update talks and presentations (8370) `Naty Clementi`_
Update Anaconda link in "Paid support" section of docs (8427) `Martin Durant`_
Fixed broken dask-gateway link in ecosystem.rst (8424) `ofirr`_
Fix CuPy doctest error (8412) `Genevieve Buckley`_
Bump Bokeh min version to 2.1.1 (8431) `Bryan Van de Ven`_
Fix following fsspec=2021.11.1 release (8428) `Martin Durant`_
Add dask/ml.py to pytest exclude list (8414) `Genevieve Buckley`_
Update gpuCI RAPIDS_VER to 22.02 (8394)
Unpin graphviz and improve package management in environment-3.7 (8411) `Julia Signell`_
- Only run gpuCI bump script daily ( dask#8404 ) Charles Blackmon-Luca
Released on November 19, 2021
Only run gpuCI bump script daily ( dask#8404 ) Charles Blackmon-Luca
Actually ignore index when asked in assert_eq ( dask#8396 ) Gabe Joseph
Ensure single-partition join divisions is tuple ( dask#8389 ) Charles Blackmon-Luca
Try to make divisions behavior clearer ( dask#8379 ) Julia Signell
Fix typo in set_index partition_size parameter description ( dask#8384 ) FredericOdermatt
Use blockwise in single_partition_join ( dask#8341 ) Gabe Joseph
Use more explicit keyword arguments ( dask#8354 ) Boaz Mohar
Fix .loc of DataFrame with nullable boolean dtype ( dask#8368 ) Marco Rossi
Parameterize shuffle implementation in tests ( dask#8250 ) Ian Rose
Remove some doc build warnings ( dask#8369 ) Boaz Mohar
Include properties in array API docs ( dask#8356 ) Julia Signell
Fix Zarr for upstream ( dask#8367 ) Julia Signell
Pin graphviz to avoid issue with windows and Python 3.7 ( dask#8365 ) Julia Signell
Import graphviz.Diagraph from top of module, not from dot ( dask#8363 ) Julia Signell
Released on November 19, 2021
Only run gpuCI bump script daily (8404) `Charles Blackmon-Luca`_
Actually ignore index when asked in assert_eq (8396) `Gabe Joseph`_
Ensure single-partition join divisions is tuple (8389) `Charles Blackmon-Luca`_
Try to make divisions behavior clearer (8379) `Julia Signell`_
Fix typo in set_index partition_size parameter description (8384) `FredericOdermatt`_
Use blockwise in single_partition_join (8341) `Gabe Joseph`_
Use more explicit keyword arguments (8354) `Boaz Mohar`_
Fix .loc of DataFrame with nullable boolean dtype (8368) `Marco Rossi`_
Parameterize shuffle implementation in tests (8250) `Ian Rose`_
Remove some doc build warnings (8369) `Boaz Mohar`_
Include properties in array API docs (8356) `Julia Signell`_
Fix Zarr for upstream (8367) `Julia Signell`_
Pin graphviz to avoid issue with windows and Python 3.7 (8365) `Julia Signell`_
Import graphviz.Diagraph from top of module, not from dot (8363) `Julia Signell`_
Patch release to update distributed dependency to version 2021.11.1 .
Released on November 8, 2021
Patch release to update distributed dependency to version 2021.11.1 .
Released on November 8, 2021
Patch release to update distributed dependency to version 2021.11.1.
- Deprecate AxisError ( dask#8305 ) crusaderky
Released on November 5, 2021
Fx required_extension behavior in read_parquet ( dask#8351 ) Richard (Rick) Zamora
Add align_dataframes to map_partitions to broadcast a dataframe passed as an arg ( dask#6628 ) Julia Signell
Better handling for arrays/series of keys in dask.dataframe.loc ( dask#8254 ) Julia Signell
Point users to Discourse ( dask#8332 ) Ian Rose
Add name_function option to to_parquet ( dask#7682 ) Matthew Powers
Get rid of environment-latest.yml and update to Python 3.9 ( dask#8275 ) Julia Signell
Require newer s3fs in CI ( dask#8336 ) James Bourbeau
Groupby Rolling ( dask#8176 ) Julia Signell
Add more ordering diagnostics to dask.visualize ( dask#7992 ) Erik Welch
Use HighLevelGraph optimizations for delayed ( dask#8316 ) Ian Rose
demo_tuples produces malformed HighLevelGraph ( dask#8325 ) crusaderky
Dask calendar should show events in local time ( dask#8312 ) Genevieve Buckley
Fix flaky test_interrupt ( dask#8314 ) crusaderky
Deprecate AxisError ( dask#8305 ) crusaderky
Fix name of cuDF in extension documentation. ( dask#8311 ) Vyas Ramasubramani
Add single eq operator (=) to parquet filters ( dask#8300 ) Ayush Dattagupta
Improve support for Spark output in read_parquet ( dask#8274 ) Richard (Rick) Zamora
Add dask.ml module ( dask#6384 ) Matthew Rocklin
CI fixups ( dask#8298 ) James Bourbeau
Make slice errors match NumPy ( dask#8248 ) Julia Signell
Fix API docs misrendering with new sphinx theme ( dask#8296 ) Julia Signell
Replace block property with blockview for array-like operations on blocks ( dask#8242 ) Davis Bennett
Deprecate file_path and make it possible to save from within a notebook ( dask#8283 ) Julia Signell
Released on November 5, 2021
Fx required_extension behavior in read_parquet (8351) `Richard (Rick) Zamora`_
Add align_dataframes to map_partitions to broadcast a dataframe passed as an arg (6628) `Julia Signell`_
Better handling for arrays/series of keys in dask.dataframe.loc (8254) `Julia Signell`_
Point users to Discourse (8332) `Ian Rose`_
Add name_function option to to_parquet (7682) `Matthew Powers`_
Get rid of environment-latest.yml and update to Python 3.9 (8275) `Julia Signell`_
Require newer s3fs in CI (8336) `James Bourbeau`_
Groupby Rolling (8176) `Julia Signell`_
Add more ordering diagnostics to dask.visualize (7992) `Erik Welch`_
Use HighLevelGraph optimizations for delayed (8316) `Ian Rose`_
demo_tuples produces malformed HighLevelGraph (8325) `crusaderky`_
Dask calendar should show events in local time (8312) `Genevieve Buckley`_
Fix flaky test_interrupt (8314) `crusaderky`_
Deprecate AxisError (8305) `crusaderky`_
Fix name of cuDF in extension documentation. (8311) `Vyas Ramasubramani`_
Add single eq operator (=) to parquet filters (8300) `Ayush Dattagupta`_
Improve support for Spark output in read_parquet (8274) `Richard (Rick) Zamora`_
Add dask.ml module (6384) `Matthew Rocklin`_
CI fixups (8298) `James Bourbeau`_
Make slice errors match NumPy (8248) `Julia Signell`_
Fix API docs misrendering with new sphinx theme (8296) `Julia Signell`_
Replace block property with blockview for array-like operations on blocks (8242) `Davis Bennett`_
Deprecate file_path and make it possible to save from within a notebook (8283) `Julia Signell`_
- Deprecate inplace argument for Dask series renaming ( dask#8136 ) Marcel Coetzee
Released on October 22, 2021
da.store to create well-formed HighLevelGraph ( dask#8261 ) crusaderky
CI: force nightly pyarrow in the upstream build ( dask#8281 ) Joris Van den Bossche
Remove chest ( dask#8279 ) James Bourbeau
Skip doctests if optional dependencies are not installed ( dask#8258 ) Genevieve Buckley
Update tmpdir and tmpfile context manager docstrings ( dask#8270 ) Daniel Mesejo-León
Unregister callbacks in doctests ( dask#8276 ) James Bourbeau
Fix typo in docs ( dask#8277 ) JoranDox
Stale label GitHub action ( dask#8244 ) Genevieve Buckley
Client-shutdown method appears twice ( dask#8273 ) German Shiklov
Add pre-commit to test requirements ( dask#8257 ) Genevieve Buckley
Refactor read_metadata in fastparquet engine ( dask#8092 ) Richard (Rick) Zamora
Support Path objects in from_zarr ( dask#8266 ) Samuel Gaist
Make nested redirects work ( dask#8272 ) Julia Signell
Set memory_usage to True if verbose is True in info ( dask#8222 ) Kinshuk Dua
Remove individual API doc pages from sphinx toctree ( dask#8238 ) James Bourbeau
Ignore whitespace in gufunc signature ( dask#8267 ) James Bourbeau
Add workflow to update gpuCI ( dask#8215 ) Charles Blackmon-Luca
DataFrame.head shouldn’t warn when there’s one partition ( dask#8091 ) Pankaj Patil
Ignore arrow doctests if pyarrow not installed ( dask#8256 ) Genevieve Buckley
Fix debugging.html redirect ( dask#8251 ) James Bourbeau
Fix null sorting for single partition dataframes ( dask#8225 ) Charles Blackmon-Luca
Fix setup.html redirect ( dask#8249 ) Florian Jetter
Run pyupgrade in CI ( dask#8246 ) crusaderky
Fix label typo in upstream CI build ( dask#8237 ) James Bourbeau
Add support for “dependent” columns in DataFrame.assign ( dask#8086 ) Suriya Senthilkumar
add NumPy array of Dask keys to Array ( dask#7922 ) Davis Bennett
Remove unnecessary dask.multiprocessing import in docs ( dask#8240 ) Ray Bell
Adjust retrieving _max_workers from Executor ( dask#8228 ) John A Kirkham
Update function signatures in delayed best practices docs ( dask#8231 ) Vũ Trung Đức
Docs reoganization ( dask#7984 ) Julia Signell
Fix df.quantile on all missing data ( dask#8129 ) Julia Signell
Add tokenize.ensure-deterministic config option ( dask#7413 ) Hristo Georgiev
Use inclusive rather than closed with pandas>=1.4.0 and pd.date_range ( dask#8213 ) Julia Signell
Add dask-gateway , Coiled, and Saturn-Cloud to list of Dask setup tools ( dask#7814 ) Kristopher Overholt
Ensure existing futures get passed as deps when serializing HighLevelGraph layers ( dask#8199 ) Jim Crist-Harif
Make sure that the divisions of the single partition merge is left ( dask#8162 ) Julia Signell
Refactor read_metadata in pyarrow parquet engines ( dask#8072 ) Richard (Rick) Zamora
Support negative drop_axis in map_blocks and map_overlap ( dask#8192 ) Gregory R. Lee
Fix upstream tests ( dask#8205 ) Julia Signell
Add support for scalar item assignment by Series ( dask#8195 ) Charles Blackmon-Luca
Add some basic examples to doc strings on dask.bag all , any , count methods ( dask#7630 ) Nathan Danielsen
Don’t have upstream report depend on commit message ( dask#8202 ) James Bourbeau
Ensure upstream CI cron job runs ( dask#8200 ) James Bourbeau
Use pytest.param to properly label param-specific GPU tests ( dask#8197 ) Charles Blackmon-Luca
Add test_set_index to tests ran on gpuCI ( dask#8198 ) Charles Blackmon-Luca
Suppress tmpfile OSError ( dask#8191 ) James Bourbeau
Use s.isna instead of pd.isna(s) in set_partitions_pre (fix cudf CI) ( dask#8193 ) Charles Blackmon-Luca
Open an issue for test-upstream failures ( dask#8067 ) Wallace Reis
Fix to_parquet bug in call to pyarrow.parquet.read_metadata ( dask#8186 ) Richard (Rick) Zamora
Add handling for null values in sort_values ( dask#8167 ) Charles Blackmon-Luca
Bump RAPIDS_VER for gpuCI ( dask#8184 ) Charles Blackmon-Luca
Dispatch walks MRO for lazily registered handlers ( dask#8185 ) Jim Crist-Harif
Configure SSHCluster instructions ( dask#8181 ) Ray Bell
Preserve HighLevelGraphs in DataFrame.from_delayed ( dask#8174 ) Gabe Joseph
Deprecate inplace argument for Dask series renaming ( dask#8136 ) Marcel Coetzee
Fix rolling for compatibility with pandas > 1.3.0 ( dask#8150 ) Julia Signell
Raise error when setitem on unknown chunks ( dask#8166 ) Julia Signell
Include divisions when doing Index.to_series ( dask#8165 ) Julia Signell
Released on October 22, 2021
da.store to create well-formed HighLevelGraph (8261) `crusaderky`_
CI: force nightly pyarrow in the upstream build (8281) `Joris Van den Bossche`_
Remove chest (8279) `James Bourbeau`_
Skip doctests if optional dependencies are not installed (8258) `Genevieve Buckley`_
Update tmpdir and tmpfile context manager docstrings (8270) `Daniel Mesejo-León`_
Unregister callbacks in doctests (8276) `James Bourbeau`_
Fix typo in docs (8277) `JoranDox`_
Stale label GitHub action (8244) `Genevieve Buckley`_
Client-shutdown method appears twice (8273) `German Shiklov`_
Add pre-commit to test requirements (8257) `Genevieve Buckley`_
Refactor read_metadata in fastparquet engine (8092) `Richard (Rick) Zamora`_
Support Path objects in from_zarr (8266) `Samuel Gaist`_
Make nested redirects work (8272) `Julia Signell`_
Set memory_usage to True if verbose is True in info (8222) `Kinshuk Dua`_
Remove individual API doc pages from sphinx toctree (8238) `James Bourbeau`_
Ignore whitespace in gufunc signature (8267) `James Bourbeau`_
Add workflow to update gpuCI (8215) `Charles Blackmon-Luca`_
DataFrame.head shouldn't warn when there's one partition (8091) `Pankaj Patil`_
Ignore arrow doctests if pyarrow not installed (8256) `Genevieve Buckley`_
Fix debugging.html redirect (8251) `James Bourbeau`_
Fix null sorting for single partition dataframes (8225) `Charles Blackmon-Luca`_
Fix setup.html redirect (8249) `Florian Jetter`_
Run pyupgrade in CI (8246) `crusaderky`_
Fix label typo in upstream CI build (8237) `James Bourbeau`_
Add support for "dependent" columns in DataFrame.assign (8086) `Suriya Senthilkumar`_
add NumPy array of Dask keys to Array (7922) `Davis Bennett`_
Remove unnecessary dask.multiprocessing import in docs (8240) `Ray Bell`_
Adjust retrieving _max_workers from Executor (8228) `John A Kirkham`_
Update function signatures in delayed best practices docs (8231) `Vũ Trung Đức`_
Docs reoganization (7984) `Julia Signell`_
Fix df.quantile on all missing data (8129) `Julia Signell`_
Add tokenize.ensure-deterministic config option (7413) `Hristo Georgiev`_
Use inclusive rather than closed with pandas>=1.4.0 and pd.date_range (8213) `Julia Signell`_
Add dask-gateway, Coiled, and Saturn-Cloud to list of Dask setup tools (7814) `Kristopher Overholt`_
Ensure existing futures get passed as deps when serializing HighLevelGraph layers (8199) `Jim Crist-Harif`_
Make sure that the divisions of the single partition merge is left (8162) `Julia Signell`_
Refactor read_metadata in pyarrow parquet engines (8072) `Richard (Rick) Zamora`_
Support negative drop_axis in map_blocks and map_overlap (8192) `Gregory R. Lee`_
Fix upstream tests (8205) `Julia Signell`_
Add support for scalar item assignment by Series (8195) `Charles Blackmon-Luca`_
Add some basic examples to doc strings on dask.bag all, any, count methods (7630) `Nathan Danielsen`_
Don't have upstream report depend on commit message (8202) `James Bourbeau`_
Ensure upstream CI cron job runs (8200) `James Bourbeau`_
Use pytest.param to properly label param-specific GPU tests (8197) `Charles Blackmon-Luca`_
Add test_set_index to tests ran on gpuCI (8198) `Charles Blackmon-Luca`_
Suppress tmpfile OSError (8191) `James Bourbeau`_
Use s.isna instead of pd.isna(s) in set_partitions_pre (fix cudf CI) (8193) `Charles Blackmon-Luca`_
Open an issue for test-upstream failures (8067) `Wallace Reis`_
Fix to_parquet bug in call to pyarrow.parquet.read_metadata (8186) `Richard (Rick) Zamora`_
Add handling for null values in sort_values (8167) `Charles Blackmon-Luca`_
Bump RAPIDS_VER for gpuCI (8184) `Charles Blackmon-Luca`_
Dispatch walks MRO for lazily registered handlers (8185) `Jim Crist-Harif`_
Configure SSHCluster instructions (8181) `Ray Bell`_
Preserve HighLevelGraphs in DataFrame.from_delayed (8174) `Gabe Joseph`_
Deprecate inplace argument for Dask series renaming (8136) `Marcel Coetzee`_
Fix rolling for compatibility with pandas > 1.3.0 (8150) `Julia Signell`_
Raise error when setitem on unknown chunks (8166) `Julia Signell`_
Include divisions when doing Index.to_series (8165) `Julia Signell`_
- Remove references to pd.Int64Index in anticipation of deprecation ( dask#8144 ) Julia Signell
Released on September 21, 2021
Fix groupby for future pandas ( dask#8151 ) Julia Signell
Remove warning filters in tests that are no longer needed ( dask#8155 ) Julia Signell
Add link to diagnostic visualize function in local diagnostic docs ( dask#8157 ) David Hoese
Add datetime_is_numeric to dataframe.describe ( dask#7719 ) Julia Signell
Remove references to pd.Int64Index in anticipation of deprecation ( dask#8144 ) Julia Signell
Use loc if needed for series get_item ( dask#7953 ) Julia Signell
Specifically ignore warnings on mean for empty slices ( dask#8125 ) Julia Signell
Skip groupby nunique test for pandas >= 1.3.3 ( dask#8142 ) Julia Signell
Implement ascending arg for sort_values ( dask#8130 ) Charles Blackmon-Luca
Replace operator.getitem ( dask#8015 ) Naty Clementi
Deprecate zero_broadcast_dimensions and homogeneous_deepmap ( dask#8134 ) SnkSynthesis
Add error if drop_index is negative ( dask#8064 ) neel iyer
Allow scheduler to be an Executor ( dask#8112 ) John A Kirkham
Handle asarray / asanyarray cases where like is a dask.Array ( dask#8128 ) Peter Andreas Entschev
Fix index_col duplication if index_col is type str ( dask#7661 ) McToel
Add dtype and order to asarray and asanyarray definitions ( dask#8106 ) Julia Signell
Deprecate dask.dataframe.Series.contains ( dask#7914 ) Julia Signell
Fix edge case with like -arrays in _wrapped_qr ( dask#8122 ) Peter Andreas Entschev
Deprecate boundary_slice kwarg: kind for pandas compat ( dask#8037 ) Julia Signell
Released on September 21, 2021
Fix groupby for future pandas (8151) `Julia Signell`_
Remove warning filters in tests that are no longer needed (8155) `Julia Signell`_
Add link to diagnostic visualize function in local diagnostic docs (8157) `David Hoese`_
Add datetime_is_numeric to dataframe.describe (7719) `Julia Signell`_
Remove references to pd.Int64Index in anticipation of deprecation (8144) `Julia Signell`_
Use loc if needed for series __get_item__ (7953) `Julia Signell`_
Specifically ignore warnings on mean for empty slices (8125) `Julia Signell`_
Skip groupby nunique test for pandas >= 1.3.3 (8142) `Julia Signell`_
Implement ascending arg for sort_values (8130) `Charles Blackmon-Luca`_
Replace operator.getitem (8015) `Naty Clementi`_
Deprecate zero_broadcast_dimensions and homogeneous_deepmap (8134) `SnkSynthesis`_
Add error if drop_index is negative (8064) `neel iyer`_
Allow scheduler to be an Executor (8112) `John A Kirkham`_
Handle asarray/asanyarray cases where like is a dask.Array (8128) `Peter Andreas Entschev`_
Fix index_col duplication if index_col is type str (7661) `McToel`_
Add dtype and order to asarray and asanyarray definitions (8106) `Julia Signell`_
Deprecate dask.dataframe.Series.__contains__ (7914) `Julia Signell`_
Fix edge case with like-arrays in _wrapped_qr (8122) `Peter Andreas Entschev`_
Deprecate boundary_slice kwarg: kind for pandas compat (8037) `Julia Signell`_
- Fewer open files ( dask#7303 ) Julia Signell
Released on September 3, 2021
Fewer open files ( dask#7303 ) Julia Signell
Add FileNotFound to expected http errors ( dask#8109 ) Martin Durant
Add DataFrame.sort_values to API docs ( dask#8107 ) Benjamin Zaitlen
Change to dask.order : be more eager at times ( dask#7929 ) Erik Welch
Add pytest color to CI ( dask#8090 ) James Bourbeau
FIX: make_people works with processes scheduler ( dask#8103 ) Dahn
Adds deep param to Dataframe copy method and restrict it to False ( dask#8068 ) João Paulo Lacerda
Fix typo in configuration docs ( dask#8104 ) Robert Hales
Update formatting in DataFrame.query docstring ( dask#8100 ) James Bourbeau
Un-xfail sparse tests for 0.13.0 release ( dask#8102 ) James Bourbeau
Add axes property to DataFrame and Series ( dask#8069 ) Jordan Jensen
Add CuPy support in da.unique (values only) ( dask#8021 ) Peter Andreas Entschev
Unit tests for sparse.zeros_like (xfailed) ( dask#8093 ) crusaderky
Add explicit like kwarg support to array creation functions ( dask#8054 ) Peter Andreas Entschev
Separate Array and DataFrame mindeps builds ( dask#8079 ) James Bourbeau
Fork out percentile_dispatch to dask.array ( dask#8083 ) GALI PREM SAGAR
Ensure filepath exists in to_parquet ( dask#8057 ) James Bourbeau
Update scheduler plugin usage in test_scheduler_highlevel_graph_unpack_import ( dask#8080 ) James Bourbeau
Add DataFrame.shuffle to API docs ( dask#8076 ) Martin Fleischmann
Order requirements alphabetically ( dask#8073 ) John A Kirkham
Released on September 3, 2021
Fewer open files (7303) `Julia Signell`_
Add FileNotFound to expected http errors (8109) `Martin Durant`_
Add DataFrame.sort_values to API docs (8107) `Benjamin Zaitlen`_
Change to dask.order: be more eager at times (7929) `Erik Welch`_
Add pytest color to CI (8090) `James Bourbeau`_
FIX: make_people works with processes scheduler (8103) `Dahn`_
Adds deep param to Dataframe copy method and restrict it to False (8068) `João Paulo Lacerda`_
Fix typo in configuration docs (8104) `Robert Hales`_
Update formatting in DataFrame.query docstring (8100) `James Bourbeau`_
Un-xfail sparse tests for 0.13.0 release (8102) `James Bourbeau`_
Add axes property to DataFrame and Series (8069) `Jordan Jensen`_
Add CuPy support in da.unique (values only) (8021) `Peter Andreas Entschev`_
Unit tests for sparse.zeros_like (xfailed) (8093) `crusaderky`_
Add explicit like kwarg support to array creation functions (8054) `Peter Andreas Entschev`_
Separate Array and DataFrame mindeps builds (8079) `James Bourbeau`_
Fork out percentile_dispatch to dask.array (8083) `GALI PREM SAGAR`_
Ensure filepath exists in to_parquet (8057) `James Bourbeau`_
Update scheduler plugin usage in test_scheduler_highlevel_graph_unpack_import (8080) `James Bourbeau`_
Add DataFrame.shuffle to API docs (8076) `Martin Fleischmann`_
Order requirements alphabetically (8073) `John A Kirkham`_
- Add ignore_metadata_file option to read_parquet ( pyarrow-dataset and fastparquet support only) ( dask#8034 ) Richard (Rick) Zamora
Released on August 20, 2021
Add ignore_metadata_file option to read_parquet ( pyarrow-dataset and fastparquet support only) ( dask#8034 ) Richard (Rick) Zamora
Add reference to pytest-xdist in dev docs ( dask#8066 ) Julia Signell
Include tz in meta from to_datetime ( dask#8000 ) Julia Signell
CI Infra Docs ( dask#7985 ) Benjamin Zaitlen
Include invalid DataFrame key in assert_eq check ( dask#8061 ) James Bourbeau
Use class when creating DataFrames ( dask#8053 ) Mads R. B. Kristensen
Use development version of distributed in gpuCI build ( dask#7976 ) James Bourbeau
Ignore whitespace when gufunc signature ( dask#8049 ) James Bourbeau
Move pandas import and percentile dispatch refactor ( dask#8055 ) GALI PREM SAGAR
Add colors to represent high level layer types ( dask#7974 ) Freyam Mehta
Upstream instance fix ( dask#8060 ) Jacob Tomlinson
Add dask.widgets and migrate HTML reprs to jinja2 ( dask#8019 ) Jacob Tomlinson
Remove wrap_func_like_safe , not required with NumPy >= 1.17 ( dask#8052 ) Peter Andreas Entschev
Fix threaded scheduler memory backpressure regression ( dask#8040 ) David Hoese
Add percentile dispatch ( dask#8029 ) GALI PREM SAGAR
Use a publicly documented attribute obj in groupby rather than private _selected_obj ( dask#8038 ) GALI PREM SAGAR
Specify module to import rechunk from ( dask#8039 ) Illviljan
Use dict to store data for {nan,}arg{min,max} in certain cases ( dask#8014 ) Peter Andreas Entschev
Fix blocksize description formatting in read_pandas ( dask#8047 ) Louis Maddox
Fix “point” -> “pointers” typo in docs ( dask#8043 ) David Chudzicki
Released on August 20, 2021
Add ignore_metadata_file option to read_parquet (pyarrow-dataset and fastparquet support only) (8034) `Richard (Rick) Zamora`_
Add reference to pytest-xdist in dev docs (8066) `Julia Signell`_
Include tz in meta from to_datetime (8000) `Julia Signell`_
CI Infra Docs (7985) `Benjamin Zaitlen`_
Include invalid DataFrame key in assert_eq check (8061) `James Bourbeau`_
Use __class__ when creating DataFrames (8053) `Mads R. B. Kristensen`_
Use development version of distributed in gpuCI build (7976) `James Bourbeau`_
Ignore whitespace when gufunc signature (8049) `James Bourbeau`_
Move pandas import and percentile dispatch refactor (8055) `GALI PREM SAGAR`_
Add colors to represent high level layer types (7974) `Freyam Mehta`_
Upstream instance fix (8060) `Jacob Tomlinson`_
Add dask.widgets and migrate HTML reprs to jinja2 (8019) `Jacob Tomlinson`_
Remove wrap_func_like_safe, not required with NumPy >= 1.17 (8052) `Peter Andreas Entschev`_
Fix threaded scheduler memory backpressure regression (8040) `David Hoese`_
Add percentile dispatch (8029) `GALI PREM SAGAR`_
Use a publicly documented attribute obj in groupby rather than private _selected_obj (8038) `GALI PREM SAGAR`_
Specify module to import rechunk from (8039) `Illviljan`_
Use dict to store data for {nan,}arg{min,max} in certain cases (8014) `Peter Andreas Entschev`_
Fix blocksize description formatting in read_pandas (8047) `Louis Maddox`_
Fix "point" -> "pointers" typo in docs (8043) `David Chudzicki`_
- Use pytest.warns instead of raises for checking parquet engine deprecation ( dask#7993 ) Joris Van den Bossche
Released on August 13, 2021
Fix to_orc delayed compute behavior ( dask#8035 ) Richard (Rick) Zamora
Don’t convert to low-level task graph in compute_as_if_collection ( dask#7969 ) James Bourbeau
Fix multifile read for hdf ( dask#8033 ) Julia Signell
Resolve warning in distributed tests ( dask#8025 ) James Bourbeau
Update to_orc collection name ( dask#8024 ) James Bourbeau
Resolve skipfooter problem ( dask#7855 ) Ross
Raise NotImplementedError for non-indexable arg passed to to_datetime ( dask#7989 ) Doug Davis
Ensure we error on warnings from distributed ( dask#8002 ) James Bourbeau
Added dict format in to_bag accessories of DataFrame ( dask#7932 ) gurunath
Delayed docs indirect dependencies ( dask#8016 ) aa1371
Add tooltips to graphviz high-level graphs ( dask#7973 ) Freyam Mehta
Close 2021 User Survey ( dask#8007 ) Julia Signell
Reorganize CuPy tests into multiple files ( dask#8013 ) Peter Andreas Entschev
Refactor and Expand Dask-Dataframe ORC API ( dask#7756 ) Richard (Rick) Zamora
Don’t enforce columns if enforce=False ( dask#7916 ) Julia Signell
Fix map_overlap trimming behavior when drop_axis is not None ( dask#7894 ) Gregory R. Lee
Mark gpuCI CuPy test as flaky ( dask#7994 ) Peter Andreas Entschev
Avoid using Delayed in to_csv and to_parquet ( dask#7968 ) Matthew Rocklin
Removed redundant check_dtypes ( dask#7952 ) gurunath
Use pytest.warns instead of raises for checking parquet engine deprecation ( dask#7993 ) Joris Van den Bossche
Bump RAPIDS_VER in gpuCI to 21.10 ( dask#7991 ) Charles Blackmon-Luca
Add back pyarrow-legacy test coverage for pyarrow>=5 ( dask#7988 ) Richard (Rick) Zamora
Allow pyarrow>=5 in to_parquet and read_parquet ( dask#7967 ) Richard (Rick) Zamora
Skip CuPy tests requiring NEP-35 when NumPy < 1.20 is available ( dask#7982 ) Peter Andreas Entschev
Add tail and head to SeriesGroupby ( dask#7935 ) Daniel Mesejo-León
Update Zoom link for monthly meeting ( dask#7979 ) James Bourbeau
Add gpuCI build script ( dask#7966 ) Charles Blackmon-Luca
Deprecate daily_stock utility ( dask#7949 ) James Bourbeau
Add distributed.nanny to configuration reference docs ( dask#7955 ) James Bourbeau
Require NumPy 1.18+ & Pandas 1.0+ ( dask#7939 ) John A Kirkham
Released on August 13, 2021
Fix to_orc delayed compute behavior (8035) `Richard (Rick) Zamora`_
Don't convert to low-level task graph in compute_as_if_collection (7969) `James Bourbeau`_
Fix multifile read for hdf (8033) `Julia Signell`_
Resolve warning in distributed tests (8025) `James Bourbeau`_
Update to_orc collection name (8024) `James Bourbeau`_
Resolve skipfooter problem (7855) `Ross`_
Raise NotImplementedError for non-indexable arg passed to to_datetime (7989) `Doug Davis`_
Ensure we error on warnings from distributed (8002) `James Bourbeau`_
Added dict format in to_bag accessories of DataFrame (7932) `gurunath`_
Delayed docs indirect dependencies (8016) `aa1371`_
Add tooltips to graphviz high-level graphs (7973) `Freyam Mehta`_
Close 2021 User Survey (8007) `Julia Signell`_
Reorganize CuPy tests into multiple files (8013) `Peter Andreas Entschev`_
Refactor and Expand Dask-Dataframe ORC API (7756) `Richard (Rick) Zamora`_
Don't enforce columns if enforce=False (7916) `Julia Signell`_
Fix map_overlap trimming behavior when drop_axis is not None (7894) `Gregory R. Lee`_
Mark gpuCI CuPy test as flaky (7994) `Peter Andreas Entschev`_
Avoid using Delayed in to_csv and to_parquet (7968) `Matthew Rocklin`_
Removed redundant check_dtypes (7952) `gurunath`_
Use pytest.warns instead of raises for checking parquet engine deprecation (7993) `Joris Van den Bossche`_
Bump RAPIDS_VER in gpuCI to 21.10 (7991) `Charles Blackmon-Luca`_
Add back pyarrow-legacy test coverage for pyarrow>=5 (7988) `Richard (Rick) Zamora`_
Allow pyarrow>=5 in to_parquet and read_parquet (7967) `Richard (Rick) Zamora`_
Skip CuPy tests requiring NEP-35 when NumPy < 1.20 is available (7982) `Peter Andreas Entschev`_
Add tail and head to SeriesGroupby (7935) `Daniel Mesejo-León`_
Update Zoom link for monthly meeting (7979) `James Bourbeau`_
Add gpuCI build script (7966) `Charles Blackmon-Luca`_
Deprecate daily_stock utility (7949) `James Bourbeau`_
Add distributed.nanny to configuration reference docs (7955) `James Bourbeau`_
Require NumPy 1.18+ & Pandas 1.0+ (7939) `John A Kirkham`_
- Add deprecation warning for top-level ucx and rmm config values ( dask#7956 ) James Bourbeau
Released on July 30, 2021
Note
This is the last release with support for NumPy 1.17 and pandas 0.25. Beginning with the next release, NumPy 1.18 and pandas 1.0 will be the minimum supported versions.
Add dask.array SVG to the HTML Repr ( dask#7886 ) Freyam Mehta
Avoid use of Delayed in to_parquet ( dask#7958 ) Matthew Rocklin
Temporarily pin pyarrow<5 in CI ( dask#7960 ) James Bourbeau
Add deprecation warning for top-level ucx and rmm config values ( dask#7956 ) James Bourbeau
Remove skips from doctests (4 of 6) ( dask#7865 ) Zhengnan Zhao
Remove skips from doctests (5 of 6) ( dask#7864 ) Zhengnan Zhao
Adds missing prepend/append functionality to da.diff ( dask#7946 ) Peter Andreas Entschev
Change graphviz font family to sans ( dask#7931 ) Freyam Mehta
Fix read-csv name - when path is different, use different name for task ( dask#7942 ) Julia Signell
Update configuration reference for ucx and rmm changes ( dask#7943 ) James Bourbeau
Add meta support to setitem ( dask#7940 ) Peter Andreas Entschev
NEP-35 support for slice_with_int_dask_array ( dask#7927 ) Peter Andreas Entschev
Unpin fastparquet in CI ( dask#7928 ) James Bourbeau
Remove skips from doctests (3 of 6) ( dask#7872 ) Zhengnan Zhao
Released on July 30, 2021
Note
This is the last release with support for NumPy 1.17 and pandas 0.25. Beginning with the next release, NumPy 1.18 and pandas 1.0 will be the minimum supported versions.
Add dask.array SVG to the HTML Repr (7886) `Freyam Mehta`_
Avoid use of Delayed in to_parquet (7958) `Matthew Rocklin`_
Temporarily pin pyarrow<5 in CI (7960) `James Bourbeau`_
Add deprecation warning for top-level ucx and rmm config values (7956) `James Bourbeau`_
Remove skips from doctests (4 of 6) (7865) `Zhengnan Zhao`_
Remove skips from doctests (5 of 6) (7864) `Zhengnan Zhao`_
Adds missing prepend/append functionality to da.diff (7946) `Peter Andreas Entschev`_
Change graphviz font family to sans (7931) `Freyam Mehta`_
Fix read-csv name - when path is different, use different name for task (7942) `Julia Signell`_
Update configuration reference for ucx and rmm changes (7943) `James Bourbeau`_
Add meta support to __setitem__ (7940) `Peter Andreas Entschev`_
NEP-35 support for slice_with_int_dask_array (7927) `Peter Andreas Entschev`_
Unpin fastparquet in CI (7928) `James Bourbeau`_
Remove skips from doctests (3 of 6) (7872) `Zhengnan Zhao`_
- Make array assert_eq check dtype ( dask#7903 ) Julia Signell
Released on July 23, 2021
Make array assert_eq check dtype ( dask#7903 ) Julia Signell
Remove skips from doctests (6 of 6) ( dask#7863 ) Zhengnan Zhao
Remove experimental feature warning from actors docs ( dask#7925 ) Matthew Rocklin
Remove skips from doctests (2 of 6) ( dask#7873 ) Zhengnan Zhao
Separate out Array and Bag API ( dask#7917 ) Julia Signell
Implement lazy Array.iter ( dask#7905 ) Julia Signell
Clean up places where we inadvertently iterate over arrays ( dask#7913 ) Julia Signell
Add numeric_only kwarg to DataFrame reductions ( dask#7831 ) Julia Signell
Add pytest marker for GPU tests ( dask#7876 ) Charles Blackmon-Luca
Add support for histogram2d in dask.array ( dask#7827 ) Doug Davis
Remove skips from doctests (1 of 6) ( dask#7874 ) Zhengnan Zhao
Add node size scaling to the Graphviz output for the high level graphs ( dask#7869 ) Freyam Mehta
Update old Bokeh links ( dask#7915 ) Bryan Van de Ven
Temporarily pin fastparquet in CI ( dask#7907 ) James Bourbeau
Add dask.array import to progress bar docs ( dask#7910 ) Fabian Gebhart
Use separate files for each DataFrame API function and method ( dask#7890 ) Julia Signell
Fix pyarrow-dataset ordering bug ( dask#7902 ) Richard (Rick) Zamora
Generalize unique aggregate ( dask#7892 ) GALI PREM SAGAR
Raise NotImplementedError when using pd.Grouper ( dask#7857 ) Ruben van de Geer
Add aggregate_files argument to enable multi-file partitions in read_parquet ( dask#7557 ) Richard (Rick) Zamora
Un- xfail test_daily_stock ( dask#7895 ) James Bourbeau
Update access configuration docs ( dask#7837 ) Naty Clementi
Use packaging for version comparisons ( dask#7820 ) Elliott Sales de Andrade
Handle infinite loops in merge_asof ( dask#7842 ) gerrymanoim
Released on July 23, 2021
Make array assert_eq check dtype (7903) `Julia Signell`_
Remove skips from doctests (6 of 6) (7863) `Zhengnan Zhao`_
Remove experimental feature warning from actors docs (7925) `Matthew Rocklin`_
Remove skips from doctests (2 of 6) (7873) `Zhengnan Zhao`_
Separate out Array and Bag API (7917) `Julia Signell`_
Implement lazy Array.__iter__ (7905) `Julia Signell`_
Clean up places where we inadvertently iterate over arrays (7913) `Julia Signell`_
Add numeric_only kwarg to DataFrame reductions (7831) `Julia Signell`_
Add pytest marker for GPU tests (7876) `Charles Blackmon-Luca`_
Add support for histogram2d in dask.array (7827) `Doug Davis`_
Remove skips from doctests (1 of 6) (7874) `Zhengnan Zhao`_
Add node size scaling to the Graphviz output for the high level graphs (7869) `Freyam Mehta`_
Update old Bokeh links (7915) `Bryan Van de Ven`_
Temporarily pin fastparquet in CI (7907) `James Bourbeau`_
Add dask.array import to progress bar docs (7910) `Fabian Gebhart`_
Use separate files for each DataFrame API function and method (7890) `Julia Signell`_
Fix pyarrow-dataset ordering bug (7902) `Richard (Rick) Zamora`_
Generalize unique aggregate (7892) `GALI PREM SAGAR`_
Raise NotImplementedError when using pd.Grouper (7857) `Ruben van de Geer`_
Add aggregate_files argument to enable multi-file partitions in read_parquet (7557) `Richard (Rick) Zamora`_
Un-xfail test_daily_stock (7895) `James Bourbeau`_
Update access configuration docs (7837) `Naty Clementi`_
Use packaging for version comparisons (7820) `Elliott Sales de Andrade`_
Handle infinite loops in merge_asof (7842) `gerrymanoim`_
- Include fastparquet in upstream CI build ( dask#7884 ) James Bourbeau
Released on July 9, 2021
Include fastparquet in upstream CI build ( dask#7884 ) James Bourbeau
Blockwise: handle non-string constant dependencies ( dask#7849 ) Mads R. B. Kristensen
fastparquet now supports new time types, including ns precision ( dask#7880 ) Martin Durant
Avoid ParquetDataset API when appending in ArrowDatasetEngine ( dask#7544 ) Richard (Rick) Zamora
Add retry logic to test_shuffle_priority ( dask#7879 ) Richard (Rick) Zamora
Use strict channel priority in CI ( dask#7878 ) James Bourbeau
Support nested dask.distributed imports ( dask#7866 ) Matthew Rocklin
Should check module name only, not the entire directory filepath ( dask#7856 ) Genevieve Buckley
Updates due to dask/fastparquet#623 ( dask#7875 ) Martin Durant
da.eye fix for chunks=-1 ( dask#7854 ) Naty Clementi
Temporarily xfail test_daily_stock ( dask#7858 ) James Bourbeau
Set priority annotations in SimpleShuffleLayer ( dask#7846 ) Richard (Rick) Zamora
Blockwise: stringify constant key inputs ( dask#7838 ) Mads R. B. Kristensen
Allow mixing dask and numpy arrays in @guvectorize ( dask#6863 ) Julia Signell
Don’t sample dict result of a shuffle group when calculating its size ( dask#7834 ) Florian Jetter
Fix scipy tests ( dask#7841 ) Julia Signell
Deterministically tokenize datetime.date ( dask#7836 ) James Bourbeau
Add sample_rows to read_csv -like ( dask#7825 ) Martin Durant
Fix typo in config.deserialize docstring ( dask#7830 ) Geoffrey Lentner
Remove warning filter in test_dataframe_picklable ( dask#7822 ) James Bourbeau
Improvements to histogramdd (for handling inputs that are sequences-of-arrays). ( dask#7634 ) Doug Davis
Make PY_VERSION private ( dask#7824 ) James Bourbeau
Released on July 9, 2021
Include fastparquet in upstream CI build (7884) `James Bourbeau`_
Blockwise: handle non-string constant dependencies (7849) `Mads R. B. Kristensen`_
fastparquet now supports new time types, including ns precision (7880) `Martin Durant`_
Avoid ParquetDataset API when appending in ArrowDatasetEngine (7544) `Richard (Rick) Zamora`_
Add retry logic to test_shuffle_priority (7879) `Richard (Rick) Zamora`_
Use strict channel priority in CI (7878) `James Bourbeau`_
Support nested dask.distributed imports (7866) `Matthew Rocklin`_
Should check module name only, not the entire directory filepath (7856) `Genevieve Buckley`_
Updates due to https://github.com/dask/fastparquet/pull/623 (7875) `Martin Durant`_
da.eye fix for chunks=-1 (7854) `Naty Clementi`_
Temporarily xfail test_daily_stock (7858) `James Bourbeau`_
Set priority annotations in SimpleShuffleLayer (7846) `Richard (Rick) Zamora`_
Blockwise: stringify constant key inputs (7838) `Mads R. B. Kristensen`_
Allow mixing dask and numpy arrays in @guvectorize (6863) `Julia Signell`_
Don't sample dict result of a shuffle group when calculating its size (7834) `Florian Jetter`_
Fix scipy tests (7841) `Julia Signell`_
Deterministically tokenize datetime.date (7836) `James Bourbeau`_
Add sample_rows to read_csv-like (7825) `Martin Durant`_
Fix typo in config.deserialize docstring (7830) `Geoffrey Lentner`_
Remove warning filter in test_dataframe_picklable (7822) `James Bourbeau`_
Improvements to histogramdd (for handling inputs that are sequences-of-arrays). (7634) `Doug Davis`_
Make PY_VERSION private (7824) `James Bourbeau`_
- layers.py compare parts_out with set(self.parts_out) ( dask#7787 ) Genevieve Buckley
Released on June 22, 2021
layers.py compare parts_out with set(self.parts_out) ( dask#7787 ) Genevieve Buckley
Make check_meta understand pandas dtypes better ( dask#7813 ) Julia Signell
Remove “Educational Resources” doc page ( dask#7818 ) James Bourbeau
Released on June 22, 2021
layers.py compare parts_out with set(self.parts_out) (7787) `Genevieve Buckley`_
Make check_meta understand pandas dtypes better (7813) `Julia Signell`_
Remove "Educational Resources" doc page (7818) `James Bourbeau`_
- Add initial deprecation utilities ( dask#7810 ) James Bourbeau
Released on June 18, 2021
Replace funding page with ‘Supported By’ section on dask.org ( dask#7817 ) James Bourbeau
Add initial deprecation utilities ( dask#7810 ) James Bourbeau
Enforce dtype conservation in ufuncs that explicitly use dtype= ( dask#7808 ) Doug Davis
Add Coiled to list of paid support organizations ( dask#7811 ) Kristopher Overholt
Small tweaks to the HTML repr for Layer & HighLevelGraph ( dask#7812 ) Genevieve Buckley
Add dark mode support to HLG HTML repr ( dask#7809 ) Jacob Tomlinson
Remove compatibility entries for old distributed ( dask#7801 ) Elliott Sales de Andrade
Implementation of HTML repr for HighLevelGraph layers ( dask#7763 ) Genevieve Buckley
Update default blockwise token to avoid DataFrame column name clash ( dask#6546 ) James Bourbeau
Use dispatch concat for merge_asof ( dask#7806 ) Julia Signell
Fix upstream freq tests ( dask#7795 ) Julia Signell
Use more context managers from the standard library ( dask#7796 ) James Bourbeau
Simplify skips in parquet tests ( dask#7802 ) Elliott Sales de Andrade
Remove check for outdated bokeh ( dask#7804 ) Elliott Sales de Andrade
More test coverage uploads ( dask#7799 ) James Bourbeau
Remove ImportError catching from dask/init.py ( dask#7797 ) James Bourbeau
Allow DataFrame.join() to take a list of DataFrames to merge with ( dask#7578 ) Krishan Bhasin
Fix maximum recursion depth exception in dask.array.linspace ( dask#7667 ) Daniel Mesejo-León
Fix docs links ( dask#7794 ) Julia Signell
Initial da.select() implementation and test ( dask#7760 ) Gabriel Miretti
Layers must implement get_output_keys method ( dask#7790 ) Genevieve Buckley
Don’t include or expect freq in divisions ( dask#7785 ) Julia Signell
A HighLevelGraph abstract layer for map_overlap ( dask#7595 ) Genevieve Buckley
Always include kwarg name in drop ( dask#7784 ) Julia Signell
Only rechunk for median if needed ( dask#7782 ) Julia Signell
Add add_(prefix|suffix) to DataFrame and Series ( dask#7745 ) tsuga
Move read_hdf to Blockwise ( dask#7625 ) Richard (Rick) Zamora
Make Layer.get_output_keys officially an abstract method ( dask#7775 ) Genevieve Buckley
Non-dask-arrays and broadcasting in ravel_multi_index ( dask#7594 ) Gabe Joseph
Fix for paths ending with “/” in parquet overwrite ( dask#7773 ) Martin Durant
Fixing calling .visualize() with filename=None ( dask#7740 ) Freyam Mehta
Generate unique names for SubgraphCallable ( dask#7637 ) Bruce Merry
Pin fsspec to 2021.5.0 in CI ( dask#7771 ) James Bourbeau
Evaluate graph lazily if meta is provided in from_delayed ( dask#7769 ) Florian Jetter
Add meta support for DatetimeTZDtype ( dask#7627 ) gerrymanoim
Add dispatch label to automatic PR labeler ( dask#7701 ) James Bourbeau
Fix HDFS tests ( dask#7752 ) Julia Signell
Released on June 18, 2021
Replace funding page with 'Supported By' section on dask.org (7817) `James Bourbeau`_
Add initial deprecation utilities (7810) `James Bourbeau`_
Enforce dtype conservation in ufuncs that explicitly use dtype= (7808) `Doug Davis`_
Add Coiled to list of paid support organizations (7811) `Kristopher Overholt`_
Small tweaks to the HTML repr for Layer & HighLevelGraph (7812) `Genevieve Buckley`_
Add dark mode support to HLG HTML repr (7809) `Jacob Tomlinson`_
Remove compatibility entries for old distributed (7801) `Elliott Sales de Andrade`_
Implementation of HTML repr for HighLevelGraph layers (7763) `Genevieve Buckley`_
Update default blockwise token to avoid DataFrame column name clash (6546) `James Bourbeau`_
Use dispatch concat for merge_asof (7806) `Julia Signell`_
Fix upstream freq tests (7795) `Julia Signell`_
Use more context managers from the standard library (7796) `James Bourbeau`_
Simplify skips in parquet tests (7802) `Elliott Sales de Andrade`_
Remove check for outdated bokeh (7804) `Elliott Sales de Andrade`_
More test coverage uploads (7799) `James Bourbeau`_
Remove ImportError catching from dask/__init__.py (7797) `James Bourbeau`_
Allow DataFrame.join() to take a list of DataFrames to merge with (7578) `Krishan Bhasin`_
Fix maximum recursion depth exception in dask.array.linspace (7667) `Daniel Mesejo-León`_
Fix docs links (7794) `Julia Signell`_
Initial da.select() implementation and test (7760) `Gabriel Miretti`_
Layers must implement get_output_keys method (7790) `Genevieve Buckley`_
Don't include or expect freq in divisions (7785) `Julia Signell`_
A HighLevelGraph abstract layer for map_overlap (7595) `Genevieve Buckley`_
Always include kwarg name in drop (7784) `Julia Signell`_
Only rechunk for median if needed (7782) `Julia Signell`_
Add add_(prefix|suffix) to DataFrame and Series (7745) `tsuga`_
Move read_hdf to Blockwise (7625) `Richard (Rick) Zamora`_
Make Layer.get_output_keys officially an abstract method (7775) `Genevieve Buckley`_
Non-dask-arrays and broadcasting in ravel_multi_index (7594) `Gabe Joseph`_
Fix for paths ending with "/" in parquet overwrite (7773) `Martin Durant`_
Fixing calling .visualize() with filename=None (7740) `Freyam Mehta`_
Generate unique names for SubgraphCallable (7637) `Bruce Merry`_
Pin fsspec to 2021.5.0 in CI (7771) `James Bourbeau`_
Evaluate graph lazily if meta is provided in from_delayed (7769) `Florian Jetter`_
Add meta support for DatetimeTZDtype (7627) `gerrymanoim`_
Add dispatch label to automatic PR labeler (7701) `James Bourbeau`_
Fix HDFS tests (7752) `Julia Signell`_
Your coding agent can read these notes before it upgrades. Set up the MCP server →