NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2137 most downloaded on PyPI
A fast & compressed ndarray library with a flexible compute engine.
Last release 3 days ago
15 Sep 2026
Release timing varies
gaps range from 9 days to 2 months
Nearly every release is documented
notes for 59 of the last 60 stable releases
2 versions withdrawn
withdrawn after publishing
5 years old
111 releases · first in 2021
One column per quarter.
Many new examples in the documentation. Now, the documentation is more complete and has a better structure. Have a look at our new docs at: https://ww
Many new examples in the documentation. Now, the documentation is more complete and has a better structure. Have a look at our new docs at: https://www.blosc.org/python-blosc2/index.html For a guide on using UDFs, check out: https://www.blosc.org/python-blosc2/reference/autofiles/lazyarray/blosc2.lazyudf.html If interested in asynchronously fetching parts of an array, take a look at: https://www.blosc.org/python-blosc2/reference/autofiles/proxy/blosc2.Proxy.afetch.html Finally, there is a new tutorial on optimizing reductions in large NDArray objects: https://www.blosc.org/python-blosc2/getting_started/tutorials/04.reductions.html Special thanks @omaech and @martaiborrar for the excellent work on the documentation and examples, and to @NumFOCUS for their support in making this possible!
New CParams, DParams and Storage dataclasses for better handling of parameters in the library. Now, you can use these dataclasses to pass parameters to the library, and get a better error handling. See here. Thanks to @martaiborra for the excellent implementation.
Better support for CParams in Proxy and C2Array instances. This allows to better propagate compression parameters from Caterva2 datasets to the Proxy and C2Array instances, improving the perception of codecs and filters used originally in datasets. Thanks to @FrancescAlted for the implementation.
Many improvements in ruff linting and code style. Thanks to @DimitriPapadopoulos for the excellent work in this area.
New ufunc support for NDArray instances. Now, you can use NumPy ufuncs on NDArray instances, and mix them with other NumPy arrays. This is a powerful feature that allows for more interoperability with NumPy.
Enhanced dtype inference, so that it mimics now more NumPy than the numexpr one. Although perfect adherence to NumPy casting conventions is not there yet, it is a big step forward towards better compatibility with NumPy.
Fix dtype for sum and prod reductions. Now, the dtype of the result of a sum or prod reduction is the same as the input array, unless the dtype is not supported by the reduction, in which case the dtype is promoted to a supported one. It is more NumPy-like now.
Many improvements on the computation of UDFs (User Defined Functions). Now, the lazy UDF computation is way more robust and efficient.
Support reductions inside queries in structured NDArrays. For example, given an array sarr with fields 'a', 'b' and 'c', the next farr = sarr["b >= c"].sum("a").compute() puts in farr the sum of the values in field 'a' for the rows that fulfills that values in fields in 'b' are larger than values in 'c' (b >= c above).
Implemented combining data filtering, as well as sorting, in structured NDArrays. For example, given an array sarr with fields 'a', 'b' and 'c', the next farr = sarr["b >= c"].indices(order="c").compute() puts in farr the indices of the rows that fulfills that values in fields in 'b' are larger than values in 'c' (b >= c above), ordered by column 'c'.
Reductions can be stored in persistent lazy expressions. Now, if you have a lazy expression that contains a reduction, the result of the reduction is preserved in the expression, so that you can reuse it later on. See https://www.blosc.org/posts/persistent-reductions/ for more information.
Many improvements in ruff linting and code style. Thanks to @DimitriPapadopoulos for the excellent work in this area.
LazyArray.eval() has been renamed to LazyArray.compute(). This avoids confusion with the eval() function in Python, and it is more in line with the Dask API.This is the main change in the API that is not backward compatible with previous beta. If you have code that still uses LazyArray.eval(), you should change it to LazyArray.compute(). Starting from this release, the API will be stable and backward compatibility will be maintained.
New reshape() function and NDArray.reshape() method allow to do efficient reshaping between NDArrays that follows C order. Only 1-dim -> n-dim is currently supported though.
New NDArray.__iter__() iterator following NumPy conventions.
Now, NDArray.__getitem__() supports (n-dim) bool arrays or sequences of integers as indices (only 1-dim for now). This follows NumPy conventions.
A new NDField.__setitem__() has been added to allow for setting values in a structured NDArray.
struct_ndarr['field'] now works as in NumPy, that is, it returns an array with the values in 'field' in the structured NDArray.
Several new constructors are available for creating NDArray instances, like arange(), linspace() and fromiter(). These constructors leverage the internal lazyudf() function and make it easier to create NDArray instances from scratch. See e.g. https://github.com/Blosc/python-blosc2/blob/main/examples/ndarray/arange-constructor.py for an example.
Structured LazyArrays received a new .indices() method that returns the indices of the elements that fulfill a condition. When combined with the new support of list of indices as key for NDArray.__getitem__(), this is useful for creating indexes for data. See https://github.com/Blosc/python-blosc2/blob/main/examples/ndarray/filter_sort_fields.py for an example.
LazyArrays received a new .sort() method that sorts the elements in the array. For example, given an array sarr with fields 'a', 'b' and 'c', the next farr = sarr["b >= c"].sort("c").compute() puts in farr the rows that fulfills that values in fields in 'b' are larger than values in 'c' (b >= c above), ordered by column 'c'.
New expr_operands() function for extracting operands from a string expression.
New validate_expr() function for validating a string expression.
New CParams, DParams and Storage dataclasses for better handling of parameters in the library. Now, you can use these dataclasses to pass parameters to the library, and get a better error handling. Thanks to @martaiborra for the excellent implementation and @omaech for revamping docs and examples to use them. See e.g. https://www.blosc.org/python-blosc2/getting_started/tutorials/02.lazyarray-expressions.html.
Much improved documentation on how to efficiently compute with compressed NDArray data. Documentation updates highlight these features and improve usability for new users. Thanks to @omaech and @martaiborra for their excellent work on the documentation and examples, and to @NumFOCUS for their support in making this possible! See https://www.blosc.org/python-blosc2/getting_started/tutorials/04.reductions.html for an example.
New remote proxy tutorial. This tutorial shows how to use the Proxy class to access remote arrays, while providing caching. https://www.blosc.org/python-blosc2/getting_started/tutorials/06.remote_proxy.html . Thanks to @omaech for her work on this tutorial.
New tutorial on "Mastering Persistent, Dynamic Reductions and Lazy Expressions". See https://www.blosc.org/posts/persistent-reductions/
Revamped documentation. Now, it is more complete and has a better structure. Thanks to Oumaima Ech Chdig (@omaech), our newcomer to the Blosc team. Al
Revamped documentation. Now, it is more complete and has a better structure. Thanks to Oumaima Ech Chdig (@omaech), our newcomer to the Blosc team. Also, thanks to NumFOCUS for their support in this task.
New Proxy class to access other arrays, while providing caching. This is useful for example when you have a big array, and you want to access a small part of it, but you want to cache the accessed data for later use. See its doc.
Lazy expressions can accept proxies as operands.
Read-ahead support for reading super-chunks from disk. This allows for overlapping reads and computations, which can be a big performance boost for some workloads.
New BLOSC_LOW_MEM envar for keeping memory under a minimum while evaluating expressions. This makes it possible to evaluate expressions on very large arrays, even if the memory is limited (at the expense of performance).
Fine tune block sizes for the internal compute engine.
Better CPU cache size guessing for linux and macOS.
Build tooling has been modernized and now uses pyproject.toml and scikit-build-core for managing dependencies and building the package. Thanks to @LecrisUT for the excellent guidance in this area.
Many code cleanup and syntax improvements in code. Thanks to @DimitriPapadopoulos.
Many new examples in the documentation. Now, the documentation is more complete and has a better structure. Have a look at our new docs at: https://www.blosc.org/python-blosc2/ For a guide on using UDFs, check out: https://www.blosc.org/python-blosc2/reference/autofiles/lazyarray/blosc2.lazyudf.html If interested in asynchronously fetching parts of an array, take a look at: https://www.blosc.org/python-blosc2/reference/autofiles/proxy/blosc2.Proxy.afetch.html Finally, there is a new tutorial on optimizing reductions in large NDArray objects: https://www.blosc.org/python-blosc2/getting_started/tutorials/04.reductions.html Special thanks @omaech and @martaiborrar for the excellent work on the documentation and examples, and to @NumFOCUS for their support in making this possible!
New CParams, DParams and Storage dataclasses for better handling of parameters in the library. Now, you can use these dataclasses to pass parameters to the library, and get a better error handling. See here. Thanks to @martaiborra for the excellent implementation.
Better support for CParams in Proxy and C2Array instances. This allows to better propagate compression parameters from Caterva2 datasets to the Proxy and C2Array instances, improving the perception of codecs and filters used originally in datasets. Thanks to @FrancescAlted for the implementation.
Many improvements in ruff linting and code style. Thanks to @DimitriPapadopoulos for the excellent work in this area.
New evaluation engine (based on numexpr) for NDArray instances. Now, you can evaluate expressions like a + b + 1 where a and b are NDArray instances.
New evaluation engine (based on numexpr) for NDArray instances. Now, you can evaluate expressions like a + b + 1 where a and b are NDArray instances. This is a powerful feature that allows for efficient computations on compressed data, and supports advanced features like reductions, filters, user-defined functions and broadcasting (still in beta). See this example.
As a consequence of the above, there are many new functions to operate with, and evaluate NDArray instances. See the function section docs for more information.
Support for NumPy 2.0.0 is here! Now, the wheels are built with NumPy 2.0.0. If you want to use NumPy 1.x, you can still use it by installing NumPy 1.23 and up.
Support for memory mapping in SChunk and NDArray instances. This allows to map super-chunks stored in disk and access them as if they were in memory. If curious, see some benchmarks here. Thanks to @JanSellner for the excellent implementation, both in the C and the Python libraries.
Internal C-Blosc2 updated to 2.15.0.
32-bit platforms are officially unsupported now. If you need support for 32-bit platforms, please use python-blosc 1.x series.
Revamped documentation. Now, the documentation is more complete and has a better structure. See here. Thanks to Oumaima Ech Chdig (@omaech), our newcomer to the Blosc team. Also, thanks to NumFOCUS for the support in this task.
New Proxy class to access other arrays, while providing caching. This is useful for example when you have a big array, and you want to access a small part of it, but you want to cache the accessed data for later use. See its doc.
Lazy expressions can accept proxies as operands.
Read-ahead support for reading super-chunks from disk. This allows for overlapping reads and computations, which can be a big performance boost for some workloads.
New BLOSC_LOW_MEM envar for keeping memory under a minimum while evaluating expressions. This makes it possible to evaluate expressions on very large arrays, even if the memory is limited (at the expense of performance).
Fine tune block sizes for the internal compute engine.
Better CPU cache size guessing for linux and macOS.
Build tooling has been modernized and now uses pyproject.toml and scikit-build-core for managing dependencies and building the package. Thanks to @LecrisUT for the excellent guidance in this area.
Many code cleanup and syntax improvements in code. Thanks to @DimitriPapadopoulos.
Updated to latest C-Blosc2 2.15.1. Fixes SIGKILL issues when using the blosc2 library in old Intel CPUs.
blosc2 library in old Intel CPUs.Deprecated LazyExpr.evaluate().
Updated to latest C-Blosc2 2.15.0.
Deprecated LazyExpr.evaluate().
Fixed _check_rc function. See https://github.com/Blosc/python-blosc2/issues/187.
Protection when platforms have just one CPU. This caused the internal number of threads to be 0, producing a division by zero.
Protection when platforms have just one CPU. This caused the internal number of threads to be 0, producing a division by zero.
Updated to latest C-Blosc2 2.14.3.
New evaluation engine (based on numexpr) for NDArray instances. Now, you can evaluate expressions like a + b + 1 where a and b are NDArray instances. This is a powerful feature that allows for efficient computations on compressed data, and supports advanced features like reductions, filters, user-defined functions and broadcasting (still in beta). See this example.
As a consequence of the above, there are many new functions to operate with, and evaluate NDArray instances. See the function section docs for more information.
Support for NumPy 2.0.0 is here! Now, the wheels are built with NumPy 2.0.0. If you want to use NumPy 1.x, you can still use it by installing NumPy 1.23 and up.
Support for memory mapping in SChunk and NDArray instances. This allows to map super-chunks stored in disk and access them as if they were in memory. If curious, see some benchmarks here. Thanks to @JanSellner for the excellent implementation, both in the C and the Python libraries.
Internal C-Blosc2 updated to 2.15.0.
32-bit platforms are officially unsupported now. If you need support for 32-bit platforms, please use python-blosc 1.x series.
Updated to latest C-Blosc2 2.14.1. This was necessary to be able to load dynamics plugins on Windows.
[EXP] New evaluation engine (based on numexpr) for NDArray instances. Now, you can evaluate expressions like a + b + 1 where a and b are NDArray insta
[EXP] New evaluation engine (based on numexpr) for NDArray instances.
Now, you can evaluate expressions like a + b + 1 where a and b
are NDArray instances. This is a powerful feature that allows for
efficient computations on compressed data. See this example to see how this works.
Thanks to @omaech for her help in the pow function.
As a consequence of the above, there are many new functions to operate with NDArray instances. See the function section in NDArray API for more information.
Support for NumPy 2.0.0 is here! Now, the wheels are built with NumPy 2.0.0rc1. Please tell us in case you see any issues with this new version.
Add **kwargs to load_tensor() function. This allows to pass additional parameters
to the deserialization function. Thanks to @jasam-sheja.
Fix vlmeta.to_dict() not honoring tuple encoding. Thanks to @ivilata.
Check that chunks/blocks computation does not allow a 0 in blocks. Thanks to @ivilata.
Many improvements in ruff rules and others. Thanks to @DimitriPapadopoulos.
Remove printing large arrays in notebooks (they use too much RAM in recent versions of Jupyter notebook).
Updated to latest C-Blosc2 2.14.0.
Updated to latest C-Blosc2 2.13.1.
Updated to latest C-Blosc2 2.13.1.
Fixed bug in b2nd.h.
Updated to latest C-Blosc2 2.13.0.
Added the filter INT_TRUNC for integer truncation.
Added some optimizations for zstd.
Now the grok library is initialized when loading the plugin from C-Blosc2.
Improved doc.
Support for slices in blosc2.get_slice_nchunks() when using SChunk
objects.
Nothing published for this version
Updated to latest C-Blosc2 2.12.0.
Updated to latest C-Blosc2 2.12.0.
Added blosc2.get_slice_nchunks() to get array of chunk
indexes needed to get a slice of a Blosc2 container.
Added grok codec plugin.
Added imported target with pkg-config to support windows.
Support for pathlib.Path objects in all the places where urlpath is used (e.g. blosc2.open()). Thanks to Marta Iborra.
Support for pathlib.Path objects in all the places where urlpath is
used (e.g. blosc2.open()). Thanks to Marta Iborra.
Included docs for SChunk.fill_special() and NDArray.dtype. Thanks
to Francesc Alted.
Upgrade to latest C-Blosc2 2.11.3. It fixes a bug preventing the use of typesize > 255 in frames. Now you can use a typesize up to 2**31-1.
Temporarily disable AVX512 support in C-Blosc2 for wheels built by CI until run-time detection works properly.
Require at least Cython 3 for building. Using previous versions worked but error handling was not correct (wheels were being built with Cython 3 anyway).
New NDArray.to_cframe() method and blosc2.ndarray_from_cframe() function for serializing and deserializing NDArrays to/from contiguous in-memory frames. Thanks to Francesc Alted.
Add an optional offset argument to blosc2.schunk.open(), to access super-chunks stored in containers like HDF5. Thanks to Ivan Vilata.
Assorted minor fixes to the blocksize/blockshape computation algorithm, avoiding some cases where it resulted in values exceeding maximum limits. Thanks to Ivan Vilata.
Updated to latest C-Blosc2 2.11.2. It adds AVX512 support for the bitshuffle filter, fixes ARM and Raspberry Pi compatibility and assorted issues.
Add python-blosc2 package definition for Guix. Thanks to Ivan Vilata.
Nothing published for this version
Support for specifying (plugable) tuner parameters in cparams. Thanks to Marta Iborra.
Support for specifying (plugable) tuner parameters in cparams. Thanks to Marta Iborra.
Re-add support for Python 3.8. Although we don't provide wheels for it, support is there (although it requires compilation time).
Avoid duplicate iteration over the same dict. Thanks to Dimitri Papadopoulos.
Fix different issues with f-strings. Thanks to Dimitri Papadopoulos.
Binary wheels for forthcoming Python 3.12 are available!
Binary wheels for forthcoming Python 3.12 are available!
Different improvements suggested by refurb and pyupgrade. Thanks to Dimitri Papadopoulos.
Updated to latest C-Blosc2 2.10.4.
Updated to latest C-Blosc2 2.10.3.
Updated to latest C-Blosc2 2.10.3.
Added openhtj2k codec plugin.
Some small fixes regarding typos.
Multithreading checks only apply to Python defined codecs and filters. Now it is possible to use multithreading with C codecs and filters plugins. See
Multithreading checks only apply to Python defined codecs and filters. Now it is possible to use multithreading with C codecs and filters plugins. See PR #127.
New support for dynamic filters registry for Python.
Now params for codec and filter plugins are correctly initialized
when using register_codec and register_filter functions.
Some fixes for Cython 3.0.0. However,compatibility with Cython 3.0.0 is not here yet, so build and install scripts are still requiring Cython<3.
Updated to latest C-Blosc2 2.10.1.
Updated to latest C-Blosc2 2.10.0.
Updated to latest C-Blosc2 2.10.0.
Use the new, fixed bytedelta filter introduced in C-Blosc2 2.10.0.
Some small fixes in tutorials.
Added a new section of tutorials for a quick get start.
Added a new section of tutorials for a quick get start.
Added a new section on how to cite Blosc.
New method interchunks_info for SChunk and NDArray classes.
This iterates through chunks for getting meta info, like decompression ratio, whether the chunk is special or not, among others. For more information on how this works see this example.
Now it is possible to register a dynamic plugin by passing None as the encoder and decoder arguments in the register_codec function.
Make shape of scalar slices NDArray objects to follow NumPy conventions. See #117.
Updated to latest C-Blosc2 2.9.3.
Nothing published for this version
Wheels are not including blosc2.pc (pkgconfig) anymore. For details see: https://github.com/Blosc/python-blosc2/pull/111 . Thanks to @bnavigator for t
Updated to latest C-Blosc2 2.9.1.
New bytedelta filter. We have blogged about this: https://www.blosc.org/posts/bytedelta-enhance-compression-toolset/. See the examples/ndarray/bytedel
New bytedelta filter. We have blogged about this: https://www.blosc.org/posts/bytedelta-enhance-compression-toolset/. See the examples/ndarray/bytedelta_filter.py for a sample script. We also have a short video on how bytedelta works: https://www.youtube.com/watch?v=5OXs7w2x6nw
The compression defaults are changed to get a better balance between compression ratio, compression speed and decompression speed. The new defaults are:
cparams.typesize = 8cparams.clevel = 1cparams.compcode = Codec.ZSTDfilters = [Filter.SHUFFLE]splitmode = SplitMode.ALWAYS_SPLITThese changes are based on the experiments performed in the blog post above.
dtype.itemsize will have preference over typesize in cparams (as it was documented).
blosc2.compressor_list(plugins=False) do not list codec plugins by default now. If you want to list plugins too, you need to pass plugins=True.
Internal C-Blosc2 updated to latest version (2.8.0).
New NDArray class for handling multidimensional arrays using compression. It includes:
New NDArray class for handling multidimensional arrays using compression. It includes:
See examples at: https://github.com/Blosc/python-blosc2/tree/main/examples/ndarray NDarray docs at: https://www.blosc.org/python-blosc2/reference/ndarray_api.html Explanatory video on why double partitioning: https://youtu.be/LvP9zxMGBng Also, see our blog on C-Blosc2 NDim counterpart: https://www.blosc.org/posts/blosc2-ndim-intro/
Internal C-Blosc2 bumped to latest 2.7.1 version.
Nothing published for this version
Add support for user-defined filters and codecs. See our blog at: https://www.blosc.org/posts/python-blosc2-pipeline/
Add support for user-defined filters and codecs. See our blog at: https://www.blosc.org/posts/python-blosc2-pipeline/
API has been frozen.
Nothing published for this version
## Changes from 0.6.4 to 0.6.5 * Add arm64 wheels for macosx.
Add arm64 wheels and remove musl builds (NumPy not having them makes the build process too long).
Use oldest-supported-numpy for maximum compatibility.
## Changes from 0.6.1 to 0.6.2 * Updated C-Blosc2 to 2.6.0.
Support for Python prefilters and postfilters. With this, you can pre-process or post-process data in super-chunks automatically. This machinery is ha
Support for Python prefilters and postfilters. With this, you can pre-process or post-process data in super-chunks automatically. This machinery is handled internally by C-Blosc2, so it is very efficient (although it cannot work in multi-thread mode due to the GIL). See the examples/ directory for different ways of using this.
Support for fillers. This is a specialization of a prefilter, and it allows to use Python functions to create new super-chunks from different kind of inputs (NumPy, SChunk instances, scalars), allowing computations among them and getting the result automatically compressed. See a sample script in the examples/ directory.
Lots of small improvements in the style, consistency and other glitches in the code. Thanks to Dimitri Papadopoulos for hist attention to detail.
No need to compile C-Blosc2 tests, benchs or fuzzers. Compilation time is much shorter now.
Added cratio, nbytes and cbytes properties to SChunk instances.
Added setters for dparams and cparams attributes in SChunk.
Nothing published for this version
Remove the testing of packing PyTorch or TensorFlow objects during wheels build.
Add msgpack as a runtime requirement
msgpack as a runtime requirementNothing published for this version
Several leaks fixed. Thanks to Christoph Gohlke.
Several leaks fixed. Thanks to Christoph Gohlke.
Internal C-Blosc2 updated to 2.3.1
Nothing published for this version
Added a new blosc2.open(urlpath, mode) function to be able to open persisted super-chunks.
Added a new blosc2.open(urlpath, mode) function to be able to open persisted super-chunks.
Added a new tutorial in notebook format (examples/tutorial-basics.ipynb) about the basics of python-blosc2.
Internal C-Blosc2 updated to 2.2.0
Internal C-Blosc updated to 2.0.4.
New SChunk class that allows to create super-chunks.
This includes the capability of storing data in 4
different ways (sparse/contiguous and in memory/on-disk),
as well as storing variable length metalayers.
Also, during the contruction of a SChunk instance,
an arbitrarily large data buffer can be given so that it is
automatically split in chunks and those are appended to the
SChunk.
See examples/schunk.py and examples/vlmeta.py for some examples.
Documentation of the new API is here: https://python-blosc2.readthedocs.io
This release is the result of a grant offered by the Python Software Foundation to Marta Iborra. A blog entry was written describing the difficulties and relevant aspects learned during the work: https://www.blosc.org/posts/python-blosc2-initial-release/
Release with C-Blosc 2.0.2 sources and binaries.
Release with C-Blosc 2.0.1 sources and binaries.
Nothing published for this version
Headers and binaries for the C-Blosc2 library are starting to being distributed inside wheels.
Headers and binaries for the C-Blosc2 library are starting to being distributed inside wheels.
Internal C-Blosc2 submodule updated to 2.0.0-rc2.
Repeating measurements 4 times in benchs so as to get more consistent figures.
Nothing published for this version
Fix some issues with PyPI packaging. Wheels are here! See: https://github.com/Blosc/python-blosc2/issues/9
Fix some issues with PyPI packaging. Wheels are here! See: https://github.com/Blosc/python-blosc2/issues/9
Nothing published for this version
Nothing published for this version
Nothing published for this version
python-blosc2 aims to leverage the new C-Blosc2 API so as to support super-chunks, serialization and all the features introduced in C-Blosc2. This is
python-blosc2 aims to leverage the new C-Blosc2 API so as to support super-chunks, serialization and all the features introduced in C-Blosc2. This is work in process and will be done incrementally in future releases.
Note: python-blosc2 is meant to be backward compatible with python-blosc data. That means that it can read data generated with python-blosc, but the opposite is not true (i.e. there is no forward compatibility).
The functions compress_ptr and decompress_ptr
are replaced by pack and unpack since
Pickle protocol 5 comes with out-of-band data.
The function pack_array is equivalent to pack,
which accepts any object with attributes itemsize
and size.
On the other hand, the function unpack doesn't
return a numpy array whereas the unpack_array
builds that array.
The blosc.NOSHUFFLE is replaced
by the blosc2.NOFILTER, but for backward
compatibility blosc2.NOSHUFFLE still exists.
A bytearray or NumPy object can be passed to
the blosc2.decompress function to store the
decompressed data.
Your coding agent can read these notes before it upgrades. Set up the MCP server →