NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #2015 most downloaded on PyPI
Probabilistic data structures for processing and searching very large datasets
Last release 3 months ago
05 Jul 2026
Release timing varies
gaps range from 2 weeks to 1.4 years
Some releases are documented
notes for 33 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
12 years old
89 releases · first in 2015
Version 2.0.0 changes the default MinHash permutation scheme to "affine32", which fixes a similarity over-estimation bias on large sets ( #212 ), halv
Version 2.0.0 changes the default MinHash permutation scheme to "affine32", which fixes a similarity over-estimation bias on large sets (#212), halves sketch memory, and speeds up updates by roughly 4x. A 64-bit "affine64" scheme is available for billion-scale sets. Hash values differ from earlier versions: rebuild persisted sketches and LSH indexes, or pass MinHash(..., scheme="legacy") to interoperate with existing data. See the MinHash documentation for details.
Full Changelog: v1.10.0...v2.0.0
One column per quarter.
Update hash function to use unsigned 32-bit integer by @davidlowryduda in #304
Full Changelog: v1.9.0...v1.10.0
Bump actions/checkout from 5 to 6 by @dependabot [bot] in #294
Redis interface with sync + introduce aioredis integration tests by @Varun0157 in #293Full Changelog: v1.8.0...v1.9.0
Fix: use of perf_counter() instead of deprecated clock() in time module by @dipeshbabu in #290
MinHashLSHDeletionSession to speed up key deletions by @Varun0157 in #272.python-version by @bhimrazy in #277prepickle to be False for Redis storage by @Varun0157 in #274buffer with memoryview for compatibility by @bhimrazy in #289Redis integration tests + resolve key type disparity across Redis and Cassandra by @Varun0157 in #284Full Changelog: v1.7.0...v1.8.0
Fix bBitMinHash NumPy pickling issue by @123epsilon in #248
uv by @bhimrazy in #257Full Changelog: v1.6.5...v1.7.0
Retrieve MinHash from LSHForest by @123epsilon in #234
Full Changelog: v1.6.4...v1.6.5
HNSW bug fixes by @ekzhu in #230
HNSW remove() point in-place. by @ekzhu in #225
HNSW as MutableMap by @ekzhu in #223
simplify reshapes by @chris-ha458 in #217
Full Changelog: v1.6.0...v1.6.1
Update MinHashLSH.query docstring detailing proximal nature of results by @micimize in https://github.com/ekzhu/datasketch/pull/199
Full Changelog: https://github.com/ekzhu/datasketch/compare/v1.5.9...v1.6.0
Create python-publish.yml by @ekzhu in https://github.com/ekzhu/datasketch/pull/191
Full Changelog: https://github.com/ekzhu/datasketch/compare/v1.5.8...v1.5.9
Add GitHub URL for PyPi by @andriyor in https://github.com/ekzhu/datasketch/pull/179
Full Changelog: https://github.com/ekzhu/datasketch/compare/v1.5.7...v1.5.8
Unable to create multiple lsh indices each one in its own keyspace - issue #171 by @ronassa in https://github.com/ekzhu/datasketch/pull/172
Full Changelog: https://github.com/ekzhu/datasketch/compare/v1.5.6...v1.5.7
Fixed broken packaging setup.py that missed experimental/aio.
Fixed broken packaging setup.py that missed experimental/aio.
minhash: Get rid of deprecation warning by @xkubov in https://github.com/ekzhu/datasketch/pull/156
redis_buffer configuration. by @QthCN in https://github.com/ekzhu/datasketch/pull/152Full Changelog: https://github.com/ekzhu/datasketch/compare/1.5.2...v1.5.4
Nothing published for this version
Performance improvement for MinHash's update method.
update_batch method for bulk update on MinHash. [See API doc].(http://ekzhu.com/datasketch/documentation.html#datasketch.MinHash.update_batch)MinHash.bulk or MinHash.generator. See API doc and pull request.MinHashLSH._H. See pull request. This leads to saving of memory/storage space used by the index.Thank you @Sinusoidal36!
Nothing published for this version
Cassandra storage layer, thank @ostefano! Now you can specify the Cassandra config just like the Redis one.
from datasketch import MinHashLSH
lsh = MinHashLSH(
threashold=0.5, num_perm=128, storage_config={
'type': 'cassandra',
'cassandra': {
'seeds': ['127.0.0.1'],
'keyspace': 'lsh_test',
'replication': {
'class': 'SimpleStrategy',
'replication_factor': '1',
},
'drop_keyspace': False,
'drop_tables': False,
}
}
)
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Now support hashfunc parameter for MinHash and HyperLogLog. The old parameter hashobj is removed.
Now support hashfunc parameter for MinHash and HyperLogLog. The old parameter hashobj is removed.
# Let's use MurmurHash3.
import mmh3
# We need to define a new hash function that outputs an integer that
# can be encoded in 32 bits.
def _hash_func(d):
return mmh3.hash32(d)
# Use this function in MinHash constructor.
m = MinHash(hashfunc=_hash_func)
Use dynamic programming to create optimal partition, allow LSH Ensemble index to adapt to any set size distribution.
Use dynamic programming to create optimal partition, allow LSH Ensemble index to adapt to any set size distribution.
Adding batch removal functionality for Async MinHashLSH
For details see Pull #70 Thanks @aastafiev for the contribution.
Add support for MongoDB replica set
Add support for MongoDB replica set
Nothing published for this version
Added Asynchronous MinHash LSH module. Thanks @aastafiev!
Fix a bug with UnorderedStorage.get_many
Fix a bug with UnorderedStorage.get_many (#56)
Test cases for checking consistency of hash value length in LSH.
Nothing published for this version
Introduced a Redis storage layer for MinHash LSH. Thanks to @ae-foster
__hash__ method for Lean MinHash.Nothing published for this version
Nothing published for this version
Added a slightly simplified version of LSH Ensemble that supports containment search with MinHash data sketches.
MinHash now uses Numpy's random number generator instead of Python's built-in random. This makes MinHash generate consistent hash values across differ
MinHash now uses Numpy's random number generator instead of Python's built-in random. This makes MinHash generate consistent hash values across different Python versions.
The side-effect is that now MinHash created before version 1.1.3 won’t work (i.e., jaccard, merge and union) correctly with those created after.
LeanMinHash is a subclass of MinHash. It uses less memory and allows faster (de)serialization. See documentation for details.
LeanMinHash is a subclass of MinHash. It uses less memory and allows faster (de)serialization. See documentation for details.serialize, deserialize, and bytesize methods from MinHash. These are supported in LeanMinHash instead.MinHash objects before this version will not be deserialized properly. To migrate see here.Nothing published for this version
Nothing published for this version
MinHash LSH Forest implementation and benchmark using synthetic data
Nothing published for this version
Fixed Issue #4 - int overflow error on Windows platform
Add remove method for LSH index - lsh.remove(key)
lsh.remove(key)key in lshAdd Weighted MinHash data sketch
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →