NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #3484 most downloaded on PyPI
An ultra-fast implementation of BM25 based on sparse matrices.
Last release 3 days ago
02 Oct 2026
Release timing varies
gaps range from 9 days to 3 months
Nearly every release is documented
notes for 37 of 38 stable releases
6 versions withdrawn
withdrawn after publishing
2 years old
61 releases · first in 2024
Add support for GIL-free concurrency by @AndreasMadsen in #198
Full Changelog: 0.3.11...0.3.12
One column per month.
Add support for Danish stopwords by @yatzima in #193
Add leave_progress and show_progress by @patryksew in #191
Full Changelog: 0.3.9...0.3.10
fix: downgrade resource module ImportError log from warning to debug on Windows by @Copilot in #187
Full Changelog: 0.3.8...0.3.9
Resolve version metadata mismatch by @xhluca in #185
Resolve version metadata mismatch by @xhluca in #185
Full Changelog: 0.3.7...0.3.8.pre1
Clarify supported corpus formats in docs by @xhluca in #183
Full Changelog: 0.3.6...0.3.7
feat: allow disabling progress when saving index by @aidanstewartthomson in #181
Full Changelog: 0.3.5...0.3.6
Use GitHub release for BEIR dataset downloads by @xhluca in #180
bm25s import from writing Windows resource fallback message to stdout by @Copilot in #179Full Changelog: 0.3.4...0.3.5
add korean stopwords supports by @emmanuel-stone in #177
Full Changelog: 0.3.3...0.3.4
Catch RuntimeError in jax import guard in selection.py by @Copilot in https://github.com/xhluca/bm25s/pull/174
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.2...0.3.3
Update BM25 homepage link to bm25-python.github.io
Update BM25 homepage link to bm25-python.github.io
This is the first release for the BM25 library, which is a high-level wrapper around bm25s aimed at quickly getting started!
This is the first release for the BM25 library, which is a high-level wrapper around bm25s aimed at quickly getting started!
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0...0.3.2
Nothing published for this version
Scipy is no longer a required dependency. scipy has been removed from install_requires in setup.py. The library now uses a pure NumPy-based CSC matrix
scipy has been removed from install_requires in setup.py. The library now uses a pure NumPy-based CSC matrix builder by default. If you need scipy's CSC builder, install it separately and pass csc_backend="scipy" to BM25(), or install via pip install bm25s[indexing].from bm25s import selection is now from bm25s import selection as selection_np internally. If you were importing selection directly from bm25s, update your imports.bm25s.high_level)A new simplified 1-line indexing and 1-line search API:
import bm25s.high_level as bm25
corpus = bm25.load("documents.csv", document_column="text")
retriever = bm25.index(corpus)
results = retriever.search(["your query"], k=5)
bm25.load() supports CSV, JSON, JSONL, and TXT files with automatic format detection.bm25.index() handles tokenization (with stemming + stopword removal) and indexing in one call.BM25Search.search() returns ranked results with document text, scores, and IDs.bm25 CLI)A new terminal CLI via the bm25 console script entry point:
bm25 index <file> — Index documents from CSV, TXT, JSON, or JSONL files.
-o to specify output directory, -c to specify text column, -u to save to user directory (~/.bm25s/indices/).bm25 search -i <index> "query" — Search an existing index.
-k for top-k, -s to save results as JSON, -u for user directory with interactive index picker.pip install bm25s[cli] for Rich-based UI, falls back to plain text).bm25s.mcp)A built-in Model Context Protocol server to expose BM25 indices as tools for LLMs:
bm25 mcp launch --index-dir <path> — Launch an MCP server with retrieve and get_info tools.pip install bm25s[mcp].compile() method on BM25 for explicit JIT compilation of both the scorer and CSC builder.auto_compile=True parameter on BM25.__init__() — automatically compiles Numba JIT functions on initialization.warmup_numba_scorer() and warmup_numba_csc() methods to pre-trigger JIT compilation with dummy data.activate_numba_csc() method — applies Numba JIT to the CSC matrix builder for faster indexing.csc_backend parameter on BM25(): choose "numpy" (default), "scipy", or "auto"._np_csc_python() — Pure NumPy implementation using packed-index argsort._np_csc_jit_ready() — Numba-compilable implementation using counting sort (linear time).BM25.load() now accepts override_params={} and **kwargs to override saved parameters at load time (e.g., change auto_compile, backend, etc.)._faketqdm fix: The fallback tqdm replacement now properly handles being called with no positional arguments (returns None instead of raising).__init__, scoring, tokenization, hf, beir, corpus) now respect the DISABLE_TQDM environment variable uniformly._compute_relevance_from_scores now wraps dtype with np.dtype() for compatibility with Numba JIT.selection_jit is None checks with a proper NUMBA_AVAILABLE boolean flag.activate_numba_scorer() now respects the NUMBA_DISABLE_JIT environment variable.| Extra | Packages | Purpose |
|---|---|---|
mcp |
mcp |
MCP server support |
cli |
rich |
Rich terminal UI for interactive index picker |
indexing |
scipy |
scipy-based CSC matrix construction |
test-numba and test-high-level CI jobs with proper thread-safety env vars (OMP_NUM_THREADS=1, etc.).coverage and report percentage.dev* branches in addition to main.claude.yml (Claude Code GitHub Action for issue/PR interaction) and claude-code-review.yml (automated PR code review).tests/core/test_core_coverage.py — 447 lines of comprehensive core module tests.tests/core/test_corpus.py, test_hf_utils.py, test_init_utils.py, test_json_functions.py, test_scoring.py, test_selection.py, test_tokenization_extended.py — Extended unit tests for core modules.tests/high_level/test_high_level.py — 121 lines testing the high-level API.tests/high_level/test_terminal.py — 647 lines testing the CLI terminal commands.tests/data/dummy.csv, dummy.jsonl, dummy.txt.examples/mcp/create_index.py — Create a test index for the MCP server.examples/mcp/verify_server.py — Verify MCP server functionality.examples/simple_load.py — Demonstrate the high-level load/index/search workflow.Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0...0.3.0.rc5
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0...0.3.0.rc5
[WIP] Add terminal cli by @xhluca in https://github.com/xhluca/bm25s/pull/157
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0.alpha3...0.3.0.alpha4
Add high level api by @xhluca in https://github.com/xhluca/bm25s/pull/154
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0.alpha2...0.3.0.alpha3
fix index_nq example by @xhluca in https://github.com/xhluca/bm25s/pull/146
Full Changelog: https://github.com/xhluca/bm25s/compare/0.3.0.alpha1...0.3.0.alpha2
Add comprehensive test coverage (93% core modules) and CI coverage reporting by @Copilot in https://github.com/xhluca/bm25s/pull/143
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.14...0.3.0.alpha1
Update README.md by @Tickloop in https://github.com/xhluca/bm25s/pull/135
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.13...0.2.14
fix merge_cqa_dupstack by @seanmacavaney in https://github.com/xhluca/bm25s/pull/132
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.12...0.2.13
Fix shifted index error when allow_empty=True and return_as="ids" for Tokenizer.tokenize by @xhluca in https://github.com/xhluca/bm25s/pull/128
allow_empty=True and return_as="ids" for Tokenizer.tokenize by @xhluca in https://github.com/xhluca/bm25s/pull/128Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.11...0.2.12
fix: raise ValueError in retrieve directly, rather than in topk by @emmanuel-stone in https://github.com/xhluca/bm25s/pull/122
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.10...0.2.11
fix: update tokenize docstring to avoid SyntaxWarning - invalid escape sequence \w by @yaminivibha in https://github.com/xhluca/bm25s/pull/124
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.9...0.2.10
fix: raise ValueError when the corpus size is less than k by @emmanuel-stone in https://github.com/xhluca/bm25s/pull/117
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.8...0.2.9
Fix error when load_vocab=False by @xhluca in https://github.com/xhluca/bm25s/pull/115
load_vocab=False by @xhluca in https://github.com/xhluca/bm25s/pull/115Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.7...0.2.8
Fix query filtering and vocabulary dict by @mossbee in https://github.com/xhluca/bm25s/pull/92 (1/2)
The behavior of tokenizers have changed wrt null token. Now, the null token will be added first to the vocab rather than at the end, as the previous approach is inconsistent with the general standard (the "" string should map to 0 in general). However, it is a backward compatible change because the tokenizers should work the same way as before, but expect the tokenizers before 0.2.7 to differ from the tokenizers in 0.2.7 and beyond in the behavior, even though both will work with the retriever object.
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.6...0.2.7
Nothing published for this version
Update corpus.py by @Restodecoca in https://github.com/xhluca/bm25s/pull/102
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.7pre2...0.2.7pre3
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.7pre1...0.2.7pre2
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.7pre1...0.2.7pre2
Fix query filtering and vocabulary dict by @xhluca in https://github.com/xhluca/bm25s/pull/96 and @mossbee in https://github.com/xhluca/bm25s/pull/92
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.6...0.2.7
Extending to Non-ASCII characters with corpora loading and saving by @IssacXid in https://github.com/xhluca/bm25s/pull/93
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.5...0.2.6
Update README.md by @xhluca in https://github.com/xhluca/bm25s/pull/83
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.4...0.2.5
Fix crash tokenizing with empty word_to_id by @mgraczyk in https://github.com/xhluca/bm25s/pull/72
Fix crash tokenizing with empty word_to_id by @mgraczyk in https://github.com/xhluca/bm25s/pull/72
Create nltk_stemmer.py by @aflip in https://github.com/xhluca/bm25s/pull/77
https://github.com/xhluca/bm25s/commit/aa31a2321250180feb8b155fec1daafd40f56182: The commit primarily focused on improving the handling of unknown tokens during the tokenization and retrieval processes, enhancing error handling, and improving the logging mechanism for better debugging.
bm25s/init.py: Added checks in the get_scores_from_ids method to raise a ValueError if max_token_id exceeds the number of tokens in the index. Enhanced handling of empty queries in _get_top_k_results method by returning zero scores for all documents.bm25s/tokenization.py: Fixed the behavior of streaming_tokenize to correctly handle the addition of new tokens and updating word_to_id, word_to_stem, and stem_to_sid.Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.3...0.2.4
PR https://github.com/xhluca/bm25s/pull/67 fixes issue #60
Tokenizer class, such as when update_vocab=True in return_as="ids" mode, which leads to unseen new token IDs being passed to retriever.retrieveFull Changelog: https://github.com/xhluca/bm25s/compare/0.2.2...0.2.3
Improve README with example of memory usage optimization
Results.merge method allowing merging list of resultsget_max_memory_usage compatible with mac osBM25.load_scores that allows loading only the scores of the objectload_vocab parameter set to True by default in BM25.load, allowing the vocabulary to not be always loaded.PR: https://github.com/xhluca/bm25s/pull/63
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.1...0.2.2
Add Tokenizer.save_vocab and Tokenizer.load_vocab methods to save/load vocabulary to a json file called vocab.tokenizer.json by default
Tokenizer.save_vocab and Tokenizer.load_vocab methods to save/load vocabulary to a json file called vocab.tokenizer.json by defaultTokenizer.save_stopwords and Tokenizer.load_stopwords methods to save/load stopwords to a json file called stopwords.tokenizer.json by defaultTokenizerHF class to allow saving/loading from huggingface hub
load_vocab_from_hub, save_vocab_to_hub, load_stopwords_from_hub, save_stopwords_to_hubNew tests and examples were added (see
examples/index_to_hf.pyandexamples/tokenizer_class.py)
*Version 0.2.0 is an exciting release! This brings a lot of new features, including numba support (over 2x faster in many cases), stopwords for 10 new
Version 0.2.0 is an exciting release! This brings a lot of new features, including numba support (over 2x faster in many cases), stopwords for 10 new languages (thank you @bm777), a new Tokenizer class (faster and more flexible), document weighting at retrieval time, a new JSON backend (orjson), improvements to utils for using BEIR, and many new examples! Hope you enjoy this new release!
See discussion here: https://github.com/xhluca/bm25s/discussions/46
The most important new feature of v0.2.0 is the addition of numba support, which only require you to install the core requirements (with pip install "bm25s[core]") or with pip install numba.
Using numba will result in a substantial speedup, so it is highly recommended if you have access to numba on your system (which should be in most cases). You can find a benchmark here.
Notably, by combining numba JIT-based scoring, numba-based top-k selection (no longer relies on jax, see discussion thread) and the new and faster bm25s.tokenization.Tokenizer (see below), we observe the following speedup on a few benchmarks, in a single-threaded setting with Kaggle CPUs:
To enable it, simply do:
import bm25s
# load corpus
# ...
retriever = bm25s.BM25(backend="numba")
# index and run retrieval
This is all you need to use numba JIT when calling the retriever.retrieve method. Note, however, that the first run might be slower, so you can warmup by passing a small query. Here are more examples:
bm25s.tokenization.Tokenizer classWith v0.2.0, we are adding the Tokenizer class, which enhances the existing features of bm25s.tokenize and makes it more flexible. Notably, it enables generator mode (stream with yield), and is much faster when tokenizing queries, if you have an existing vocabulary. Also, you can specify your own splitter function, which is no longer locked to a regex pattern.
You can find more information here:
examples/tokenizer_class.pyhelp(bm25s.tokenization.Tokenizer)Stopwords for 10 languages (from NLTK) were added by @bm777 in https://github.com/xhluca/bm25s/pull/33
orjson is now supported as a JSON backend, as it is faster than ujson and is currently supported.
BM25.retrieve now supports a weight_mask array, which applies a weight (binary or float) on each of the document retrieved. This is useful, for example, if you want to use a binary mask to hide certain documents deemed irrelevant.
orjson replaces ujson as a core dependencyjax[cpu] is no longer a core dependency, but a selection dependency now. Be careful to not use backend_selection='jax' if you don't have it installed!numba is a new core dependency, allowing you to directly use the backend='numba' when initializing a retriever.pytrec_eval is a new evaluation dependency, which is useful if you want to use the evaluation function in bm25s.utils.beir which is copied from the BEIR dataset.Here's an example of how to leverage numba speedups using the alternative method of activing numba scorer and choosing the backend_selection manually. It is not recommended to use this method unless you speicfically want to have more control over how the backend is activated.
import os
import Stemmer
import bm25s.hf
def main(repo_name="xhluca/bm25s-fiqa-index"):
queries = [
"Is chemotherapy effective for treating cancer?",
"Is Cardiac injury is common in critical cases of COVID-19?",
]
retriever = bm25s.hf.BM25HF.load_from_hub(
repo_name, load_corpus=False, mmap=False
)
# Tokenize the queries
stemmer = Stemmer.Stemmer("english")
queries_tokenized = bm25s.tokenize(queries, stemmer=stemmer)
# Retrieve the top-k results
retriever.activate_numba_scorer()
results = retriever.retrieve(queries_tokenized, k=3, backend_selection="numba")
# show first results
result = results.documents[0]
print(f"First score (# 1 result):{results.scores[0, 0]}")
print(f"First result (# 1 result):\n{result[0]}")
if __name__ == "__main__":
main()
Again, this method is only recommended if you want to have more control.
WARNING: it will not do well with multithreading. For the full example, see retrieve_with_numba_advanced.py
In this release, we add the Tokenizer class. Please see readme section on tokenization and examples/tokenizer_class.py for more details.
In this release, we add the Tokenizer class. Please see readme section on tokenization and examples/tokenizer_class.py for more details.
This is the final version of the numba improvements:
This is the final version of the numba improvements:
Full Changelog: https://github.com/xhluca/bm25s/compare/data...0.2.0rc7
This is a pretty exciting pre-release! It is a major new feature for the v0.2.0 that will come out. I hope you get to try this and share your thoughts
This is a pretty exciting pre-release! It is a major new feature for the v0.2.0 that will come out. I hope you get to try this and share your thoughts in the discussions!
texts argument in tokenize function and replace time.time() with time.monotonic()` by @dantetemplar in https://github.com/xhluca/bm25s/pull/44Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.10...0.2.0rc6
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.0rc3...0.2.0rc5
Full Changelog: https://github.com/xhluca/bm25s/compare/0.2.0rc3...0.2.0rc5
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Add bibtex to the auto-generated readme for huggingface
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.9...0.1.10
Allow retrieve() to take a tuple as input
Fix bug when tqdm is not available in bm25s.utils.corpus
Fix bug when tqdm is not available in bm25s.utils.corpus
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.6...0.1.7
Improve readme generated when saving to huggingface by @xhluca in https://github.com/xhluca/bm25s/pull/6
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.5...0.1.6
Fix readme template example when saving to hub.
Fix readme template example when saving to hub.
[hf] Add "library_name" metadata to avoid confusion about what the primary library is by @tomaarsen in https://github.com/xhluca/bm25s/pull/2
hf] Add "library_name" metadata to avoid confusion about what the primary library is by @tomaarsen in https://github.com/xhluca/bm25s/pull/2compat] Allow for local install on Windows by @tomaarsen in https://github.com/xhluca/bm25s/pull/3Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.3...0.1.4
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.2...0.1.3
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.2...0.1.3
Fix speed issue with in-memory corpus
Fix speed issue with in-memory corpus
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.0...0.1.1
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.0...0.1.1
Nothing published for this version
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.0rc1...0.1.0rc3
Full Changelog: https://github.com/xhluca/bm25s/compare/0.1.0rc1...0.1.0rc3
Full Changelog: https://github.com/xhluca/bm25s/commits/0.1.0rc2
Full Changelog: https://github.com/xhluca/bm25s/commits/0.1.0rc2
Full Changelog: https://github.com/xhluca/bm25s/compare/0.0.1rc1...0.1.0rc1
Full Changelog: https://github.com/xhluca/bm25s/compare/0.0.1rc1...0.1.0rc1
Your coding agent can read these notes before it upgrades. Set up the MCP server →