NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #745 most downloaded on PyPI
Industrial-strength Natural Language Processing (NLP) in Python
Last release 1 months ago
24 Aug 2026
Release timing varies
gaps range from 1 weeks to 6 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
9 versions withdrawn
withdrawn after publishing
12 years old
218 releases · first in 2015
Nothing published for this version
Nothing published for this version
Nothing published for this version
One column per quarter.
Its trusted-publisher configuration no longer exists on PyPI, so it fails on every release tag. Publishing is handled by a separate release process.
Remove publish_pypi workflow
Its trusted-publisher configuration no longer exists on PyPI, so it fails
on every release tag. Publishing is handled by a separate release process.
Typer removed click as a dependency in favour of vendoring it, but spaCy imports from Click, so the requirement needs to be added.
Typer removed click as a dependency in favour of vendoring it, but spaCy imports from Click, so the requirement needs to be added.
Fix spacy download failing in environments where pip is not on PATH but is available as a Python module (e.g., some virtual environments and container
spacy download failing in environments where pip is not on PATH but is available as a Python module (e.g., some virtual environments and containers)v3.8.14: Bug fix for model downloading in environments without pip on PATH
Compare
The v3.8.12 release didn't update the confection pin, which meant that if you did an upgrade-install models wouldn't load.
The v3.8.12 release didn't update the confection pin, which meant that if you did an upgrade-install models wouldn't load.
v3.8.13: Pin confection to new version
Compare
I've taken some lengths to explain this because migrating off a dependency after breaking changes can be a sensitive topic. I want to stress that the…
Use confection v1.3 and Thinc v8.3.13, which implement custom validation logic in place of Pydantic, allowing us to properly adopt Pydantic v2 and provide full Python 3.14 support.
Our dependency tree used Pydantic v1 in unusual ways, and relied on behaviours that Pydantic v2 reformed. In the time since Pydantic v2 was released there were a few attempts to migrate over to it, but the task has been complicated by the fact that the confection library has a fairly tangled implementation and I had reduced availability for open-source work in 2024 and 2025.
Specifically, our library confection provides the extensible configuration system we use in spaCy and Thinc. The config system allows you to refer to values that will be supplied by arbitrary functions, that e.g. define some neural network model or its sublayers. The functionality in confection is complicated because we aggressively prioritised user experience in the specification, even if it required increased implementation complexity.
Confection's original implementation built a dynamic Pydantic v1 schema for function-supplied values ("promises"). We validate the schema before calling any promises, and then validate the schema again after calling all the promises and substituting in their values. The variable-interpolation system adds further difficulties to the implementation, and we have to do it all subclassing the Python built-in configparser, which ties us to implementation choices I'd do differently if I had a clean slate.
Here's one summary of Pydantic v1-specific behaviours that the migration to v2 particularly difficult for us. This particular summary was produced during a session with Claude Code Opus 4.6, so nuances of it might be wrong. The full history of attempts at doing this spans over different refactors separated by a few months at a time, so I don't have a full record of all the things that I struggled with. It's possible some details of this summary are incorrect though.
The core problem we kept hitting: Pydantic v2 compiles validation schemas upfront and has much stricter immutability. The whole session has been a series of workarounds for this:
1. Schema mutation — v1 let you mutate __fields__ in place; v2 needs model_rebuild() which loses forward ref namespaces, or create_model subclasses which don't propagate to parent schemas.
2. model_dump vs dict — v2 converts dataclasses to dicts, breaking resolved objects. Needed a custom _model_to_dict helper.
3. model_construct drops extras — v2 silently drops fields with extra="forbid", needed manual workarounds.
4. Strict coercion — v2 coerces ndarray to List[Floats1d] via iteration, needed strict=True.
5. Forward refs — Every schema with TYPE_CHECKING imports needs model_rebuild() with the right namespace, and that breaks when confection re-rebuilds later.
In order to adjust for behavioural differences like this, I'd refactored confection to build the different versions of the schema in multiple passes, instead of building all the representations together as we'd been doing. However this refactor itself had problems, further complicating the migration.
I've now bitten the bullet and rolled back the refactor I'd been attempting of confection, and instead replaced the Pydantic validation with custom logic. This allows Confection to remove Pydantic as a dependency entirely.~ Update: Actually I went back and got the refactor working. All much nicer now.
I've taken some lengths to explain this because migrating off a dependency after breaking changes can be a sensitive topic. I want to stress that the changes Pydantic made from v1 to v2 are very good, and I greatly appreciate them as a user of FastAPI in our services. It would be very bad for the ecosystem if Pydantic pinned themselves to exactly matching the behaviours they had in v1 just to avoid breaking support for the sort of thing we'd been doing. Instead users who were relying on those behaviours like us should just find some way to adapt --- either vendor the v1 version we need, or change our behaviours, or implement an alternative. I would have liked to do this sooner but we've ultimately gone with the third option.
Add wheels for Python 3.11, 3.12, 3.13 and 3.14 for Windows ARM. Windows ARM wheels for Python 3.10 and earlier are not available in numpy, so aren't
Add wheels for Python 3.11, 3.12, 3.13 and 3.14 for Windows ARM. Windows ARM wheels for Python 3.10 and earlier are not available in numpy, so aren't provided.
Windows arm needs to be disabled at the ci level, so remove this skip…
Windows arm needs to be disabled at the ci level, so remove this skip…
… selector
v3.8.10: Fix missing Python 3.14 wheels
Compare
Add wheels for Python 3.14
Add wheels for Python 3.14
Fix deprecation warnings from click imports
typer-slim to reduce dependency footprintOther dependencies in spaCy's tree have also been updated to widen the numpy compatibility pin, which should reduce installation problems for some users.
v3.8.8: Fix deprecation warnings, update requirements, drop 3.9
Compare
__getattr__ import shims have been added to the previous locations of these functions to prevent backwards incompatibilities.
In order to support Python 3.13, spaCy is now compiled with Cython 3. This brings a change to the way types are handled at runtime (Cython 3 uses the from __future__ import annotations semantics, which stores types as strings at runtime. This difference caused problems for components registered within Cython files, as we rely on building Pydantic models from factory function signatures to do validation.
To support Python 3.13 we therefore create a new module, spacy.pipeline.factories, which contains the factory function implementations. __getattr__ import shims have been added to the previous locations of these functions to prevent backwards incompatibilities.
As well as moving the factories, the new implementation avoids import-time side-effects, by moving the actual calls to the decorator inside a function, which is executed once when the Language class is initialised.
A matching change has been made to the catalogue registry decorators. A new module spacy.registrations has been created that performs all the catalogue registrations. Moving these registrations away from the functions prevents these decorators from running at import time. This change was not necessary for the Python 3.13 support, but it means we no longer rely on any import-time side-effects, which will allow us to improve spaCy's import time and therefore CLI execution time. The change also makes maintenance easier as it's easier to find the implementations of different registry functions (this may help library users as well).
v3.8.7: Python 3.13 support, Cython 3, centralize registry entries
Compare
Restores support for wheels for ARM platforms, while correctly noting compatibility range.
Restores support for wheels for ARM platforms, while correctly noting compatibility range.
v3.8.6: Restore wheels, remove Python 3.13 compatibility
Compare
Nothing published for this version
Nothing published for this version
Fix bug in memory zones when non-transient strings were added to the StringStore inside a memory zone. This caused a bug in the morphological analyser
Fix bug in memory zones when non-transient strings were added to the StringStore inside a memory zone. This caused a bug in the morphological analyser that caused string not found errors when applied during a memory zone.
Support a new context manager method Language.memory_zone(), to allow long-running services to avoid growing memory usage from cached entries in the V
Support a new context manager method Language.memory_zone(), to allow long-running services to avoid growing memory usage from cached entries in the Vocab or StringStore. Once the memory zone block ends, spaCy will evict Vocab and StringStore entries that were added during the block, freeing up memory. Doc objects created inside a memory zone block should not be accessed outside the block.
The current implementation disables population of the tokenizer cache inside the memory zone, resulting in some performance impact. The performance difference will likely be negligible if you're running a full pipeline, but if you're only running the tokenizer, it'll be much slower. If this is a problem, you can mitigate it by warming the cache first, by processing the first few batches of text without creating a memory zone. Support for memory zones in the tokenizer will be added in a future update.
The Language.memory_zone() context manager also checks for a memory_zone() method on pipeline components, so that components can perform similar memory management if necessary. None of the built-in components currently require this.
If you component needs to add non-transient entries to the StringStore or Vocab, you can pass the allow_transient=False flag to the Vocab.add() or StringStore.add() components.
Example usage:
import spacy
import json
from pathlib import Path
from typing import Iterator
from collections import Counter
import typer
from spacy.util import minibatch
def texts(path: Path) -> Iterator[str]:
with path.open("r", encoding="utf8") as file_:
for line in file_:
yield json.loads(line)["text"]
def main(jsonl_path: Path) -> None:
nlp = spacy.load("en_core_web_sm")
counts = Counter()
batches = minibatch(texts(jsonl_path), 1000)
for i, batch in enumerate(batches):
print("Batch", i)
with nlp.memory_zone():
for doc in nlp.pipe(batch):
for token in doc:
counts[token.text] += 1
for word, count in counts.most_common(100):
print(count, word)
if __name__ == "__main__":
typer.run(main)
Numpy 2.0 isn't binary-compatible with numpy v1, so we need to build against one or the other. This release isolates the dependency change and has no other changes, to make things easier if the dependency change causes problems.
This dependency change was previously attempted in version 3.7.6, but dependencies within the v3.7 family of models resulted in some conflicts, and some packages depending on numpy v1 were incompatible with v3.7.6. I've therefore removed the 3.7.6 release and replaced it with this one, which increments the minor version.
I've also made a change to the way models are packaged to make it easier to release more quickly. Previously spaCy models specified a versioned requirement on spacy itself. This meant that there was no way to increment the spaCy version and have it work with the existing models, because the models would specify they were only compatible with spacy>=3.7.0,<3.8.0. We have a compatibility table that allows spacy to see which models are compatible, but the models themselves can't know which future versions of spaCy they work with.
I've therefore added a flag --require-parent/--no-require-parent to the spacy package CLI, which controls where the parent package (e.g. spaCy) should be listed as a requirement of the model. --require-parent is the default for v3.8, but this will change to --no-require-parent by default in v4. I've set --no-require-parent for the v3.8 models, so that further changes can be published that don't impact the models, without retraining the models or forcing users to redownload them.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Sanitize direct download for spacy download (#13313).
spacy download (#13313).typing-extensions<5.0.0 for Python < 3.8 (#13516).use_gold_ents behaviour for EntityLinker.MorphAnalysis (#13433).@danieldk, @honnibal, @ines, @JoeSchiff, @nokados, @Paillat-dev, @rmitsch, @schorfma, @strickvl, @svlandeg, @ynx0
Improve NumPy 2.0 compatibility (#13103).
TextCatReduce.v1 layer for text classification (#13181).TextCatParametricAttention.v1 layer for text classification (#13201).build module for creating model packages by default (#13109).benchmark speed command (#13247).Language.pipe.Doc.Tokenizer.explain for special cases with whitespace.SparseLinear layer.trf_data examples and the transformer pipeline design section.@adrianeboyd, @danieldk, @evornov, @honnibal, @ines, @lise-brinck, @ridge-kimani, @rmitsch, @shadeMe, @svlandeg
Nothing published for this version
Update __all__ fields (#13063).
__all__ fields (#13063).spacy.cli.project API.Any comparisons for Token and Span.spacy-llm including Azure OpenAI, PaLM, and Mistral support.@adrianeboyd, @honnibal, @ines, @rmitsch, @svlandeg
Revert lazy loading of CLI module for spacy.info to fix availability of spacy.cli following import spacy (#13040).
spacy.info to fix availability of spacy.cli following import spacy (#13040).@adrianeboyd, @honnibal, @ines, @svlandeg
spacy project has a few backwards incompatibilities due to the transition to the standalone library Weasel, which is not as tightly coupled to spaCy.…
This release drops support for Python 3.6 and adds support for Python 3.12.
spacy project commands should run as before, just now they're using Weasel under the hood.transformers extra to spacy-transformers v1.3 (#13025).--spans-key option for CLI evaluation with spacy benchmark accuracy (#12981).spacy.info (#12962).spacy.training.example (#12801).Language.replace_listeners: Pass the replaced listener and the tok2vec pipe to the callback in order to support spacy-curated-transformers (#12785).tqdm with disable=None to disable output in non-interactive environments (#12979).The transformer-based trf pipelines have been updated to use our new Curated Transformers library through the Thinc model wrappers and pipeline component from spaCy Curated Transformers.
ray extra.spacy project has a few backwards incompatibilities due to the transition to the standalone library Weasel, which is not as tightly coupled to spaCy. Weasel produces warnings when it detects older spaCy-specific settings in your environment or project config.
spacy_version configuration key has been dropped.check_requirements configuration key has been dropped due to the deprecation of pkg_resources.SPACY_CONFIG_OVERRIDES environment variable is no longer checked. You can set configuration overrides using WEASEL_CONFIG_OVERRIDES.SPACY_PROJECT_USE_GIT_VERSION environment variable has been dropped.@adrianeboyd, @bdura, @connorbrinton, @danieldk, @davidberenstein1957, @denizcodeyaa, @eltociear, @evornov, @honnibal, @ines, @jmyerston, @koaning, @magdaaniol, @pdhall99, @ringohoffman, @rmitsch, @senisioi, @shadeMe, @svlandeg, @vinbo8, @wjbmattingly
Allow Pydantic v2 using transitional v1 support (#12888).
find-function CLI for finding locations of registered functions (#12757).spacy[cuda12x] for cupy-cuda12x (#12890).init config and train CLI (#12173).distutils to setuptools/sysconfig (#12853).<br> tags in displaCy.@adrianeboyd, @afriedman412, @arplusman, @bdura, @connorbrinton, @honnibal, @ines, @it176131, @pmbaumgartner, @rmitsch, @shadeMe, @svlandeg, @thomashacker, @victorialslocum, @x-tabdeveloping
NEW: `span_finder` pipeline component to identify overlapping, unlabeled spans (#12507).
span_finder pipeline component to identify overlapping, unlabeled spans (#12507).spacy evaluate --per-component, Language.evaluate(per_component=True) and Scorer.score(per_component=True) (#12540).spancat_singlelabel in spacy debug data CLI (#12749).PhraseMatcher and SpanGroup (#12642, #12714).SpanGroup spans come from the current doc.We have added new pipelines for Slovenian that use the trainable lemmatizer and floret vectors.
| Package | UPOS | Parser LAS | NER F |
|---|---|---|---|
sl_core_news_sm |
96.9 | 82.1 | 62.9 |
sl_core_news_md |
97.6 | 84.3 | 73.5 |
sl_core_news_lg |
97.7 | 84.3 | 79.0 |
sl_core_news_trf |
99.0 | 91.7 | 90.0 |
The English pipelines have been updated to improve handling of contractions with various apostrophes and to lemmatize "get" as a passive auxiliary.
The Danish pipeline da_core_news_trf has been updated to use vesteinn/DanskBERT with performance improvements across the board.
SpanGroup spans are now required to be from the same doc. When initializing a SpanGroup, there is a new check to verify that all added spans refer to the current doc. Without this check, it was possible to run into string store or other errors.@adrianeboyd, @bdura, @danieldk, @davidberenstein1957, @diyclassics, @essenmitsosse, @honnibal, @ines, @isabelizimm, @jmyerston, @kadarakos, @KennethEnevoldsen, @khursani8, @ljvmiranda921, @rmitsch, @shadeMe, @svlandeg, @tomaarsen, @victorialslocum, @vin-ivar, @ZiadAmerr
Nothing published for this version
Extend Typer support to v0.9 (#12631).
@adrianeboyd, @bdura, @honnibal, @ines, @svlandeg
Huge speed improvements for spancat, in particular on GPU (~10x-30x faster) (#12577).
spancat, in particular on GPU (~10x-30x faster) (#12577).>+, >-, >++, >--) for the dependency matcher (#12528).doc.spans for displaCy output in spacy benchmark accuracy / spacy evaluate (#12575).MorphAnalysis.get(default=) argument for user-provided default values similar to dict (#12545).#egg from download URLs due to future deprecation in pip.@adrianeboyd, @andyjessen, @bdura, @davidberenstein1957, @diyclassics, @honnibal, @ines, @kadarakos, @KennethEnevoldsen, @ljvmiranda921, @moxley01, @royashcenazi, @svlandeg, @tanloong, @victorialslocum
Add support for floret vectors in spacy pretrain (#12435).
spacy pretrain (#12435).model-last.bin for spacy pretrain (#12459).Span input for displacy.parse_deps (#12477).cupy install extras.Span.sents.spancat_singlelabel.Span.sents when the final sentence is the last token in a Doc.Span.kb_id and Span.id strings in Doc and DocBin serialization.@adrianeboyd, @BLKSerene, @honnibal, @ines, @kadarakos, @prajakta-1527, @rmitsch, @shadeMe, @sloev, @svlandeg, @thomashacker, @willfrey
💥 We'd love to hear more about your experience with spaCy! Take our survey here.
💥 We'd love to hear more about your experience with spaCy! Take our survey here.
spancat_singlelabel pipeline component for multi-class and non-overlapping span classification. The spancat_singlelabel component predicts at most one label for each suggested span and adds a new setting allow_overlap to restrict the output to non-overlapping spans (#11365).transformer + CNN for efficient GPU textcat with spacy init config (#11900).spacy debug data (#11419).>+, >-, <+, <-) (#12334).spacy.PlainTextCorpusReader.v1 for plain text input (#12122).alignment_mode and span_id to Span.char_span() (#12145, #12196).top_k>1 in trainable lemmatizer.test_cli_find_threshold() test more robust.registry.find().Matcher patterns with extension attributes.grc to languages with lexeme norms in spacy-lookups-data.KnowledgeBase instances configurable.auto_select_port.InMemoryLookupKB.is_empty.Lexeme.orth and Lexeme.lower.PretrainVectors.pkg_resources.@adrianeboyd, @andyjessen, @danieldk, @essenmitsosse, @honnibal, @ines, @itssimon, @kadarakos, @kwhumphreys, @ljvmiranda921, @pmbaumgartner, @polm, @richardpaulhudson, @rmitsch, @shadeMe, @svlandeg, @tanloong, @thomashacker, @victorialslocum
Patch a security vulnerability in extracting tar files (#11746).
apply CLI command to annotate new documents with a trained pipeline (#11376).benchmark CLI command to benchmark pipelines. The new benchmark speed subcommand measures the speed of a pipeline, the benchmark accuracy subcommand is a new alias for evaluate (#11902).find-threshold CLI command to identify an optimal threshold for classification models (#11280).FUZZY Matcher operator for fuzzy matches based on Levenshtein edit distance. In addition, the FUZZY and REGEX operators are now supported in combination with IN/NOT_IN. (#11359).typer v0.7.x (#11720), mypy 0.990 (#11801) and typing_extensions v4.4.x (#12036).spacy.ConsoleLogger.v3 with expanded progress tracking (#11972).textcat with spacy.textcat_scorer.v2 (#11696 and #11971) and spacy.textcat_multilabel_scorer.v2 (#11820).InMemoryLookupKB (#11268).before_update callback that is invoked at the start of each training step (#11739).SpanGroup (#11380).displacy.serve when the default port is in use (#11948).tok2vec version (#11618).tok2vec or transformer layer.textcat.Vocab.to_disk respects the exclude setting for lookups and vectors.SpanGroup and Span objects.The following changes may require you to update code that is using the relevant functionality:
textcat or textcat_multilabel model - ensure that values are 0.0 or 1.0 as explained in the docs.KnowledgeBase is now an abstract class, you should call the constructor of the new InMemoryLookupKB instead when you want to use spaCy's default KB implementation. If you've written a custom KB that inherits from KnowledgeBase, you'll need to implement its abstract methods, or alternatively inherit from InMemoryLookupKB instead.The following changes may influence the output of your language pipeline or trained models:
pymorphy3 (#11345, #11811).tok2vec defaults in all components (#11618).textcat and textcat_multilabel components (#11698).textcat and textcat_multilabel to fix a bug related to threshold for textcat and to make it possible to score multiple textcat/textcat_multilabel components in a single pipeline with custom scorers. If no custom scorers are used, the cat_p/r/f scores will now only reflect the final component's labels and performance (#11696, #11820).token_acc score to report the intended measure (# correct tokens / # predicted tokens, the same as in spaCy v2). The token_acc scores for v3.5 will be lower for the same performance because they were incorrectly inflated in v3.0-v3.4. The token_p/r/f scores should remain unchanged (#12073).The following functionality will be changed in the near future - so it's best to start updating your scripts now to make them more generic:
master branch to main.IS_SPACE as a tok2vec feature for tagger and morphologizer components to improve tagging of non-whitespace vs. whitespace tokens.spacy-transformers v1.2, which uses the exact alignment from tokenizers for fast tokenizers instead of the heuristic alignment from spacy-alignments. For all trained pipelines except ja_core_news_trf, the alignments between spaCy tokens and transformer tokens may be slightly different. More details about the spacy-transformers changes in the v1.2.0 release notes.biluo_to_iob and iob_to_biluo functions.@aaronzipp, @adrianeboyd, @albertvillanova, @ArchiDevil, @cfuerbachersparks, @damian-romero, @danieldk, @darigovresearch, @DSLituiev, @essenmitsosse, @gremur, @honnibal, @ines, @jmyerston, @JosPolfliet, @kadarakos, @koaning, @kwhumphreys, @ljvmiranda921, @MarcoGorelli, @orglce, @pmbaumgartner, @polm, @richardpaulhudson, @rmitsch, @ryndaniels, @shadeMe, @svlandeg, @thomashacker, @TrellixVulnTeam, @wannaphong, @zhiiw, @zrpxx
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
spancat for docs with zero suggestions.smart_open requirement and update deprecated options.spacy init config --gpu for environments without spacy-transformers.EditTreeLemmatizer.@adrianeboyd, @danieldk, @honnibal, @ines, @polm, @svlandeg
Extend Typer support to v0.7.x (#11720).
EntityLinker.Doc.to_json() for attributes set by getters.pipeline_package.load().spacy project requirements checks for unsupported specifiers and requirements lines.spacy.load(disable=) that could enable currently disabled components.@aaronzipp, @adrianeboyd, @honnibal, @ines, @polm, @rmitsch, @ryndaniels, @svlandeg, @thomashacker
NEW: Luganda language support (#10847).
spacy.ConsoleLogger.v2 optionally saves training logs to JSONL (#11214).DependencyMatcher to include matching parents or children to the left or the right of the node (#10371).cuda11x and cuda-autodetect (using cupy-wheel) (#11279).Doc.to_json() and Doc.from_json() (#11125).enable and disable options for spacy.load() more consistent (#11459).disable/enclude/exclude for spacy.load() (#11406).--url flag for spacy info to print the direct download URL for a pipeline (#11175).spacy project CLI (#11226).spacy debug data CLI for spancat data (#11504).spacy_version in spacy package metadata (#11552).spacy project assets (#11458).spacy pretrain command (#11210).natto-py for the ko extra (#11222).This release includes updated English pipelines for spaCy v3.4 with improved NER performance. The updates in en_core_web_* v3.4.1 address issues related to training from data with partial named entity annotation, which led to lower NER recall in English pipeline versions v3.0.0–v3.4.0. In particular, entities that appear in the sections of the OntoNotes training data without NER annotation were not predicted consistently by the earlier pipeline versions, such as names and places that are frequent in the Biblical sections, e.g., "David" and "Egypt" (see #7493).
Use spacy download to update your English pipelines to the newest version. If you'd prefer to keep using an earlier version, you can specify the version directly with e.g. spacy download -d en_core_web_sm-3.4.0. You can check that you are using the new version (v3.4.1) with spacy validate:
NAME SPACY VERSION
en_core_web_md >=3.4.0,<3.5.0 3.4.1 ✔
SetPredicate.Doc.__init__.pymorphy2_lookup lemmatizer mode for Russian and Ukrainian.Doc type, an error will now be raised (#11424).spacy.models_and_pipes_with_nvtx_range.v1 callback.Example API documentation.displacy docs.spacy project dvc.spacy-wordnet.initialize() function for pipeline components.@adrianeboyd, @bdura, @danieldk, @diyclassics, @DSLituiev, @GabrielePicco, @honnibal, @ines, @JulesBelveze, @kadarakos, @ljvmiranda921, @ninjalu, @pmbaumgartner, @polm, @radandreicristian, @richardpaulhudson, @rmitsch, @shadeMe, @stefawolf, @svlandeg, @thomashacker, @tobiusaolo, @tzussman , @yasufumy
Fix issue #11137: Fix compatibility with CuPy v9.x.
@adrianeboyd, @danieldk, @honnibal, @ines, @lll-lll-lll-lll, @Lucaterre, @MaartenGr, @mr-bjerre, @polm, @radenkovic
Support for mypy 0.950+ and pydantic v1.9 (#10786).
{n,m} operator for Matcher patterns (#10981).saxpy/sgemm provided by the Ops implementation in order to use Accelerate through thinc-apple-ops (#10773).Example.get_aligned_parse and Example.get_aligned (#10952).StringStore lookups (#10938).spacy project clone to try both main and master branches by default (#10843).init_config_cli (#10788).debug data (#10960).TrainablePipe components (#10965).SPACY_NUM_BUILD_JOBS to specify the number of build jobs to run in parallel with pip (#11073).We have added new pipelines for Croatian that use the trainable lemmatizer and floret vectors.
| Package | UPOS | Parser LAS | NER F |
|---|---|---|---|
hr_core_news_sm |
96.6 | 77.5 | 76.1 |
hr_core_news_md |
97.3 | 80.1 | 81.8 |
hr_core_news_lg |
97.5 | 80.4 | 83.0 |
🙏 Special thanks to @gtoffoli for help with the new pipelines!
The English pipelines have new word vectors:
| Package | Model Version | TAG | Parser LAS | NER F |
|---|---|---|---|---|
en_core_news_md |
v3.3.0 | 97.3 | 90.1 | 84.6 |
en_core_news_md |
v3.4.0 | 97.2 | 90.3 | 85.5 |
en_core_news_lg |
v3.3.0 | 97.4 | 90.1 | 85.3 |
en_core_news_lg |
v3.4.0 | 97.3 | 90.2 | 85.6 |
All CNN pipelines have been extended to add whitespace augmentation.
Doc.has_vector, distinguish 0-vectors and missing vectors in similarity warnings.get_array_module in textcat.Doc.has_vector now matches Token.has_vector and Span.has_vector: it returns True if at least one token in the doc has a vector rather than checking only whether the vocab contains vectors.@adrianeboyd, @danieldk, @ericholscher, @gorarakelyan, @honnibal, @ines, @jademlc, @kadarakos, @KennethEnevoldsen, @koaning, @Lucaterre, @maxTarlov, @philipvollet, @pmbaumgartner, @polm, @richardpaulhudson, @rmitsch, @sadovnychyi, @shadeMe, @shen-qin, @single-fingal, @svlandeg, @victorialslocum, @Zackere
Remove #egg from download URLs due to future deprecation in pip.
This bug fix release is primarily to address Pydantic incompatibility with typing_extensions>=4.6.0.
spancat, in particular on GPU (~10x-30x faster) (#12577).typing_extensions requirement due to Pydantic incompatibility with typing_extensions>=4.6.0.#egg from download URLs due to future deprecation in pip.@adrianeboyd, @honnibal, @ines, @kadarakos, @svlandeg
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
precomputable_biaffine by avoiding concatenation.spancat for docs with zero suggestions.smart_open requirement and update deprecated options.spacy init config --gpu for environments without spacy-transformers.EditTreeLemmatizer.@adrianeboyd, @danieldk, @honnibal, @ines, @polm, @svlandeg
Add the SpanRuler component. This component saves a list of matched spans to [Doc.spans[spans_key]](https://spacy.io/api/doc#spans).
Doc.spans[spans_key].Doc objects.debug data.Doc objects.SpanGroup objects that share the same name within one SpanGroups container.walk_head_nodes to avoid acquiring the GIL.StringStore.__getitem__ return type dependent on its parameter type.PhraseMatcher.SpanGroups.setdefault to also support Iterable[SpanGroup] as the default.ROOT is in the glossary.Doc.has_annotation and Matcher.Doc inputs passed to Language.pipe().Doc.Before this release, a validation bug allowed the configuration of a pipeline component to override the name of the pipeline itself through the name attribute. For example, the following pipeline component:
[components.transformer]
factory = "transformer"
name = "custom_transformer_name"
would be registered erroneously as custom_transformer_name. Such overrides are now ignored and a warning is emitted (#10779). From spaCy v3.3.1 onwards, this component will be registered as transformer.
@adrianeboyd, @danieldk, @freddyheppell, @honnibal, @ines, @kadarakos, @ldorigo, @ljvmiranda921, @maxTarlov, @pmbaumgartner, @polm, @pypae, @richardpaulhudson, @rmitsch, @shadeMe, @single-fingal, @svlandeg
Improved speeds for many components, see speed benchmarks for trained pipelines:
spacy.Tagger.v2 to speed up inference for the tagger, morphologizer, senter and trainable lemmatizer (#10197).Ragged with faster AlignmentArray in Example for training (#10319).Matcher speed (#10659).Doc.spans (#10250).spacy init config -p trainable_lemmatizer or using the quickstart.thinc v8.0.14+ and thinc-bigendian-ops.spacy debug diff-config.SpanCategorizer.set_candidates for debugging span suggesters.spancat and trainable_lemmatizer components.v3.3 introduces trained pipelines for Finnish, Korean and Swedish which feature the trainable lemmatizer and floret vectors. Due to the use Bloom embeddings and subwords, the pipelines have compact vectors with no out-of-vocabulary words.
| Package | Language | UPOS | Parser LAS | NER F |
|---|---|---|---|---|
fi_core_news_sm |
Finnish | 92.5 | 71.9 | 75.9 |
fi_core_news_md |
Finnish | 95.9 | 78.6 | 80.6 |
fi_core_news_lg |
Finnish | 96.2 | 79.4 | 82.4 |
ko_core_news_sm |
Korean | 86.1 | 65.6 | 71.3 |
ko_core_news_md |
Korean | 94.7 | 80.9 | 83.1 |
ko_core_news_lg |
Korean | 94.7 | 81.3 | 85.3 |
sv_core_news_sm |
Swedish | 95.0 | 75.9 | 74.7 |
sv_core_news_md |
Swedish | 96.3 | 78.5 | 79.3 |
sv_core_news_lg |
Swedish | 96.3 | 79.1 | 81.1 |
🙏 Special thanks to @aajanki, @thiippal (Finnish) and Elena Fano (Swedish) for their help with the new pipelines!
The new trainable lemmatizer is used for Danish, Dutch, Finnish, German, Greek, Italian, Korean, Lithuanian, Norwegian, Polish, Portuguese, Romanian and Swedish.
| Model | v3.2 Lemma Acc | v3.3 Lemma Acc |
|---|---|---|
da_core_news_md |
84.9 | 94.8 |
de_core_news_md |
73.4 | 97.7 |
el_core_news_md |
56.5 | 88.9 |
fi_core_news_md |
- | 86.2 |
it_core_news_md |
86.6 | 97.2 |
ko_core_news_md |
- | 90.0 |
lt_core_news_md |
71.1 | 84.8 |
nb_core_news_md |
76.7 | 97.1 |
nl_core_news_md |
81.5 | 94.0 |
pl_core_news_md |
87.1 | 93.7 |
pt_core_news_md |
76.7 | 96.9 |
ro_core_news_md |
81.8 | 95.5 |
sv_core_news_md |
- | 95.5 |
Scorer.score_cats for missing labels._ value for UPOS in CoNLL-U converter.Span attributes consistently."spans" to the output of doc.to_json.Matcher handling for all special cases.Example to align whitespace annotation.Tok2Vec for empty batches.rehearse.Vectors.n_keys for floret vectors.meta in util.load_model_from_config.Example.get_matching_ents.Tokenizer.explain.KoreanTokenizer tag map.init vectors.Tagger architecture, edit your configs to switch from spacy.Tagger.v1 to spacy.Tagger.v2 and then run init fill-config.<, <=, >, >=) now take all span attributes into account (start, end, label, and KB ID) so spans may be sorted in a slightly different order (#9956).Doc.from_docs now includes Doc.tensor by default and supports excludes with an exclude argument in the same format as Doc.to_bytes. The supported exclude fields are spans, tensor and user_data.@aajanki, @adrianeboyd, @apjanco, @bdura, @BramVanroy, @danieldk, @danmysak, @davidberenstein1957, @DuyguA, @fonfonx, @gremur, @HaakonME, @harmbuisman, @honnibal, @ines, @internaut, @jfainberg, @jnphilipp, @jsnfly, @kadarakos, @koaning, @ljvmiranda921, @martinjack, @mgrojo, @nrodnova, @ofirnk, @orglce, @pepemedigu, @philipvollet, @pmbaumgartner, @polm, @richardpaulhudson, @ryndaniels, @SamEdwardes, @Schero1994, @shadeMe, @single-fingal, @svlandeg, @thebugcreator, @thomashacker, @umaxfun, @y961996
Remove #egg from download URLs due to future deprecation in pip.
This bug fix release is primarily to address Pydantic incompatibility with typing_extensions>=4.6.0.
spancat, in particular on GPU (~10x-30x faster) (#12577).typing_extensions requirement due to Pydantic incompatibility with typing_extensions>=4.6.0.#egg from download URLs due to future deprecation in pip.@adrianeboyd, @honnibal, @ines, @kadarakos, @svlandeg
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
spancat for docs with zero suggestions.smart_open requirement and update deprecated options.spacy init config --gpu for environments without spacy-transformers.@adrianeboyd, @honnibal, @ines, @polm, @svlandeg
Fix issue #10564: Restrict supported Click versions as a workaround for incompatibilities between Click v8.1.0 and Typer v0.4.0.
@adrianeboyd, @honnibal, @ines
Fix issue #10324: Fix Tok2Vec for empty batches.
Tok2Vec for empty batches.@adrianeboyd, @honnibal, @ines
Improved parser and ner speeds on long documents (see technical details in #10019).
parser and ner speeds on long documents (see technical details in #10019).spancat components in debug data.ENT_IOB as a Matcher token pattern key.ENT_IOB.debug data.Lexeme.rank.spacy project.Doc.from_docs() for empty docs.debug data for components with custom names.Underscore and DependencyMatcher and improve types in Language, Matcher and PhraseMatcher.Tokenizer.explain when infixes appear as prefixes.spancat initialization.IS_SENT_END in Doc.has_annotation.spacy package.PhraseMatcher.Dockerfile for repeatable website builds and easier local development.@adrianeboyd, @antonpibm, @ColleterVi, @danieldk, @DuyguA, @ezorita, @HaakonME, @honnibal, @ines, @jboynyc, @KennethEnevoldsen, @ljvmiranda921, @mrshu, @pmbaumgartner, @polm, @ramonziai, @richardpaulhudson, @ryndaniels, @svlandeg, @thiippal, @thomashacker, @yoavxyoav
NEW: doc_cleaner component for removing doc.tensor,doc._._trf_data or other Doc attributes at the end of the pipeline to reduce size of output docs.
doc_cleaner component for removing doc.tensor,doc._._trf_data or other Doc attributes at the end of the pipeline to reduce size of output docs.ENT_ID and ENT_KB_ID to Matcher pattern attributes.kb_id for entities in displaCy from Doc input.Span.sents property for spans spanning over more than one sentence.EntityRuler.remove to remove patterns by id.Tagger neg_prefix configurable.Language.pipe in Language.evaluate for more efficient processing.JsonlCorpus path optional again.spancat for empty docs and zero suggestions..jsonl paths in EntityRuler.Scorer.score_spans to handle predicted docs with missing annotation.parser from reference parse rather than aligned example.tagger and morphologizer.init_tok2vec after pretraining, batch contract for listeners.eng-spacysentiment: Sentiment analysis for English.@adrianeboyd, @danieldk, @DuyguA, @honnibal, @ines, @ljvmiranda921, @narayanacharya6, @nrodnova, @Pantalaymon, @polm, @richardpaulhudson, @svlandeg, @thiippal, @Vishnunkumar
NEW: Registered scoring functions for each component in the config.
nlp() and nlp.pipe() accept Doc input, which simplifies setting custom tokenization or extensions before processing.overwrite config settings for entity_linker, morphologizer, tagger, sentencizer and senter.extend config setting for morphologizer for whether existing feature types are preserved.spacy.blank() including IETF language tags, for example fra for French and zh-Hans for Chinese.spacy-loggers for additional loggers.sudachipy are annotated as Token.morph features.morph_micro_p/r/f scores for morphological features from Scorer.score_morph_per_feat().LIKE_URL attribute includes the tokenizer URL pattern.--n-save-epoch option for spacy pretrain.ja_core_news_trf, thanks to @hiroshi-matsuda-rit and the spaCy Japanese community!tok2vec feature, improving the performance for many components, especially fine-grained tagging and sentence segmentation.Token.pos and Token.morph.For more details, see the New in v3.2 usage guide.
Language.pipe(as_tuples=True) for multiprocessing with custom error handlers.Tokenizer.Tokenizer, prefixes are now removed before suffix matches are applied, which may lead to minor differences in the output. In particular, the default tokenization of °[cfk]. is now ° c . instead of ° c. for most languages.ChineseTokenizer, JapaneseTokenizer, KoreanTokenizer, ThaiTokenizer and VietnameseTokenizer require Vocab rather than Language in __init__.DocBin, user data is now always serialized according to the store_user_data option, see #9190.pipelines/floret_vectors_demo: basic floret vector training and importing.pipelines/floret_fi_core_demo: Finnish UD+NER vector and pipeline training, comparing standard vs. floret vectors.pipelines/floret_ko_ud_demo: Korean UD vector and pipeline training, comparing standard vs. floret vectors.@adrianeboyd, @Avi197, @baxtree, @BramVanroy, @cayorodriguez, @DuyguA, @fgaim, @honnibal, @ines, @Jette16, @jimregan, @polm, @rspeer, @rumeshmadhusanka, @svlandeg, @syrull, @thomashacker
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
spancat for docs with zero suggestions.smart_open requirement and update deprecated options.spacy init config --gpu for environments without spacy-transformers.@adrianeboyd, @honnibal, @ines, @polm, @svlandeg
Fix issue #10564: Restrict supported Click versions as a workaround for incompatibilities between Click v8.1.0 and Typer v0.4.0.
@adrianeboyd, @honnibal, @ines
Fix issue #9593: Use metaclass to subclass errors for easier pickling.
spancat for empty docs and zero suggestions.Lexeme.rank.Tok2Vec for empty batches.@adrianeboyd, @BramVanroy, @brucewlee, @danieldk, @honnibal, @ines, @ljvmiranda921, @polm, @svlandeg, @vgautam, @xxyzz
NEW: Binary wheels for Python 3.10.
AppleOps: pip install spacy[apple].spacy.models_with_nvtx_range.v1.mypy integration in the CI and many type fixes across the code base.Protocol classes in ty.py to define behavior of pipeline components.displacy.spacy project assets .train function to run the training from Python scripts just like the spacy train CLI.spacy-transformers>=1.1.0 with improved IO.thinc>=8.0.11 with improved gradient clipping.KnowledgeBase.set_entities.DocBin constructor.spacy project title.DependencyMatcher.textcat and textcat_multilabel configurations.Doc object creation.convert CLI..pyi files in the distributed package.deplacy: CUI-based dependency visualizeripymarkup: Visualizations for NER and syntax treesPhruzzMatcher: Find fuzzy matchesspacy-huggingface-hub: Push spaCy pipelines to the Hugging Face HubspaCyOpenTapioca: Entity Linking on Wikidataspacy-clausie: Clause-based information extraction system@adrianeboyd, @connorbrinton, @danieldk, @DuyguA, @honnibal, @ines, @Jette16, @ljvmiranda921, @mjvallone, @philipvollet, @polm, @rspeer, @ryndaniels, @shigapov, @svlandeg, @thomashacker
The v3 of `WandbLogger` now supports optional run_name and entity parameters.
v3 of WandbLogger now supports optional run_name and entity parameters.pos values for a Doc or Token.Matcher callbacks.config in create_pipe.typer 0.4 to provide support for both Click 7 and Click 8.spacy project workflows.repo and path arguments in spacy project.epoch_resume in spacy pretrain.spacy-legacy in spacy package dependency detection.spacy package.StringStore and the Vocab.@adrianeboyd, @davidefiocco, @davidstrouk, @filipematos95, @honnibal, @ines, @j-frei, @Joozty, @kwhumphreys, @mjhajharia, @mylibrar, @polm, @rspeer, @shigapov, @svlandeg, @thomashacker
NEW: Provide scores for the SpanCategorizer predictions.
SpanCategorizer predictions..pyi stub files.spacy package.INTERSECTS operator for the Matcher.spacy project push and pull commands.Span.as_doc calls.da transformer is now the same as the one from the trained pipelines (Maltehb/danish-bert-botxo).debug data runs correctly with a custom tokenizer.ISSUBSET and ISSUPERSET in schema and docs.no_skip value for spacy project run.ConsoleLogger flush after each logging line.exclude when serializing the vocab.allow_overlap default for span categorizer scoring._SP.@adrianeboyd, @bbieniek, @DuyguA, @ezorita, @HLasse, @honnibal, @ines, @kabirkhan, @kevinlu1248, @ldorigo, @Ledenel, @nsorros, @polm, @svlandeg, @swfarnsworth, @themrmax, @thomashacker
Alpha tokenization support for Ancient Greek.
noun_chunk iterator for Dutch.black & flake8 as pre-commit hooks.spacy.ngram_range_suggester.v1 for suggesting a range of n-gram sizes for the spancat component.ru and uk multiprocessing (with spawn).meta information with spacy package.replace_pipe takes disabled components into account.@adrianeboyd, @honnibal, @ines, @jmyerston, @julien-talkair, @KennethEnevoldsen, @mariosasko, @mylibrar, @polm, @rynoV, @svlandeg, @thomashacker, @yohasebe
NEW: Trained pipelines for Catalan and a new transformer-based pipeline for Danish.
SpanCategorizer component for labeling arbitrary and potentially overlapping spans of text.[training.annotating_components] config setting.EntityRecognizer with known incorrect span annotations.README.md based on the meta in spacy package.For more details, see the New in v3.1 usage guide.
| Package | Language | UPOS | Parser LAS | NER F |
|---|---|---|---|---|
ca_core_news_sm |
Catalan | 98.2 | 87.4 | 79.8 |
ca_core_news_md |
Catalan | 98.3 | 88.2 | 84.0 |
ca_core_news_lg |
Catalan | 98.5 | 88.4 | 84.2 |
ca_core_news_trf |
Catalan | 98.9 | 93.0 | 91.2 |
da_core_news_trf |
Danish | 98.0 | 85.0 | 82.9 |
spacy_version in your model package meta to ">=3.0.0,<3.2.0". If you run into degraded performance, retrain your pipeline with v3.1.spacy init fill-config to update a v3.0 config for v3.1.[initialize.vectors].warnings.filterwarnings or the new helper method spacy.errors.filter_warning(action, error_msg='') to manage warnings.For more information, see Notes on upgrading from v3.0.
spacy ray command works.debug data.EntityLinker robust for nO=None.minn is not set.debug model for transformers.ENT_KB_ID in ner annotation.Doc.from_docs() for all empty docs.textcat with listener.ENT_ID and NORM to DocBin strings.Span.as_doc.Span attrs writable.debug data for textcat.DocBin is too large.to/from_bytes for KnowledgeBase and EntityLinker.Span.get_lca_matrix.attrs.IDS.spacy.batch_by_words.v1.EntityRuler: ent_ids returns None for phrases.EntityRuler.Doc.Span.lemma_.Example.from_dict.Language.pipe return values.Doc.from_docs.textcat with <2 labels.@aajanki, @adrianeboyd, @bodak, @bryant1410, @dhruvrnaik, @explosion-bot, @fhopp, @frascuchon, @graue70, @gtoffoli, @honnibal, @ines, @jacopofar, @jenojp, @jhroy, @jklaise, @juliensalinas, @kevinlu1248, @ldorigo, @mathcass, @meghanabhange, @michael-k, @narayanacharya6, @NirantK, @nsorros, @polm, @sevdimali, @svlandeg, @themrmax, @xadrianzetx, @yohasebe, @ZeeD
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
This bug fix release is primarily to avoid deprecation warnings and future incompatibility with NumPy v1.24+.
smart_open requirement and update deprecated options.spacy init config --gpu for environments without spacy-transformers.@adrianeboyd, @honnibal, @ines, @polm, @svlandeg
Your coding agent can read these notes before it upgrades. Set up the MCP server →