NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #829 most downloaded on PyPI
Industrial-strength Natural Language Processing (NLP) in Python
Last release 1 months ago
24 Aug 2026
Release timing varies
gaps range from 1 weeks to 6 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
9 versions withdrawn
withdrawn after publishing
12 years old
218 releases · first in 2015
Fix issue #10324: Fix Tok2Vec for empty batches.
Tok2Vec for empty batches.@adrianeboyd, @danieldk, @honnibal, @ines
Alpha tokenization support for Azerbaijani.
debug data.load_lookups return type and docstring.EntityLinker robust for nO=None.minn is not set.debug model for transformers.ENT_KB_ID in ner annotation.Matcher(as_spans) on spans.Doc.from_docs() for all empty docs.textcat with listener.ENT_ID and NORM to DocBin strings.Span.as_doc.Span attrs writable.debug data for textcat.DocBin is too large.to/from_bytes for KnowledgeBase and EntityLinker.Span.get_lca_matrix.attrs.IDS.spacy.batch_by_words.v1.EntityRuler: ent_ids returns None for phrases.EntityRuler.pymorphy2 requirement to pymorphy2 mode in Russian and Ukrainian lemmatizers.Doc.Span.lemma_.JsonlReader path optional.Example.from_dict.Doc.from_docs.textcat with <2 labels.@adrianeboyd, @bodak, @bryant1410, @dhruvrnaik, @fhopp, @frascuchon, @graue70, @ines, @jenojp, @jhroy, @jklaise, @juliensalinas, @meghanabhange, @michael-k, @narayanacharya6, @polm, @sevdimali, @svlandeg, @ZeeD
One column per quarter.
New assemble CLI command for assembling a pipeline from a config without training.
assemble CLI command for assembling a pipeline from a config without training.Matcher to align matched tokens with matcher patterns.spacy.WandbLogger.v2.Scorer.score_spans to support overlapping and unlabeled spans.debug data for new v3 components.vocab kwarg to spacy.load.--code usage in CLI commands.upstream check in pretraining.callbacks entry points.doc.spans in Doc.from_docs().DependencyMatcher on spans.__add__ method for PRFScore.Span.as_doc and Doc.from_docs.replace_listeners in configs.StaticVectors.Lemmatizer.Tokenizer.explain for special cases in v3.Example.from_dict for sent starts.pipe and multiprocessing.Thanks to @alvaroabascar, @armsp, @AyushExel, @BramVanroy, @broaddeep, @bryant1410, @bsweileh, @dpalmasan, @Findus23, @graue70, @jaidevd, @koaning, @langdonholmes, @m0canu1, @meghanabhange, @paoloq, @plison, @richardpaulhudson, @SamEdwardes, @Stannislav for the pull requests and contributions!
Fix related to issue #7075: Update thinc requirement for Jupyter notebook GPU warning
thinc requirement for Jupyter notebook GPU warningAllow sourcing disabled components in config.
Doc.spans in Example.from_dict.init config.tok2vec pretraining and pretrain command work as expected again.Span.sent.conll converter option.n_sents to entity linker and fix config handling and I/O.evaluate CLI.UkrainianLemmatizer.Sentencizer to use Pipe API.SpanGroup import from spacy.tokens.bg and bn.spans weakref in Doc.copy.is_cython_func for additional imported code.spacy.orth_variants.v1 and spacy.lower_case.v1 augmenters work as expected.EntityRuler.labels alphabetically.textcat_multilabel component.Vocab.get_noun_chunks.Thanks to @MartinoMensio, @SergeyShk, @R1j1t, @palandlom, @dardoria, @Tocic, @clippered, @graue70, @koaning and @jankrepl for the pull requests and contributions!
Fix issue #7035, #7056: Fix parser transition bug that could lead to incorrect sentence fragments.
init fill-config.Thanks @MartinoMensio for the pull request!
NEW: Base support for Setswana.
PhraseMatcher can now also be run on Span objects.project.yml: a section env defines environment variable names that can be used in commands. The project run command now also supports CLI overrides, e.g. --vars.batch_size 128.init config CLI.noun_chunks when pickling Vocab.vocab forward in spacy.blank.is_same_func works correctly for classes in component decorator.spacy evaluate printer.Doc in batch.Thanks to @peter-exos, @KoichiYasuoka, @tarskiandhutch, @reneoctavio, @melonwater211, @mapmeld and @Shumie82 for the pull requests and contributions.
Fix issue #6883: Fix bug in transformer training for Cannot get dimension 'nO' for model 'transformer': value unset.
Cannot get dimension 'nO' for model 'transformer': value unset.New in v3.0: New features, backwards incompatibilities and migration guide.
📣 NEW: Want to make the transition from spaCy v2 to spaCy v3 as smooth as possible for you and your organization? We're now offering commercial migration support for your spaCy pipelines! We've put a lot of work into making it easy to upgrade your existing code and training workflows – but custom projects may always need some custom work, especially when it comes to taking advantage of the new capabilities. Details & application →
For the smoothest updating process, we recommend starting with a fresh virtual environment.
pip install -U spacy
SentenceRecognizer, Morphologizer, Lemmatizer, AttributeRuler and Transformer.DependencyMatcher for matching patterns within the dependency parse using Semgrex operators.Matcher.SpanGroup for efficiently storing collections of potentially overlapping spans via the Doc.spans.To download a trained pipeline, you can use the spacy download command. See the training documentation for details on how to train your own pipelines on your data.
| Name | Language | POS | TAG | LAS | UAS | NER | Sent | Size | |
|---|---|---|---|---|---|---|---|---|---|
da_core_news_lg v3.0.0 |
Danish | 0.97 | 0.97 | 0.78 | 0.82 | 0.82 | 0.88 | 547 MB | 📖 |
da_core_news_md v3.0.0 |
Danish | 0.96 | 0.96 | 0.78 | 0.82 | 0.81 | 0.86 | 47 MB | 📖 |
da_core_news_sm v3.0.0 |
Danish | 0.95 | 0.95 | 0.76 | 0.81 | 0.72 | 0.86 | 17 MB | 📖 |
de_core_news_lg v3.0.0 |
German | 0.98 | 0.98 | 0.91 | 0.93 | 0.85 | 0.95 | 546 MB | 📖 |
de_core_news_md v3.0.0 |
German | 0.98 | 0.98 | 0.91 | 0.93 | 0.84 | 0.95 | 47 MB | 📖 |
de_core_news_sm v3.0.0 |
German | 0.98 | 0.97 | 0.90 | 0.92 | 0.82 | 0.94 | 18 MB | 📖 |
de_dep_news_trf v3.0.0 |
German | 0.99 | 0.99 | 0.95 | 0.96 | n/a | 0.98 | 393 MB | 📖 |
el_core_news_lg v3.0.0 |
Greek | 0.97 | 0.94 | 0.85 | 0.88 | 0.80 | 1.00 | 544 MB | 📖 |
el_core_news_md v3.0.0 |
Greek | 0.96 | 0.93 | 0.84 | 0.87 | 0.79 | 1.00 | 42 MB | 📖 |
el_core_news_sm v3.0.0 |
Greek | 0.94 | 0.91 | 0.81 | 0.85 | 0.72 | 1.00 | 12 MB | 📖 |
en_core_web_lg v3.0.0 |
English | n/a | 0.97 | 0.90 | 0.92 | 0.86 | 0.89 | 742 MB | 📖 |
en_core_web_md v3.0.0 |
English | n/a | 0.97 | 0.90 | 0.92 | 0.85 | 0.89 | 44 MB | 📖 |
en_core_web_sm v3.0.0 |
English | n/a | 0.97 | 0.90 | 0.92 | 0.84 | 0.89 | 13 MB | 📖 |
en_core_web_trf v3.0.0 |
English | n/a | 0.98 | 0.94 | 0.95 | 0.90 | 0.89 | 438 MB | 📖 |
es_core_news_lg v3.0.0 |
Spanish | 0.99 | 0.98 | 0.88 | 0.91 | 0.90 | 1.00 | 547 MB | 📖 |
es_core_news_md v3.0.0 |
Spanish | 0.99 | 0.98 | 0.88 | 0.91 | 0.90 | 1.00 | 46 MB | 📖 |
es_core_news_sm v3.0.0 |
Spanish | 0.98 | 0.97 | 0.87 | 0.90 | 0.89 | 1.00 | 17 MB | 📖 |
es_dep_news_trf v3.0.0 |
Spanish | 0.99 | 0.98 | 0.93 | 0.95 | n/a | 0.97 | 395 MB | 📖 |
fr_core_news_lg v3.0.0 |
French | 0.98 | 0.95 | 0.86 | 0.90 | 0.82 | 0.88 | 546 MB | 📖 |
fr_core_news_md v3.0.0 |
French | 0.97 | 0.94 | 0.85 | 0.89 | 0.81 | 0.87 | 45 MB | 📖 |
fr_core_news_sm v3.0.0 |
French | 0.96 | 0.93 | 0.84 | 0.88 | 0.79 | 0.85 | 16 MB | 📖 |
fr_dep_news_trf v3.0.0 |
French | 0.99 | 0.96 | 0.92 | 0.94 | n/a | 0.94 | 381 MB | 📖 |
it_core_news_lg v3.0.0 |
Italian | 0.98 | 0.97 | 0.88 | 0.91 | 0.89 | 0.97 | 545 MB | 📖 |
it_core_news_md v3.0.0 |
Italian | 0.97 | 0.97 | 0.88 | 0.91 | 0.87 | 0.97 | 44 MB | 📖 |
it_core_news_sm v3.0.0 |
Italian | 0.97 | 0.97 | 0.86 | 0.90 | 0.85 | 0.97 | 16 MB | 📖 |
ja_core_news_lg v3.0.0 |
Japanese | 0.96 | 0.97 | 0.90 | 0.92 | 0.72 | 0.98 | 531 MB | 📖 |
ja_core_news_md v3.0.0 |
Japanese | 0.96 | 0.97 | 0.90 | 0.92 | 0.72 | 0.99 | 41 MB | 📖 |
ja_core_news_sm v3.0.0 |
Japanese | 0.96 | 0.97 | 0.90 | 0.92 | 0.64 | 0.99 | 12 MB | 📖 |
lt_core_news_lg v3.0.0 |
Lithuanian | 0.96 | 0.89 | 0.68 | 0.75 | 0.80 | 0.82 | 545 MB | 📖 |
lt_core_news_md v3.0.0 |
Lithuanian | 0.95 | 0.86 | 0.67 | 0.74 | 0.79 | 0.83 | 44 MB | 📖 |
lt_core_news_sm v3.0.0 |
Lithuanian | 0.91 | 0.82 | 0.59 | 0.68 | 0.74 | 0.79 | 15 MB | 📖 |
mk_core_news_lg v3.0.0 |
Macedonian | 0.93 | n/a | 0.51 | 0.68 | 0.76 | 0.73 | 312 MB | 📖 |
mk_core_news_md v3.0.0 |
Macedonian | 0.93 | n/a | 0.51 | 0.67 | 0.75 | 0.73 | 44 MB | 📖 |
mk_core_news_sm v3.0.0 |
Macedonian | 0.92 | n/a | 0.47 | 0.62 | 0.70 | 0.62 | 18 MB | 📖 |
nb_core_news_lg v3.0.0 |
Norwegian | 0.97 | 0.97 | 0.87 | 0.89 | 0.85 | 0.94 | 547 MB | 📖 |
nb_core_news_md v3.0.0 |
Norwegian | 0.97 | 0.97 | 0.87 | 0.90 | 0.85 | 0.93 | 44 MB | 📖 |
nb_core_news_sm v3.0.0 |
Norwegian | 0.97 | 0.97 | 0.85 | 0.88 | 0.77 | 0.93 | 15 MB | 📖 |
nl_core_news_lg v3.0.0 |
Dutch | 0.96 | 0.95 | 0.82 | 0.87 | 0.77 | 0.87 | 546 MB | 📖 |
nl_core_news_md v3.0.0 |
Dutch | 0.96 | 0.95 | 0.82 | 0.87 | 0.74 | 0.87 | 45 MB | 📖 |
nl_core_news_sm v3.0.0 |
Dutch | 0.95 | 0.93 | 0.80 | 0.85 | 0.72 | 0.86 | 16 MB | 📖 |
pl_core_news_lg v3.0.0 |
Polish | 0.97 | 0.98 | 0.84 | 0.89 | 0.85 | 0.99 | 584 MB | 📖 |
pl_core_news_md v3.0.0 |
Polish | 0.97 | 0.98 | 0.84 | 0.89 | 0.84 | 0.98 | 84 MB | 📖 |
pl_core_news_sm v3.0.0 |
Polish | 0.95 | 0.98 | 0.79 | 0.86 | 0.80 | 0.98 | 55 MB | 📖 |
pt_core_news_lg v3.0.0 |
Portuguese | 0.97 | 0.90 | 0.86 | 0.90 | 0.91 | 0.95 | 551 MB | 📖 |
pt_core_news_md v3.0.0 |
Portuguese | 0.97 | 0.90 | 0.86 | 0.90 | 0.90 | 0.95 | 49 MB | 📖 |
pt_core_news_sm v3.0.0 |
Portuguese | 0.97 | 0.89 | 0.85 | 0.89 | 0.88 | 0.92 | 21 MB | 📖 |
ro_core_news_lg v3.0.0 |
Romanian | 0.96 | 0.97 | 0.84 | 0.89 | 0.77 | 0.96 | 546 MB | 📖 |
ro_core_news_md v3.0.0 |
Romanian | 0.96 | 0.97 | 0.85 | 0.89 | 0.76 | 0.96 | 44 MB | 📖 |
ro_core_news_sm v3.0.0 |
Romanian | 0.96 | 0.96 | 0.82 | 0.87 | 0.72 | 0.97 | 15 MB | 📖 |
ru_core_news_lg v3.0.0 |
Russian | 0.99 | 0.99 | 0.95 | 0.96 | 0.95 | 1.00 | 491 MB | 📖 |
ru_core_news_md v3.0.0 |
Russian | 0.99 | 0.99 | 0.95 | 0.96 | 0.94 | 1.00 | 41 MB | 📖 |
ru_core_news_sm v3.0.0 |
Russian | 0.99 | 0.99 | 0.95 | 0.96 | 0.95 | 1.00 | 16 MB | 📖 |
xx_ent_wiki_sm v3.0.0 |
MultiLanguage | n/a | n/a | n/a | n/a | 0.82 | n/a | 14 MB | 📖 |
xx_sent_ud_sm v3.0.0 |
MultiLanguage | n/a | n/a | n/a | n/a | n/a | 0.86 | 10 MB | 📖 |
zh_core_web_lg v3.0.0 |
Chinese | n/a | 0.90 | 0.66 | 0.71 | 0.71 | 0.75 | 577 MB | 📖 |
zh_core_web_md v3.0.0 |
Chinese | n/a | 0.90 | 0.65 | 0.70 | 0.70 | 0.76 | 76 MB | 📖 |
zh_core_web_sm v3.0.0 |
Chinese | n/a | 0.90 | 0.64 | 0.70 | 0.69 | 0.75 | 47 MB | 📖 |
zh_core_web_trf v3.0.0 |
Chinese | n/a | 0.92 | 0.73 | 0.77 | 0.75 | 0.65 | 398 MB | 📖 |
💬 TAG: Part-of-speech tags (fine-grained tags, i.e.
Token.tag_) POS: Part-of-speech tags (coarse-grained tags, i.e.Token.pos_) UAS: Unlabelled dependencies (parser). LAS: Labelled dependencies (parser). NER: Named entities (F-score). Sent: Sentence segmentation. Size: Model file size (zipped archive).
For more info on how to migrate from spaCy v2.x, see the detailed migration guide.
link command and shortcut names are now deprecated. There can be many different trained pipelines and not just one "English model", so you should always use the full package name like en_core_web_sm explicitly.meta.json is now only used to provide meta information like the package name, author, license and labels. It's not used to construct the processing pipeline anymore. This is all defined in the config.cfg, which also includes all settings used to train the pipeline.train, pretrain and debug data commands now only take a config.cfg.Language.add_pipe now takes the string name of the component factory instead of the component function.@Language.component or @Language.factory decorator.Language.update, Language.evaluate and TrainablePipe.update methods now all take batches of Example objects instead of Doc and GoldParse objects, or raw text and a dictionary of annotations.begin_training methods have been renamed to initialize and now take a function that returns a sequence of Example objects to initialize the model instead of a list of tuples.Matcher.add and PhraseMatcher.add now only accept a list of patterns as the second argument (instead of a variable number of arguments). The on_match callback becomes an optional keyword argument.Doc flags like Doc.is_parsed or Doc.is_tagged have been replaced by Doc.has_annotation.spacy.gold module has been renamed to spacy.training.PRON_LEMMA symbol and -PRON- as an indicator for pronoun lemmas has been removed.TAG_MAP and MORPH_RULES in the language data have been replaced by the more flexible AttributeRuler.Lemmatizer is now a standalone pipeline component and doesn't provide lemmas by default or switch automatically between lookup and rule-based lemmas. You can now add it to your pipeline explicitly and set its mode on initialization.| Removed | Replacement |
|---|---|
Language.disable_pipes |
Language.select_pipes, Language.disable_pipe, Language.enable_pipe |
Language.begin_training, Pipe.begin_training, ... |
Language.initialize, Pipe.initialize, ... |
Doc.is_tagged, Doc.is_parsed, ... |
Doc.has_annotation |
GoldParse |
Example |
GoldCorpus |
Corpus |
KnowledgeBase.load_bulk, KnowledgeBase.dump |
KnowledgeBase.from_disk, KnowledgeBase.to_disk |
Matcher.pipe, PhraseMatcher.pipe |
not needed |
gold.offsets_from_biluo_tags, gold.spans_from_biluo_tags, gold.biluo_tags_from_offsets |
training.biluo_tags_to_offsets, training.biluo_tags_to_spans, training.offsets_to_biluo_tags |
spacy init-model |
spacy init vectors |
spacy debug-data |
spacy debug data |
spacy profile |
spacy debug profile |
spacy link, util.set_data_path, util.get_data_path |
not needed, symlinks are deprecated |
The following deprecated methods, attributes and arguments were removed in v3.0. Most of them have been deprecated for a while and many would previously raise errors. Many of them were also mostly internals. If you've been working with more recent versions of spaCy v2.x, it's unlikely that your code relied on them.
| Removed | Replacement |
|---|---|
Doc.tokens_from_list |
Doc.__init__ |
Doc.merge, Span.merge |
Doc.retokenize |
Token.string, Span.string, Span.upper, Span.lower |
Span.text, Token.text |
Language.tagger, Language.parser, Language.entity |
Language.get_pipe |
keyword-arguments like vocab=False on to_disk, from_disk, to_bytes, from_bytes |
exclude=["vocab"] |
n_threads argument on Tokenizer, Matcher, PhraseMatcher |
n_process |
verbose argument on Language.evaluate |
logging (DEBUG) |
SentenceSegmenter hook, SimilarityHook |
user hooks, Sentencizer, SentenceRecognizer |
This release is brought to you by @honnibal, @ines, @svlandeg and @adrianeboyd. Thanks to @AMArostegui, @BramVanroy, @Cristianasp, @DeNeutoy, @DuyguA, @Jan-711, @KKsharma99, @KeshavG-lb, @KoichiYasuoka, @MartinoMensio, @Nuccy90, @PluieElectrique, @SamEdwardes, @Stannislav, @abchapman93, @alexcombessie, @alvaroabascar, @baranitharan2020, @bittlingmayer, @bjascob, @borijang, @borijang, @bratao, @bryant1410, @buriy, @chopeen, @danielvasic, @delzac, @dhruvrnaik, @erip, @florijanstamenkovic, @forest1988, @gandersen101, @garethsparks, @graue70, @guadiromero, @hertelm, @hiroshi-matsuda-rit, @holubvl3, @idoshr, @jabortell, @jbesomi, @jenojp, @jganseman, @jgutix, @jmargeta, @jumasheff, @kuk, @leyendecker, @lizhe2004, @lorenanda, @mahnerak, @mikeizbicki, @myavrum, @nipunsadvilkar, @oculusrepairo, @ophelielacroix, @rahul1990gupta, @rameshhpathak, @rasyidf, @revuel, @richardliaw, @robertsipek, @snsten, @solarmist, @tamuhey, @thomasbird, @tiangolo, @tilusnet, @timgates42, @vha14, @walterhenry, @wannaphong, @werew, @yosiasz and @zaibacu for the pull requests and contributions!
This release addresses future compatibility with NumPy v1.24+.
This release addresses future compatibility with NumPy v1.24+.
@adrianeboyd, @honnibal, @ines, @svlandeg
Updates and binary wheels for Python 3.10 and 3.11.
@adrianeboyd, @honnibal, @ines
Fix issue #8286: Fix spacy download.
spacy download.Add noun chunk iterator for Danish.
token_match and url_match for the tokenizer.Matcher.IS_SENT_START in the PhraseMatcher.Span.text for empty spans.Doc.char_span alignment_mode handling.--no-cache-dir when downloading models.Span.get_lca_matrix.Thanks to @alexcombessie, @AMArostegui, @bryant1410, @Cristianasp, @garethsparks, @jenojp, @jganseman, @jumasheff, @lorenanda, @ophelielacroix, @thomasbird, @timgates42, @tupui and @yosiasz for the pull requests and contributions.
Modify blis and numpy build dependencies to simplify source installations.
blis and numpy build dependencies to simplify source installations.cupy v8+ in combination with thinc v7.4.5.NORM on token in retokenizer.SPACY as a Matcher attribute.nlp.max_length check to nlp.pipe through nlp.make_doc..pipe methods to Chinese, Japanese, Korean and Thai tokenizers.EntityRuler.--use-chars from train CLI.Thanks to @KoichiYasuoka for the pull requests and contributions.
Fix issue #6446: Restore cleanup_beam method.
cleanup_beam method.Thanks to @jabortell for the pull requests and contributions.
NEW: Add alpha support for Macedonian and Sanskrit.
sys.argv exists.ent_id_ to strings serialized with Doc.pretrain.EntityRenderer to support break lines (after last entity).EntityRuler.Doc.char_span to snap to token boundaries.Span index boundary checks.Matcher pattern node by quantifier.TextCategorizer and Tok2Vec.on_match callback and exclude empty match lists from results for DependencyMatcher.beam_parse (requires thinc>=7.4.3).thinc>=7.4.3).Matcher.Thanks to @abchapman93, @baranitharan2020, @bittlingmayer, @bjascob, @borijang, @BramVanroy, @chopeen, @danielvasic, @delzac, @DuyguA, @erip, @florijanstamenkovic, @graue70, @hiroshi-matsuda-rit, @holubvl3, @idoshr, @jgutix, @KKsharma99, @leyendecker, @lizhe2004, @MartinoMensio, @nipunsadvilkar, @Nuccy90, @oculusrepairo, @rahul1990gupta, @rasyidf, @robertsipek, @SamEdwardes, @snsten, @solarmist, @Stannislav, @tamuhey, @tilusnet, @vha14, @wannaphong, @zaibacu for the pull requests and contributions.
Nothing published for this version
Improve Korean tokenizer speed.
Thanks to @graue70, @mikeizbicki, @jbesomi, @gandersen101 and @DeNeutoy for the pull requests and contributions.
NEW: Add alpha support for Nepali.
Token.is_oov and Lexeme.is_oov.ENT_KB_ID to Doc serialization.is_base_form to language settings.Thanks to @myavrum, @mahnerak, @rameshhpathak, @hiroshi-matsuda-rit, @PluieElectrique, @hertelm and @alvaroabascar for the pull requests and contributions.
> ⚠️ This version of spaCy requires downloading new models. You can use the `spacy validate` command to find out which models need updating, and print
⚠️ This version of spaCy requires downloading new models. You can use the
spacy validatecommand to find out which models need updating, and print update instructions. If you've been training your own models, you'll need to retrain them with the new version.
Matcher to match on both Doc and Span objects.Token.is_sent_end property.pkuseg alongside jieba for Chinese.fugashi to sudachipy for Japanese.gold.align.spacy-lookups-data.Span objects.Vectors.from_glove.exclusive_classes in textcat ensemble.Vocab.set_vector.None values in gold fields.GoldParse initialization when the number of tokens has changed.cupy-cuda extra dependencies.Vectors.resize to work with cupy.unittest warnings when saving a model.token_match patterns.__init__.py files to language data tests.!=.TokenC.sent_start values for Matcher.--n-save_every.max(uint64) for OOV lexeme rank.most_similar for vectors with unused rows.Span.similarity that could trigger TypeError.'d (would/had) in English.OOV_RANK.Token.sent_start for Span.sent.ErrorsWithCodes().__class__ return value.spacy validate command to find out which models need updating, and print update instructions.spacy-lookups-data, which now includes both the lemmatization tables (as in v2.2) and the normalization tables (new in v2.3). If you're using pretrained models, nothing changes, because the relevant tables are included in the model packages.ADP_DET for French "au", which maps to UPOS ADP based on the head "à". This increases the accuracy of the models by improving the alignment between spaCy's tokenization and Universal Dependencies multi-word tokens used for contractions.warnings. Instead of setting SPACY_WARNING_IGNORE, use the warnings filters to manage warnings.bin/wiki_entity_linking scripts for Wikipedia to projects repo.🔥 ICYMI: We recently updated the free and interactive spaCy course to include translations for German (with German NLP examples), Spanish (with Spanish NLP examples) and Japanese, as well as videos for English and German. Translations for Chinese (with Chinese NLP examples), French (with French NLP examples) and Russian coming soon!
| Model | Language | Version | Vectors |
|---|---|---|---|
zh_core_web_sm |
Chinese | 2.3.0 | 𐄂 |
zh_core_web_md |
Chinese | 2.3.0 | ✓ |
zh_core_web_lg |
Chinese | 2.3.0 | ✓ |
da_core_news_sm |
Danish | 2.3.0 | 𐄂 |
da_core_news_md |
Danish | 2.3.0 | ✓ |
da_core_news_lg |
Danish | 2.3.0 | ✓ |
nl_core_news_sm |
Dutch | 2.3.0 | 𐄂 |
nl_core_news_md |
Dutch | 2.3.0 | ✓ |
nl_core_news_lg |
Dutch | 2.3.0 | ✓ |
en_core_web_sm |
English | 2.3.0 | 𐄂 |
en_core_web_md |
English | 2.3.0 | ✓ |
en_core_web_lg |
English | 2.3.0 | ✓ |
fr_core_news_sm |
French | 2.3.0 | 𐄂 |
fr_core_news_md |
French | 2.3.0 | ✓ |
fr_core_news_lg |
French | 2.3.0 | ✓ |
de_core_news_sm |
German | 2.3.0 | 𐄂 |
de_core_news_md |
German | 2.3.0 | ✓ |
de_core_news_lg |
German | 2.3.0 | ✓ |
el_core_news_sm |
Greek | 2.3.0 | 𐄂 |
el_core_news_md |
Greek | 2.3.0 | ✓ |
el_core_news_lg |
Greek | 2.3.0 | ✓ |
it_core_news_sm |
Italian | 2.3.0 | 𐄂 |
it_core_news_md |
Italian | 2.3.0 | ✓ |
it_core_news_lg |
Italian | 2.3.0 | ✓ |
ja_core_news_sm |
Japanese | 2.3.0 | 𐄂 |
ja_core_news_md |
Japanese | 2.3.0 | ✓ |
ja_core_news_lg |
Japanese | 2.3.0 | ✓ |
lt_core_news_sm |
Lithuanian | 2.3.0 | 𐄂 |
lt_core_news_md |
Lithuanian | 2.3.0 | ✓ |
lt_core_news_lg |
Lithuanian | 2.3.0 | ✓ |
nb_core_news_sm |
Norwegian Bokmål | 2.3.0 | 𐄂 |
nb_core_news_md |
Norwegian Bokmål | 2.3.0 | ✓ |
nb_core_news_lg |
Norwegian Bokmål | 2.3.0 | ✓ |
pl_core_news_sm |
Polish | 2.3.0 | 𐄂 |
pl_core_news_md |
Polish | 2.3.0 | ✓ |
pl_core_news_lg |
Polish | 2.3.0 | ✓ |
pt_core_news_sm |
Portuguese | 2.3.0 | 𐄂 |
pt_core_news_md |
Portuguese | 2.3.0 | ✓ |
pt_core_news_lg |
Portuguese | 2.3.0 | ✓ |
ro_core_news_sm |
Romanian | 2.3.0 | 𐄂 |
ro_core_news_md |
Romanian | 2.3.0 | ✓ |
ro_core_news_lg |
Romanian | 2.3.0 | ✓ |
es_core_news_sm |
Spanish | 2.3.0 | 𐄂 |
es_core_news_md |
Spanish | 2.3.0 | ✓ |
es_core_news_lg |
Spanish | 2.3.0 | ✓ |
xx_ent_wiki_sm |
Multi-language | 2.3.0 | 𐄂 |
Thanks to @mabraham, @sloev, @pinealan, @pmbaumgartner, @Baciccin, @nlptechbook, @guerda, @Tiljander, @nikhilsaldanha, @tommilligan, @Jacse, @leicmi, @YohannesDatasci, @mirfan899, @koaning, @umarbutler, @chopeen, @paoloq, @thomasthiebaud, @sebastienharinck, @elben10, @laszabine, @Mlawrence95, @sabiqueqb, @punitvara, @michael-k, @louisguitton, @vondersam, @thoppe, @vishnupriyavr, @ilivans and @osori for the pull requests and contributions.
🙏 Special thanks to everyone who helped us develop and test the new models: @lixiepeng, @lingvisa and @howl-anderson (Chinese), @hvingelby (Danish), @hiroshi-matsuda-rit and @polm (Japanese), @ryszardtuora (Polish) and @avramandrei and @dumitrescustefan (Romanian).
Nothing published for this version
NEW: Add Span.char_span method.
Span.char_span method.--tag-map-path argument to debug-data and train commands.add_lemma option to displacy dependency visualizer.IDX as an attribute available via Doc.to_array.EntityRuler.python-mecab3 with fugashi for Japanese.tok2vec parameters to train command.TransitionSystem.HEAD for is_parsed in Doc.from_array.SHAPE docs and examples.HEAD field in CoNLL-U format to be an underscore.Vocab.set_entities in the KnowledgeBase more robust.el, es and pt.lr_edges until Doc.sents are correct.disabled when calling from_disk during load.Matcher.EntityLinker example.Doc.cats in serialization of Doc and DocBin.EntityLinker.predict.pyproject.toml.ENT_ID..pyx and .pxd files in the distribution.Language.evaluate.Model was already defined.Sentencizer.pipe for empty Doc.Span.__eq__ and Span.__hash__.srsly pin.get_doc test utility.IS_SENT_START to SENT_START for Matcher.Doc.from_array.merge_entities.Tokenizer.to_disk and Tokenizer.from_disk.Doc.is_ flags for empty Docs.Thanks to @polm, @mmaybeno, @jarib, @questoph, @aajanki, @mr-bjerre, @Tclack88, @thiagola92, @tamuhey, @Olamyy, @AlJohri, @iechevarria, @iurshina, @lineality, @pbadeer, @BramVanroy, @kabirkhan, @ceteri, @omri374, @maknotavailable, @onlyanegg, @drndos, @ju-sh, @nlptechbook, @chkoar, @Jan-711, @MisterKeefe, @bryant1410, @mirfan899, @dhpollack and @mabraham for the pull requests and contributions!
NEW: Tokenizer.explain method to see which rule or pattern was matched. `python tok_exp = nlp.tokenizer.explain("(don't)") assert [t[0] for t in tok_e
Tokenizer.explain method to see which rule or pattern was matched.tok_exp = nlp.tokenizer.explain("(don't)")
assert [t[0] for t in tok_exp] == ["PREFIX", "SPECIAL-1", "SPECIAL-2", "SUFFIX"]
assert [t[1] for t in tok_exp] == ["(", "do", "n't", ")"]
Scorer.las_per_type (labelled depdencency scores per label).debug-data if no dev docs are available.as_tuples=True in Language.pipe work with multiprocessing.on_match in DependencyMatcher.Retokenizer.split.conllu2json converter when -n > 1.Language.evaluate for components without .pipe method.EntityRuler is deserialized correctly from disk.Tagger or TextCategorizer.Vectors.find return keys in correct order.Thanks to @yash1994, @walterhenry, @prilopes, @f11r, @questoph, @erip, @richardpaulhudson and @GuiGel for the pull requests and contributions.
Nothing published for this version
NEW: Support multiprocessing in nlp.pipe via the n_process argument (Python 3 only).
nlp.pipe via the n_process argument (Python 3 only).spacy-lookups-data.debug-data for low sentences per doc ratio.convert and debug-data CLI.EntityRuler ID resolution 2× faster and support "id" in patterns to set Token.ent_id.displacy for RTL languages.thinc_gpu_ops for simpler GPU install.spacy pretrain.Language.disable_pipes API, which will become
the default in the future. The method can now also take a list of component names as its first argument (instead of a variable number of arguments).- disabled = nlp.disable_pipes("tagger", "parser")
+ disabled = nlp.disable_pipes(["tagger", "parser"])
Matcher.add and PhraseMatcher.add API, which will become the default in the future. The patterns are now the second argument and a list (instead of a variable number of arguments). The on_match callback becomes an optional keyword argument.patterns = [[{"TEXT": "Google"}, {"TEXT": "Now"}], [{"TEXT": "GoogleNow"}]]
- matcher.add("GoogleNow", None, *patterns)
+ matcher.add("GoogleNow", patterns)
- matcher.add("GoogleNow", on_match, *patterns)
+ matcher.add("GoogleNow", patterns, on_match=on_match)
gold.align behind a feature flag. The new alignment may produce backwards-incompatible results, so it won't be enabled by default before v3.0.import spacy.gold
spacy.gold.USE_NEW_ALIGN = True
nlp.pipe.thinc_gpu_ops for simpler GPU install.Vectors.most_similar.spacy-lookups-data.URL_PATTERN and handling in tokenizer.PhraseMatcher.vocab consistent with Matcher.vocab.pkg_resources and handling of entry points.batch_size when sorting similar vectors.ner_jsonl2json converter.on_match callback is executed in PhraseMatcher.PhraseMatcher.remove for overlapping patterns.Vectors.most_similar.gold.docs_to_json documentation.cats to GoldParse.from_annot_tuples in Scorer.stdout.PhraseMatcher.add arguments.Vectors.most_similar returns 1.0 for identical vectors.None iteration error in entity linking script.Parser sample construction of GoldParse.DocBin.GoldParse is initialized correctly with misaligned tokens.lemma_rules, lemma_index, lemma_exc and lemma_lookup of the Language.Defaults have now been removed to prevent confusion (e.g. if users add rules that then have no effect). The only place lemmatization tables are stored and can be modified at runtime is via nlp.vocab.lookups.- nlp.Defaults.lemma_lookup["spaCies"] = "spaCy"
+ lemma_lookup = nlp.vocab.lookups.get_table("lemma_lookup")
+ lemma_lookup["spaCies"] = "spaCy"
Thanks to @tamuhey, @PeterGilles, @akornilo, @danielkingai2, @ghollah, @pberba, @gustavengstrom, @ju-sh, @kabirkhan, @ZhuoruLin, @nipunsadvilkar and @neelkamath for the pull requests and contributions.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Make Vectors.most_similar return the top most similar vectors instead of only one.
Vectors.most_similar return the top most similar vectors instead of only one.DocBin with attributes.Vectors.most_similar.Thanks to @bintay and @svlandeg for the pull requests and contributuons.
> ⚠️ This version of spaCy requires downloading new models. You can use the `spacy validate` command to find out which models need updating, and print
⚠️ This version of spaCy requires downloading new models. You can use the
spacy validatecommand to find out which models need updating, and print update instructions. If you've been training your own models, you'll need to retrain them with the new version.
EntityLinker and KnowledgeBase API to train and access entity linking models, plus scripts to train your own Wikidata models.PhraseMatcher and improved phrase matching algorithm.DocBin class to efficiently serialize collections of Doc objects.spacy train and get textcat results via the Scorer.debug-data command to validate your training and development data, get useful stats, and find problems like invalid entity annotations, cyclic dependencies, low data labels and more.Lookups class using Bloom filters that allows storing, accessing and serializing large dictionaries via vocab.lookups.spacy train via the --orth-variant-level flag, which defines the percentage of occurrences of some tokens subject to replacement during training.nlp.pipe_labels (labels assigned by pipeline components) and include "labels" in nlp.meta.spacy_displacy_colors entry point to allow packages to add entity colors to displacy.template config option in displacy to customize entity HTML template.Doc.retokenize.subtok errors.Matcher retry loop that'd cause problems with ? operator.displacy.PhraseMatcher.remove method.pos and tag are correctly serialized.PhraseMatcher returns multiple matches for identical rules.biluo_tags_from_offsets.debug-data.? operator at the end of pattern.E069..html files as UTF-8 in evaluate command.force flag on set_extension.tqdm bug that'd remove text color from terminal output..pos/.tag distinction more clear in the docs.--vectors-loc documentation.Parser.tok2vec property.as_tuples and return_matches in Matcher.pipe.PhraseMatcher with very large lists to miss matches.spacy validate command to find out which models need updating, and print update instructions.spacy-lookups-data, which is not installed by default. If you're using pre-trained models, nothing changes, because the tables are now included in the model packages. If you want to use the lemmatizer for other languages that don't yet have pre-trained models (e.g. Turkish or Croatian) or start off with a blank model that contains lookup data (e.g. spacy.blank("en")), you'll need to explicitly install spaCy plus data via pip install spacy[lookups]. The data will be registered automatically via entry points.Vocab and serialized with it. This means that serialized objects (nlp, pipeline components, vocab) will now include additional data, and models written to disk will include additional files.Lemmatizer class is now initialized with an instance of Lookups containing the rules and tables, instead of dicts as separate arguments. This makes it easier to share data tables and modify them at runtime. This is mostly internals, but if you've been implementing a custom Lemmatizer, you'll need to update your code.spacy download command does not set the --no-deps pip argument anymore by default, meaning that model package dependencies (if available) will now be also downloaded and installed. If spaCy (which is also a model dependency) is not installed in the current environment, e.g. if a user has built from source, --no-deps is added back automatically to prevent spaCy from being downloaded and installed again from pip.biluo_tags_from_offsets converter is now stricter and will raise an error if entities are overlapping (instead of silently skipping them). If your data contains invalid entity annotations, make sure to clean it and resolve conflicts. You can now also use the new debug-data command to find problems in your data.ent_iob value set, it won't be reset to an "unset" state and will always have at least O assigned. list(doc.ents) now actually keeps the annotations on the token level consistent, instead of resetting O to an empty string.Sentencizer has been extended and now includes more characters common in various languages. This also means that the results it produces may change, depending on your text. If you want the previous behaviour with limited characters, set punct_chars=[".", "!", "?"] on initialization.PhraseMatcher algorithm was rewritten from scratch and it's now 10× faster. The rewrite also resolved a few subtle bugs with very large terminology lists. So if you were matching large lists, you may see slightly different results – however, the results should now be fully correct. See #4309 for details on this change.Serbian language class (introduced in v2.1.8) incorrectly used the language code rs instead of sr. This has now been fixed, so Serbian is now available via spacy.lang.sr."sources" in the meta.json have changed from a list of strings to a list of dicts. This is mostly internals, but if your code used nlp.meta["sources"], you might have to update it.| Model | Language | Version | UAS | LAS | POS | NER F | Vec | Size |
|---|---|---|---|---|---|---|---|---|
en_core_web_sm |
English | 2.2.0 | 91.61 | 89.71 | 97.03 | 85.07 | 𐄂 | 11 MB |
en_core_web_md |
English | 2.2.0 | 91.65 | 89.77 | 97.14 | 86.10 | ✓ | 91 MB |
en_core_web_lg |
English | 2.2.0 | 91.98 | 90.16 | 97.21 | 86.30 | ✓ | 789 MB |
de_core_news_sm |
German | 2.2.0 | 90.75 | 88.63 | 96.29 | 83.11 | 𐄂 | 14 MB |
de_core_news_md |
German | 2.2.0 | 91.26 | 89.36 | 96.44 | 83.42 | ✓ | 214 MB |
es_core_news_sm |
Spanish | 2.2.0 | 90.20 | 87.05 | 96.79 | 89.45 | 𐄂 | 15 MB |
es_core_news_md |
Spanish | 2.2.0 | 90.89 | 87.94 | 97.03 | 89.86 | ✓ | 74 MB |
pt_core_news_sm |
Portuguese | 2.2.0 | 89.53 | 86.07 | 79.96 | 87.97 | 𐄂 | 20 MB |
fr_core_news_sm |
French | 2.2.0 | 87.27 | 84.28 | 94.38 | 82.77 | 𐄂 | 14 MB |
fr_core_news_md |
French | 2.2.0 | 88.82 | 86.07 | 95.15 | 82.82 | ✓ | 84 MB |
it_core_news_sm |
Italian | 2.2.0 | 90.79 | 86.94 | 96.06 | 86.29 | 𐄂 | 13 MB |
nl_core_news_sm |
Dutch | 2.2.0 | 76.79 | 69.53 | 90.10 | 68.79 | 𐄂 | 14 MB |
el_core_news_sm |
Greek | 2.2.0 | 84.40 | 80.98 | 94.41 | 71.88 | 𐄂 | 10 MB |
el_core_news_md |
Greek | 2.2.0 | 87.96 | 84.88 | 96.38 | 77.59 | ✓ | 126 MB |
nb_core_news_sm |
Norwegian | 2.2.0 | 89.02 | 86.49 | 95.72 | 83.99 | 𐄂 | 12 MB |
lt_core_news_sm |
Lithuanian | 2.2.0 | 59.87 | 48.00 | 74.02 | 76.58 | 𐄂 | 12 MB |
xx_ent_wiki_sm |
Multi | 2.2.0 | - | - | - | 79.88 | 𐄂 | 3 MB |
💬 UAS: Unlabelled dependencies (parser). LAS: Labelled dependencies (parser). POS: Part-of-speech tags (fine-grained tags, i.e.
Token.tag_). NER F: Named entities (F-score). Vec: Model contains word vectors. Size: Model file size (zipped archive).
sources listed in the meta.json of pre-trained models with more details on the training corpora and include more information in the models directory.debug-data, EntityLinker, KnowledgeBase and Lookups.Thanks to @ICLRandD, @phiedulxp, @ajrader, @RyanZHe, @jenojp, @yanaiela, @isaric, @mrdbourke, @avramandrei, @Pavle992, @chkoar, @wannaphongcom, @BreakBB, @b1uec0in, @mihaigliga21, @tamuhey, @euand, @Hazoom, @SeanBE, @esemeniuc, @zqianem, @ajkl, @jaydeepborkar, @EarlGreyT and @er-raoniz for the pull requests and contributions.
Special thanks to our spaCy team @svlandeg and @adrianeboyd for the bug fixes and new features, @polm for the Bloom filters implementation and data compression and @yvespeirsman, @lemontheme, @jarib, @miktoki and @rokasramas for the help and resources for the new models.
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
> This is a small maintenance update that backports a bug fix for a memory leak that'd occur in long-running parsing processes. It's intended for user
This is a small maintenance update that backports a bug fix for a memory leak that'd occur in long-running parsing processes. It's intended for users who can't or don't yet want to upgrade to spaCy v2.2 (e.g. because it requires retraining all the models). If you're able to upgrade, you shouldn't use this version and instead install the latest v2.2.
NEW: Alpha tokenization support for Serbian
PhraseMatcher work as expected for NORM attribute.Matcher attributes.validate option to EntityRuler.pyproject.toml.Language.evaluate.Matcher attribute docs.Thanks to @akornilo, @mirfan899, @veer-bains, @seppeljordan, @Pavle992, @svlandeg, @jenojp and @adrianeboyd for the pull requests and contributions.
Add Token.tensor and Span.tensor attributes.
Token.tensor and Span.tensor attributes.(text, annotations) instead of only (doc, gold) for nlp.evaluate."lang_factory" setting in model meta.json (see #4031)."requirements" in meta.json to define packages for setup's install_requires.Pipe base class methods and make them less presumptuous.Span.as_doc that could cause segfault.TextCategorizer.predict with empty Doc.Span.sent docs.init-model command if there's no vocab.lang of nlp and nlp.vocab stay consistent.Token.similarity and Span.similarity when called via hook.gold.align helper.Thanks to @sorenlind, @pmbaumgartner, @svlandeg, @FallakAsad, @BreakBB, @adrianeboyd, @polm, @b1uec0in, @mdaudali and @ejarkm for the pull requests and contributions.
Nothing published for this version
Fix issue #3958: Fix order of symbols that caused tag maps to be out-of-sync.
NEW: Base language data for Marathi and Korean (via `mecab-ko`, `mecab-ko-dic` and `natto-py`).
mecab-ko, mecab-ko-dic and natto-py).spacy pretrain.id property to EntityRuler patterns.Doc.is_sentenced always return True for single-token docs.Scorer.EntityRuler settings correctly.E024 error message for incorrect GoldParse.ngram parameter in text classifier.Language.replace_pipe.lex_id.Doc.is_sentenced default to True for single-token Docs.--seed option in spacy pretrain.PhraseMatcher arguments to EntityRuler.Matcher returns correct match IDs when used with operators.spacy pretrain.pretrain to prevent unintended overwriting of weight files."v.s." to English tokenizer exceptions.Doc.count_by work as expected.user_data when copying doc in displaCy.Tokenizer initialization docs.Scorer, Language.evaluate and gold.docs_to_json.Thanks to @BreakBB, @ujwal-narayan, @estr4ng7d, @maknotavailable, @ramananbalakrishnan, @nipunsadvilkar, @NirantK, @munozbravo, @intrafindBreno, @Azagh3l, @jarib, @tokestermw, @polm, @skrcode, @kabirkhan, @demongolem, @elbaulp, @clarus, @BramVanroy, @rokasramas, @askhogan, @khellan, @kognate, @cedar101 and @yash1994 for the pull requests and contributions.
NEW: util.filter_spans helper to filter duplicates and overlaps from a list of Span objects.
util.filter_spans helper to filter duplicates and overlaps from a list of Span objects.--n-save-every to spacy pretrain and rename --nr-iter to --n-iter for consistency.--return-scores flag to spacy evaluate to return a dict.--n-early-stopping option to spacy train to define maximum number of iterations without dev accuracy improvements.spacy convert correctly default to json."settings" or "title" required in displaCy data.GoldParse.__del__ is a string.DependencyParser.predict docs.jupyter=False to override Jupyter mode in displacy..iob converter.jsonschema pin.spacy.explain.Doc.retokenize.Language.update docs."text" in spacy pretrain optional when "tokens" is provided.Token.prob and Lexeme.prob docs.rmtree and copytree with strings in spacy train.--base-model argument in spacy train docs.Thanks to @svlandeg, @wannaphongcom, @Bharat123rox, @DuyguA, @SamuelLKane, @graus, @HiromuHota, @jeannefukumaru, @ivigamberdiev, @socool, @yvespeirsman, @lemontheme, @Dobita21, @w4nderlust, @pierremonico, @bryant1410, @celikomer, @xssChauhan, @kowaalczyk, @BreakBB, @fizban99, @tokestermw, @bjascob, @pickfire, @yaph, @amitness, @henry860916, @d5555, @BramVanroy, @F0rge1cE, @richardpaulhudson, @ldorigo, @aaronkub and @devforfu for the pull requests and contributions.
Allow customizing punctuation characters in sentencizer and make it serializable.
"bow" architecture for TextCategorizer, to do faster bag-of-words text classification.train_textcat.py example.Token.is_sent_start correctly."ensemble" TextClassifier architecture that prevented the unigram bag-of-words submodel from working properly.Thanks to @chkoar for the pull request!
Fix issue #3356: Fix handling of unicode ranges in regular expressions on Python 2.
wasabi to better handle non-UTF-8 terminals.label argument in Span.__init__.tag_map in line with UD Treebank.--init-tok2vec argument to train_textcat.py example.Raise error if user is running a narrow unicode build.
ud_train, ud_evaluate and other UD scripts from CLI to /bin in repo only.spacy pretrain by implementing cosine loss.Doc.vector and Doc.vector_norm work as expected on GPU.spacy.cli.Thanks to @mhham and @Bharat123Rox for the pull requests!
Nothing published for this version
This replaces the spacy vocab command, which is now deprecated.
⚠️ This version of spaCy requires downloading new models. You can use the
spacy validatecommand to find out which models need updating, and print update instructions. If you've been training your own models, you'll need to retrain them with the new version.
spacy pretrain command. This pre-trains the CNN using BERT's cloze task. A new trick we're calling Language Modelling with Approximate Outputs is used to apply the pre-training to smaller models. The pre-training outputs CNN and embedding weights that can be used in spacy train, using the new -t2v argument.TextCategorizer, and allow setting exclusive_classes and architecture arguments on initialization.EntityRecognizer.labels property.blis kernel for matrix multiplication. Parallelisation should be performed at the task level, e.g. by running more containers.French by ~30%.Vocab.writing_system (populated via the language data) to expose settings like writing direction.pretrain command for ULMFit/BERT/Elmo-like pretraining (see #2931).ud-train command, to train and evaluate using the CoNLL 2017 shared task data.spacy download.download command to pip to customise installation.train command by letting GoldCorpus stream data, instead of loading into memory.init-model command, including support for lexical attributes and word-vectors, using a variety of formats. This replaces the spacy vocab command, which is now deprecated.train command.train command.Matcher (see #1971).Doc.retokenize context manager for merging and splitting tokens more efficiently.PhraseMatcher to match on token attributes other than ORTH, e.g. LOWER (for case-insensitive matching) or even POS or TAG.ujson, msgpack, msgpack-numpy, pickle, cloudpickle and dill with our own package srsly to centralise dependencies and allow binary wheels.Doc.to_json() method which outputs data in spaCy's training format. This will be the only place where the format is hard-coded (see #2932).EntityRuler component to make it easier to build rule-based NER and combinations of statistical and rule-based systems.gold.spans_from_biluo_tags helper that returns Span objects, e.g. to overwrite the doc.ents..similarity method is called with empty vectors or without word vectors.Matcher and add return_matches keyword argument to Matcher.pipe to yield (doc, matches) tuples instead of only Doc objects, and as_tuples to add context to the Doc objects.Token.is_stop and Lexeme.is_stop case-insensitive."TEXT" as an alternative to "ORTH" in Matcher patterns.black for auto-formatting .py source and optimse codebase using flake8. You can now run flake8 spacy and it should return no errors or warnings. See CONTRIBUTING.md for details.Token.conjuncts.Doc.retokenize() context manager.Span.as_doc return a copy, not a view.regex with re and speed up tokenization.Animacy_inan and add Animacy_nhum.TextCategorizer.POS but not TAG.Language subclasses via entry points.it_core_news_sm model.relcl dependency label to symbols.Doc.tensor when merging spans.Matcher engine to support regex, extension attributes and rich comparison.Token.pos_ writeable.displacy support for RTL languages.TextCategorizer and GoldParse API docs.Doc.get_lca_matrix.Matcher's ? quantifier.KeyError in Vectors.most_similar.'sentencizer' as built-in sentence boundary component name.displacy NER visualization and correct API docs.NORM a Token attribute instead of a Lexeme attribute to allow setting context-specific norms in tokenizer exceptions.EntityRecognizer.add_label.like_num work with prefixed numbers.Token or Span are pickled.Retokenizer.split method to split one token into several.doc[0].is_sent_start == True.B, L or U.nlp in Japanese (MeCab).Doc.is_tagged in Doc.from_array.Span to take unicode value for label argument.Doc.to_array.vectors.name correctly when exporting model via CLI.Japanese.Token.subtree and Span.subtree.PhraseMatcher pickling and make __len__ consistent.Token.sent work as expected without the parser.Token.pos_.numpy directly for similarity.#egg fragments in direct downloads.Doc.from_array consistent with Doc.to_array.conllu converters.spacy validate command to find out which models need updating, and print update instructions.blis for faster platform-independent matrix multiplication, v2.1.x currently doesn't work on Python 2.7 on Windows. We expect this to be corrected in the future.Matcher API is fully backwards compatible, its algorithm has changed to fix a number of bugs and performance issues. This means that the Matcher in v2.1.x may produce different results compared to the Matcher in v2.0.x.Doc.merge and Span.merge methods still work, but you may notice that they now run slower when merging many objects in a row. That's because the merging engine was rewritten to be more reliable and to support more efficient merging in bulk. To take advantage of this, you should rewrite your logic to use the Doc.retokenize context manager and perform as many merges as possible together in the with block.- doc[1:5].merge()
- doc[6:8].merge()
+ with doc.retokenize() as retokenizer:
+ retokenizer.merge(doc[1:5])
+ retokenizer.merge(doc[6:8])
to_disk, from_disk, to_bytes and from_bytes now support a single exclude argument to provide a list of string names to exclude. The docs have been updated to list the available serialization fields for each class. The disable argument on the Language serialization methods has been renamed to exclude for consistency.- nlp.to_disk("/path", disable=["parser", "ner"])
+ nlp.to_disk("/path", exclude=["parser", "ner"])
- data = nlp.tokenizer.to_bytes(vocab=False)
+ data = nlp.tokenizer.to_bytes(exclude=["vocab"])
.pos value for several common English words has changed, due to corrections to long-standing mistakes in the English tag map (see #593, #3311).n_threads on the .pipe methods is now deprecated, as the v2.x models cannot release the global interpreter lock. (Future versions may introduce a n_process argument for parallel inference via multiprocessing.)Doc.print_tree method is not deprecated in favour of a unified Doc.to_json method, which outputs data in the same format as the expected JSON training data.'sentencizer' – the name 'sbd' is deprecated.- sentence_splitter = nlp.create_pipe('sbd')
+ sentence_splitter = nlp.create_pipe('sentencizer')
is_sent_start attribute of the first token in a Doc now correctly defaults to True. It previously defaulted to None.spacy train command now lets you specify a comma-separated list of pipeline component names, instead of separate flags like --no-parser to disable components. This is more flexible and also handles custom components out-of-the-box.- $ spacy train en /output train_data.json dev_data.json --no-parser
+ $ spacy train en /output train_data.json dev_data.json --pipeline tagger,ner
spacy init-model command now uses a --jsonl-loc argument to pass in a a newline-delimited JSON (JSONL) file containing one lexical entry per line instead of a separate --freqs-loc and --clusters-loc.- $ spacy init-model en ./model --freqs-loc ./freqs.txt --clusters-loc ./clusters.txt
+ $ spacy init-model en ./model --jsonl-loc ./vocab.jsonl
it_core_news_sm is now correctly licensed under CC BY-NC-SA 3.0, and all English and German models are now published under the MIT license.| Model | Language | Version | UAS | LAS | POS | NER F | Vec | Size |
|---|---|---|---|---|---|---|---|---|
en_core_web_sm |
English | 2.1.0 | 91.5 | 89.7 | 96.8 | 85.9 | 𐄂 | 10 MB |
en_core_web_md |
English | 2.1.0 | 91.8 | 90.0 | 96.9 | 86.6 | ✓ | 90 MB |
en_core_web_lg |
English | 2.1.0 | 91.8 | 90.1 | 97.0 | 86.6 | ✓ | 788 MB |
de_core_news_sm |
German | 2.1.0 | 90.7 | 88.6 | 96.3 | 83.1 | 𐄂 | 10 MB |
de_core_news_md |
German | 2.1.0 | 91.2 | 89.4 | 96.6 | 83.8 | ✓ | 210 MB |
es_core_news_sm |
Spanish | 2.1.0 | 90.4 | 87.3 | 96.9 | 89.5 | 𐄂 | 10 MB |
es_core_news_md |
Spanish | 2.1.0 | 91.0 | 88.2 | 97.2 | 89.7 | ✓ | 69 MB |
pt_core_news_sm |
Portuguese | 2.1.0 | 89.1 | 85.9 | 80.4 | 88.9 | 𐄂 | 12 MB |
fr_core_news_sm |
French | 2.1.0 | 87.6 | 84.7 | 94.5 | 82.6 | 𐄂 | 14 MB |
fr_core_news_md |
French | 2.1.0 | 89.1 | 86.4 | 95.3 | 83.1 | ✓ | 82 MB |
it_core_news_sm |
Italian | 2.1.0 | 91.0 | 87.3 | 95.8 | 86.1 | 𐄂 | 10 MB |
nl_core_news_sm |
Dutch | 2.1.0 | 83.7 | 77.6 | 91.6 | 87.0 | 𐄂 | 10 MB |
el_core_news_sm |
Greek | 2.1.0 | 84.4 | 80.6 | 94.6 | 71.6 | 𐄂 | 10 MB |
el_core_news_md |
Greek | 2.1.0 | 88.3 | 85.0 | 96.6 | 81.1 | ✓ | 126 MB |
xx_ent_wiki_sm |
Multi | 2.1.0 | - | - | - | 81.3 | 𐄂 | 3 MB |
💬 UAS: Unlabelled dependencies (parser). LAS: Labelled dependencies (parser). POS: Part-of-speech tags (fine-grained tags, i.e.
Token.tag_). NER F: Named entities (F-score). Vec: Model contains word vectors. Size: Model file size (zipped archive).
Although it looks pretty much the same, we've rebuilt the entire documentation using Gatsby and MDX. It's now an even faster progressive web app and allows us to write all content entirely in Markdown, without having to compromise on easy-to-use custom UI components. We're hoping that the Markdown source will make it even easier to contribute to the documentation. For more details, check out the styleguide and source.
While converting the pages to Markdown, we've also fixed a bunch of typos, improved the existing pages and added some new content:
Matcher, PhraseMatcher and the new EntityRuler, and write powerful components to combine statistical models and rules.Doc using the new retokenize context manager and merge spans into single tokens and split single tokens into multiple.EntityRulerSentenceSegmenterThanks to @DuyguA, @giannisdaras, @mgogoulos, @louridas, @skrcode, @gavrieltal, @svlandeg, @jarib, @alvaroabascar, @kbulygin, @moreymat, @mirfan899, @ozcankasal, @willprice, @alvations, @amperinet, @retnuh, @Loghijiaha, @DeNeutoy, @gavrieltal, @boena, @BramVanroy, @pganssle, @foufaster, @adrianeboyd, @maknotavailable, @pierremonico, @lauraBaakman, @juliamakogon, @Gizzio, @Abhijit-2592, @akki2825, @grivaz, @roshni-b, @mpuig, @mikelibg, @danielkingai2, @adrienball and @Poluglottos for the pull requests and contributions.
NEW: Alpha tokenization support for Catalan.
regex pin to harmonise dependencies with conda.msgpack pin.pytest 4.0.is_ascii documentation.Vocab.prune_vectors did not use batch_size.Span.ents was added.msgpack pin.Thanks to @mpuig, @ALSchwalm, @bpben, @svlandeg and @wxv for the pull requests and contributions.
Nothing published for this version
Nothing published for this version
Make max_length of input text inclusive.
max_length of input text inclusive.doc.ents.displacy arcs would receive the same IDs in Jupyter notebooks, causing weirdly positioned arc labels.'\n'.Thanks to @digest0r, @BramVanroy, @grivaz, @wannaphongcom, @mikelibg, @danielhers, @frascuchon, @mauryaland and @cicorias for the pull requests and contributions.
Nothing published for this version
Nothing published for this version
Fix msgpack-numpy pin, which could affect serialization on Python 2.7.
msgpack-numpy pin, which could affect serialization on Python 2.7.Nothing published for this version
Improve version compatibility to support wheels for all spaCy dependencies maintained by us: `thinc`, `cymem`, `preshed` and `murmurhash`.
thinc, cymem, preshed and murmurhash.spacy[cuda], spacy[cuda90], spacy[cuda91], spacy[cuda92] or spacy[cuda10], which will install cupy and thinc_gpu_ops.spacy.prefer_gpu() and spacy.require_gpu() functions.Nothing published for this version
Nothing published for this version
NEW: Pre-built wheels and up to 10 times faster installation! This release starts the journey towards pre-built wheels for all of spaCy's dependencies
explosion/wheelwright.Span.ents property for consistency with Doc.ents.--verbose option to spacy train to output more details for debugging.numpy warning.FAC to spacy.explain glossary.getoption() in conftest.py.Thanks to @DimaBryuhanov, @kororo, @AndriyMulyar, @katarkor, @giannisdaras, @bphi, @vikaskyadav, @sammous, @EmilStenstrom, @howl-anderson, @ohenrik, @aashishg, @aryaprabhudesai, @steve-prod, @njsmith, @aniruddha-adhikary, @pzelasko, @mbkupfer, @sainathadapa, @tyburam, @grivaz, @filipecaixeta, @aongko, @free-variation, @mauryaland, @pmj642, @keshan, @darindf, @charlax, @phojnacki, @skrcode, @jacopofar, @Cinnamy and @JKhakpour for the pull requests and contributions!
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →