NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #829 most downloaded on PyPI
Industrial-strength Natural Language Processing (NLP) in Python
Last release 1 months ago
24 Aug 2026
Release timing varies
gaps range from 1 weeks to 6 months
Nearly every release is documented
notes for 55 of the last 60 stable releases
9 versions withdrawn
withdrawn after publishing
12 years old
218 releases · first in 2015
Nothing published for this version
Nothing published for this version
Nothing published for this version
Fix issue #2226: Use correct, non-deprecated merge syntax in merge_ents.
One column per quarter.
We had to release another update to the v2.0.x branch of spaCy to resolve a dependency issue, so we decided to also include and/or backport a bunch of features and fixes that were originally intended for v2.1.0 (see here for the nightly version).
Token.sent property that returns the sentence Span the token is part of.remove_extension method on Doc, Token and Span.Doc.is_sentenced property that returns True if sentence boundaries have been applied.SPACY_WARNING_IGNORE environment variable.--silent option to info command.download command to pip and check if model is already installed before downloading it.README section on tests and dependencies.Doc.noun_chunks_iterator isn't None before calling it.cli.info and add silent option.spacy package command message.download command and use requests library again.merge_ents.tensor=False or sentiment=False.remove_extension method on Doc, Token and Span.collapse_phrases option to displaCy visualizer.KeyError by renaming SP to _SP.attrs argument in Doc.retokenize and allow ints/unicode.displacy.render.Matcher examples and add section on using pipeline components.displacy arrows to only point in one direction.Span objects in displacy.render.msgpack-python to msgpack to hopefully prevent conda from downloading a two-year-old spaCy version when installing with latest the Anaconda distribution.Doc.is_tagged is set correctly when using Language.pipe.merge_noun_chunks factory that would return None if Doc wasn't parsed.pathlib backport on Python 2 only.Matcher Explorer demo – create token patterns interactively, test them against your text and copy-paste the Python pattern code.Thanks to @mollerhoj, @howl-anderson, @pktippa, @skrcode, @miroli, @ivyleavedtoadflax, @5hirish, @therealronnie, @alexvy86, @mn3mos, @polm, @knoxdw, @bellabie, @mauryaland, @LRAbbade, @janimo, @vishnumenon, @tzano, @cclauss, @armsp, @aristorinjuang, @BigstickCarpet, @idealley, @ansgar-t, @mpszumowski, @91ns, @msklvsk, @himkt, @DanielRuf, @nathanathan, @GolanLevy, @nipunsadvilkar, @cjhurst, @aliiae, @mirfan899, @ohenrik, @btrungchi, @kleinay, @DuyguA, @stefan-it, @Eleni170, @datascouting, @tjkemp, @x-ji, @giannisdaras, @kororo and @katarkor for the pull requests and contributions.
Nothing published for this version
Nothing published for this version
NEW: Alpha Vietnamese support with tokenization via Pyvi.
asserts have been replaced with descriptive errors. See #2163 for implementation details, and let us know if you have any suggestions for errors and warnings in #2164!six, html5lib, ftfy and requests.beam_update_prob entry in nlp.parser.cfg. The default value is 0.5, so 50% of beam updates will be done as greedy updates.Token.ent_iob after Doc.merge(), and ensure consistency in Doc.ents.Vocab.set_vector.Token.set_extension work as expected.six and html5lib and prevent dependency conflict with TensorFlow / Keras.Language.to_bytes and pickling in Thinc.Matcher docs.set_extension if getter and setter are specified or if default=None, and add error if setter is specified with no getter.CITATION for assigning a DOI via Zenodo.Thanks to @jimregan, @justindujardin, @trungtv, @katrinleinweber and @skrcode for the pull requests and contributions.
Nothing published for this version
Improve language data for Turkish and Croatian.
merge_entities and merge_noun_chunks to allow models to specify those components as part of their pipeline.merge_entities = nlp.create_pipe('merge_entities')
nlp.add_pipe(merge_entities, after='ner')
noun_chunks failure caused by typo.Token.lemma always returns a hash value.Vectors class.Thanks to @thomasopsomer, @alldefector, @DuyguA, @dejanmarich, @justindujardin, @calumcalder, @SebastinSanty, @iann0036, @doug-descombaz and @willismonroe for the pull requests and contributions.
Nothing published for this version
Fix issue #2015: Pin msgpack-python to 0.5.4 to avoid conflict with new msgpack release.
msgpack-python to 0.5.4 to avoid conflict with new msgpack release.NEW: Lexical attribute IS_CURRENCY via Token.is_currency for currency symbols.
IS_CURRENCY via Token.is_currency for currency symbols.noun_chunks syntax iterator for Norwegian.get_beam_parse method in ArcEager.Matcher in favour of the new and improved API (#1971) coming in v2.1.0.from_disk are closed.init_model command.begin_training.html5lib in setup.py to prevent six error.Matcher docs to only include ORTH and LOWER.Matcher example.Language.pipe.random.seed globally in CLI commands.match_id and improve example.Thanks to @ohenrik, @tokestermw, @azarezade, @piratos, @mhaddy, @pktippa, @mdcclv, @oxinabox, @SThomasP, @DuyguA, @emulbreh, @ursachec and @enerrio for the pull requests and contributions.
Fix issue #1919: Fix missing config property in parser when resuming training.
Fix issue #1712, #1813: Don't raise deprecation warning in property.
Matcher bugs and behaviour of * and + operators.Vectors.resize on Python 3.march=native.LIKE_URL doesn't return True for email addresses.nlp.to_disk in spacy train command.Span.vocab property.token.sent_start.--no-deps to avoid conda errors.download and validate commands exit correctly.pretrained_dims setting from cfg.TextCategorizer documentation.None objects.LIKE_NUM case-insensitive.Chinese language class.l_edge and r_edge correctly for non-projective parses.set_vector add word to vocab.Matcher operators.spacy train.Vocab.__contains__ work with ints.Matcher.spacy init_model command.Thanks to @cbilgili, @melanuria, @mpuels, @IsaacHaze, @sorenlind, @Bri-Will, @d99kris, @mdda, @kimfalk, @benjaminp, @zqhZY, @avinashrubird, @nirdesh37, @kwhumphreys, @fucking-signup, @wrathagom, @pbnsilva, @savkov, @matatusko, @GregDubbin, @avadhpatel, @azarezade, @ohenrik, @azarezade, @thomasopsomer, @Kimahriman and @hassanshamim for the pull requests and contributions.
Nothing published for this version
Add spacy init-model command to create a model directory from raw data (similar to the spacy model command in v1.x).
spacy init-model command to create a model directory from raw data (similar to the spacy model command in v1.x).None.Thanks to @mpuels and @GreenRiverRUS for the pull requests and contributions.
Nothing published for this version
Alpha support for Russian via pymorphy2.
offsets_from_biluo_tags helper to convert BILUO notation to entity offsets.POS instead of TAG by default in displaCy, to prevent visualisation issues in languages with long combined tags (e.g. Italian or Dutch).Span.noun_chunks.Span.orth_ == Span.text.entity_relations.py example Python 2 compatible and fix French test.nlp.add_pipe when using after.spacy package.StringStore cleanup.python -m spacy for CLI commands again to prevent issues on Windows etc.Thanks to @ligser, @pavillet, @yuukos, @GreenRiverRUS, @MartinoMensio, @raphael0202, @tokestermw, @fsonntag, @cclauss, @bdewilde, @markulrich, @sorenlind, @hugovk, @atomobianco, @twerkmeister, @mkdynamic and @jimregan for the pull requests and contributions.
Nothing published for this version
Require Thinc v6.10.1 to fix GPU installation fix and beam parsing.
v6.10.1 to fix GPU installation fix and beam parsing.KeyError from cleaning up strings during Language.pipe (work in progress).Doc.to_disk and Doc.from_disk.util.minibatch work correctly.Tokenizer (partially addresses performance regression in #1371 and #1508).python -m to CLI commands to ensure cross-platform compatibility.Thanks to @MathiasDesch, @mcsalgado, @Wahib, @ligser, @abhi18av, @DuyguA, @KMLDS and @yogendrasoni for the pull requests and contributions.
Nothing published for this version
Fix issue #1507, #1512, #1513, #1514, #1516: Improve new documentation and list of backwards incompatibilities.
train_textcat.py example.Vectors.resize work as expected.Thanks to @danielhers and @abhi18av for the pull requests.
Nothing published for this version
Fix syntax error in language data examples that prevented conda build.
Nothing published for this version
Fix issue #1243: Resolve undefined names in deprecated functions.
We're very excited to finally introduce spaCy v2.0. The new version gets spaCy up to date with the latest deep learning technologies and makes it much easier to run spaCy in scalable cloud computing workflows. We've fixed over 60 bugs (every open bug!), including several long-standing issues, trained 13 neural network models for 7+ languages and added alpha tokenization support for 8 new languages. We also re-wrote almost all of the usage guides, API docs and code examples.
pip install -U spacy
conda install -c conda-forge spacy
Vectors class for managing word vectors, plus trainable document vectors and contextual similarity via convolutional neural networks.Doc, Token and Span via Doc._, Token._ and Span._.MultiLanguage class (xx).PhraseMatcher for matching large terminology lists as Doc objects, plus revised Matcher API.validate, vocab and evaluate, plus entry point for spacy command to use instead of python -m spacy.spaCy v2.0 comes with 13 new convolutional neural network models for 7+ languages. The models have been designed and implemented from scratch specifically for spaCy. A novel bloom embedding strategy with subword features is used to support huge vocabularies in tiny tables.
All core models include part-of-speech tags, dependency labels and named entities. Small models include only context-specific token vectors, while medium-sized and large models ship with word vectors. For more details, see the models directory or try our new model comparison tool.
| Name | Language | Features | Size |
|---|---|---|---|
en_core_web_sm |
English | Tagger, parser, entities | 35 MB |
en_core_web_md |
English | Tagger, parser, entities, vectors | 115 MB |
en_core_web_lg |
English | Tagger, parser, entities, vectors | 812 MB |
en_vectors_web_lg |
English | Vectors | 627 MB |
de_core_news_sm |
German | Tagger, parser, entities | 36 MB |
es_core_news_sm |
Spanish | Tagger, parser, entities | 35 MB |
es_core_news_md |
Spanish | Tagger, parser, entities, vectors | 93 MB |
pt_core_news_sm |
Portuguese | Tagger, parser, entities | 36 MB |
fr_core_news_sm |
French | Tagger, parser, entities | 37 MB |
fr_core_news_md |
French | Tagger, parser, entities, vectors | 106 MB |
it_core_news_sm |
Italian | Tagger, parser, entities | 34 MB |
nl_core_news_sm |
Dutch | Tagger, parser, entities | 34 MB |
xx_ent_wiki_sm |
Multi-language | Entities | 33MB |
You can download a model by using its name or shortcut. To load a model, use spacy.load(), or import it as a module and call its load() method:
spacy download en_core_web_sm
import spacy
nlp = spacy.load('en_core_web_sm')
import en_core_web_sm
nlp = en_core_web_sm.load()
spaCy v2.0's new neural network models bring significant improvements in accuracy, especially for English Named Entity Recognition. The new en_core_web_lg model makes about 25% fewer mistakes than the corresponding v1.x model and is within 1% of the current state-of-the-art (Strubell et al., 2017). The v2.0 models are also cheaper to run at scale, as they require under 1 GB of memory per process.
| Model | spaCy | Type | UAS | LAS | NER F | POS | Size |
|---|---|---|---|---|---|---|---|
en_core_web_sm-2.0.0 |
v2.x | neural | 91.7 | 89.8 | 85.3 | 97.0 | 35MB |
en_core_web_md-2.0.0 |
v2.x | neural | 91.7 | 89.8 | 85.9 | 97.1 | 115MB |
en_core_web_lg-2.0.0 |
v2.x | neural | 91.9 | 90.1 | 85.9 | 97.2 | 812MB |
en_core_web_sm-1.1.0 |
v1.x | linear | 86.6 | 83.8 | 78.5 | 96.6 | 50MB |
en_core_web_md-1.2.1 |
v1.x | linear | 90.6 | 88.5 | 81.4 | 96.7 | 1GB |
| Model | spaCy | Type | UAS | LAS | NER F | POS | Size |
|---|---|---|---|---|---|---|---|
es_core_news_sm-2.0.0 |
v2.x | neural | 89.8 | 86.8 | 88.7 | 96.9 | 35MB |
es_core_news_md-2.0.0 |
v2.x | neural | 90.2 | 87.2 | 89.0 | 97.8 | 93MB |
es_core_web_md-1.1.0 |
v1.x | linear | 87.5 | n/a | 94.2 | 96.7 | 377MB |
For more details of the other models, see the models directory and model comparison tool.
Doc objects.ROOT objects.SP tag.Doc, Token and Span.train command finally works correctly if used without dev_data.Spans without keyword arguments.== works as expected on tokens.Token.nbor raises IndexError correctly."*" ends the match pattern.For the complete table and more details, see the guide on what's new in v2.0.
Note that the old v1.x models are not compatible with spaCy v2.0.0. If you've trained your own models, you'll have to re-train them to be able to use them with the new version. For a full overview of changes in v2.0, see the documentation and guide on migrating from spaCy 1.x.
The Language.pipe method allows spaCy to batch documents, which brings a significant performance advantage in v2.0. The new neural networks introduce some overhead per batch, so if you're processing a number of documents in a row, you should use nlp.pipe and process the texts as a stream.
docs = nlp.pipe(texts)
# BAD: docs = (nlp(text) for text in texts)
To make usage easier, there's now a boolean as_tuples keyword argument, that lets you pass in an iterator of (text, context) pairs, so you can get back an iterator of (doc, context) tuples.
spacy.load() is now only intended for loading models – if you need an empty language class, import it directly instead, e.g. from spacy.lang.en import English. If the model you're loading is a shortcut link or package name, spaCy will expect it to be a model package, import it and call its load() method. If you supply a path, spaCy will expect it to be a model data directory and use the meta.json to initialise a language class and call nlp.from_disk() with the data path.
nlp = spacy.load('en')
nlp = spacy.load('en_core_web_sm')
nlp = spacy.load('/model-data')
nlp = English().from.disk('/model-data')
# OLD: nlp = spacy.load('en', path='/model-data')
All built-in pipeline components are now subclasses of Pipe, fully trainable and serializable, and follow the same API. Instead of updating the model and telling spaCy when to stop, you can now explicitly call begin_training, which returns an optimizer you can pass into the update function. While update still accepts sequences of Doc and GoldParse objects, you can now also pass in a list of strings and dictionaries describing the annotations. This is the recommended usage, as it removes one layer of abstraction from the training.
optimizer = nlp.begin_training()
for itn in range(1000):
for texts, annotations in train_data:
nlp.update(texts, annotations, sgd=optimizer)
nlp.to_disk('/model')
spaCy's serialization API is now consistent across objects. All containers and pipeline components have .to_disk(), .from_disk(), .to_bytes() and .from_bytes() methods.
nlp.to_disk('/model')
nlp.vocab.to_disk('/vocab')
# OLD: nlp.save_to_directory('/model')
Models can now define their own processing pipelines as a list of strings, mapping to component names. Components receive a Doc, modify it and return it to be processed by the next component in the pipeline. You can add custom components to nlp.pipeline and create extensions to add custom attributes, properties and methods to the Doc, Token and Span objects.
nlp = spacy.load('en')
my_component = MyComponent()
nlp.add_pipe(my_component, before='tagger')
Doc.set_extension('my_attr', default=True)
doc = nlp(u"This is a text.")
assert doc._.my_attr
This release is brought to you by @honnibal and @ines. Thanks to @Gregory-Howard, @luvogels, @Ferdous-Al-Imran, @uetchy, @akYoung, @kengz, @raphael0202, @ardeego, @yuvalpinter, @dvsrepo, @frascuchon, @oroszgy, @v3t3a, @Tpt, @thinline72, @jarle, @jimregan, @nkruglikov, @delirious-lettuce, @geovedi, @wannaphongcom, @h4iku, @IamJeffG, @binishkaspar, @ramananbalakrishnan, @jerbob92, @mayukh18, @abhi18av and @uwol for the pull requests and contributions. Also thanks to everyone who submitted bug reports and took the spaCy user survey – your feedback made a big difference!
Fix issue #2112: Avoid import pip to ensure compatibility with pip v9.0.2 which deprecated this usage. See pypa/pip#5081 for more details.
import pip to ensure compatibility with pip v9.0.2 which deprecated this usage. See pypa/pip#5081 for more details.Thanks to @mdcclv for the pull request!
> ⚠️ Important note: This is a bridge release that gets the current state of the v1.x branch published. Stay tuned for v2.0.
⚠️ Important note: This is a bridge release that gets the current state of the v1.x branch published. Stay tuned for v2.0.
Doc.get_lca_matrix().Doc.to_scalar and Doc.to_array.Thanks to @raphael0202, @gideonite, @delirious-lettuce, @polm, @kevinmarsh, @IamJeffG, @Vimos, @ericzhao28, @galaxyh, @hscspring, @wannaphongcom, @Wellan89, @kokes, @mdcclv, @ameyuuno, @ramananbalakrishnan, @Demfier, @johnhaley81, @mayukh18 and @jnothman for the pull requests and contributions.
Thanks to all of you for 5,000 stars on GitHub, the valuable feedback in the user survey and testing [spaCy v2.0 alpha](https://alpha.spacy.io/docs/us
Thanks to all of you for 5,000 stars on GitHub, the valuable feedback in the user survey and testing spaCy v2.0 alpha. We're working hard on getting the new version ready and can't wait to release it. In the meantime, here's a new release for the 1.x branch that fixes a variety of outstanding bugs and adds capabilities for new languages.
💌 P.S.: If you haven't gotten your hands on a set of spaCy stickers yet, you can still do so – send us a DM with your address on Twitter or Gitter, and we'll mail you some!
python -m spacy download es
nlp = spacy.load('es')
doc = nlp(u'Esto es una frase.')
Parser and EntityRecognizer, using the drop keyword argument to the update() method.spacy.explain(). For example, spacy.explain('NORP') will return "Nationalities or religious or political groups".Language.parse_tree method to generate POS tree for all sentences in a Doc.Lexeme API.spacy.explain().SP symbol to tag map.flush_cache method to tokenizer.Doc.sents iterator when customised with generator.requirements.txt.requests dependency.Span.noun_chunks.six and its dependencies that occasionally caused spaCy to fail.package command that caused error when printing error messages.Thanks to @kengz, @luvogels, @Ferdous-Al-Imran, @uetchy, @akYoung, @pasupulaphani, @dvsrepo, @raphael0202, @yuvalpinter, @frascuchon, @kootenpv, @oroszgy, @bartbroere, @ianmobbs, @garfieldnate, @polm, @callumkift, @swierh, @val314159, @lgenerknol and @jsparedes for the contributions!
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incred
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incredibly helpful in improving the library. As we're getting closer to v2.0 we hope you'll take a few minutes to fill out the survey, to help us understand how you're using the library, and how it can be better.
shortcuts.json to allow adding new ones without updating spaCy.python -m spacy download fr_depvec_web_lg
import fr_depvec_web_lg
nlp = fr_depvec_web_lg.load()
doc = nlp(u'Parlez-vous Français?')
train command is used without dev_data.Span hashable.Thanks to @raphael0202 and @julien-c for the contributions!
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incred
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incredibly helpful in improving the library. As we're getting closer to v2.0 we hope you'll take a few minutes to fill out the survey, to help us understand how you're using the library, and how it can be better.
convert command now uses Python 2/3 compatible json.dumps.regex library for non-latin characters to simplify punctuation rules.SPACE to Spanish tag map.train command now works correctly if used without dev_data.Language.save_to_directory() now converts strings to pathlib paths.pos_tag.py examples.Language.save_to_directory() method to API docs.Thanks to @dvsrepo, @beneyal and @oroszgy for the pull requests!
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incred
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incredibly helpful in improving the library. As we're getting closer to v2.0 we hope you'll take a few minutes to fill out the survey, to help us understand how you're using the library, and how it can be better.
Language.save_to_directory() method to make it easier to save user-trained models.spacy.compat module to handle platform and Python version compatibility.package command to read from existing meta.json and supply custom location to meta file.spacy.cli.spacy.load() now prints warning if no model is found.token.lemma and token.lemma_ attributes writeable.spacy.compat to handle compatibility.Thanks to @tsohil and @oroszgy for the pull requests!
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incred
We've been delighted to see spaCy growing so much over the last few months. Before the v1.0 release, we asked for your feedback, which has been incredibly helpful in improving the library. As we're getting closer to v2.0 we hope you'll take a few minutes to fill out the survey, to help us understand how you're using the library, and how it can be better.
convert and model commands to convert files to spaCy's JSON format for training, and initialise a new model and its data directory.package command now works correctly and doesn't fail when creating files.EntityRecognizer transition bug.label keyword argument is now handled correctly in doc.merge()./ infixes are now split by the tokenizer.invalid switch error.setup.py and __init__.py.regex module to avoid back-tracking on URL regex.convert and model commands.--no-cache-dir error resulting from outdated pip version and file name shadowing model problem.Thanks to @ericzhao28, @Gregory-Howard, @kinow, @jreeter, @mamoit, @kumaranvpl and @dvsrepo for the pull requests!
NEW: Alpha tokenization for Hebrew.
train and package commands to train a model and convert it to a Python package.download command.mlink to create symlinks in Python 2 on Windows.--no-cache-dir when downloading models via pip.spacy.info.package and train commands.Thanks to @raphael0202, @pavlin99th, @iddoberger and @solresol for the pull requests!
Success message in link is now displayed correctly when using local paths.
link is now displayed correctly when using local paths.beam_parser.Fix issue #892: Data now downloads and installs correctly on system Python.
This will be the last major release before v2.0, which will introduce a few breaking changes to allow native deep learning integration. If you're usin…
To increase transparency and make it easier to use spaCy with your own models, all data is now available as direct downloads, organised in individual releases. spaCy v1.7 also supports installing and loading models as Python packages. You can now choose how and where you want to keep the data files, and set up "shortcut links" to load models by name from within spaCy. For more info on this, see the new models documentation.
# out-of-the-box: download best-matching default model
python -m spacy download en
# download best-matching version of specific model for your spaCy installation
python -m spacy download en_core_web_md
# pip install .tar.gz archive from path or URL
pip install /Users/you/en_core_web_md-1.2.0.tar.gz
pip install https://github.com/explosion/spacy-models/releases/download/en_core_web_md-1.2.0/en_core_web_md-1.2.0.tar.gz
# set up shortcut link to load installed package as "en_default"
python -m spacy link en_core_web_md en_default
# set up shortcut link to load local model as "my_amazing_model"
python -m spacy link /Users/you/data my_amazing_model
nlp1 = spacy.load('en')
nlp2 = spacy.load('en_core_web_md')
nlp3 = spacy.load('my_amazing_model')
en_core_web_md v1.2.0). The German model is still valid and will be linked to the de shortcut automatically.sputnik is now deprecated. For now, we will keep maintaining our download server to support the python -m spacy.{en|de}.download all command in older versions, but it will soon re-route to download the models from GitHub instead.spacy/en and the WordNet data previously stored in corpora/en has been removed. This should not affect your code, unless you have added functionality that relies on these data files.This will be the last major release before v2.0, which will introduce a few breaking changes to allow native deep learning integration. If you're using spaCy in production, don't forget to pin your dependencies:
# requirements.txt
spacy>=1.7.0,<2.0.0
# setup.py
install_requires=['spacy>=1.7.0,<2.0.0']
's are now lemmatized correctly.Span class now has lower_ and upper_ properties.-PRON-.en_core_web_sm) is now available.--force.NUM_WORDS are now recognised correctly as like_number.matcher and make sure open patterns are closed at doc end.ujson and plac are now specified correctly.conda-forge.cythonize.load_vectors() now accepts arbitrary space characters as word tokens.token_match regex.token.idx now matches original index when text contains newlines.spacy-models, including the latest model releases.help wanted (easy) issue label for contribution requests suitable for beginners.This release is brought to you by @honnibal and @ines. Thanks to @magnusburton, @jktong, @JasonKessler, @sudowork, @oiwah, @raphael0202, @latkins, @ematvey, @Tpt, @wallinm1, @knub, @wehlutyk, @vaulttech, @nycmonkey, @jondoughty, @aniruddha-adhikary, @badbye, @shuvanon, @rappdw, @ericzhao28, @juanmirocks and @rominf for the pull requests!
Updated token exception handling mechanism to allow the usage of arbitrary functions as token exception matchers.
richcmp method to Token.She are now handled correctly.Token is now hashable.were and Were are now excluded correctly from contractions.Doc object manually.Thanks to @oroszgy, @magnusburton, @guyrosin and @danielhers and for the pull requests!
Nothing published for this version
NEW: Alpha support for Swedish tokenization.
language_data package in the setup.py.vec_path declaration that was failing if add_vectors was set.Vocab to load without serializer_freqs.Thanks to @oroszgy, @magnusburton, @jmizgajski, @aikramer2, @fnorf and @bhargavvader for the pull requests!
NEW: Alpha support for Dutch tokenization.
token.ent_iob_ return unicode.TOKENIZER_EXCEPTIONS.Morphology class now supplies tag map value for the special space tag if it's missing.spacy.en.English() loads the Glove vector data if available. Previously was inconsistent with behaviour of spacy.load('en').TOKENIZER_EXCEPTIONS with unicode apostrophe (’).STOP_WORDS.No changes to the public, documented API, but the previously undocumented language data and model initialisation processes have been refactored and reorganised. If you were relying on the bin/init_model.py script, see the new spaCy Developer Resources repo. Code that references internals of the spacy.en or spacy.de packages should also be reviewed before updating to this version.
Thanks to @dafnevk, @jvdzwaan, @RvanNieuwpoort, @wrvhage, @jaspb, @savvopoulos and @davedwards for the pull requests!
Fix issue #605: accept argument to Matcher now rejects matches as expected.
Span.sentiment attribute.Span.noun_chunks iterator (thanks @pokey).--data-path be specified when running download.py scripts (thanks @ExplodingCabbage).PhraseMatcher to work with new Matcher (thanks @sadovnychyi).accept argument to Matcher now rejects matches as expected.Vocab.load() now works with string paths, as well as Path objects.Language class now used as expected.Tokenizer special-case rules now support arbitrary token attributes.Thanks to @pokey, @ExplodingCabbage, @souravsingh, @sadovnychyi, @manojsakhwar, @TiagoMRodrigues, @savkov, @pspiegelhalter, @chenb67, @kylepjohnson, @YanhaoYang, @tjrileywisc, @dechov, @wjt, @jsmootiv and @blarghmatey for the pull requests!
NEW: Support Chinese tokenization, via Jieba.
--force argument on download command now operates correctly.Matcher now rejects empty patterns.token.tag and token.tag_ setters.Matcher to sometimes segfault.Nothing published for this version
Nothing published for this version
Nothing published for this version
Rename new pipeline keyword argument of spacy.load() to create_pipeline.
pipeline keyword argument of spacy.load() to create_pipeline.vectors keyword argument of spacy.load() to add_vectors.vocab.resize_vectors() method, to support changing to vectors of different dimensionality.ent_iob attribute was incorrect after setting entities via doc.entsNothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →