NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4698 most downloaded on PyPI
John Snow Labs Spark NLP is a natural language processing library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines, that scale easily in a distributed environment.
Last release 11 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
196 releases · first in 2018
Thank you again for all your feedback and questions in our Slack channel. Such feedback from users and contributors (thank you Stuart Lynn @sllynn) he
Thank you again for all your feedback and questions in our Slack channel. Such feedback from users and contributors (thank you Stuart Lynn @sllynn) helped to find several python module bugs. We also fixed and improved OCR support towards extracting page coordinates and fixed NerDL evaluator from Python
One column per quarter.
This short release is to address a few uncovered issues in the previous 2.2.0 release. Thank you all for quick feedback.
This short release is to address a few uncovered issues in the previous 2.2.0 release. Thank you all for quick feedback.
Last time, following a release candidate schedule proved to be a quite effective method to avoid silly bugs right after release! Fortunately, there we
Last time, following a release candidate schedule proved to be a quite effective method to avoid silly bugs right after release! Fortunately, there were no breaking bugs by carefully testing releases alongside the community, which ended up in various pull requests. This huge release features OCR based coordinate highlighting, BERT embeddings refactor and tuning, more tools for accuracy evaluation in python, and much more. We welcome your feedback in our Slack channels, as always!
withCoverageColumn and overallCoverage offer metric analysisincludeConfidence param that enables confidence scores on prediction metadataenableOutputLog outputs training metric logs to filepoolingLayer allows for polling layer selectionfullAnnotate to retrieve fully information of AnnotationsWe are so glad to present the release candidate of this new release. Last time, following a release candidate schedule proved to be a quite effective
We are so glad to present the release candidate of this new release. Last time, following a release candidate schedule proved to be a quite effective method to avoid silly bugs right after release! Fortunately, there were no breaking bugs by carefully testing releases alongside the community, which ended up in various pull requests. This huge release features OCR based coordinate highlighting, BERT embeddings refactor and tuning, more tools for accuracy evaluation in python, and much more. We welcome your feedback in our Slack channels, as always!
withCoverageColumn and overallCoverage offer metric analysisincludeConfidence param that enables confidence scores on prediction metadatapoolingLayer allows for polling layer selectionfullAnnotate to retrieve fully information of AnnotationsWe are so glad to present the release candidate of this new release. Last time, following a release candidate schedule proved to be a quite effective
We are so glad to present the release candidate of this new release. Last time, following a release candidate schedule proved to be a quite effective method to avoid silly bugs right after release! Fortunately, there were no breaking bugs by carefully testing releases alongside the community, which ended up in various pull requests. This huge release features OCR based coordinate highlighting, BERT embeddings refactor and tuning, more tools for accuracy evaluation in python, and much more. We welcome your feedback in our Slack channels, as always!
withCoverageColumn and overallCoverage offer metric analysisincludeConfidence param that enables confidence scores on prediction metadatapoolingLayer allows for polling layer selectionfullAnnotate to retrieve fully information of AnnotationsWe are so glad to present the first release candidate of this new release. Last time, following a release candidate schedule allowed us to move from 2
We are so glad to present the first release candidate of this new release. Last time, following a release candidate schedule allowed us to move from 2.1.0 straight to 2.2.0! Fortunately, there were no breaking bugs by carefully testing releases alongside the community, which ended up in various pull requests. This huge release features OCR based coordinate highlighting, BERT embeddings refactor and tuning, more tools for accuracy evaluation in python, and much more. We welcome your feedback in our Slack channels, as always!
withCoverageColumn and overallCoverage offer metric analysisThank you so much for your feedback on slack. This release is to extend life length of the 2.1.x release, with important bugfixes from upstream
Thank you so much for your feedback on slack. This release is to extend life length of the 2.1.x release, with important bugfixes from upstream
Thank you for following up with release candidates. This release is backwards breaking because two basic annotators have been redesigned. The tokenize
Thank you for following up with release candidates. This release is backwards breaking because two basic annotators have been redesigned.
The tokenizer now has easier to customize params and simplified exception management.
DocumentAssembler trimAndClearNewLiens was redesigned into a cleanupMode for further control over the cleanup process.
Tokenizer now supports pretrained models, meaning you'll be capable of accessing any of our language-based Tokenizers.
Another big introduction is the eval module. An optional Spark NLP sub-module that provides evaluation scripts, to
make it easier when looking to measure your own models are against a validation dataset, now using MLFlow.
Some work also began on metrics during training, starting now with the NerDLApproach.
Finally, we'll have Scaladocs ready for easy library reference.
Thank you for your feedback in our Slack channels.
Particular thanks to @csnardi for fixing a bug in one of the release candidates.
cleanupMode allows the user to decide what kind of cleanup to apply to sourcesetTrainValidationPropFixed Tokenizer missing pretrained() functions
Release candidate #2 for 2.1.0
pretrained() functionsThis is a pre-release for 2.1.0. The tokenizer has been revamped, and some of the DocumentAssembler defaults changed. For this reason, many pipelines
This is a pre-release for 2.1.0. The tokenizer has been revamped, and some of the DocumentAssembler defaults changed. For this reason, many pipelines and models may now change their accuracies and performance. Old tokenizer default rules will be translated in a new english specific pretrained Tokenizer. NerDLApproach will now report metrics if setTrainValidationProp has been set, as well as confidence scores reporting in spell checkers. DependencyParser output has been reviewed and fixed a bunch of other bugs in the embeddings scope. Please feedback and bugs, and remember, this is a pre-release, so not yet intended for production use. Join Slack!
setTrainValidationPropspark-nlp-eval evaluation model with multiple scripts that help users evaluate their models and pipelines. To be improved.This release fixes a bug in embeddingsRef param causing embeddings not to be loadable when setIncludeEmbeddings was set to false
This release fixes a bug in embeddingsRef param causing embeddings not to be loadable when setIncludeEmbeddings was set to false
This release fixes a few tiny but meaningful issues that prevent from new trained models having internal compatibility issues.
This release fixes a few tiny but meaningful issues that prevent from new trained models having internal compatibility issues.
This release addresses bugs related to cluster support, improving error messages and fixing various potential bugs depending on the cluster configurat
This release addresses bugs related to cluster support, improving error messages and fixing various potential bugs depending on the cluster configuration, such as Kryo Serialization or non default FS systems
Following after 2.0.5 release (read notes below), this release fixes a bug when disabling useContrib param in NerDLApproach on non-windows OS.
Following after 2.0.5 release (read notes below), this release fixes a bug when disabling useContrib param in NerDLApproach on non-windows OS.
Following the 2.0.5 (read notes below), this release fixes a bug when disabling contrib param in NerDLApproach on non-windows OS
========
This release bumps Spark NLP by default to Apache Spark 2.4.3. Spark has been undergoing testing with Scala 2.12 and they are back in 2.11 now, so thi
This release bumps Spark NLP by default to Apache Spark 2.4.3. Spark has been undergoing testing with Scala 2.12 and they are back in 2.11 now, so this should be a working release. In this version, we fixed a series of Pretrained models, as well as focused on improving the flexibility of NerDL annotator, which is, if not, the most popular one based on user feedback. Users can point to graphs they create without having to re-compile the library, graph options as well whether to use Tensorflow contrib is now user defined. Particular thanks to @CyborgDroid because of reporting important and well-reported bugs that helped us improve Spark NLP. Thank you for reporting issues and feedback, and we always welcome more. Join us on Slack!
We are excited about the Spark NLP workshop (spark-nlp-workshop repository) being so useful for many users. Now we also made a step forward by moving
We are excited about the Spark NLP workshop (spark-nlp-workshop repository) being so useful for many users. Now we also made a step forward by moving the website's documentation to an easy to maintain Jekyll template with Markdown. Spark NLP library received key bug fixes on this release. Thanks to the community for reporting issues on GitHub. Much more to come, as always.
We are excited about Spark NLP workshop (spark-nlp-workshop repository) being so useful for many users. Now we also made a step forward by moving website's documentation to an easy to maintain Wiki!. Spark NLP library received key bug fixes on this release. Thanks to the community for reporting issues on GitHub. Much more to come, as always.
========
Short after 2.0.2, a hotfix release was made to address two bugs that prevented users from using pretrained tensorflow models in clusters. Please read
Short after 2.0.2, a hotfix release was made to address two bugs that prevented users from using pretrained tensorflow models in clusters. Please read release notes for 2.0.2 to catch up!
Thank you for joining us in this exciting Spark NLP year!. We continue to make progress towards a better performing library, both in speed and in accu
Thank you for joining us in this exciting Spark NLP year!. We continue to make progress towards a better performing library, both in speed and in accuracy. This release focuses strongly in the quality and stability of the library, making sure it works well in most cluster environments and improving the compatibility across systems. Word Embeddings continue to be improved for better performance and lower memory blueprint. Context Spell Checker continues to receive enhancements in concurrency and usage of spark. Finally, tensorflow based annotators have been significantly improved by refactoring the serialization design. Help us with feedback and we'll welcome any issue reports!
Thanks for following up after our 2.0.0 release!. This release covers a few holes left by the immense 2.0.0 release, to address high priority issues f
Thanks for following up after our 2.0.0 release!. This release covers a few holes left by the immense 2.0.0 release, to address high priority issues found after release. More importantly, the library should now behave correctly when using Spark cluster modes, and memory and CPU utilization should be reduced to normal levels after some serious profiling of Serialization revealed a bunch of problems. Aside from performance and resource management improvements, we include an OCR dependency handler in start() function as well as improve the support of GPU for NER Deep Learning models. Finally, check out our spark-nlp-workshop repo, it has cool features!
Thank you for following up with the biggest changelog ever on Spark NLP: Spark NLP 2.0.0! Where to begin? We have no less than 50 Pull Requests merged
Thank you for following up with the biggest changelog ever on Spark NLP: Spark NLP 2.0.0! Where to begin? We have no less than 50 Pull Requests merged this time. Most importantly, we become the first library to have a production ready implementation of BERT embeddings. Along with this interesting deep learning and context based embeddings algorithm, here is a quick overview of new things:
This release is meant to push downstream a few improvements from 2.0.x to the 1.8.x branch, mostly with the objective of keeping the stable branch lin
This release is meant to push downstream a few improvements from 2.0.x to the 1.8.x branch, mostly with the objective of keeping the stable branch line stable, and solving a few serious issues that were pending. This makes 1.8.4 an ideal version for stable deployments.
We're glad to announce a new release for Spark NLP. This one calls the attention of the community who contributed immensely towards reporting bugs and
We're glad to announce a new release for Spark NLP. This one calls the attention of the community who contributed immensely towards reporting bugs and feedback to the library. This release focuses in various bugfixes around DeepSentenceDetector and also python deserialization of some specific pipelines. It also improves the DeepSentenceDetector allowing further fine-tuning and customization. Then, we have embeddings that are being cached in the models folder, and further improvements towards accessing them through S3 storage. Finally, we have made serious improvements in noteoboks and documentation around the library. Special thanks to @Tshimanga and @haimco10 for very interesting contributions. See you on Slack!
This release potentially targets to improve performance and resource usage in some pipelines that use word embeddings, it also comes together with a v
This release potentially targets to improve performance and resource usage in some pipelines that use word embeddings, it also comes together with a very interesting autorotation feature in OCR, and a couple of new annotators to solve particular needs, including the ChunkTokenizer or a Param to limit sentence lengths. Finally, we are starting to organize our multilingual store of models and data for training models. Check the examples for some italian notebooks!. Thanks again to all community for such quick feedback all the time.
This hotfix version of Spark-NLP improves framework support by adding Maven coordinates for OCR and allowing S3 retrieval of files. We also included c
This hotfix version of Spark-NLP improves framework support by adding Maven coordinates for OCR and allowing S3 retrieval of files. We also included code for generating Graphs for NerDL and also for creating your own metadata files for a private model downloader. As new features, we are including a new experimental machine learning based sentence detector, which uses NER for bounds detections. Aside from this, we are including a few bug fixes and OCR improvements. Enjoy! and thanks again for community contributions!
This release is huge! Spark-NLP made the leap into Spark 2.4.0, even with the challenge of not having everyone yet on board there (i.e. Zeppelin doesn
This release is huge! Spark-NLP made the leap into Spark 2.4.0, even with the challenge of not having everyone yet on board there (i.e. Zeppelin doesn't yet support it). In this version we release three new NLP annotators. Two for dependency parsing processes and one for contextual deep learning based spell checking. We also significantly improved OCR functionality, fine-tuning capabilities and general output performance, particularly on tesseract. Finally, there's plenty of bug fixes and improvements in the word embeddings field, along with performance boosts and reduced disk IO. Feel free to shoot us with any feedback you have! Particularly on your Spark 2.4.x experience.
This hotfix release focuses on fixing word-embeddings cluster problems on some frameworks such as Databricsk, while keeping 1.7.x performance benefits
This hotfix release focuses on fixing word-embeddings cluster problems on some frameworks such as Databricsk, while keeping 1.7.x performance benefits. Various YARN based clusters have been tested, databricks cloud among them to test this hotfix. Aside of that, multiple improvements have been commited towards a better support of PySpark-NLP, fixing diverse technical issues in the API that help consistency in Annotator's super classes. Finally, PIP installation has been made easier with a SparkNLP class that creates SparkSession automatically, for those who are learning Python Spark on their local computers. Thanks to all the community for reporting issues.
Failed to get broadcast_6_piece0 of broadcast_6 causing pretrained models not work in cluster frameworks (thanks @EnricoMi)Quick release with another hotfix, due to a new found bug when deserializing word embeddings in a distributed fs. Also introduces changes in applicati
Quick release with another hotfix, due to a new found bug when deserializing word embeddings in a distributed fs. Also introduces changes in application.conf reader in order to allow run-time changes. Also introduces renaming from EmbeddingsHelper API.
Thanks to our slack community (Bryan Wilkinson, @maziyarpanahi, @apiltamang), a few bugs been pointed out very quickly from 1.7.0 release. This hotfix
Thanks to our slack community (Bryan Wilkinson, @maziyarpanahi, @apiltamang), a few bugs been pointed out very quickly from 1.7.0 release. This hotfix fixes an embeddings deserialization issue when cache_pretrained is located on a distributed filesystem. Also, fixes some path resolution in Windows OS. Thanks to Maziyar, .gitattributes been added in order to identify proper languages in GitHub. Finally, 1.7.1 adds a missing annotator from 1.7.0 Chunk2Doc, which converts CHUNK types into DOCUMENT types, for further retokenization or other annotations.
Having multiple annotators that use the same word embeddings set, may result in huge pipelines, driver memory and storage consumption. Since now on, e
Having multiple annotators that use the same word embeddings set, may result in huge pipelines, driver memory and storage consumption. Since now on, embeddings may be shared and reutilized across annotators making the process much more efficient. Also, thanks to @apiltamang, we now better support path resolution for Windows implementations.
Memory and storage saving by allowing annotators with embeddings through params 'includeEmbeddings' and 'embeddingsRef' to allow them to set whether they should be included when saved, or referenced by id from other annotators. EmbeddingsHelper class allows embeddings management
Thanks to @apiltamang for improving URI path support for Windows Servers
Embeddings interfaces and method names completely refactored, hopefully simplified and easier to understand
This release includes a new annotator for de-identification of sensitive information. It uses CHUNK annotations, meaning its accuracy will depend on p
This release includes a new annotator for de-identification of sensitive information. It uses CHUNK annotations, meaning its accuracy will depend on previous annotators on the pipeline. Also, OCR capabilities have been improved in the OCR module. In terms of broken stuff, we've fixed a few annoying bugs on SymmetricDelete and SentenceDetector explode feature. Finally, pip is now part of the official repositories, meaning you can install it just as any other module. It also includes jars and we've added a SparkNLP class which creates SparkSession easily for you. Thanks again for all community contribution in issues, feedback and comments in GitHub and in Slack.
In this release, we focused on reviewing out streaming performance, buy measuring our amount of sentences processed by second, through a LightPipeline
In this release, we focused on reviewing out streaming performance, buy measuring our amount of sentences processed by second, through a LightPipeline. We increased Norvig Spell Checker by more than 300% by disabling DoubleVariants and improving algorithm orders. It is now reported capable of 42K sentences per second. Symmetric Delete Spell checker is more accurate, although it has been reported to process 2K sentences per second. NerCRF has been reported to process 300 hundred sentences per second, while NerDL can do twice fast (about 700 sentences per second). Vivekn Sentiment Analysis was improved and is now capable to processing 100K sentences per sentence (before it was below 500). Finally, SentenceDetector performance was improved by a 40% from ~30K rows processed per second to ~40K. But, we have now enabled Abbreviation processing by default which reduces final speed to 22K rows per second with a negative net but better accuracy. Again, thanks for the community for helping with feedback. We welcome everyone asking questions or giving feedback in our Slack channel or reporting issues on Github.
Your coding agent can read these notes before it upgrades. Set up the MCP server →