NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4698 most downloaded on PyPI
John Snow Labs Spark NLP is a natural language processing library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines, that scale easily in a distributed environment.
Last release 11 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
196 releases · first in 2018
We are pleased to release Spark NLP 🚀 3.2.2! This release comes with accessible Models Hub to our community to host their models and pipelines for fre
We are pleased to release Spark NLP 🚀 3.2.2! This release comes with accessible Models Hub to our community to host their models and pipelines for free, new RoBERTa and XLM-RoBERTa Sentence Embeddings, over 40 new models and pipelines in 20+ languages, bug fixes, and more
As always, we would like to thank our community for their feedback, questions, and feature requests.
Serve Your Spark NLP Models for Free! You can host and share your Spark NLP models & pipelines publicly with everyone to reuse them with one line of code!
We are opening Models Hub to everyone to upload their models and pipelines, showcase their work, and share them with others.
Please visit the following page for more information: https://modelshub.johnsnowlabs.com/
elmo as a default poolingLayer in ElmoEmbeddingsSpark NLP 3.2.2 comes with new Turkish text classifier pipelines, Expert BERT Word and Sentence embeddings such as wiki books and PubMed, new BERT model for 17 Indian languages, and Sentence Detection models for 15 new languages.
| Name | Build | Lang |
|---|---|---|
| classifierdl_berturk_cyberbullying_pipeline | 3.1.3 | tr |
| classifierdl_bert_news_pipeline | 3.1.3 | de |
| classifierdl_electra_questionpair_pipeline | 3.2.0 | en |
| classifierdl_bert_news_pipeline | 3.2.0 | tr |
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | ner_conll_elmo | 3.2.2 | en |
| NerDLModel | ner_conll_albert_base_uncased | 3.2.2 | en |
| NerDLModel | ner_conll_albert_large_uncased | 3.2.2 | en |
| NerDLModel | ner_conll_xlnet_base_cased | 3.2.2 | en |
| Model | Name | Build | Lang |
|---|---|---|---|
| BertEmbeddings | bert_muril | 3.2.0 | xx |
| BertEmbeddings | bert_wiki_books_sst2 | 3.2.0 | en |
| BertEmbeddings | bert_wiki_books_squad2 | 3.2.0 | en |
| BertEmbeddings | bert_wiki_books_qqp | 3.2.0 | en |
| BertEmbeddings | bert_wiki_books_qnli | 3.2.0 | en |
| BertEmbeddings | bert_wiki_books_mnli | 3.2.0 | en |
| BertEmbeddings | bert_wiki_books | 3.2.0 | en |
| BertEmbeddings | bert_pubmed_squad2 | 3.2.0 | en |
| BertEmbeddings | bert_pubmed | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books_sst2 | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books_squad2 | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books_qqp | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books_qnli | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books_mnli | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_wiki_books | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_pubmed_squad2 | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_pubmed | 3.2.0 | en |
| BertSentenceEmbeddings | sent_bert_muril | 3.2.0 | xx |
Yiddish, Ukrainian, Telugu, Tamil, Somali, Sindhi, Russian, Punjabi, Nepali, Marathi, Malayalam, Kannada, Indonesian, Gujrati, Bosnian
| Model | Name | Build | Lang |
|---|---|---|---|
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | yi |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | uk |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | te |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | ta |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | so |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | sd |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | ru |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | pa |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | ne |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | mr |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | ml |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | kn |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | id |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | gu |
| SentenceDetectorDLModel | sentence_detector_dl | 3.2.0 | bs |
The complete list of all 3700+ models & pipelines in 200+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==3.2.2
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.2
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.2
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.2.2
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.2.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.2.2</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.2.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.2.2</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.2.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.2.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.2.2.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.2.2.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.2.2.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.2.2.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.2.2.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.2.2.jar
One column per quarter.
Fix unsupported model error in pretrained function for LongformerEmbeddings, BertForTokenClassification, and DistilBertForTokenClassification https://
unsupported model error in pretrained function for LongformerEmbeddings, BertForTokenClassification, and DistilBertForTokenClassification https://github.com/JohnSnowLabs/spark-nlp/issues/5947Python
#PyPI
pip install spark-nlp==3.2.1
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.1
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.1
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.2.1
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.2.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.2.1</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.2.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.2.1</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.2.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.2.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.2.1.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.2.1.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.2.1.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.2.1.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.2.1.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.2.1.jar
NEW: Introducing new Spark NLP configurations via spark.conf() by deprecating application.conf usage. You can easily change Spark NLP configurations i…
We are very excited to release Spark NLP 🚀 3.2.0! This is a big release with new Longformer models for long documents, BertForTokenClassification & DistilBertForTokenClassification for existing or fine-tuned models on HuggingFace, GraphExctraction & GraphFinisher to find relevant relationships between words, support for multilingual Date Matching, new Pydoc for Python APIs, and so many more!
As always, we would like to thank our community for their feedback, questions, and feature requests.
Longformer is a transformer model for long documents. Longformer is a BERT-like model started from the RoBERTa checkpoint and pretrained for MLM on long documents. It supports sequences of length up to 4,096.We have trained two NER models based on Longformer Base and Large embeddings:
| Model | Accuracy | F1 Test | F1 Dev |
|---|---|---|---|
| ner_conll_longformer_base_4096 | 94.75% | 90.09 | 94.22 |
| ner_conll_longformer_large_4096 | 95.79% | 91.25 | 94.82 |
BertForTokenClassification can load BERT Models with a token classification head on top (a linear layer on top of the hidden-states output) e.g. for Named-Entity-Recognition (NER) tasks. This annotator is compatible with all the models trained/fine-tuned by using BertForTokenClassification or TFBertForTokenClassification in HuggingFace 🤗DistilBertForTokenClassification can load BERT Models with a token classification head on top (a linear layer on top of the hidden-states output) e.g. for Named-Entity-Recognition (NER) tasks. This annotator is compatible with all the models trained/fine-tuned by using DistilBertForTokenClassification or TFDistilBertForTokenClassification in HuggingFace 🤗NerDLModel and creates a dependency tree that describes how the entities relate to each other. For that, a triple store format is used. Nodes represent the entities and the edges represent the relations between those entities. The graph can then be used to find relevant relationships between wordsapplication.conf usage. You can easily change Spark NLP configurations in SparkSession. For more examples please vistit Spark NLP Configurationlog_folder Spark NLP config and outputLogsPath param in NerDLApproach, ClassifierDlApproach, MultiClassifierDlApproach, and SentimentDlApproach annotatorscache_folder, log_folder, and cluster_tmp_dir to sparknlp.start() function to set Spark NLP configurationsSpark NLP 3.2.0 comes with new LongformerEmbeddings, BertForTokenClassification, and DistilBertForTokenClassification annotators.
| Model | Name | Build | Lang |
|---|---|---|---|
| LongformerEmbeddings | longformer_base_4096 | 3.2.0 | en |
| LongformerEmbeddings | longformer_large_4096 | 3.2.0 | en |
New NER models for CoNLL (4 entities) and OntoNotes (18 entities) trained by using BERT, RoBERTa, DistilBERT, XLM-RoBERTa, and Longformer Embeddings:
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | ner_ontonotes_roberta_base | 3.2.0 | en |
| NerDLModel | ner_ontonotes_roberta_large | 3.2.0 | en |
| NerDLModel | ner_ontonotes_distilbert_base_cased | 3.2.0 | en |
| NerDLModel | ner_conll_bert_base_cased | 3.2.0 | en |
| NerDLModel | ner_conll_distilbert_base_cased | 3.2.0 | en |
| NerDLModel | ner_conll_roberta_base | 3.2.0 | en |
| NerDLModel | ner_conll_roberta_large | 3.2.0 | en |
| NerDLModel | ner_conll_xlm_roberta_base | 3.2.0 | en |
| NerDLModel | ner_conll_longformer_base_4096 | 3.2.0 | en |
| NerDLModel | ner_conll_longformer_large_4096 | 3.2.0 | en |
New BERT and DistilBERT fine-tuned for the Named Entity Recognition (NER) in English, Persian, Spanish, Swedish, and Turkish:
| Model | Name | Build | Lang |
|---|---|---|---|
| BertForTokenClassification | bert_base_token_classifier_conll03 | 3.2.0 | en |
| BertForTokenClassification | bert_large_token_classifier_conll03 | 3.2.0 | en |
| BertForTokenClassification | bert_base_token_classifier_ontonote | 3.2.0 | en |
| BertForTokenClassification | bert_large_token_classifier_ontonote | 3.2.0 | en |
| BertForTokenClassification | bert_token_classifier_parsbert_armanner | 3.2.0 | fa |
| BertForTokenClassification | bert_token_classifier_parsbert_ner | 3.2.0 | fa |
| BertForTokenClassification | bert_token_classifier_parsbert_peymaner | 3.2.0 | fa |
| BertForTokenClassification | bert_token_classifier_turkish_ner | 3.2.0 | tr |
| BertForTokenClassification | bert_token_classifier_spanish_ner | 3.2.0 | es |
| BertForTokenClassification | bert_token_classifier_swedish_ner | 3.2.0 | sv |
| BertForTokenClassification | bert_base_token_classifier_few_nerd | 3.2.0 | en |
| DistilBertForTokenClassification | distilbert_base_token_classifier_few_nerd | 3.2.0 | en |
| DistilBertForTokenClassification | distilbert_base_token_classifier_conll03 | 3.2.0 | en |
| DistilBertForTokenClassification | distilbert_base_token_classifier_ontonotes | 3.2.0 | en |
| DistilBertForTokenClassification | distilbert_token_classifier_persian_ner | 3.2.0 | fa |
The complete list of all 3700+ models & pipelines in 200+ languages is available on Models Hub.
Import hundreds of models in different languages to Spark NLP
| Spark NLP | HuggingFace Notebooks | Colab |
|---|---|---|
| LongformerEmbeddings | HuggingFace in Spark NLP - Longformer | |
| BertForTokenClassification | HuggingFace in Spark NLP - BertForTokenClassification | |
| DistilBertForTokenClassification | HuggingFace in Spark NLP - DistilBertForTokenClassification |
You can visit Import Transformers in Spark NLP for more info
New Multilingual DateMatcher and MultiDateMatcher
| Spark NLP | Jupyter Notebooks |
|---|---|
| MultiDateMatcher | Date Matcher in English |
| MultiDateMatcher | Date Matcher in French |
| MultiDateMatcher | Date Matcher in German |
| MultiDateMatcher | Date Matcher in Italian |
| MultiDateMatcher | Date Matcher in Portuguese |
| MultiDateMatcher | Date Matcher in Spanish |
| GraphExtraction | Graph Extraction Intro |
| GraphExtraction | Graph Extraction |
| GraphExtraction | Graph Extraction Explode Entities |
The use of application.conf has been deprecated in Spark NLP 3.2.0 release. You can set those configurations via Spark Conf during SparkSession creation. For the full list and examples please visit the Spark NLP Configuration.
Python
#PyPI
pip install spark-nlp==3.2.0
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.2.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.2.0
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.2.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.2.0
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.2.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.2.0
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.2.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.2.0</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.2.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.2.0</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.2.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.2.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.2.0.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.2.0.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.2.0.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.2.0.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.2.0.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.2.0.jar
We are pleased to release Spark NLP 🚀 3.1.3! In this release, we bring notebooks to easily import models for BERT and ALBERT models from TF Hub into S
We are pleased to release Spark NLP 🚀 3.1.3! In this release, we bring notebooks to easily import models for BERT and ALBERT models from TF Hub into Spark NLP, new multilingual NER models for 40 languages with a fine-tuned XLM-RoBERTa model, and new state-of-the-art document/sentence embeddings models for English and 100+ languages!
As always, we would like to thank our community for their feedback, questions, and feature requests.
We have trained multilingual NER models by using the entire XTREME (40 languages) and WIKINER (8 languages).
Multilingual Named Entity Recognition:
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | ner_xtreme_xlm_roberta_xtreme_base | 3.1.3 | xx |
| NerDLModel | ner_xtreme_glove_840B_300 | 3.1.3 | xx |
| NerDLModel | ner_wikiner_xlm_roberta_base | 3.1.3 | xx |
| NerDLModel | ner_wikiner_glove_840B_300 | 3.1.3 | xx |
| NerDLModel | ner_mit_movie_simple_distilbert_base_cased | 3.1.3 | en |
| NerDLModel | ner_mit_movie_complex_distilbert_base_cased | 3.1.3 | en |
| NerDLModel | ner_mit_movie_complex_bert_base_cased | 3.1.3 | en |
Fine-tuned XLM-RoBERTa base model by randomly masking 15% of XTREME dataset:
| Model | Name | Build | Lang |
|---|---|---|---|
| XlmRoBertaEmbeddings | xlm_roberta_xtreme_base | 3.1.3 | xx |
New Universal Sentence Encoder trained with CMLM (English & 100+ languages):
The models extend the BERT transformer architecture and that is why we use them with BertSentenceEmbeddings.
| Model | Name | Build | Lang |
|---|---|---|---|
| BertSentenceEmbeddings | sent_bert_use_cmlm_en_base | 3.1.3 | en |
| BertSentenceEmbeddings | sent_bert_use_cmlm_en_large | 3.1.3 | en |
| BertSentenceEmbeddings | sent_bert_use_cmlm_multi_base | 3.1.3 | xx |
| BertSentenceEmbeddings | sent_bert_use_cmlm_multi_base_br | 3.1.3 | xx |
We used BERT base, large, and the new Universal Sentence Encoder trained with CMLM extending the BERT transformer architecture to train ClassifierDL with News dataset:
(120k training examples - 10 Epochs - 512 max sequence - Nvidia Tesla P100)
| Model | Accuracy | F1 | Duration |
|---|---|---|---|
| tfhub_use | 0.90 | 0.89 | 10 min |
| tfhub_use_lg | 0.91 | 0.90 | 24 min |
| sent_bert_base_cased | 0.92 | 0.90 | 35 min |
| sent_bert_large_cased | 0.93 | 0.91 | 75 min |
| sent_bert_use_cmlm_en_base | 0.934 | 0.91 | 36 min |
| sent_bert_use_cmlm_en_large | 0.945 | 0.92 | 72 min |
The complete list of all 3700+ models & pipelines in 200+ languages is available on Models Hub.
| Spark NLP | TF Hub Notebooks |
|---|---|
| BertEmbeddings | TF Hub in Spark NLP - BERT |
| BertSentenceEmbeddings | TF Hub in Spark NLP - BERT Sentence |
| AlbertEmbeddings | TF Hub in Spark NLP - ALBERT |
Python
#PyPI
pip install spark-nlp==3.1.3
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.3
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.3
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.1.3
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.1.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.1.3</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.1.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.1.3</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.1.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.1.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.1.3.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.1.3.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.1.3.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.1.3.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.1.3.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.1.3.jar
========
We are pleased to release Spark NLP 🚀 3.1.2! We have a new and much-improved XLNet annotator with support for HuggingFace 🤗 models in Spark NLP. We ma
We are pleased to release Spark NLP 🚀 3.1.2! We have a new and much-improved XLNet annotator with support for HuggingFace 🤗 models in Spark NLP. We managed to make XlnetEmbeddings almost 5x times faster on GPU compare to prior releases!
As always, we would like to thank our community for their feedback, questions, and feature requests.
ContextSpellChecker
ViveknSentimentApproach
RegexMatcher
WordSegmenterApproach
ViveknSentimentApproach
PerceptronApproach
Introducing a new batch annotation technique implemented in Spark NLP 3.1.2 for XlnetEmbeddings annotator to radically improve prediction/inferencing performance. From now on the batchSize for these annotators means the number of rows that can be fed into the models for prediction instead of sentences per row. You can control the throughput when you are on accelerated hardware such as GPU to fully utilize it.
We have migrated XlnetEmbeddings to TensorFlow v2, the earlier models prior to 3.1.2 won't work after this release.
We have already updated the models and uploaded them on Models Hub. You can use pretrained() that takes care of it automatically or please make sure you download the new models manually.
Python
#PyPI
pip install spark-nlp==3.1.2
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.2
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.2
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.1.2
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.1.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.1.2</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.1.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.1.2</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.1.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.1.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.1.2.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.1.2.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.1.2.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.1.2.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.1.2.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.1.2.jar
We have migrated XlnetEmbeddings to TensorFlow v2, the earlier models prior to 3.1.2 won't work after this release.
We have already updated the models and uploaded them on Models Hub. You can use pretrained() that takes care of it automatically or please make sure you download the new models manually.
========
We are pleased to release Spark NLP 🚀 3.1.1! We have a new and much-improved ALBERT annotator with support for HuggingFace 🤗 models in Spark NLP. We m
We are pleased to release Spark NLP 🚀 3.1.1! We have a new and much-improved ALBERT annotator with support for HuggingFace 🤗 models in Spark NLP. We managed to make AlbertEmbeddings almost 7x times faster on GPU compare to prior releases!
As always, we would like to thank our community for their feedback, questions, and feature requests.
sparknlp.start(). Thanks to PySpark 3.x, this is now possible with sparknlp.start(real_time_output=True) to have the outputs of Spark NLP (such as metrics during training) right in your Jupyter, Colab, and Kaggle notebooks.day after tomorrow or day before yesterday https://github.com/JohnSnowLabs/spark-nlp/pull/5706logger inside session on some setup https://github.com/JohnSnowLabs/spark-nlp/pull/5715init_all_tables https://github.com/JohnSnowLabs/spark-nlp/pull/5715Introducing a new batch annotation technique implemented in Spark NLP 3.1.1 for AlbertEmbeddings annotator to radically improve prediction/inferencing performance. From now on the batchSize for these annotators means the number of rows that can be fed into the models for prediction instead of sentences per row. You can control the throughput when you are on accelerated hardware such as GPU to fully utilize it.
(Performed on a Databricks cluster)
| Spark NLP 2.x/3.0.x vs. 3.1.1 | CPU | GPU |
|---|---|---|
| ALBERT Base | 22% | 340% |
| Albert Large | 20% | 770% |
We will update this benchmark table in future pre-releases.
We have migrated AlbertEmbeddings to TensorFlow v2, the earlier models prior to 3.1.1 won't work after this release. We have already updated the models and uploaded them on Models Hub. You can use pretrained() that takes care of it automatically or please make sure you download the new models manually.
Python
#PyPI
pip install spark-nlp==3.1.1
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.1
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.1
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.1.1
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.1.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.1.1</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.1.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.1.1</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.1.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.1.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.1.1.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.1.1.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.1.1.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.1.1.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.1.1.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.1.1.jar
We are very excited to release Spark NLP 🚀 3.1.0! This is one of our biggest releases with lots of models, pipelines, and groundworks for future featu
We are very excited to release Spark NLP 🚀 3.1.0! This is one of our biggest releases with lots of models, pipelines, and groundworks for future features that we are so proud to share it with our community.
Spark NLP 3.1.0 comes with over 2600+ new pretrained models and pipelines in over 200+ languages, new DistilBERT, RoBERTa, and XLM-RoBERTa annotators, support for HuggingFace 🤗 (Autoencoding) models in Spark NLP, and extends support for new Databricks and EMR instances.
As always, we would like to thank our community for their feedback, questions, and feature requests.
bert-base-uncased, runs 60% faster while preserving over 95% of BERT’s performancessaved_model feature in HuggingFace within a few lines of codes and import any BERT, DistilBERT, RoBERTa, and XLM-RoBERTa models to Spark NLP. We will work on the remaining annotators and extend this support to the rest with each release - For more information please visit this discussionTokenizer or RegexTokenizer and generates token pieces, encodes, and decodes the resultsSpark NLP 3.1.0 comes with over 2600+ new pretrained models and pipelines in over 200 languages available for Windows, Linux, and macOS users.
| Model | Name | Build | Lang |
|---|---|---|---|
| BertEmbeddings | bert_base_dutch_cased | 3.1.0 | nl |
| BertEmbeddings | bert_base_german_cased | 3.1.0 | de |
| BertEmbeddings | bert_base_german_uncased | 3.1.0 | de |
| BertEmbeddings | bert_base_italian_cased | 3.1.0 | it |
| BertEmbeddings | bert_base_italian_uncased | 3.1.0 | it |
| BertEmbeddings | bert_base_turkish_cased | 3.1.0 | tr |
| BertEmbeddings | bert_base_turkish_uncased | 3.1.0 | tr |
| BertEmbeddings | chinese_bert_wwm | 3.1.0 | zh |
| BertEmbeddings | bert_base_chinese | 3.1.0 | zh |
| DistilBertEmbeddings | distilbert_base_cased | 3.1.0 | en |
| DistilBertEmbeddings | distilbert_base_uncased | 3.1.0 | en |
| DistilBertEmbeddings | distilbert_base_multilingual_cased | 3.1.0 | xx |
| RoBertaEmbeddings | roberta_base | 3.1.0 | en |
| RoBertaEmbeddings | roberta_large | 3.1.0 | en |
| RoBertaEmbeddings | distilroberta_base | 3.1.0 | en |
| XlmRoBertaEmbeddings | xlm_roberta_base | 3.1.0 | xx |
| XlmRoBertaEmbeddings | twitter_xlm_roberta_base | 3.1.0 | xx |
| Model | Name | Build | Lang |
|---|---|---|---|
| MarianTransformer | Chinese to Vietnamese | 3.1.0 | xx |
| MarianTransformer | Chinese to Ukrainian | 3.1.0 | xx |
| MarianTransformer | Chinese to Dutch | 3.1.0 | xx |
| MarianTransformer | Chinese to English | 3.1.0 | xx |
| MarianTransformer | Chinese to Finnish | 3.1.0 | xx |
| MarianTransformer | Chinese to Italian | 3.1.0 | xx |
| MarianTransformer | Yoruba to English | 3.1.0 | xx |
| MarianTransformer | Yapese to French | 3.1.0 | xx |
| MarianTransformer | Waray to Spanish | 3.1.0 | xx |
| MarianTransformer | Ukrainian to English | 3.1.0 | xx |
| MarianTransformer | Hindi to Urdu | 3.1.0 | xx |
| MarianTransformer | Italian to Ukrainian | 3.1.0 | xx |
| MarianTransformer | Italian to Icelandic | 3.1.0 | xx |
Import hundreds of models in different languages to Spark NLP
| Spark NLP | HuggingFace Notebooks |
|---|---|
| BertEmbeddings | HuggingFace in Spark NLP - BERT |
| BertSentenceEmbeddings | HuggingFace in Spark NLP - BERT Sentence |
| DistilBertEmbeddings | HuggingFace in Spark NLP - DistilBERT |
| RoBertaEmbeddings | HuggingFace in Spark NLP - RoBERTa |
| XlmRoBertaEmbeddings | HuggingFace in Spark NLP - XLM-RoBERTa |
The complete list of all 3700+ models & pipelines in 200+ languages is available on Models Hub.
3.1.x release. You can either use MarianTransformer.pretrained(MODEL_NAME) and it will automatically download the compatible model or you can visit Models Hub to download the compatible models for offline use via MarianTransformer.load(PATH)Python
#PyPI
pip install spark-nlp==3.1.0
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.1.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.1.0
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.1.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.1.0
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.1.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.1.0
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.1.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.1.0</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.1.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.1.0</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.1.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.1.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.1.0.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.1.0.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.1.0.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.1.0.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.1.0.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.1.0.jar
We are glad to release Spark NLP 3.0.3! We have added some new features to our T5 Transformer annotator to help with longer and more accurate text gen
We are glad to release Spark NLP 3.0.3! We have added some new features to our T5 Transformer annotator to help with longer and more accurate text generation, trained some new multi-lingual models and pipelines in Farsi, Hebrew, Korean, and Turkish, and fixed some bugs in this release.
As always, we would like to thank our community for their feedback, questions, and feature requests.
top_p or higher are kept for generationnext friday or next Friday https://github.com/JohnSnowLabs/spark-nlp/pull/2848New multilingual models and pipelines for Farsi, Hebrew, Korean, and Turkish
| Model | Name | Build | Lang |
|---|---|---|---|
| ClassifierDLModel | classifierdl_bert_news | 3.0.2 | tr |
| UniversalSentenceEncoder | tfhub_use_multi | 3.0.0 | xx |
| UniversalSentenceEncoder | tfhub_use_multi_lg | 3.0.0 | xx |
| Pipeline | Name | Build | Lang |
|---|---|---|---|
| PretrainedPipeline | recognize_entities_dl | 3.0.0 | fa |
| PretrainedPipeline | explain_document_lg | 3.0.2 | he |
| PretrainedPipeline | explain_document_lg | 3.0.2 | ko |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==3.0.3
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.3
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.3
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.0.3
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.0.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.0.3</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.0.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.0.3</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.0.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.0.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.3.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.3.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.0.3.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.0.3.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.0.3.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.0.3.jar
========
We are glad to release Spark NLP 3.0.2! We have added some new features, improvements, trained some new multi-lingual models, and fixed some bugs in t
We are glad to release Spark NLP 3.0.2! We have added some new features, improvements, trained some new multi-lingual models, and fixed some bugs in this release.
As always, we would like to thank our community for their feedback, questions, and feature requests.
# NerDLModel and NerCrfModel before 3.0.2
[[named_entity, 0, 4, B-LOC, [word -> Japan, confidence -> 0.9998], []]
# Now in Spark NLP 3.0.2
[[named_entity, 0, 4, B-LOC, [B-LOC -> 0.9998, I-ORG -> 0.0, I-MISC -> 0.0, I-LOC -> 0.0, I-PER -> 0.0, B-MISC -> 0.0, B-ORG -> 1.0E-4, word -> Japan, O -> 0.0, B-PER -> 0.0], []]
[chunk, 30, 41, Barack Obama, [entity -> PERSON, sentence -> 0, chunk -> 0, confidence -> 0.94035]
New multilingual models for Afrikaans, Welsh, Maltese, Tamil, and Vietnamese
| Model | Name | Build | Lang |
|---|---|---|---|
| PerceptronModel | pos_afribooms | 3.0.0 | af |
| LemmatizerModel | lemma | 3.0.0 | cy |
| LemmatizerModel | lemma | 3.0.0 | mt |
| LemmatizerModel | lemma | 3.0.0 | af |
| LemmatizerModel | lemma | 3.0.0 | ta |
| LemmatizerModel | lemma | 3.0.0 | vi |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==3.0.2
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.2
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.2
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.0.2
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.0.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.0.2</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.0.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.0.2</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.0.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.0.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.2.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.2.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.0.2.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.0.2.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.0.2.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.0.2.jar
# Previously in NerDLModel and NerCrfModel
[[named_entity, 0, 4, B-LOC, [word -> Japan, confidence -> 0.9998], []]
# In Spark NLP 3.0.2
[[named_entity, 0, 4, B-LOC, [B-LOC -> 0.9998, I-ORG -> 0.0, I-MISC -> 0.0, I-LOC -> 0.0, I-PER -> 0.0, B-MISC -> 0.0, B-ORG -> 1.0E-4, word -> Japan, O -> 0.0, B-PER -> 0.0], []]
[chunk, 30, 37, john, [entity -> PERSON, sentence -> 0, chunk -> 0, confidence -> 0.44035]
========
We are glad to release Spark NLP 3.0.1! We have made some improvements, added 1 line bash script to set up Google Colab and Kaggle kernel for Spark NL
We are glad to release Spark NLP 3.0.1! We have made some improvements, added 1 line bash script to set up Google Colab and Kaggle kernel for Spark NLP 3.x, and improved our Models Hub filtering to help our community to have easier access to over 1300 pretrained models and pipelines in over 200+ languages.
As always, we would like to thank our community for their feedback, questions, and feature requests.
Python
#PyPI
pip install spark-nlp==3.0.1
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.1
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.1
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.0.1
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.0.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.0.1</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.0.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.0.1</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.0.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.0.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.1.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.1.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.0.1.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.0.1.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.0.1.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.0.1.jar
We are very excited to release Spark NLP 3.0.0! This has been one of the biggest releases we have ever done and we are so proud to share this with our
We are very excited to release Spark NLP 3.0.0! This has been one of the biggest releases we have ever done and we are so proud to share this with our community.
Spark NLP 3.0.0 extends the support for Apache Spark 3.0.x and 3.1.x major releases on Scala 2.12 with both Hadoop 2.7. and 3.2. We will support all 4 major Apache Spark and PySpark releases of 2.3.x, 2.4.x, 3.0.x, and 3.1.x helping the community to migrate from earlier Apache Spark versions to newer releases without being worried about Spark NLP support.
As always, we would like to thank our community for their feedback, questions, and feature requests.
spark-nlp and spark-nlp-gpu will be compatible only with Apache Spark 3.x and Scala 2.12)spark-nlp-spark24 and spark-nlp-gpu-spark24)spark-nlp-spark23 and spark-nlp-gpu-spark23)spark24=True)memory="16G")Introducing a new batch annotation technique implemented in Spark NLP 3.0.0 for NerDLModel, BertEmbeddings, and BertSentenceEmbeddings annotators to radically improve prediction/inferencing performance. From now on the batchSize for these annotators means the number of rows that can be fed into the models for prediction instead of sentences per row. You can control the throughput when you are on accelerated hardware such as GPU to fully utilize it.
(Performed on a Databricks cluster)
| Spark NLP 3.0.0 vs. 2.7.x | PySpark 3.x on CPU | PySpark 3.x on GPU |
|---|---|---|
| BertEmbeddings (bert-base) | +10% | +550% (6.6x) |
| BertEmbeddings (bert-large) | +12%. | +690% (7.9x) |
| NerDLModel | +185% | +327% (4.2x) |
There are only 6 annotators that are not compatible to be used with both Scala 2.11 (Apache Spark 2.3 and Apache Spark 2.4) and Scala 2.12 (Apache Spark 3.x) at the same time. You can either train and use them on Apache Spark 2.3.x/2.4.x or train and use them on Apache Spark 3.x.
The rest of our models/pipelines can be used on all Apache Spark and Scala major versions without any issue.
We have already retrained and uploaded all the exiting pretrained for Part of Speech and WordSegmenter models in Apache Spark 3.x and Scala 2.12. We will continue doing this as we see existing models which are not compatible with Apache Spark 3.x and Scala 2.12.
NOTE: You can always use the .pretrained() function which seamlessly will find the compatible and most recent models to download for you. It will download and extract them in your HOME DIRECTORY ~/cached_pretrained/.
More info: https://github.com/JohnSnowLabs/spark-nlp/discussions/2562
Starting Spark NLP 3.0.0 release we no longer publish any artifacts on spark-packages and we continue to host all the artifacts only Maven Repository.
Python
#PyPI
pip install spark-nlp==3.0.0
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.0
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.0
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.0.0
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.0.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.0.0</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.0.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.0.0</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.0.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.0.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.0.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.0.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark24-assembly-3.0.0.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark24-assembly-3.0.0.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.0.0.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-spark23-assembly-3.0.0.jar
Nothing published for this version
Nothing published for this version
We are very pleased to share the very first publicly available pre-release of Spark NLP 3.0.0 with our community! This pre-release expands the support
We are very pleased to share the very first publicly available pre-release of Spark NLP 3.0.0 with our community! This pre-release expands the support for Apache Spark 3.0.x and 3.1.x major releases on Scala 2.12 with both Hadoop 2.7. and 3.2.
Spark NLP 3.0.0 will support all 4 major Apache Spark and PySpark releases of 2.3.x, 2.4.x, 3.0.x, and 3.1.x helping the community to migrate from earlier Apache Spark versions to newer releases without being worried about Spark NLP support.
We have already made more than 7 pre-releases to have Spark NLP 3.0.0 ready for the community to test. We will continue the improvements, introducing new features, and keep releasing our release candidates until the end of March when we release our final Spark NLP 3.0.0! 😊
As always, we would like to thank our community for their feedback, questions, and feature requests.
spark-nlp and spark-nlp-gpu will be compatible only with Apache Spark 3.x and Scala 2.12)spark-nlp-spark24 and spark-nlp-gpu-spark24)spark-nlp-spark23 and spark-nlp-gpu-spark23)spark24=True)memory="16G")Introducing a new batch annotation technique implemented in Spark NLP 3.0.0 for NerDLModel, BertEmbeddings, and BertSentenceEmbeddings annotators to radically improve prediction/inferencing performance. From now on the batchSize for these annotators means the number of rows that can be fed into the models for prediction instead of sentences per row. You can control the throughput when you are on accelerated hardware such as GPU to fully utilize it.
(Performed on a Databricks cluster)
| Spark NLP 3.0.0 vs. 2.7.x | PySpark 2.4.x on CPU | PySpark 2.4.x on GPU | PySpark 3.x on CPU | PySpark 3.x on GPU |
|---|---|---|---|---|
| NerDLModel | 2.2x | 2.6x | 2.8x | 4.3x |
We will update this benchmark table in future pre-releases.
There are only 5 annotators that are not compatible with both Scala 2.11 (Apache Spark 2.3 and Apache Spark 2.4) and Scala 2.12 (Apache Spark 3.x). You can either train and use them on Apache Spark 2.3.x/2.4.x or train and use them on Apache Spark 3.x. The rest of our models/pipelines can be used on all Apache Spark and Scala major versions.
We have already retrained and uploaded all the exiting pretrained Part of Speech and WordSegmenter models for Apache Spark 3.x and Scala 2.12. We will continue doing this as we see existing models which are not compatible with Apache Spark 3.x and Scala 2.12.
NOTE: You can always use the .pretrained() function which seamlessly will find the compatible and most recent models to download for you. It will download and extract them in your HOME DIRECTORY ~/cached_pretrained/.
Spark NLP 3.0.0 doesn't require any migration guide, however, if you are planning to migrate from Apache Spark 2.3.x or 2.4.x to 3.x the following guides will help you in your process:
Python
#PyPI
pip install spark-nlp==3.0.0-rc8
Spark Packages
spark-nlp on Apache Spark 3.0.x and 3.1.x (Scala 2.12 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:3.0.0-rc8
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:3.0.0-rc8
spark-nlp on Apache Spark 2.4.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark24_2.11:3.0.0-rc8
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark24_2.11:3.0.0-rc8
spark-nlp on Apache Spark 2.3.x (Scala 2.11 only):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:3.0.0-rc8
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:3.0.0-rc8
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu-spark23_2.11:3.0.0-rc8
Maven
spark-nlp on Apache Spark 3.0.x and 3.1.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>3.0.0-rc8</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>3.0.0-rc8</version>
</dependency>
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark24_2.11</artifactId>
<version>3.0.0-rc8</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark24_2.11</artifactId>
<version>3.0.0-rc8</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>3.0.0-rc8</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>3.0.0-rc8</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.0-rc8.jar
GPU on Apache Spark 3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.0-rc8.jar
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-3.0.0-rc8.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-3.0.0-rc8.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-3.0.0-rc8.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-3.0.0-rc8.jar
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
We are glad to release Spark NLP 2.7.5 release! Starting this release we no longer ship Hadoop AWS and AWS Java SDK dependencies. This change allows u
We are glad to release Spark NLP 2.7.5 release! Starting this release we no longer ship Hadoop AWS and AWS Java SDK dependencies. This change allows users to avoid any conflicts in AWS environments and also results in more EMR 5.x versions support.
As always, we would like to thank our community for their feedback, questions, and feature requests.
Python
#PyPI
pip install spark-nlp==2.7.5
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.5
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.5
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.5
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.5</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.5</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.5.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.5.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.5.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.5.jar
We are glad to release Spark NLP 2.7.4 release! This release comes with a few bug fixes, enhancements, and 4 new pretrained models.
We are glad to release Spark NLP 2.7.4 release! This release comes with a few bug fixes, enhancements, and 4 new pretrained models.
As always, we would like to thank our community for their feedback, questions, and feature requests.
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | bengali_cc_300d | 2.7.3 | bn |
| WordEmbeddingsModel | bengaliner_cc_300d | 2.7.3 | bn |
| NerDLModel | nerdl_snips_100d | 2.7.3 | en |
| ClassifierDLModel | classifierdl_use_snips | 2.7.3 | en |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Starting today, we have moved all of the Fat JARs hosted on our S3 to the auxdata.johnsnowlabs.com/public/jars/ location. We have also fixed the links in the previous releases.
Python
#PyPI
pip install spark-nlp==2.7.4
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.4
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.4
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.4
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.4</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.4</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.4.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.4.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.4.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.4.jar
We are glad to release Spark NLP 2.7.3 release! This release comes with a couple of bug fixes, enhancements, and 20+ pretrained models and pipelines i
We are glad to release Spark NLP 2.7.3 release! This release comes with a couple of bug fixes, enhancements, and 20+ pretrained models and pipelines including support for Bengali Named Entity Recognition, Hindi Word Embeddings, and state-of-the-art transformer based OntoNotes models and pipelines!
As always, we would like to thank our community for their feedback, questions, and feature requests.
This release comes with support for Bengali Named Entity Recognition and Hindi Word Embeddings. We are also announcing the release of 18 new state-of-the-art transformer based OntoNotes models and pipelines! These models are trained by using Transformers pretrained models such as BERT Tiny, BERT Mini, BERT Small, BERT Medium, BERT Base, BERT Large, ELECTRA Small, ELECTRA Base, and ELECTRA Large.
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | ner_jifs_glove_840B_300d | 2.7.0 | bn |
| WordEmbeddingsModel | hindi_cc_300d | 2.7.0 | hi |
| Model | Name | Build | Lang |
|---|---|---|---|
| NerDLModel | onto_small_bert_L2_128 | 2.7.0 | en |
| NerDLModel | onto_small_bert_L4_256 | 2.7.0 | en |
| NerDLModel | onto_small_bert_L4_512 | 2.7.0 | en |
| NerDLModel | onto_small_bert_L8_512 | 2.7.0 | en |
| NerDLModel | onto_bert_base_cased | 2.7.0 | en |
| NerDLModel | onto_bert_large_cased | 2.7.0 | en |
| NerDLModel | onto_electra_small_uncased | 2.7.0 | en |
| NerDLModel | onto_electra_base_uncased | 2.7.0 | en |
| NerDLModel | onto_electra_large_uncased | 2.7.0 | en |
| Pipeline | Build | Lang |
|---|---|---|
| onto_recognize_entities_bert_tiny | 2.7.0 | en |
| onto_recognize_entities_bert_mini | 2.7.0 | en |
| onto_recognize_entities_bert_small | 2.7.0 | en |
| onto_recognize_entities_bert_medium | 2.7.0 | en |
| onto_recognize_entities_bert_base | 2.7.0 | en |
| onto_recognize_entities_bert_large | 2.7.0 | en |
| onto_recognize_entities_electra_small | 2.7.0 | en |
| onto_recognize_entities_electra_base | 2.7.0 | en |
| onto_recognize_entities_electra_large | 2.7.0 | en |
| SYSTEM | YEAR | LANGUAGE | ONTONOTES |
|---|---|---|---|
| Spark NLP v2.7 | 2021 | Python/Scala/Java/R | 90.0 (test F1) 92.5 (dev F1) |
| spaCy 3.0 (RoBERTa) | 2020 | Python | 89.7 (dev F1) |
| Stanza (StanfordNLP) | 2020 | Python | 88.8 (dev F1) |
| Flair | 2018 | Python | 89.7 |
Python
#PyPI
pip install spark-nlp==2.7.3
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.3
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.3
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.3
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.3</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.3.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.3.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.3.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.3.jar
We are glad to release Spark NLP 2.7.2 release! This release comes with a couple of bug fixes, enhancements, and 25+ pretrained models and pipelines i
We are glad to release Spark NLP 2.7.2 release! This release comes with a couple of bug fixes, enhancements, and 25+ pretrained models and pipelines in Amharic, Bengali, Bhojpuri, Japanese, and Korean languages.
As always, we would like to thank our community for their feedback, questions, and feature requests.
The 2.7.x release comes with over 720+ new pretrained models and pipelines available for Windows, Linux, and macOS users.
| Model | Name | Build | Lang |
|---|---|---|---|
| SentimentDLModel | sentimentdl_glove_imdb | 2.7.1 | en |
| SentimentDLModel | sentimentdl_use_imdb | 2.7.1 | en |
| SentimentDLModel | sentimentdl_use_twitter | 2.7.1 | en |
| ClassifierDLMode | classifierdl_use_spam | 2.7.1 | en |
| ClassifierDLModel | classifierdl_use_sarcasm | 2.7.1 | en |
| ClassifierDLModel | classifierdl_use_fakenews | 2.7.1 | en |
| ClassifierDLModel | classifierdl_use_emotion | 2.7.1 | en |
| ClassifierDLModel | classifierdl_use_cyberbullying | 2.7.1 | en |
| MultiClassifierDLModel | multiclassifierdl_use_toxic_sm | 2.7.1 | en |
| MultiClassifierDLModel | multiclassifierdl_use_toxic | 2.7.1 | en |
| MultiClassifierDLModel | multiclassifierdl_use_e2e | 2.7.1 | en |
Some of the new models for Amharic, Bengali, Bhojpuri, Japanese, and Korean languages
| Model | Name | Build | Lang |
|---|---|---|---|
| SentimentDLModel | sentiment_jager_use | 2.7.1 | th |
| SentimentDLModel | sentimentdl_urduvec_imdb | 2.7.1 | ur |
| LemmatizerModel | lemma | 2.7.0 | am |
| LemmatizerModel | lemma | 2.7.0 | bn |
| LemmatizerModel | lemma | 2.7.0 | bh |
| LemmatizerModel | lemma | 2.7.0 | ja |
| LemmatizerModel | lemma | 2.7.0 | ko |
| PerceptronModel | pos_ud_att | 2.7.0 | am |
| PerceptronModel | pos_ud_bhtb | 2.7.0 | bh |
| PerceptronModel | pos_msri | 2.7.0 | bn |
| PerceptronModel | pos_lst20 | 2.7.0 | th |
| WordSegmenterModel | wordseg_best | 2.7.0 | th |
| NerDLModel | ner_lst20_glove_840B_300d | 2.7.0 | th |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==2.7.2
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.2
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.2
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.2
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.2</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.2.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.2.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.2.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.2.jar
We are glad to release Spark NLP 2.7.1 towards making 2.7 stable release! This release comes with 3 new optimized T5 models, 2 new TREC pipelines, a f
We are glad to release Spark NLP 2.7.1 towards making 2.7 stable release! This release comes with 3 new optimized T5 models, 2 new TREC pipelines, a few bug fixes, and other improvements. We highly recommend all users to upgrade to 2.7.1 for more stability while paying attention to the backward compatibility notice.
As always, we would like to thank our community for their feedback, questions, and feature requests.
The 2.7.x release comes with over 720+ new pretrained models and pipelines available for Windows, Linux, and macOS users.
| Model | Name | Build | Lang |
|---|---|---|---|
| T5Transformer | t5_small | 2.7.1 | en |
| T5Transformer | t5_base | 2.7.1 | en |
| T5Transformer | google_t5_small_ssm_nq | 2.7.1 | en |
| Model | Name | Build | Lang |
|---|---|---|---|
| ClassifierDL | classifierdl_use_trec6 | 2.7.1 | en |
| ClassifierDL | classifierdl_use_trec50 | 2.7.1 | en |
| ClassifierDL | classifierdl_use_trec6_pipeline | 2.7.1 | en |
| ClassifierDL | classifierdl_use_trec50_pipeline | 2.7.1 | en |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==2.7.1
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.1
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.1
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.1
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.1</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.1.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.1.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.1.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.1.jar
We are very excited to release Spark NLP 2.7.0! This has been one of the biggest releases we have ever done that we are so proud to share it with our
We are very excited to release Spark NLP 2.7.0! This has been one of the biggest releases we have ever done that we are so proud to share it with our community!
In this release, we are bringing support to state-of-the-art Seq2Seq and Text2Text transformers. We have developed annotators for Google T5 (Text-To-Text Transfer Transformer) and MarianMNT for Neural Machine Translation with over 646 pretrained models and pipelines.
This release also comes with a refactored and brand new models for language detection and identification. They are more accurate, faster, and support up to 375 languages.
The 2.7.0 release has over 720+ new pretrained models and pipelines while extending our support of multi-lingual models to 192+ languages such as Chinese, Japanese, Korean, Arabic, Persian, Urdu, and Hebrew.
As always, we would like to thank our community for their feedback and support.
The 2.7.0 release comes with over 720+ new pretrained models and pipelines available for Windows, Linux, and macOS users.
| Model | Name | Build | Lang |
|---|---|---|---|
| T5Transformer | google_t5_small_ssm_nq |
2.7.0 | en |
| T5Transformer | t5_small |
2.7.0 | en |
| MarianTransformer | opus_mt_en_aav |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_af |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_afa |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_alv |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ar |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_az |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bat |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bcl |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bem |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ber |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bg |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bi |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bnt |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_bzs |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ca |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ceb |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cel |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_chk |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cpf |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cpp |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_crs |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cs |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cus |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_cy |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_da |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_de |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_dra |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ee |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_efi |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_el |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_eo |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_es |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_et |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_eu |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_euq |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_fi |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_fiu |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_fj |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_fr |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ga |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gaa |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gem |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gil |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gl |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gmq |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gmw |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_grk |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_guw |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_gv |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ha |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_he |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_hi |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_hil |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ho |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ht |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_hu |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_hy |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_id |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ig |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_iir |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ilo |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_inc |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ine |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_is |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_iso |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_it |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_itc |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_jap |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_kg |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_kj |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_kqn |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_kwn |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_kwy |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lg |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ln |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_loz |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lu |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lua |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lue |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lun |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_luo |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_lus |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_map |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mfe |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mg |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mh |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mk |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mkh |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ml |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mos |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mr |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mt |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_mul |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ng |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_nic |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_niu |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_nl |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_nso |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_ny |
2.7.0 | xx |
| MarianTransformer | opus_mt_en_nyk |
2.7.0 | xx |
| Models | Name | Build | Lang |
|---|---|---|---|
| WordSegmenterModel | wordseg_weibo |
2.7.0 | zh |
| WordSegmenterModel | wordseg_pku |
2.7.0 | zh |
| WordSegmenterModel | wordseg_msra |
2.7.0 | zh |
| WordSegmenterModel | wordseg_large |
2.7.0 | zh |
| WordSegmenterModel | wordseg_ctb9 |
2.7.0 | zh |
| PerceptronModel | pos_ud_gsd |
2.7.0 | zh |
| PerceptronModel | pos_ctb9 |
2.7.0 | zh |
| NerDLModel | ner_msra_bert_768d |
2.7.0 | zh |
| NerDLModel | ner_weibo_bert_768d |
2.7.0 | zh |
| Models | Name | Build | Lang |
|---|---|---|---|
| StopWordsCleaner | stopwords_ar |
2.7.0 | ar |
| LemmatizerModel | lemma |
2.7.0 | ar |
| PerceptronModel | pos_ud_padt |
2.7.0 | ar |
| WordEmbeddingsModel | arabic_w2v_cc_300d |
2.7.0 | ar |
| NerDLModel | aner_cc_300d |
2.7.0 | ar |
| Models | Name | Build | Lang |
|---|---|---|---|
| StopWordsCleaner | stopwords_fa |
2.7.0 | fa |
| LemmatizerModel | lemma |
2.7.0 | fa |
| PerceptronModel | pos_ud_perdt |
2.7.0 | fa |
| WordEmbeddingsModel | persian_w2v_cc_300d |
2.7.0 | fa |
| NerDLModel | personer_cc_300d |
2.7.0 | fa |
The complete list of all 1100+ models & pipelines in 192+ languages is available on Models Hub.
Python
#PyPI
pip install spark-nlp==2.7.0
#Conda
conda install -c johnsnowlabs spark-nlp==2.7.0
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.7.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.7.0
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.7.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.7.0
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.7.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.7.0</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.7.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.7.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.7.0.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.7.0.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.7.0.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.7.0.jar
Nothing published for this version
We are glad to release Spark NLP 2.6.5! This release comes with a few bug fixes before we move to a new major version.
We are glad to release Spark NLP 2.6.5! This release comes with a few bug fixes before we move to a new major version.
As always, we would like to thank our community for their feedback, questions, and feature requests.
Python
#PyPI
pip install spark-nlp==2.6.5
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.5
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.5
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.5
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.5</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.5</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.5.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.5.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.6.5.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.5.jar
We are glad to release Spark NLP 2.6.4! This release comes with a few bug fixes before we move to a new major version.
We are glad to release Spark NLP 2.6.4! This release comes with a few bug fixes before we move to a new major version.
As always, we would like to thank our community for their feedback, questions, and feature requests.
Python
#PyPI
pip install spark-nlp==2.6.4
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.4
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.4
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.4
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.4</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.4</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.4.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.4.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/sjars/park-nlp-spark23-assembly-2.6.4.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.4.jar
We are glad to release Spark NLP 2.6.3! This release comes with a refactored NerDLApproach that allows users to train their NerDL on any size of the C
We are glad to release Spark NLP 2.6.3! This release comes with a refactored NerDLApproach that allows users to train their NerDL on any size of the CoNLL file regardless of the memory limitations. We also have some bug fixes and improvements in the 2.6.3 release.
Spark NLP has a new and improved Website for its documentation and models. We have been moving our 330+ pretrained models and pipelines into Models Hubs and we would appreciate your feedback! :)
As always, we would like to thank our community for their feedback, questions, and feature requests.
Python
#PyPI
pip install spark-nlp==2.6.3
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.3
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.3
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.3
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.3</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.3.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.3.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.6.3.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.3.jar
Nothing published for this version
Nothing published for this version
Nothing published for this version
DeepSentenceDetector is deprecated in favor of SentenceDetectorDL
We are glad to release Spark NLP 2.6.2! This release comes with a brand new SentenceDetectorDL (SDDL) that is based on a general-purpose neural network model for sentence boundary detection with higher accuracy. In addition, we are releasing 12 new and improved BioBERT models for BertEmbeddings and BertSentenceEembeddings used for sequence and text classifications.
Spark NLP has a new and improved Website for its documentation and models. We have been moving our 330+ pretrained models and pipelines into Models Hubs and we would appreciate your feedback! :)
As always, we would like to thank our community for their feedback, questions, and feature requests.
| Model | Name | Build | Lang |
|---|---|---|---|
| BertEmbeddings | biobert_pubmed_base_cased |
2.6.2 | en |
| BertEmbeddings | biobert_pubmed_large_cased |
2.6.2 | en |
| BertEmbeddings | biobert_pmc_base_cased |
2.6.2 | en |
| BertEmbeddings | biobert_pubmed_pmc_base_cased |
2.6.2 | en |
| BertEmbeddings | biobert_clinical_base_cased |
2.6.2 | en |
| BertEmbeddings | biobert_discharge_base_cased |
2.6.2 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_base_cased |
2.6.2 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_large_cased |
2.6.2 | en |
| BertSentenceEmbeddings | sent_biobert_pmc_base_cased |
2.6.2 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_pmc_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_clinical_base_cased |
2.6.2 | en |
| BertSentenceEmbeddings | sent_biobert_discharge_base_cased |
2.6.2 | en |
The complete list of all 330+ models & pipelines in 46+ languages is available here.
Python
#PyPI
pip install spark-nlp==2.6.2
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.2
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.2
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.2
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.2</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.2.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.2.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.6.2.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.2.jar
========
We are glad to release Spark NLP 2.6.1! This release comes with new Portuguese BERT models, a notebook to demonstrate how to import any BERT models to
We are glad to release Spark NLP 2.6.1! This release comes with new Portuguese BERT models, a notebook to demonstrate how to import any BERT models to Spark NLP, and a fix for ClassifierDL which was introduced in the 2.6.0 release that resulted in low accuracy during training.
As always, we would like to thank our community for their feedback, questions, and feature requests.
| Model | Name | Build | Lang |
|---|---|---|---|
| BertEmbeddings | bert_portuguese_base_cased |
2.6.0 | pt |
| BertEmbeddings | bert_portuguese_large_cased |
2.6.0 | pt |
The complete list of all 330+ models & pipelines in 46+ languages is available here.
Python
#PyPI
pip install spark-nlp==2.6.1
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.1
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.1
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.1
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.1</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.1.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.1.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.6.1.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.1.jar
========
We are very excited to finally release Spark NLP 2.6.0! This has been one of the biggest releases we have ever made and we are so proud to share it wi
We are very excited to finally release Spark NLP 2.6.0! This has been one of the biggest releases we have ever made and we are so proud to share it with our community!
This release comes with a brand new MultiClassifierDL for multi-label text classification, BertSentenceEmbeddings with 42 models, unsupervised keyword extractions annotator, and adding 28 new pretrained Transformers such as Small BERT, CovidBERT, ELECTRA, and the state-of-the-art language-agnostic BERT Sentence Embedding model(LaBSE).
The 2.6.0 release has over 110 new pretrained models, pipelines, and Transformers with extending full support for Danish, Finnish, and Swedish languages.
This release comes with over 100+ new pretrained models and pipelines available for Windows, Linux, and macOS users.
The complete list of all 330+ models & pipelines in 46+ languages is available here.
| Model | Name | Build | Lang |
|---|---|---|---|
| BertEmbeddings | electra_small_uncased |
2.6.0 | en |
| BertEmbeddings | electra_base_uncased |
2.6.0 | en |
| BertEmbeddings | electra_large_uncased |
2.6.0 | en |
| BertEmbeddings | covidbert_large_uncased |
2.6.0 | en |
| BertEmbeddings | small_bert_L2_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L4_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L6_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L8_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L10_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L12_128 |
2.6.0 | en |
| BertEmbeddings | small_bert_L2_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L4_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L6_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L8_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L10_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L12_256 |
2.6.0 | en |
| BertEmbeddings | small_bert_L2_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L4_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L6_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L8_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L10_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L12_512 |
2.6.0 | en |
| BertEmbeddings | small_bert_L2_768 |
2.6.0 | en |
| BertEmbeddings | small_bert_L4_768 |
2.6.0 | en |
| BertEmbeddings | small_bert_L6_768 |
2.6.0 | en |
| BertEmbeddings | small_bert_L8_768 |
2.6.0 | en |
| BertEmbeddings | small_bert_L10_768 |
2.6.0 | en |
| BertEmbeddings | small_bert_L12_768 |
2.6.0 | en |
| BertEmbeddings | bert_finnish_cased |
2.6.0 | fi |
| BertEmbeddings | bert_finnish_uncased |
2.6.0 | fi |
| BertSentenceEmbeddings | sent_bert_finnish_cased |
2.6.0 | fi |
| BertSentenceEmbeddings | sent_bert_finnish_uncased |
2.6.0 | fi |
| BertSentenceEmbeddings | sent_electra_small_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_electra_base_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_electra_large_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_bert_base_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_bert_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_bert_large_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_bert_large_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_large_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_pmc_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_pubmed_pmc_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_clinical_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_biobert_discharge_base_cased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_covidbert_large_uncased |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L2_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L4_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L6_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L8_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L10_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L12_128 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L2_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L4_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L6_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L8_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L10_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L12_256 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L2_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L4_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L6_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L8_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L10_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L12_512 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L2_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L4_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L6_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L8_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L10_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_small_bert_L12_768 |
2.6.0 | en |
| BertSentenceEmbeddings | sent_bert_multi_cased |
2.6.0 | xx |
| BertSentenceEmbeddings | labse |
2.6.0 | xx |
| Pipeline | Name | Build | Lang |
|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.6.0 | da |
| Explain Document Medium | explain_document_md |
2.6.0 | da |
| Explain Document Large | explain_document_lg |
2.6.0 | da |
| Entity Recognizer Small | entity_recognizer_sm |
2.6.0 | da |
| Entity Recognizer Medium | entity_recognizer_md |
2.6.0 | da |
| Entity Recognizer Large | entity_recognizer_lg |
2.6.0 | da |
| Pipeline | Name | Build | Lang |
|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.6.0 | fi |
| Explain Document Medium | explain_document_md |
2.6.0 | fi |
| Explain Document Large | explain_document_lg |
2.6.0 | fi |
| Entity Recognizer Small | entity_recognizer_sm |
2.6.0 | fi |
| Entity Recognizer Medium | entity_recognizer_md |
2.6.0 | fi |
| Entity Recognizer Large | entity_recognizer_lg |
2.6.0 | fi |
| Pipeline | Name | Build | Lang |
|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.6.0 | sv |
| Explain Document Medium | explain_document_md |
2.6.0 | sv |
| Explain Document Large | explain_document_lg |
2.6.0 | sv |
| Entity Recognizer Small | entity_recognizer_sm |
2.6.0 | sv |
| Entity Recognizer Medium | entity_recognizer_md |
2.6.0 | sv |
| Entity Recognizer Large | entity_recognizer_lg |
2.6.0 | sv |
Python
#PyPI
pip install spark-nlp==2.6.0
#Conda
conda install -c johnsnowlabs spark-nlp==2.6.0
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.6.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.6.0
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.6.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.6.0
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.6.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.6.0</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.6.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.6.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.6.0.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.6.0.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.6.0.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.6.0.jar
Nothing published for this version
Nothing published for this version
Nothing published for this version
We are excited to release Spark NLP 2.5.5 with 28 new pretrained models for Lemma and POS in 14 languages, bug fixes, new notebooks, and more!
We are excited to release Spark NLP 2.5.5 with 28 new pretrained models for Lemma and POS in 14 languages, bug fixes, new notebooks, and more!
As always, we would like to thank our community for their feedback, questions, and feature requests.
Example:
ner_model = NerDLModel.pretrained('onto_100')
print(ner_model.getClasses())
#['O', 'B-CARDINAL', 'B-EVENT', 'I-EVENT', 'B-WORK_OF_ART', 'I-WORK_OF_ART', 'B-ORG', 'B-DATE', 'I-DATE', 'I-ORG', 'B-GPE', 'B-PERSON', 'B-PRODUCT', 'B-NORP', 'B-ORDINAL', 'I-PERSON', 'B-MONEY', 'I-MONEY', 'I-GPE', 'B-LOC', 'I-LOC', 'I-CARDINAL', 'B-FAC', 'I-FAC', 'B-LAW', 'I-LAW', 'B-TIME', 'I-TIME', 'B-PERCENT', 'I-PERCENT', 'I-NORP', 'I-PRODUCT', 'B-QUANTITY', 'I-QUANTITY', 'B-LANGUAGE', 'I-ORDINAL', 'I-LANGUAGE', 'X']
| Model | Name | Build | Lang |
|---|---|---|---|
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | br |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | ca |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | da |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | ga |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | hi |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | hy |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | eu |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | mr |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | yo |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | la |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | lv |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | sl |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | gl |
| LemmatizerModel (Lemmatizer) | lemma |
2.5.5 | id |
| PerceptronModel (POS UD) | pos_ud_keb |
2.5.5 | br |
| PerceptronModel (POS UD) | pos_ud_ancora |
2.5.5 | ca |
| PerceptronModel (POS UD) | pos_ud_ddt |
2.5.5 | da |
| PerceptronModel (POS UD) | pos_ud_idt |
2.5.5 | ga |
| PerceptronModel (POS UD) | pos_ud_hdtb |
2.5.5 | hi |
| PerceptronModel (POS UD) | pos_ud_armtdp |
2.5.5 | hy |
| PerceptronModel (POS UD) | pos_ud_bdt |
2.5.5 | eu |
| PerceptronModel (POS UD) | pos_ud_ufal |
2.5.5 | mr |
| PerceptronModel (POS UD) | pos_ud_ytb |
2.5.5 | yo |
| PerceptronModel (POS UD) | pos_ud_llct |
2.5.5 | la |
| PerceptronModel (POS UD) | pos_ud_lvtb |
2.5.5 | lv |
| PerceptronModel (POS UD) | pos_ud_ssj |
2.5.5 | sl |
| PerceptronModel (POS UD) | pos_ud_treegal |
2.5.5 | gl |
| PerceptronModel (POS UD) | pos_ud_gsd |
2.5.5 | id |
Languages: Armenian, Basque, Breton, Catalan, Danish, Galician, Hindi, Indonesian, Irish, Latin, Latvian, Marathi, Slovenian, Yoruba
Python
#PyPI
pip install spark-nlp==2.5.5
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.5
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.5.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.5.5
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.5.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.5.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.5.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.5.5
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.5.5</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.5.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.5.5</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.5.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.5.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.5.5.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.5.5.jar
We are excited to release Spark NLP 2.5.4 with the full support of Apache Spark 2.3.x, adding 43 new pre-trained models for stop words cleaning, suppo
We are excited to release Spark NLP 2.5.4 with the full support of Apache Spark 2.3.x, adding 43 new pre-trained models for stop words cleaning, supporting 26 new languages, a new RegexTokenizer annotator and more!
As always, we would like to thank our community for their feedback, questions, and feature requests.
spark23 to start() function to start the session for Apache Spark 2.3.x| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_af |
2.5.4 | af |
Download |
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_ar |
2.5.4 | ar |
Download |
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_hy |
2.5.4 | hy |
Download |
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_eu |
2.5.4 | eu |
Download |
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_bn |
2.5.4 | bn |
Download |
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| StopWordsCleaner | stopwords_br |
2.5.4 | br |
Download |
Python
#PyPI
pip install spark-nlp==2.5.4
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.4
Spark
spark-nlp on Apache Spark 2.4.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.5.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.11:2.5.4
spark-nlp on Apache Spark 2.3.x:
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.5.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23_2.11:2.5.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.5.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark23-gpu_2.11:2.5.4
Maven
spark-nlp on Apache Spark 2.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.11</artifactId>
<version>2.5.4</version>
</dependency>
spark-nlp on Apache Spark 2.3.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-spark23_2.11</artifactId>
<version>2.5.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu-spark23_2.11</artifactId>
<version>2.5.4</version>
</dependency>
FAT JARs
CPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.4.jar
GPU on Apache Spark 2.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.4.jar
CPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-assembly-2.5.4.jar
GPU on Apache Spark 2.3.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-spark23-gpu-assembly-2.5.4.jar
We are very happy to release Spark NLP 2.5.3 with 5 new pre-trained ClassifierDL models for multi-class text classification. There are also bug-fixes
We are very happy to release Spark NLP 2.5.3 with 5 new pre-trained ClassifierDL models for multi-class text classification. There are also bug-fixes and other enhancements introduced in this release which were reported and requested by Spark NLP users.
As always, we thank our community for their feedback, questions, and feature requests.
We have added 5 new pre-trained ClassifierDL models for multi-class text classification.
| Model | Name | Build | Lang | Description | Offline |
|---|---|---|---|---|---|
| ClassifierDLModel | classifierdl_use_spam |
2.5.3 | en |
Detect if a message is spam or not | Download |
| ClassifierDLModel | classifierdl_use_fakenews |
2.5.3 | en |
Classify if a news is fake or real | Download |
| ClassifierDLModel | classifierdl_use_emotion |
2.5.3 | en |
Detect Emotions in TweetsDetect Emotions in Tweets | Download |
| ClassifierDLModel | classifierdl_use_cyberbullying |
2.5.3 | en |
Classify if a tweet is bullying | Download |
| ClassifierDLModel | classifierdl_use_sarcasm |
2.5.3 | en |
Identify sarcastic tweets | Download |
Python
#PyPI
pip install spark-nlp==2.5.3
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.3
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.3
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.3
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.3</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.3.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.3.jar
We are very happy to release Spark NLP 2.5.2 with a new state-of-the-art LanguageDetectorDL annotator to detect and identify up to 20 languages. There
We are very happy to release Spark NLP 2.5.2 with a new state-of-the-art LanguageDetectorDL annotator to detect and identify up to 20 languages. There are also bug-fixes and other enhancements introduced in this release which were reported and requested by Spark NLP users.
As always, we thank our community for their feedback, questions, and feature requests.
We have added 4 new LanguageDetectorDL models and pipelines to detect and identify up to 20 languages:
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| LanguageDetectorDL | ld_wiki_7 |
2.5.2 | xx |
Download |
| LanguageDetectorDL | ld_wiki_20 |
2.5.2 | xx |
Download |
| Pipeline | Name | Build | Lang | Offline |
|---|---|---|---|---|
| LanguageDetectorDL | detect_language_7 |
2.5.2 | xx |
Download |
| LanguageDetectorDL | detect_language_20 |
2.5.2 | xx |
Download |
Python
#PyPI
pip install spark-nlp==2.5.2
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.2
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.2
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.2
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.2</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.2.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.2.jar
We are very excited to extend Spark NLP support to 6 new BERT models for medical and clinical documents. We have also updated our documentation for 2.
We are very excited to extend Spark NLP support to 6 new BERT models for medical and clinical documents. We have also updated our documentation for 2.5.x releases, notebooks in our workshop, and made some enhancements in this release.
As always, we thank our community for their feedback and questions in our Slack channel.
We have added 6 new BERT models for medical and clinical purposes. The 4 BERT pre-trained models are from BioBERT and the other 2 are coming from ClinicalBERT models:
| Model | Name | Build | Lang | Offline |
|---|---|---|---|---|
| BertEmbeddings | biobert_pubmed_base_cased |
2.5.0 | en |
Download |
| BertEmbeddings | biobert_pubmed_large_cased |
2.5.0 | en |
Download |
| BertEmbeddings | biobert_pmc_base_cased |
2.5.0 | en |
Download |
| BertEmbeddings | biobert_pubmed_pmc_base_cased |
2.5.0 | en |
Download |
| BertEmbeddings | biobert_clinical_base_cased |
2.5.0 | en |
Download |
| BertEmbeddings | biobert_discharge_base_cased |
2.5.0 | en |
Download |
Python
#PyPI
pip install spark-nlp==2.5.1
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.1
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.1
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.1
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.1</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.1.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.1.jar
When we started planning for Spark NLP 2.5.0 release a few months ago the world was a different place!
When we started planning for Spark NLP 2.5.0 release a few months ago the world was a different place!
We have been blown away by the use of Natural Language Processing for early outbreak detections, question-answering chatbot services, text analysis of medical records, monitoring efforts to minimize the virus spread, and many more.
In that spirit, we are honored to announce Spark NLP 2.5.0 release! Witnessing the world coming together to fight coronavirus has driven us to deliver perhaps one of the biggest releases we have ever made.
As always, we thank our community for their feedback, bug reports, and contributions that made this release possible.
Spark NLP 2.5.0 comes with 87 new pretrained models and pipelines in 14 new languages available for all Windows, Linux, and macOS users. We added new languages such as Dutch, Norwegian. Polish, Portuguese, Bulgarian, Czech, Greek, Finnish, Hungarian, Romanian, Slovak, Swedish, Turkish, and Ukrainian.
The complete list of 160+ models & pipelines in 22+ languages is available here.
| Pipeline | Name | Build | lang | Description | Offline |
|---|---|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.5.0 | nl |
Download | |
| Explain Document Medium | explain_document_md |
2.5.0 | nl |
Download | |
| Explain Document Large | explain_document_lg |
2.5.0 | nl |
Download | |
| Entity Recognizer Small | entity_recognizer_sm |
2.5.0 | nl |
Download | |
| Entity Recognizer Medium | entity_recognizer_md |
2.5.0 | nl |
Download | |
| Entity Recognizer Large | entity_recognizer_lg |
2.5.0 | nl |
Download |
| Pipeline | Name | Build | lang | Description | Offline |
|---|---|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.5.0 | no |
Download | |
| Explain Document Medium | explain_document_md |
2.5.0 | no |
Download | |
| Explain Document Large | explain_document_lg |
2.5.0 | no |
Download | |
| Entity Recognizer Small | entity_recognizer_sm |
2.5.0 | no |
Download | |
| Entity Recognizer Medium | entity_recognizer_md |
2.5.0 | no |
Download | |
| Entity Recognizer Large | entity_recognizer_lg |
2.5.0 | no |
Download |
| Pipeline | Name | Build | lang | Description | Offline |
|---|---|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.5.0 | pl |
Download | |
| Explain Document Medium | explain_document_md |
2.5.0 | pl |
Download | |
| Explain Document Large | explain_document_lg |
2.5.0 | pl |
Download | |
| Entity Recognizer Small | entity_recognizer_sm |
2.5.0 | pl |
Download | |
| Entity Recognizer Medium | entity_recognizer_md |
2.5.0 | pl |
Download | |
| Entity Recognizer Large | entity_recognizer_lg |
2.5.0 | pl |
Download |
| Pipeline | Name | Build | lang | Description | Offline |
|---|---|---|---|---|---|
| Explain Document Small | explain_document_sm |
2.5.0 | pt |
Download | |
| Explain Document Medium | explain_document_md |
2.5.0 | pt |
Download | |
| Explain Document Large | explain_document_lg |
2.5.0 | pt |
Download | |
| Entity Recognizer Small | entity_recognizer_sm |
2.5.0 | pt |
Download | |
| Entity Recognizer Medium | entity_recognizer_md |
2.5.0 | pt |
Download | |
| Entity Recognizer Large | entity_recognizer_lg |
2.5.0 | pt |
Download |
Python
#PyPI
pip install spark-nlp==2.5.0
#Conda
conda install -c johnsnowlabs spark-nlp==2.5.0
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.0
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.5.0
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.5.0</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.5.0.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.5.0.jar
Nothing published for this version
Nothing published for this version
We are very excited to extend Spark NLP support to 6 new Databricks runtimes and add support to Cloudera and EMR YARN cluster-mode. As always, we than
We are very excited to extend Spark NLP support to 6 new Databricks runtimes and add support to Cloudera and EMR YARN cluster-mode. As always, we thank our community for their feedback and questions in our Slack channel.
Python
#PyPI
pip install spark-nlp==2.4.5
#Conda
conda install -c johnsnowlabs spark-nlp==2.4.5
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.5
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.5
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.4.5</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.4.5.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.4.5.jar
We are very excited to release the very first multi-class text classifier in Spark NLP v2.4.4! We have built a generic ClassifierDL annotator that use
NOTE: ClassifierDL is an experimental feature in 2.4.4 before it becomes stable in 2.4.5 release. We have worked hard to aim for simplicity and we are looking forward to your feedback as always. We will add more examples by the upcoming days:
Models:
| Model | name | language |
|---|---|---|
| LemmatizerModel (Lemmatizer) | lemma |
ru |
| PerceptronModel (POS UD) | pos_ud_gsd |
ru |
| NerDLModel | wikiner_6B_100 |
ru |
| NerDLModel | wikiner_6B_300 |
ru |
| NerDLModel | wikiner_840B_300 |
ru |
Pipelines:
| Pipeline | name | language |
|---|---|---|
| Explain Document (Small) | explain_document_sm |
ru |
| Explain Document (Medium) | explain_document_md |
ru |
| Explain Document (Large) | explain_document_lg |
ru |
| Entity Recognizer (Small) | entity_recognizer_sm |
ru |
| Entity Recognizer (Medium) | entity_recognizer_md |
ru |
| Entity Recognizer (Large) | entity_recognizer_lg |
ru |
Evaluation:
wikiner_6B_100 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 97.76% | 88.85% | 88.55% | 88.70 |
wikiner_6B_300 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 97.78% | 89.09% | 88.51% | 88.80 |
wikiner_840B_300 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 97.85% | 89.85% | 89.11% | 89.48 |
import com.johnsnowlabs.nlp.pretrained.PretrainedPipeline
val pipeline = PretrainedPipeline("explain_document_sm", lang="ru")
val testData = spark.createDataFrame(Seq(
(1, "Пик распространения коронавируса и вызываемой им болезни Covid-19 в Китае прошел, заявил в четверг агентству Синьхуа официальный представитель Госкомитета по гигиене и здравоохранению КНР Ми Фэн.")
)).toDF("id", "text")
val annotation = pipeline.transform(testData)
annotation.show()
Python
#PyPI
pip install spark-nlp==2.4.4
#Conda
conda install -c johnsnowlabs spark-nlp==2.4.4
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.4
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.4
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.4.4</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.4.4.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.4.4.jar
NOTE: ClassifierDL is an experimental feature in 2.4.4 release. We have worked hard to aim for simplicity and we are looking forward to your feedback as always.
ClassifierDL. This annotator can train any dataset from 2 up to 50 classes.========
Nothing published for this version
This minor release fixes a bug on our Python side that was introduced in 2.4.2 release. As always, we thank our community for their feedback and quest
This minor release fixes a bug on our Python side that was introduced in 2.4.2 release. As always, we thank our community for their feedback and questions in our Slack channel.
NOTE: We highly recommend our Python users to update to 2.4.3 release.
pip install spark-nlp==2.4.3
conda install -c johnsnowlabs spark-nlp==2.4.3
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.3
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.4.3</version>
</dependency>
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.4.3.jar GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.4.3.jar
This minor release fixes a few bugs in some of our annotators reported by our community. As always, we thank our community for their feedback and ques
This minor release fixes a few bugs in some of our annotators reported by our community. As always, we thank our community for their feedback and questions in our Slack channel.
pip install spark-nlp==2.4.2
conda install -c johnsnowlabs spark-nlp==2.4.2
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.2
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.4.2</version>
</dependency>
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.4.2.jar GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.4.2.jar
This minor release fixes a few bugs in some of the annotators reported by our community. As always, we thank our community for their feedback on our S
This minor release fixes a few bugs in some of the annotators reported by our community. As always, we thank our community for their feedback on our Slack channel.
Models:
| Model | name | language |
|---|---|---|
| LemmatizerModel (Lemmatizer) | lemma |
es |
| PerceptronModel (POS UD) | pos_ud_gsd |
es |
| NerDLModel | wikiner_6B_100 |
es |
| NerDLModel | wikiner_6B_300 |
es |
| NerDLModel | wikiner_840B_300 |
es |
Pipelines:
| Pipeline | name | language |
|---|---|---|
| Explain Document (Small) | explain_document_sm |
es |
| Explain Document (Medium) | explain_document_md |
es |
| Explain Document (Large) | explain_document_lg |
es |
| Entity Recognizer (Small) | entity_recognizer_sm |
es |
| Entity Recognizer (Medium) | entity_recognizer_md |
es |
| Entity Recognizer (Large) | entity_recognizer_lg |
es |
Evaluation:
wikiner_6B_100 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 98.35% | 88.97% | 88.64% | 88.80 |
wikiner_6B_300 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 98.38% | 89.42% | 89.03% | 89.22 |
wikiner_840B_300 with conlleval.pl
| Accuracy | Precision | Recall | F1-Score |
|---|---|---|---|
| 98.46% | 89.74% | 89.43% | 89.58 |
import com.johnsnowlabs.nlp.pretrained.PretrainedPipeline
val pipeline = PretrainedPipeline("explain_document_sm", lang="es")
val testData = spark.createDataFrame(Seq(
(1, "Ésta se convertiría en una amistad de por vida, y Peleo, conociendo la sabiduría de Quirón , más adelante le confiaría la educación de su hijo Aquiles."),
(2, "Durante algo más de 200 años el territorio de la actual Bolivia constituyó la Real Audiencia de Charcas, uno de los centros más prósperos y densamente poblados de los virreinatos españoles.")
)).toDF("id", "text")
val annotation = pipeline.transform(testData)
annotation.show()
More info on pre-trained models and pipelines
We are very excited to finally release Spark NLP v2.4.0! This has been one of the largest releases we have ever made since the inception of the librar
We are very excited to finally release Spark NLP v2.4.0! This has been one of the largest releases we have ever made since the inception of the library! The new release of Spark NLP 2.4.0 has been migrated to TensorFlow 1.15.0 which takes advantage of the latest deep learning technologies and pre-trained models.
Python
#PyPI
pip install spark-nlp==2.4.0
#Conda
conda install -c johnsnowlabs spark-nlp==2.4.0
Spark
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.0
PySpark
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.11:2.4.0
Maven
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.11</artifactId>
<version>2.4.0</version>
</dependency>
FAT JARs
CPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-2.4.0.jar
GPU: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-2.4.0.jar
Storage. Allows any annotator to have it's own distributed local index databaseSpark NLP 2.4.0 comes with new models including Universal Sentence Encoder, BERT, and Elmo models from TF Hub. In addition, our multilingual pipelines are now available for Windows as same as Linux and macOS users.
| Models | Name |
|---|---|
| UniversalSentenceEncoder | tf_use |
| UniversalSentenceEncoder | tf_use_lg |
| BertEmbeddings | bert_large_cased |
| BertEmbeddings | bert_large_uncased |
| BertEmbeddings | bert_base_cased |
| BertEmbeddings | bert_base_uncased |
| BertEmbeddings | bert_multi_cased |
| ElmoEmbeddings | elmo |
| NerDLModel | onto_100 |
| NerDLModel | onto_300 |
| Pipelines | Name | Language |
|---|---|---|
| Explain Document Large | explain_document_lg |
fr |
| Explain Document Medium | explain_document_md |
fr |
| Entity Recognizer Large | entity_recognizer_lg |
fr |
| Entity Recognizer Medium | entity_recognizer_md |
fr |
| Explain Document Large | explain_document_lg |
de |
| Explain Document Medium | explain_document_md |
de |
| Entity Recognizer Large | entity_recognizer_lg |
de |
| Entity Recognizer Medium | entity_recognizer_md |
de |
| Explain Document Large | explain_document_lg |
it |
| Explain Document Medium | explain_document_md |
it |
| Entity Recognizer Large | entity_recognizer_lg |
it |
| Entity Recognizer Medium | entity_recognizer_md |
it |
Example:
# Import Spark NLP
from sparknlp.base import *
from sparknlp.annotator import *
from sparknlp.pretrained import PretrainedPipeline
import sparknlp
# Start Spark Session with Spark NLP
# If you already have a SparkSession (Zeppelin, Databricks, etc.)
# you can skip this
spark = sparknlp.start()
# Download a pre-trained pipeline
pipeline = PretrainedPipeline('explain_document_md', lang='fr')
# Your testing dataset
text = """
Emmanuel Jean-Michel Frédéric Macron est le fils de Jean-Michel Macron, né en 1950, médecin, professeur de neurologie au CHU d'Amiens4 et responsable d'enseignement à la faculté de médecine de cette même ville5, et de Françoise Noguès, médecin conseil à la Sécurité sociale.
"""
# Annotate your testing dataset
result = pipeline.annotate(text)
# What's in the pipeline
list(result.keys())
# result:
# ['entities', 'lemma', 'document', 'pos', 'token', 'ner', 'embeddings', 'sentence']
# Check the results
result['entities']
# entities:
# ['Emmanuel Jean-Michel Frédéric Macron', 'Jean-Michel Macron', "CHU d'Amiens4", 'Françoise Noguès', 'Sécurité sociale']
Please note that in 2.4.0 we have added storageRef parameter to our WordEmbeddogs. This means every WordEmbeddingsModel will now have storageRef which is also bound to NerDLModel trained by that embeddings.
This assures users won't use a NerDLModel with a wrong WordEmbeddingsModel.
Example:
val embeddings = new WordEmbeddings()
.setStoragePath("/tmp/glove.6B.100d.txt", ReadAs.TEXT)
.setDimension(100)
.setStorageRef("glove_100d") // Use or save this WordEmbeddings with storageRef
.setInputCols("document", "token")
.setOutputCol("embeddings")
If you save theWordEmbeddings model the storageRef will be glove_100d. If you ever train any NerDLApproach the glove_100d will bind to that NerDLModel.
If you have already WordEmbeddingsModels saved from earlier versions, you either need to re-save them with storageRed or you can manually add this param in their metadata/. The same advice works for the NerDLModel from earlier versions.
We are very excited to finally release Spark NLP v2.4.0! This has been one of the largest releases we have ever made since the inception of the library!
The new release of Spark NLP 2.4.0 has been migrated to TensorFlow 1.15.0 which takes advantage of the latest deep learning technologies and pre-trained models.
As always, thanks to the community for the feedback and questions in our Slack channel.
Please beware as this release breaks backwards compatibility with previously saved models, particularly on Tensorflow and Embeddings, aside from code-breaking changes in the API.
We will be working in our documentation to enhance the learning curve.
Storage. Allows any annotator to have it's own distributed local index database========
Nothing published for this version
This minor release fixes a bug in ChunkEmbeddings causing an out of boundaries exception in some scenarios. We also switch to maven coordinates as def
This minor release fixes a bug in ChunkEmbeddings causing an out of boundaries exception in some scenarios. We also switch to maven coordinates as default source for start() function since spark-packages has not been responsive on their package approval process. Thank you all for your consistent feedback.
We would like to thank you all for your valuable feedback via our Slack and our GitHub repositories. Spark NLP 2.3.4 was a very stable and rock-solid
We would like to thank you all for your valuable feedback via our Slack and our GitHub repositories.
Spark NLP 2.3.4 was a very stable and rock-solid release. However, we wanted to release 2.3.5 to fix the few remaining minor bugs before moving to our bigger release 2.4.0!
We would like to thank you all for your valuable feedback via our Slack channels and our GitHub repositories.
Spark NLP 2.3.4 is a very stable and rock-solid release. However, we wanted to fix the few remaining minor bugs before moving to our bigger release 2.4.0!
========
Thank you, as always, for the feedback given at Slack and our repos. The most important part of this release is how we internally started organizing m
Thank you, as always, for the feedback given at Slack and our repos. The most important part of this release is how we internally started organizing models. We'll be deploying our model news in https://github.com/JohnSnowLabs/spark-nlp-models . The model's repo will be kept up to date.
As for this release, it improves various internal API functionalities, allowing for positive side-effects across the library. As an important enhancement, we have added user UDFs and functions for both Scala and Python users to be able to easily manipulate annotations on DataFrames. Finally, we have fixed various bugs in embeddings metadata to make sure we provide accurate offsetting information for other annotators to consume it successfully.
map_annotations and filter_by_annotationsFixed bad deprecated OCR and SpellChecker python classpath
We are very glad to announce this release, it actually ended up much bigger than we expected. Thanks to the community feedback, we arranged many bugfixes. We also spent some times and started building models for the TextMatcher, so it got various improvements and bugfixes when dealing with empty sentences or cleaned up tokens. We also added UDF ready functions in Python to easily deal with Annotations. Finally, we fixed a few bugs when loading models from disk. Thank you very much for constant feedback on Slack.
mergeOverlapping allows for handling overlapping output chunks when matching entities share keywordsmap_annotations, map_annotations_strict, map_annotations_col, filter_by_annotations_col and explode_annotations_col functions to python side. Allows dealing with Annotations easily.This release addresses multiple bug fixes and some enhancements regarding memory consumption in our BertEmbeddings.
This release addresses multiple bug fixes and some enhancements regarding memory consumption in our BertEmbeddings.
This release addresses multiple bug fixes and some enhancements regarding memory consumption in BertEmbeddings annotator. Thanks for your feedback and reports!
========
This quick release addresses a bug in Lemmatizer loading/pretrained function causing it not to work in 2.3.0. We took the chance to include a feature
This quick release addresses a bug in Lemmatizer loading/pretrained function causing it not to work in 2.3.0. We took the chance to include a feature which did not make it for base 2.3.0 and slightly changed protected variables for better Java API, also including a pretrained compatible function with Java. Thanks for the quick issue feedback again!
This quick release addresses a bug in Lemmatizer loading/pretrained function causing it not to work in 2.3.0. We took the chance to include a feature which did not make it for base 2.3.0 and slightly changed protected variables for better Java API, also including a pretrained compatible function with Java. Thanks for the quick issue feedback again!
========
…credential profiles. Unfortunately, we have deprecated Eval and OCR due to internal patents in some of the latest improvements John Snow Labs has cont…
Thanks for your contributions and feedback on Slack. This amazing release comes with many new features in the scope of the embeddings, allowing pipeline builders to retrieve embeddings for specific bodies of texts in any form given, from sentences to chunks or n-grams. We also worked a lot on making sure Spark NLP in Java works as intended. Finally, we improved the AWS profile's compatibility for frameworks that utilize multiple credential profiles. Unfortunately, we have deprecated Eval and OCR due to internal patents in some of the latest improvements John Snow Labs has contributed to.
Chunker, NGramGenerator, or NerConverter outputsYour coding agent can read these notes before it upgrades. Set up the MCP server →