NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4698 most downloaded on PyPI
John Snow Labs Spark NLP is a natural language processing library built on top of Apache Spark ML. It provides simple, performant & accurate NLP annotations for machine learning pipelines, that scale easily in a distributed environment.
Last release 11 days ago
23 Sep 2026
Ships fairly regularly
a new release about every 2 weeks
Nearly every release is documented
notes for 60 of the last 60 stable releases
Nothing withdrawn
no release was ever pulled
8 years old
196 releases · first in 2018
📢 Spark NLP 7.0.0: Welcome, Apache Spark 4! New support for Spark 4.x alongside Spark 3.x
Spark NLP is taking a major step forward with 7.0.0: support for Apache Spark 4.x and Scala 2.13, while preserving the Spark 3.x / Scala 2.12 line. This is more than a version-number update. The Spark 4 work covers build profiles, annotation and UDF compatibility, annotator behavior, Python package selection, CI, and deployment guidance. If your team is moving to Spark 4, Spark NLP now has a dedicated path for you; teams staying on Spark 3 retain their own build line.
And there is more to explore: compare embedding pairs, rerank text with a cross-encoder, write expressive token-matching rules, summarize long documents, or translate documents with GGUF models. These additions make 7.0.0 a significant update for both existing Spark NLP pipelines and new Spark 4 projects. 🎉
PairwiseVectorSimilarity and CrossEncoder. Score pairs of embeddings with cosine, dot product, or negative Euclidean distance, or use an ONNX cross-encoder to score a document/query pair together for reranking workflows. PairwiseVectorSimilarity / Spark 4 counterpart; CrossEncoder / Spark 4 counterpart.RuleBasedMatcher. Match token spans using rules over text and optional linguistic annotations, with CHUNK output for downstream pipeline stages. Spark 3 PR / Spark 4 PR.Summarization estimator offers LLM, encoder–decoder, and extractive modes; DocumentTranslator combines document reading, sentence segmentation, and GGUF-based translation in a pipeline stage. Summarization / Spark 4 counterpart; DocumentTranslator / Spark 4 counterpart.ZeroShotNer decoding, model runtime handling, and package selection. These are targeted changes, not a claim of universal performance or compatibility improvements.The defining change in 7.0.0 is an additional Spark 4 / Scala 2.13 support line. This required more than recompilation: the branch updates annotation/UDF encoder paths and annotators, introduces distinct Spark 4 build profiles, adjusts Python Maven-coordinate selection, and adds Spark 4 CI and setup documentation. Spark 3 / Scala 2.12 remains a separate line; this release is about adding a choice, not replacing one. Spark 4 encoder work · annotator fixes · build profiles · CI work.
The regular Spark 4 / Scala 2.13 profile compiles against Spark 4.0.1. A second spark400 profile targets unpatched Spark 4.0.0, which differs in the binary shape of Spark's Param class. According to the project's Spark 4 migration guide, Databricks Runtime 17.3 identifies as Spark 4.0.0 but already includes the relevant fix: it needs the regular _2.13 artifact, not the -spark400_2.13 artifact. A NoSuchMethodError involving Param construction is a sign to check this choice. Select by your actual runtime, not its version string alone.
The Python resolver now maps release-form Spark 3.x versions to Scala 2.12 artifacts, plain Spark 4.0.0 to -spark400_2.13, and later Spark 4.x release-form versions to _2.13. It also retains CPU, GPU, Apple Silicon, and AArch64 naming. For the patched Databricks 4.0.0 case, explicitly follow the guide's regular-artifact instruction rather than assuming automatic 4.0.0 routing will recognize the patch. Automatic coordinate routing does not prove runtime compatibility with every future Spark release.
If you are planning a Spark 4 rollout, start with the artifact-selection guide and the installation guidance below. Your existing Spark 3 workloads continue to require Scala 2.12 artifacts.
Retrieval workflows often need different scoring tools at different stages. 7.0.0 adds two complementary options: score already-computed embeddings with PairwiseVectorSimilarity, or score an already-paired query and passage jointly with CrossEncoder. Neither is an index or an automatic corpus search; both can slot into retrieval and reranking pipelines you control.
PairwiseVectorSimilarity consumes two SENTENCE_EMBEDDINGS columns on the same row and emits a score for each embedding pair as VECTOR_SIMILARITY. Choose cosine (default), dot product, or Euclidean scoring; the Euclidean option returns negative L2 distance, so higher scores consistently mean more similar vectors. With multiple embeddings in either column, it scores every cross-pair rather than automatically retrieving nearest neighbors across a corpus.CrossEncoder jointly encodes two DOCUMENT inputs per row with a single-logit ONNX model and emits a relevance CATEGORY score. This is a reranking/scoring stage for already paired text, not a vector-index or corpus-search implementation.RuleBasedMatcher accepts DOCUMENT and a base TOKEN column, with optional annotation columns for rule attributes such as POS, lemma, NER, dependency and metadata. Rules may be expressed as JSON/JSONL patterns; matched spans become CHUNK annotations. A follow-up fix preserves the base token column when optional attributes are inferred automatically.Two new document-oriented components bring generation and transformation closer to the Spark pipeline: Summarization for configurable summary methods and DocumentTranslator for file-to-translation workflows. They address distinct tasks and have different model requirements.
Summarization / SummarizationModel offer llm, encoder_decoder, and extractive methods with controls for summary length and long documents. fit() resolves a method-specific default model; plan for model download and runtime requirements. The implementation documents limitations: it is a DataFrame-level stage rather than a LightPipeline stage, and a fitted model should not be used for concurrent transform calls on the same instance. LLM input text can also influence its own prompt; this workflow does not eliminate prompt-injection risk.DocumentTranslator reads supported document inputs, segments them into sentences, translates with a GGUF model via llama.cpp, and returns translated DOCUMENT annotations. Translation has substantial model/context requirements; increasing decoding batch size reduces the per-sentence context budget. A later branch commit also corrects its Scala implementation and Java example; include the final fix in the release candidate.addFile collision affecting models with external ONNX data, plus configurable verbosity for fallback logging. #14834, #14839.silicon build variant includes arm64 OpenVINO runtime work. #14824, #14825.ZeroShotNer decoding on that path. #14827, #14828, #14846, #14855.NorvigSweetingModel. #14819, #14820.RuleBasedMatcher attribute inference. #14812, #14815.addFile collision. #14792, #14858, #14834.DocumentTranslator correction #14878. It is not part of the Spark 3/master tag; the main release narrative does not depend on this follow-up's inclusion in any particular published binary.Choose the artifact for your Spark and Scala runtime. Spark 3.x uses Scala 2.12; Spark 4.x uses Scala 2.13. Plain, unpatched Spark 4.0.0 needs the dedicated -spark400_2.13 artifact. The regular _2.13 artifact is for Spark 4.0.1 and the other supported Spark 4 versions listed in the installation guide. The patched Databricks Spark 4.0.0 runtime needs the regular _2.13 artifact; select that coordinate explicitly, following the Spark 4 artifact-selection guide. Do not pair a Spark 3 runtime with a Scala 2.13 artifact or vice versa.
PyPI has spark-nlp==7.0.0. Install the matching PySpark runtime in the same environment:
# Spark 3.x / Scala 2.12
pip install spark-nlp==7.0.0 pyspark==3.5.1
# Plain Spark 4.0.0 / Scala 2.13 (dedicated spark400 artifact)
pip install spark-nlp==7.0.0 pyspark==4.0.0
# Spark 4.0.1 / Scala 2.13 (regular Spark 4 artifact)
pip install spark-nlp==7.0.0 pyspark==4.0.1These are alternative environments, not three commands to run in one environment. sparknlp.start() chooses a Maven coordinate from PySpark's version for ordinary distributions. For the patched Databricks 4.0.0 runtime, select the regular _2.13 JAR using your platform's package configuration rather than relying on the default 4.0.0 resolver rule.
CPU examples (choose the line matching your Spark runtime):
# Spark 3.x / Scala 2.12
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:7.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:7.0.0
# Plain Spark 4.0.0 / Scala 2.13
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-spark400_2.13:7.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-spark400_2.13:7.0.0
# Spark 4.0.1 and the other validated Spark 4 versions / Scala 2.13
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.13:7.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.13:7.0.0GPU, Apple Silicon, and Linux AArch64 Maven artifacts are also published for each lane. Their base names are spark-nlp-gpu, spark-nlp-silicon, and spark-nlp-aarch64: append _2.12 for Spark 3; append -spark400_2.13 for plain Spark 4.0.0; append _2.13 for the regular Spark 4 lane. For example, GPU on plain Spark 4.0.0 is com.johnsnowlabs.nlp:spark-nlp-gpu-spark400_2.13:7.0.0.
For CPU, choose one artifact ID from this table in your project's Maven dependency:
| Spark runtime | Scala | Maven artifact ID |
|---|---|---|
| Spark 3.x | 2.12 | spark-nlp_2.12 |
| Plain Spark 4.0.0 | 2.13 | spark-nlp-spark400_2.13 |
| Spark 4.0.1 and the other validated Spark 4 versions; patched Databricks 4.0.0 | 2.13 | spark-nlp_2.13 |
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.13</artifactId>
<version>7.0.0</version>
</dependency>The XML example uses the regular Spark 4 lane. Replace only the artifactId for Spark 3 or plain Spark 4.0.0. For hardware variants, use the same suffix rules shown under Spark Packages.
Spark 3.x / Scala 2.12 fat assemblies are directly under public/jars/ (not public/jars/scala-2.12/):
| Spark 3 variant | Scala 2.12 fat assembly |
|---|---|
| CPU | spark-nlp-assembly-7.0.0.jar |
| GPU | spark-nlp-gpu-assembly-7.0.0.jar |
| Apple Silicon | spark-nlp-silicon-assembly-7.0.0.jar |
| AArch64 | spark-nlp-aarch64-assembly-7.0.0.jar |
Spark 4 / Scala 2.13 fat assemblies are instead under public/jars/scala-2.13/. Choose the spark400 column for plain Spark 4.0.0 and the regular spark4 column for Spark 4.0.1 and the supported later Spark 4 versions (or the patched Databricks 4.0.0 runtime):
| Variant | Plain Spark 4.0.0 (spark400) |
Regular Spark 4 (spark4) |
|---|---|---|
| CPU | spark-nlp-spark400-assembly-7.0.0.jar | spark-nlp-assembly-7.0.0.jar |
| GPU | spark-nlp-gpu-spark400-assembly-7.0.0.jar | spark-nlp-gpu-assembly-7.0.0.jar |
| Apple Silicon | spark-nlp-silicon-spark400-assembly-7.0.0.jar | spark-nlp-silicon-assembly-7.0.0.jar |
| AArch64 | spark-nlp-aarch64-spark400-assembly-7.0.0.jar | spark-nlp-aarch64-assembly-7.0.0.jar |
Fat JARs are not Maven JARs. The S3 files above are fat assemblies; Maven publishes separate regular JARs. For Spark 3 CPU, the Maven JAR is distinct from the S3 fat assembly. For plain Spark 4.0.0 CPU, the Maven JAR is distinct from the S3 fat assembly. None of these pairs is byte-identical or interchangeable by filename. The S3 prefix is part of the lane selection: spark-nlp-assembly-7.0.0.jar under public/jars/ is Spark 3; the same basename under public/jars/scala-2.13/ is regular Spark 4.
The 7.0.0 Git tag is on the Spark 3/master line; Spark 4 is published from its separate branch. Paired PRs below identify the work in each release lane rather than implying both lanes share a single source commit.
PairwiseVectorSimilarity #14787, RuleBasedMatcher #14788, CrossEncoder #14795, Summarization #14794, and DocumentTranslator #14790.NorvigSweetingModel intersection #14819, multi-row optimized-annotator batching #14827, and batched ZeroShotNer decoding #14846.PairwiseVectorSimilarity #14806, RuleBasedMatcher #14807, DocumentTranslator #14810, CrossEncoder #14811, and Summarization #14799.ZeroShotNer #14855 fixes; ONNX CUDA recovery #14858 and arm64 OpenVINO runtime work #14825.addFile collision handling #14834, configurable fallback logging #14839, and Spark 4 integration work for the healthcare library #14818 (not an independently verified healthcare-product support statement).DocumentTranslator correction #14878; this link documents branch work, not a claim about a specific binary's contents.One column per quarter.
Spark 3.x Lane
Spark 4.x Lane
=======
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
Nothing published for this version
📢 Spark NLP 6.4.2: BM25 Retrieval, SaT Sentence Detection, and Model Conversion Workflows
Spark NLP 6.4.2 is a feature and reliability release focused on strengthening retrieval, sentence segmentation, model conversion, and release infrastructure. This release introduces BM25Approach and BM25Model for Okapi BM25 lexical retrieval, adds SentenceDetectorSaTModel for transformer-based sentence boundary detection with SaT / Segment any Text models, and expands model conversion examples with new Hugging Face to Spark NLP notebooks for ONNX, OpenVINO, and GGUF workflows.
In addition, this release improves production robustness by properly closing resource streams and file systems during local resource copying, which avoids leak warnings that can affect dependent pipelines. It also updates jsl-llamacpp to 2.0.3.
BM25Approach + BM25Model: New Okapi BM25 lexical retrieval annotators implemented as a fit-once, query-many Estimator/Model pair. They complement semantic retrieval components such as DocumentSimilarityRanker with a fast, corpus-statistics-based lexical ranking option.SentenceDetectorSaTModel: New sentence detector based on SaT / Segment any Text transformer models. It uses ONNX-backed XLM-R SentencePiece models to produce sentence spans from per-token boundary probabilities and supports long documents through overlapping token windows.Spark NLP 6.4.2 introduces BM25 lexical document retrieval through a two-stage Spark ML design. This implements SPARKNLP-1145 and adds a purely lexical ranking option that complements semantic retrieval components such as DocumentSimilarityRanker.
BM25Approach scans the tokenized corpus during fit() and learns corpus-level statistics: document count, per-term document frequency, average document length, and inverse document frequency.BM25Model reuses those learned statistics at transform time to score each document against a query.This design is necessary because BM25 scores depend on corpus-level statistics that are only known after fitting over the full dataset. Once fitted, the model can be reused across many queries without rescanning the corpus.
Key capabilities:
BM25_RANKINGS annotations.bm25_score, num_query_terms_matched, query, and doc_len metadata.k1, b, minDocFreq, and caseSensitive parameters.setQuery(...) for raw query strings.setQueryTokens(...) for pre-analyzed query terms when the corpus pipeline includes normalization, stemming, lemmatization, or other token transformations.caseSensitive read-only on fitted BM25Model to avoid corrupting scores after the IDF vocabulary has been built.from sparknlp.base import DocumentAssembler
from sparknlp.annotator import Tokenizer, StopWordsCleaner, BM25Approach
from pyspark.ml import Pipeline
from pyspark.sql.functions import col, explode
corpus = spark.createDataFrame([
(1, "Apples are a great source of dietary fiber and vitamin C."),
(2, "Machine learning uses neural networks and statistical models."),
(3, "Vitamin C deficiency can affect the immune system."),
], ["id", "text"])
document = DocumentAssembler() \
.setInputCol("text") \
.setOutputCol("document")
tokenizer = Tokenizer() \
.setInputCols(["document"]) \
.setOutputCol("token")
stop_words = StopWordsCleaner() \
.setInputCols(["token"]) \
.setOutputCol("clean_token") \
.setCaseSensitive(False)
bm25 = BM25Approach() \
.setInputCols(["clean_token"]) \
.setOutputCol("bm25_rankings") \
.setK1(1.2) \
.setB(0.75) \
.setMinDocFreq(1) \
.setCaseSensitive(False)
pipeline = Pipeline(stages=[document, tokenizer, stop_words, bm25])
model = pipeline.fit(corpus)
# Re-query the fitted model without recomputing corpus statistics.
model.stages[-1].setQuery("vitamin C health")
result = model.transform(corpus)
result.select(col("id"), explode(col("bm25_rankings")).alias("ranking")) \
.select(
col("id"),
col("ranking.metadata")["bm25_score"].cast("double").alias("bm25_score"),
col("ranking.metadata")["num_query_terms_matched"].cast("int").alias("terms_matched"),
) \
.orderBy(col("bm25_score").desc()) \
.show(truncate=False)For pipelines that transform tokens before BM25 training, use setQueryTokens(...) with query tokens produced by the same analysis pipeline. This avoids query/document analyzer asymmetry and ensures query terms match the learned IDF vocabulary.
See the new BM25 Retrieval notebook for a complete example.
This release adds SentenceDetectorSaTModel, a new transformer-based sentence segmentation annotator built around SaT / Segment any Text models.
SentenceDetectorSaTModel predicts sentence boundaries from per-token probabilities using an ONNX-backed XLM-R SentencePiece model. It supports documents longer than a single model window by slicing text into overlapping sub-word token windows, merging the boundary probabilities, and projecting them back to character spans.
Key capabilities:
DOCUMENT.DOCUMENT.segment-any-text/sat-12l-sm and segment-any-text/sat-12l.sat_12l_sm with language xx.blockSize and stride.satBatchSize.hat and uniform.minSentenceLength and maxSentenceLength.loadSavedModel(...) when the model folder contains model.onnx and assets/sentencepiece.bpe.model.from sparknlp.base import DocumentAssembler
from sparknlp.annotator import SentenceDetectorSaTModel
from pyspark.ml import Pipeline
assembler = DocumentAssembler() \
.setInputCol("text") \
.setOutputCol("document")
sentence_detector = SentenceDetectorSaTModel.pretrained() \
.setInputCols(["document"]) \
.setOutputCol("sentence")
pipeline = Pipeline().setStages([assembler, sentence_detector])
data = spark.createDataFrame([["This is a sentence. This is another one."]]).toDF("text")
result = pipeline.fit(data).transform(data)
result.selectExpr("explode(sentence.result) as sentence").show(truncate=False)A new example notebook demonstrates Hugging Face ONNX usage with SentenceDetectorSaTModel:
Spark NLP 6.4.2 adds three end-to-end model-conversion notebooks under examples/python/model-conversion/:
These notebooks make it easier to validate and document model import paths across the main runtime families used by Spark NLP: ONNX, OpenVINO, and GGUF.
This release also includes several infrastructure and maintenance improvements:
ResourceHelper now closes copied resource streams and associated file systems, avoiding stream/file-system leak warnings that could affect dependent pipelines even when the underlying issue appears as a warning rather than a hard error.jsl-llamacpp upgrade: The jsl-llamacpp dependency is bumped from 2.0.0 to 2.0.3.sbt/setup-sbt@v1 disk-cache behavior in Spark 3.3, 3.4, and 3.5 build jobs to avoid unrelated hashFiles(...) cache hashing failures before the Spark NLP build starts.ResourceHelper by closing the copied resource and the associated file system.similarity/__init__.py issue in Python API wiring.sbt/setup-sbt disk cache.Validation added or updated in this release includes:
BM25TestSpec coverage for statistics learning, ranking, query reuse, save/load, exact-score verification, non-default parameters, minDocFreq pruning, analyzer symmetry, and parameter-range rejection.bm25_test.py coverage for fit/query/reuse, save/load, parameter propagation, exact score checks, setQueryTokens(...), and read-only caseSensitive behavior.SentenceDetectorSaTSpec coverage for the new SaT sentence detector.sentence_detector_sat_test.py coverage for the Python wrapper.pip install spark-nlp==6.4.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.2<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.4.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.4.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.4.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.4.2</version>
</dependency>https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-6.4.2.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-6.4.2.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-6.4.2.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-6.4.2.jarBM25Approach, BM25Model, Scala/Python APIs, docs, tests, and example notebook #14776 by @AbdullahMubeenAnwarjsl-llamacpp to 2.0.3 #14777 by @danilojslSentenceDetectorSaTModel sentence segmentation annotator with Scala/Python APIs, tests, and ONNX example notebook #14782 by @ahmedlone127Full Changelog: 6.4.1...6.4.2
=======
Nothing published for this version
Nothing published for this version
📢 Spark NLP 6.4.1: Late Chunking, Multimodal Embeddings, and Smarter Pretrained Workflows
Spark NLP 6.4.1 is a feature-rich follow-up release that expands retrieval and multimodal capabilities while improving pretrained model loading and production robustness. This release introduces two powerful new annotators: LateChunkEmbeddings for context-aware chunk embeddings and BiEncoderMultimodalEmbeddings for dual-encoder text-image retrieval workflows.
In addition, VectorDBConnector now supports multimodal image indexing, pretrained model resolution becomes smarter with preferred engine selection, and document ingestion workflows gain important reliability fixes for HTML readers and cloud-backed pretrained caches.
For a deeper walkthrough of the major additions in this release, see our Medium post: Spark NLP 6.4.1: Context-Aware Retrieval, Engine Control, and Multimodal Indexing.
LateChunkEmbeddings: A new annotator that applies the late chunking technique so chunk embeddings are computed from token embeddings generated over the full document context rather than embedding each chunk independently.BiEncoderMultimodalEmbeddings: A new multimodal dual-encoder annotator for generating aligned text and image embeddings for retrieval, search, and indexing workflows.VectorDBConnector: VectorDBConnector now supports indexing image embeddings in addition to text embeddings, enabling multimodal vector database pipelines..pretrained() resolution: Pretrained model loading now supports a preferred engine parameter and resolves duplicate model names more reliably by filtering first by annotator class.Spark NLP 6.4.1 introduces LateChunkEmbeddings, a new annotator based on the Late Chunking approach. Unlike traditional chunk embedding approaches that embed each chunk in isolation, this annotator assumes the upstream embedding model has already encoded the full document in a single pass. It then pools token embeddings corresponding to each chunk span, allowing every chunk embedding to retain broader document context.
This is especially valuable for:
A follow-up improvement in this same release adds sentenceAwareFiltering, which is enabled by default. This restricts pooling to token embeddings within the chunk’s sentence boundary, reducing noise from overlapping tokens across adjacent sentences.
late_chunk = LateChunkEmbeddings() \
.setInputCols(["document", "chunk", "token", "embeddings"]) \
.setOutputCol("chunk_embeddings")See the LateChunkEmbeddings example notebook for a full walkthrough and practical usage examples. You can also read our Medium article for a broader explanation of why late chunking improves retrieval quality over naive chunk-first embedding pipelines.
This release adds BiEncoderMultimodalEmbeddings, a new annotator for multimodal dual-encoder workflows. It accepts aligned DOCUMENT and IMAGE annotations and emits two embedding outputs:
<outputCol>_doc_embeddings<outputCol>_image_embeddingsThis enables use cases such as:
The first supported implementation targets dual-encoder ONNX exports in the Ops-MM / Qwen2VL-style architecture.
mm = BiEncoderMultimodalEmbeddings.pretrained() \
.setInputCols(["vision_pair_doc", "vision_pair_image"]) \
.setOutputCol("mm")To see this in action, check out the BiEncoderMultimodalEmbeddings + Pinecone RAG notebook, which demonstrates multimodal retrieval and indexing in a complete end-to-end workflow. Our Medium post also covers how these multimodal embeddings fit into retrieval and vector search pipelines.
VectorDBConnector now supports multimodal image indexing through a new modalityMode parameter.
Supported modes:
text (default): expects DOCUMENT + SENTENCE_EMBEDDINGSimage: expects IMAGE + SENTENCE_EMBEDDINGSIn image mode, Spark NLP augments metadata with image-specific fields such as origin, width, height, and channel count, and generates deterministic vector IDs derived from the image origin path for stable re-indexing.
This makes it possible to build end-to-end pipelines that extract images, compute embeddings, and store them directly in vector databases such as Pinecone.
vectorDB = VectorDBConnector() \
.setInputCols(["image_assembler", "image_embeddings"]) \
.setOutputCol("vectordb_result") \
.setProvider("pinecone") \
.setIndexName("my-multimodal-index") \
.setModalityMode("image")For a complete multimodal retrieval example, see the BiEncoderMultimodalEmbeddings + Pinecone RAG notebook, which demonstrates how multimodal embeddings and vector indexing work together in an end-to-end RAG pipeline.
Spark NLP 6.4.1 improves .pretrained() behavior in two important ways:
.pretrained() now accepts an engine parameter, allowing users to prefer a specific engine when multiple versions of a model are available.If no engine is specified, Spark NLP uses this priority order:
This makes model loading more predictable and helps users better control runtime behavior across supported backends.
This release fixes several issues affecting pretrained resource caching when cache_pretrained points to cloud-backed storage such as S3, GCS, or Azure Blob Storage.
Improvements include:
wasbs:// cache paths through Hadoop FSspark.hadoop.* params into Hadoop configurationAdditionally, sparknlp.start() now includes a skip_sparknlp_maven option for developer workflows that need to run Spark NLP using a custom local JAR.
HTMLReader receives important reliability and metadata improvements in this release:
paragraph_indexparagraph_ypage_yThese additions improve downstream layout-aware document understanding and help preserve ordering and spatial context in HTML ingestion workflows.
This release also improves LLMEntityExtractor usability and robustness:
getFewShotExamples handling for multiple input cases, particularly in notebook and Colab-style environmentsA metadata cleanup pass also ensures Markdown model cards use engine tags more consistently by aligning engine tags with the declared engine field where available.
ResourceDownloaderwasbs:// cache handlingHTMLReaderLLMEntityExtractor few-shot example handlingpip install spark-nlp==6.4.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.1<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.4.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.4.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.4.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.4.1</version>
</dependency>https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-6.4.1.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-6.4.1.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-6.4.1.jarhttps://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-6.4.1.jarllama_cpp_in_Spark_NLP_LLMEntityExtractor.ipynb #14756 by @AbdullahMubeenAnwarVectorDBConnector #14760 by @AbdullahMubeenAnwarGetFewShotExamples handling in LLMEntityExtractor #14762 by @ahmedlone127LateChunkEmbeddings annotator #14764 by @AbdullahMubeenAnwarBiEncoderMultimodalEmbeddings #14767 by @danilojslsentenceAwareFiltering param #14772 by @danilojslFull Changelog: 6.4.0...6.4.1
=======
Nothing published for this version
Nothing published for this version
Nothing published for this version
📢 Spark NLP 6.4.0: LLM-Powered Entity Extraction, Expanded Document Readers, and Pipeline Robustness
Spark NLP 6.4.0 is a feature-rich release focused on bringing LLM-powered extraction into standard NLP pipelines, dramatically expanding the document reader ecosystem, and hardening production ingestion workflows. This release introduces LLMEntityExtractor, a new annotator that leverages on-device GGUF LLMs to extract named entities from text using flexible few-shot prompting, no fine-tuning required. We also introduce DocumentTitleSplitter, a semantic chunker for reader-based pipelines, and extend the document reader stack with native support for four new file formats: ODT, EPUB, TSV, and RTF. Finally, reader pipelines gain significant production robustness improvements including graceful error handling for HTML and image fetch failures.
LLMEntityExtractor: Use any GGUF-based LLM to extract named entities from text with configurable few-shot examples, no model fine-tuning needed. Returns results as standard CHUNK annotations compatible with the rest of the Spark NLP ecosystem.DocumentTitleSplitter: A new semantic annotator that splits element-level reader output into coherent chunks using title boundaries, designed to work alongside Reader2Doc in element mode across Word, PDF, HTML, PowerPoint, ODT, and Markdown files.HTMLReader and Reader2Image now degrade gracefully on unreachable URLs and broken remote images rather than failing the entire Spark job critical for large-scale batch processing and ETL workflows.Teams that need flexible entity extraction without model fine-tuning can now use LLMEntityExtractor to leverage on-device GGUF models (via AutoGGUFModel) for NER-style extraction. The annotator uses structured prompts (optionally enhanced with few-shot examples) to identify and extract entities from input text, returning them as standard CHUNK annotations.
Key capabilities:
AutoGGUFModelFewShotExample objects.pretrained() and .loadSavedModel()llm_extractor = LLMEntityExtractor.pretrained() \
.setInputCols(["document"]) \
.setOutputCol("entities")See the LLMEntityExtractor example notebook for full usage including few-shot configuration.
Reader pipelines that use Reader2Doc in element mode produce structured output with elementType metadata (e.g., Title, NarrativeText). DocumentTitleSplitter closes a key gap by turning this element sequence into semantic, title-bounded chunks that downstream annotators like sentence embedders and LLMs can consume coherently.
Key parameters:
setSplitOnPageBreak(bool): Also enforce chunk boundaries on page breakssetJoinDelimiter(str): Delimiter used when joining narrative text within a chunksetMaxChunkLength(int): Overflow splitting for oversized title sectionsSupported document types include Word, PDF, HTML, PowerPoint, ODT, and Markdown any format for which Reader2Doc extracts title metadata.
splitter = DocumentTitleSplitter() \
.setInputCols(["document"]) \
.setOutputCol("title_chunks")See the DocumentTitleSplitter example notebook for full usage.
A new dedicated ODTReader brings full parity with WordReader for LibreOffice and OpenOffice documents. It supports:
A new EpubReader enables direct ingestion of e-book content with full structural preservation:
Tab-separated files are now supported natively alongside the existing CSV reader, using the same ingestion APIs without any preprocessing or conversion step.
A new RTFReader adds native ingestion of legacy .rtf files common in enterprise document repositories and archived templates directly into Reader2Doc, and ReaderAssembler pipelines.
Large-scale document ingestion pipelines no longer need to worry about a single unreachable URL or broken remote image crashing the entire Spark job. Two targeted improvements address this:
HTMLReader: Unreachable URLs and timeouts are now handled gracefully by returning a structured fallback artifact instead of failing the task. Configurable via ignoreUrlErrors (default: true).Reader2Image: Remote image fetch failures now produce image-level error artifacts when setIgnoreExceptions(true) (the default), instead of aborting the task.Fail-fast behavior remains available for strict workloads:
# Image strict mode
Reader2Image().setIgnoreExceptions(False)Python API improvement: Reader2Doc and Reader2Table are now re-exported directly from sparknlp.reader, simplifying imports:
# Before
from sparknlp.reader.reader2doc import Reader2Doc
# Now
from sparknlp.reader import Reader2Doc, Reader2TableDocumentation has been added with instructions for loading and caching pretrained Spark NLP models using Databricks Unity Catalog Volumes, enabling teams running on Databricks to take advantage of centralized model storage within the Unity Catalog ecosystem.
pip install spark-nlp==6.4.0spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.4.0spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.4.0spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.4.0spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.4.0<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.4.0</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.4.0</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.4.0</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.4.0</version>
</dependency>LLMEntityExtractor annotator #14735 by @ahmedlone127DocumentTitleSplitter annotator #14751 by @danilojslFull Changelog: 6.3.3...6.4.0
=======
Nothing published for this version
…from 4.1.2 to 5.4.1 ( poi-ooxml-full ) to avoid deprecated dependencies.
Spark NLP 6.3.3 is a feature-packed release aimed at practitioners building modern NLP and multimodal document pipelines. This release introduces ModernBertEmbeddings for dramatically faster and more memory-efficient text embeddings, VectorDBConnector to close the gap between embedding pipelines and vector search infrastructure.
Moreover, we introduce a new suite of document-understanding annotators: LayoutAlignerForVision and LayoutAlignerForText which together enable coherent end-to-end multimodal pipelines over complex documents like PDFs and PowerPoint files. To gain an in-depth walkthrough on how to build use your own pipelines and documents, please see our Medium blog post at Efficient Document Ingestion with Layout Aware Annotators: A Case Study on Mixed-Type Documents.
In addition, a new MultiColumnAssembler can merge multiple annotation column into one and LightPipeline also gains metadata column support for more powerful batch inference workflows.
ModernBertEmbeddings: A state-of-the-art encoder that is 8x faster and uses 5x less memory than traditional BERT, with native support for sequences up to 8192 tokens which is ideal for long-document NLP tasks.VectorDBConnector: Bridges Spark NLP embedding pipelines with external vector databases (initially Pinecone), so teams building semantic search and RAG systems no longer need custom glue code to store embeddings.LayoutAlignerForVision and LayoutAlignerForText: New annotators for multimodal document pipelines that spatially align text and images from complex documents, giving downstream Vision-Language Models (VLMs) the coherent context they need to produce better results.MultiColumnAssembler: Closes a common pipeline gap when ReaderAssembler splits document content across multiple columns (text, table, image captions). This annotator merges them back into a single column that downstream annotators like AutoGGUFVisionModel expect.LightPipeline with metadata column support for richer, context-aware inference workflows.Teams working with long documents, code, or large-scale embedding workloads will benefit from ModernBertEmbeddings, which brings the latest generation of bidirectional encoder models to Spark NLP. Based on the paper Smarter, Better, Faster, Longer, ModernBERT was trained on 2 trillion tokens with a native sequence length of up to 8,192 tokens (eight times the limit of classic BERT) enabling faster, cheaper embeddings for longer sequences without truncation.
"modernbert-base" (English)WORD_EMBEDDINGS outputembeddings = ModernBertEmbeddings.pretrained() \
.setInputCols(["document", "token"]) \
.setOutputCol("modernbert_embeddings")See the ModernBertEmbeddings notebook for extended examples, including how to import custom HuggingFace ModernBERT models via ONNX.
For teams building semantic search, retrieval-augmented generation (RAG), or similarity-based recommendation systems, manually bridging Spark NLP with a vector database has historically required custom integration code. VectorDBConnector eliminates this gap by letting you store embeddings from any Spark NLP embedding annotator directly into a vector database as part of the pipeline. It initially supports Pinecone, with more providers planned.
vectorDB = VectorDBConnector() \
.setInputCols(["document", "sentence_embeddings"]) \
.setOutputCol("vectordb_result") \
.setProvider("pinecone") \
.setIndexName("my-semantic-index") \
.setNamespace("production") \
.setIdColumn("doc_id") \
.setMetadataColumns(["text", "category"]) \
.setBatchSize(100)The Pinecone API key is configured via spark.jsl.settings.vectordb.api_key. See the VectorDBConnector Pinecone Demo notebook for a full walkthrough.
When processing rich documents like PDFs or PowerPoint presentations, text and images are spatially interleaved. For example, A chart sits next to the paragraph it illustrates, a diagram is surrounded by its explanation. Without layout awareness, VLMs operating on extracted content lose this spatial context entirely. LayoutAlignerForVision and LayoutAlignerForText solve this problem for teams building multimodal document intelligence pipelines.
LayoutAlignerForVision takes document chunks and images extracted by ReaderAssembler and aligns each image with its spatially nearby text paragraphs based on actual page coordinates. It produces three output columns <outputCol>_doc, <outputCol>_image, and <outputCol>_prompt, ready to be fed directly into a VLM (e.g. AutoGGUFVisionModel) for captioning or question answering.
Key parameters:
setMaxDistance(int): Maximum vertical distance (px) for image-paragraph alignmentsetIncludeContextWindow(bool): Include neighboring paragraphs as context for floating imagessetAddNeighborText(bool): Include aligned text in the prompt outputsetImageCaptionBasePrompt(str): Customize the captioning prompt sent to downstream VLMssetNeighborTextCharsWindow(int): Include surrounding text characters as prompt contextsetExplodeDocs(bool): Emit one output row per aligned doc/image pairLayoutAlignerForText takes the VLM-generated image captions produced after LayoutAlignerForVision and weaves them back into the document's text flow, replacing raw image placeholders with meaningful captions and re-computing begin/end offsets so the resulting document is coherent for downstream NLP tasks.
Key parameters:
setJoinDelimiter(str) – Delimiter used to join rebuilt text segmentssetExplodeElements(bool) – Emit one output row per aligned text elementFor more extended examples and walkthroughs, see refer to the notebook Spark NLP LayoutAligners for Document Understanding and our Medium blog post Efficient Document Ingestion with Layout Aware Annotators: A Case Study on Mixed-Type Documents.
When using ReaderAssembler to process documents such as PDFs or PPTX files, content is extracted into separate typed columns: document_text, document_table, and image-related outputs. However, many downstream annotators expect a single input column. Previously, bridging this split required custom Spark transformations. MultiColumnAssembler fills this gap directly within Spark NLP pipelines.
It merges any number of DOCUMENT-type annotation columns into a single output column, preserving all annotation metadata and adding a source_column key to track provenance. Annotations can optionally be sorted by their begin offset using setSortByBegin(True).
multiColumnAssembler = MultiColumnAssembler() \
.setInputCols(["document_text", "document_table"]) \
.setOutputCol("merged_document")Key parameters:
setInputCols([...]) – List of DOCUMENT-type annotation columns to mergesetOutputAsAnnotatorType(str) – Override the output annotator type (default: "document")setSortByBegin(bool) – Sort merged annotations by begin position (default: False)Note: Columns using the
AnnotationImageschema (i.e., IMAGE-typed columns fromReaderAssembler) are not supported.
See the Merging Annotation Columns notebook for a full walkthrough.
Users running inference with LightPipeline on data that carries additional context — such as document source, language, or category — previously had no way to pass that context through alongside the text. LightPipeline now supports passing metadata columns alongside text inputs in both annotate() and fullAnnotate(), enabling richer, context-aware inference for applications like routing, filtering, and conditional processing.
New supported call signatures:
fullAnnotate(text: str, metadata: dict[str, list[str]])fullAnnotate(texts: list[str], metadata: list[dict])fullAnnotate(texts: list[str], metadata: dict[str, list[str]]) (columnar format)annotate()Metadata can be passed as a keyword argument or as a positional trailing argument:
result = light_pipeline.fullAnnotate(
"U.N. official Ekeus heads for Baghdad.",
metadata={"source": ["news_article"]}
)This feature is also surfaced through PretrainedPipeline.annotate() and PretrainedPipeline.fullAnnotate().
poi-ooxml-full) to avoid deprecated dependencies.pip install spark-nlp==6.3.3spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.3spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.3spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.3spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.3Supported on Apache Spark 3.x.
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.3.3</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.3.3</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.3.3</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.3.3</version>
</dependency>Full Changelog: 6.3.2...6.3.3
=======
Nothing published for this version
Nothing published for this version
📢 Spark NLP 6.3.2: Scala 2.13 Support, Layout-Aware Images, and Enhanced LightPipeline Tracking
Spark NLP 6.3.2 is a foundational release that introduces official support for Scala 2.13, alongside important improvements in document layout understanding and lightweight inference workflows.
This release improves long-term model portability through JSON-based serialization, enriches document image extraction with spatial metadata, and enhances LightPipeline with document ID tracking and output filtering.
Reader2Image for HTML, DOCX, and PPTX documents.LightPipeline with document ID propagation and output column filtering for better batch inference workflows.Spark NLP now supports Scala 2.13 with this release! This will enable you to run your Spark NLP pipelines on Spark versions that run on Scala 2.13, such as used by Databricks and Dataproc. See our Installation Instructions for Scala 2.13 on how to use it with our project.
There are some things you have to consider when using the Scala 2.13 version
spark-nlp_2.12 to spark-nlp_2.13.SPARK_HOME environment variable to a Spark Scala 2.13 installation, or install PySpark from the official Spark archives.DependencyParserModel or TextMatcherModel from Scala 2.12 into Scala 2.13, you will need to manually export them again with the latest version. See the notebook Reader2ImageThe Reader2Image annotator now extracts spatial image coordinates from rich document formats, adding layout awareness to image annotations.
x, y, width, heightThis enables:
LightPipelineLightPipeline now supports passing document IDs together with text inputs, improving traceability in batch and production inference scenarios.
Key capabilities:
fullAnnotate(ids, texts)annotate(ids, texts)doc_id)output_cols parameter to restrict returned annotation typesBenefits:
Existing LightPipeline usage remains unchanged and backward compatible.
pip install spark-nlp==6.3.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.2spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.2Supported on on Apache Spark 3.x.
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.3.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.3.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.3.2</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.3.2</version>
</dependency>Full Changelog: 6.3.1...6.3.2
=======
Nothing published for this version
Nothing published for this version
Nothing published for this version
📢 Spark NLP 6.3.1: LLM Backend Upgrade and Document Processing Improvements
Spark NLP 6.3.1 focuses on strengthening distributed local LLM inference by upgrading the jsl-llamacpp backend to a newer llama.cpp release, while also delivering important improvements in document structure handling and metadata consistency.
This enables you to use the latest LLMs and embeddings compatible with llama.cpp and perform advanced ingestion of tables and images.
jsl-llamacpp backend to llama.cpp tag b7247, bringing upstream performance improvements, stability fixes, and expanded model compatibility for local LLM inference.Reader2X annotator capabilities with structural position metadata for tables and images and integration with AutoGGUFVisionModel
The jsl-llamacpp backend has been upgraded to llama.cpp tag b7247, applying upstream fixes and enabling the use of the latest LLMs. These benefit distributed LLM workloads in Spark NLP and affects the annotators AutoGGUFModel, AutoGGUFEmbeddings, AutoGGUFVisionModel, AutoGGUFReranker:
llama.cpp for offline LLM inference within Spark NLP pipelinesgpt-oss, Qwen3 and embeddinggemma.Previously, our document parsers (HTMLReader, XMLReader, WordReader, PowerPointReader, ExcelReader) relied heavily on positional or page-based coordinates for layout metadata.
However, non-PDF formats such as HTML, XML, DOC(X), PPT(X), and XLS(X) do not have fixed pages
To ensure deterministic element referencing and structural traceability across all document types, we needed to adopt a unified DOM-like metadata model.
This change standardizes metadata extraction so every element can be uniquely identified and re-located within its source document, independent of visual layout.
These additions enable layout-aware downstream processing and more precise filtering especially for HTML and rich document formats.
Previously, you could use Reader2Image to ingest images from various file formats into Spark NLP. However, processing was limited to Spark NLP native VLM implementations (such as Qwen2VLTransformer).
Reader2Image now supports interoperability with our llama.cpp backend with AutoGGUFVisionModel by introducing flexible handling of encoded vs. decoded image bytes and optional prompt output.
useEncodedImageBytes to control whether the image result stores:
true: Encoded (compressed) file bytes for models like AutoGGUFVisionModelfalse: Decoded pixel matrix for models such as Qwen2VLTransformerAutoGGUFVisionModel.Added official documentation and instructions for setting up and running Spark NLP on Microsoft Fabric, simplifying configuration and improving developer onboarding on the platform. You can see them at Spark NLP - Installation
DocumentAssembler outputs when using LightPipeline.ResourceDownloader could fail under certain conditions.BertEmbeddings models with non-standard output tensor names.pip install spark-nlp==6.3.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.3.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.3.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.3.1spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.3.1Supported on on Apache Spark 3.x.
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.3.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.3.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.3.1</version>
</dependency><dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.3.1</version>
</dependency>Full Changelog: 6.3.0...6.3.1
=======
Nothing published for this version
[SPARKNLP-1318] Reader2Image Integration with AutoGGUFVisionModel #14705
None
=======
📢 Spark NLP 6.2.3: Further Improvements for NerDL
Spark NLP 6.2.3 introduces targeted improvements to training performance and stability of NerDLApproach and bug fixes for CamemBertForTokenClassification.
NerDLApproach now uses new internal data-loading behavior, and improving training speed and preventing out-of-memory errors.
Enhanced NerDLApproach training performance through threaded data loading and optimized partitioning.
Significant performance improvements for training of NerDLApproach:
Threaded Data Loading: When enabling the memory optimizer (setEnableMemoryOptimizer(true)), data can now be pre-fetched through a threaded data loader. By default, it is disabled but can be tuned by using:
.setPrefetchBatches(int)By tuning this parameter (for example 20 batches), you can get training time reductions of about 10%.
Optimized Partitioning Strategy: NerDLApproach now applies optimized dataframe partitioning when using the memory optimizer (setEnableMemoryOptimizer(true)) by default, improving parallelization efficiency during training and preventing out-of-memory errors.
For manual tuning of the input data frames, this behavior can be disabled with:
.setOptimizePartitioning(false)pip install spark-nlp==6.2.3CPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.3GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.3Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.3AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.3<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.2.3</version>
</dependency>Full Changelog: 6.2.2...6.2.3
None
=======
📢 Spark NLP 6.2.2: Bugfix Release
Spark NLP 6.2.2 brings bug fixes to WordEmbeddings and NerDLApproach logging.
WordEmbeddings would duplicate input tokens in the outputpip install spark-nlp==6.2.2CPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.2GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.2Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.2AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.2<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.2.2</version>
</dependency>Full Changelog: 6.2.0...6.2.2
None
=======
Nothing published for this version
📢 Spark NLP 6.2.1: Enhanced hierarchical document processing and training optimizations
Spark NLP 6.2.1 brings significant improvements to document ingestion with expanded hierarchical support, XML processing enhancements, and optimizations for NerDL training. This release builds on the foundation of 6.2.0, continuing to focus on structure-awareness, flexibility, and performance for production NLP pipelines.
Reader2DocBuilding on the HTMLReader hierarchical features introduced in 6.2.0, this release extends structured element tracking to additional document formats:
Reader2Doc now supports hierarchical processing for PDF, Microsoft Word, and Markdown files
Each extracted element includes:
element_id: Unique UUID identifier per elementparent_id: References the parent element's ID for logical document structureEnables tree-like navigation and contextual understanding of document hierarchy:
Chapter 1
├── Narrative Text A
├── Narrative Text B
Chapter 2
├── Paragraph C
Supports advanced use cases including hierarchical retrieval, graph-based indexing, and multi-level document analysis
Metadata propagation ensures downstream annotators maintain structural relationships
Significant performance improvements for training of NerDLApproach:
setEnableMemoryOptimizer(true) with maxEpoch > 1, input datasets are automatically cached to improve training speedNerDLGraphChecker now populates TensorFlow graph metadata that NerDLApproach can reuse, reducing redundant computations during training initializationWith all these improvements you can expect up half the memory consumption and training time on RAM constrained environments (when using setEnableMemoryOptimizer(true)). For larger distributed datasets, the effect will be more pronounced.
Single Document Output by Default: Reader2Doc now creates single document annotations per file by default, providing more expected behavior when processing large documents
\n by default, configurable via new setJoinString(string) parameter for custom separatorsImproved Tag Handling: XML reader now ignores empty tags without text content, reducing noise in parsed output
Enhanced content type handling for application/xml documents
XML Tag Attribute Extraction: New setExtractTagAttributes(attributes: list[str]) parameter enables extraction of XML attribute values. Example:
<bookstore>
<book category="children">
<title lang="en">Harry Potter</title>
<author>J K. Rowling</author>
<year>2005</year>
<price>29.99</price>
</book>
<book category="web">
<title lang="en">Learning XML</title>
<author>Erik T. Ray</author>
<year>2003</year>
<price>39.95</price>
</book>
</bookstore>We can extract category and lang values with the Reader2Doc Config
reader2doc = Reader2Doc() \
.setContentType("application/xml") \
.setContentPath("../src/test/resources/reader/xml/test.xml") \
.setOutputCol("document") \
.setExtractTagAttributes(["category", "lang"])Resulting in
children
en
Harry Potter
J K. Rowling
2005
29.99
web
en
Learning XML
Erik T. Ray
2003
39.95
pip install spark-nlp==6.2.1CPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.1GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.1Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.1AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.1<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.2.1</version>
</dependency>Full Changelog: 6.2.0...6.2.1
=======
Spark NLP 6.2.0 introduces key upgrades across entity extraction, document normalization, HTML reading, and GGUF-based models. To recap, since the rel
Spark NLP 6.2.0 introduces key upgrades across entity extraction, document normalization, HTML reading, and GGUF-based models. To recap, since the releases of Spark NLP 6.1 you can:
AutoGGUFRerankerReader2Doc: streamlines the process of loading and integrating diverse file formats (PDFs, Word, Excel, PowerPoint, HTML, Text, Email, Markdown) directly into Spark NLP pipelines with a unified and flexible interface.Reader2Table: streamlines tabular data extraction from multiple document formats with seamless pipeline integration.Reader2Image: extract structured image content from various document typesSpark NLP release 6.2.0 further focuses on automation, structure-awareness, and resource efficiency, making pipelines easier to configure, manage, and extend.
EntityRulerModelautoMode parameter to enable predefined regex entity groups ("network_entities", "communication_entities", "media_entities", "email_entities", "all_entities").extractEntities parameter to filter entities within auto modes.DocumentNormalizerpresetPattern and autoMode parameters to apply built-in text cleaning patterns."light_clean", "document_clean", "social_clean", "html_clean", and "full_auto".Together, these additions significantly reduce boilerplate setup for common text extraction and normalization workflows.
element_id and parent_id metadata fields for each parsed HTML element.title → paragraph → link) for hierarchical retrieval and contextual reasoning.For AutoGGUFModel, AutoGGUFVision, AutoGGUFEmbeddings, AutoGGUFReranker
close() method to explicitly release llama.cpp model resources, preventing memory retention in long-running sessions.setRemoveThinkingTag(tag: String) parameter to remove internal <think>...</think> sections from model outputs.
(?s)<$tag>.+?</$tag>pip install spark-nlp==6.2.0
CPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.2.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.2.0
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.2.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.2.0
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.2.0</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.5...6.2.0
=======
Spark NLP 6.1.5 focuses on improving data ingestion reliability and pipeline flexibility. This release enhances reader components with better fault to
Spark NLP 6.1.5 focuses on improving data ingestion reliability and pipeline flexibility. This release enhances reader components with better fault tolerance, broader input support, and introduces a new ReaderAssembler annotator for streamlined integration. Several key fixes also improve model loading and stability in distributed environments.
ReaderAssembler Annotator: Unify multiple reader annotators into one configurable component for simpler and cleaner ingestion pipelines.ReaderAssembler Annotator
A new meta-annotator that unifies Reader2X components (e.g., Reader2Doc, Reader2Image, Reader2Table) under a single interface.
Support for String Input Columns in Readers (SPARKNLP-1291) Spark NLP readers only supported inputs via file paths. That means if you already had a DataFrame with text content (say from another pipeline or a preliminary load), you had to write it to disk just to let the reader ingest it. This adds friction and overhead, especially in streaming or in-memory pipelines.
With this change, you can:
Fault-Tolerant XML Reader The XML reader now skips malformed XML fragments (e.g., mismatched tags, missing closures, invalid characters) instead of failing the job. Enhanced error handling ensures more resilient ingestion of imperfect real-world data.
FeaturesFallbackReader that caused duplicate loading or missing model files when calling .pretrained() on GGUF-based annotators such as AutoGGUFModel and rerankers, especially in Databricks environments.pip install spark-nlp==6.1.5
CPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.5
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.5
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.5
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.1.5</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.1.5</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.1.5</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.4...6.1.5
=======
We are excited to announce the release of Spark NLP 6.1.4! This version introduces a powerful new annotator, Reader2Image, which extends Spark NLP’s u
We are excited to announce the release of Spark NLP 6.1.4!
This version introduces a powerful new annotator, Reader2Image, which extends Spark NLP’s universal ingestion capabilities to embedded images across a wide range of document formats. With this release, Spark NLP users can now seamlessly integrate text and image processing in the same pipeline, unlocking new opportunities for vision-language modeling (VLM), multimodal search, and document understanding.
Reader2Image Annotator: Extract and structure image content directly from documents like PDFs, Word, PowerPoint, Excel, HTML, Markdown, and Email files.Reader2Image Annotator
A new multimodal annotator designed to parse image content embedded in structured documents. Supported formats include:
Output Fields:
This enables seamless integration with vision-language models (VLMs), multimodal embeddings, and downstream Spark NLP annotators, all within the same distributed pipeline.
pip install spark-nlp==6.1.4
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.4
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.4
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.4
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.4</version>
</dependency>
spark-nlp-gpu_2.12:6.1.4spark-nlp-silicon_2.12:6.1.4spark-nlp-aarch64_2.12:6.1.4Reader2Image Annotator (#14658) by @danilojslFull Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.3...6.1.4
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.3...6.1.4
None
=======
We are pleased to announce Spark NLP 6.1.3, introducing a new graph validation annotator for NER training, enhancements to Reader2Doc for flexible doc
We are pleased to announce Spark NLP 6.1.3, introducing a new graph validation annotator for NER training, enhancements to Reader2Doc for flexible document handling, and a new ranking finisher for AutoGGUFReranker outputs. This release focuses on improving training robustness, document processing flexibility, and retrieval ranking capabilities.
NerDLGraphChecker:
A new annotator that validates whether a suitable NerDL graph is available for a given training dataset before embeddings or training start. This helps avoid wasted computation in custom training scenarios. (Link to notebook)
NerDLApproach annotators.Reader2Doc Enhancements:
New configuration options provide more control over output formatting:
outputAsDocument: Concatenates all sentences into a single document.excludeNonText: Filters out non-textual elements (e.g., tables, images) from the document.AutoGGUFRerankerFinisher:
A finisher for processing AutoGGUFReranker outputs, adding advanced ranking and filtering capabilities (Link to notebook):
None.
pip install spark-nlp==6.1.3
spark-nlp on Apache Spark 3.0.x–3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.3
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.3
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.3
spark-nlp:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.1.3</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.1.3</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.1.3</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.2...6.1.3
None
=======
Nothing published for this version
We are excited to announce Spark NLP 6.1.2, enhancing AutoGGUF model support and introduces a brand new reranking annotator based on llama.cpp LLMs. T
We are excited to announce Spark NLP 6.1.2, enhancing AutoGGUF model support and introduces a brand new reranking annotator based on llama.cpp LLMs. This release also brings fixes for AutoGGUFVision model and improvements for CUDA compatibility of AutoGGUF models.
New AutoGGUFReranker annotator for advanced LLM-based reranking in information retrieval and retrieval-augmented generation (RAG) pipelines.
AutoGGUFReranker
A new annotator for reranking candidate results using AutoGGUF-based LLM embeddings. This enables more accurate ranking in retrieval pipelines, benefiting applications such as search, RAG, and question answering. (Link to notebook)AutoGGUFVisionModel.save for AutoGGUF models now supports more file protocols.pip install spark-nlp==6.1.2
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.2
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.2
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.2
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.1.2</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.1.2</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.1.2</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.1...6.1.2
=======
Nothing published for this version
Robust Model Loading: Pretrained AutoGGUF-based annotators now load despite the inclusion of deprecated parameters, ensuring broader compatibility.
We are thrilled to announce Spark NLP 6.1.1, a focused release that delivers significant performance improvements and enhanced functionality for large language models and universal data ingestion. This release continues our commitment to providing state-of-the-art AI capabilities within the native Spark ecosystem, with optimized inference performance and expanded multimodal support.
AutoGGUFModel and AutoGGUFEmbeddings deliver improvements for large language model workflows on GPU.AutoGGUFVisionModel annotator is back with full functionality and latest SOTA VLMs, enabling sophisticated vision-language processing capabilities.Reader2Table annotator streamlines tabular data extraction from multiple document formats with seamless pipeline integration.AutoGGUFModel Performance: We improved the inference of llama.cpp models and achieved a 10% performance increase for AutoGGUFModel on GPU.AutoGGUFVisionModel: The multimodal vision model annotator is fully operational again, enabling powerful vision-language processing capabilities. Users can now process images alongside text for comprehensive multimodal AI applications while using the latest SOTA vision-language models.AutoGGUFModel can now seamlessly load the language model components from pretrained AutoGGUFVisionModel instances, providing greater flexibility in model deployment and usage. (Link to notebook)| Annotator | Default pretrained model |
|---|---|
| AutoGGUFModel | Phi_4_mini_instruct_Q4_K_M_gguf |
| AutoGGUFEmbeddings | Qwen3_Embedding_0.6B_Q8_0_gguf |
| AutoGGUFVisionModel | Qwen2.5_VL_3B_Instruct_Q4_K_M_gguf |
Reader2Table Annotator: This powerful new annotator provides a streamlined interface for extracting and processing tabular data from various document formats (Link to notebook). It offers:
None
pip install spark-nlp==6.1.1
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.1
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.1
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.1
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.1.1</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.1.1</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.1.1</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.1.0...6.1.1
None
=======
We are excited to announce Spark NLP 6.1.0, another milestone for building scalable, distributed AI pipelines! This major release significantly enhanc
We are excited to announce Spark NLP 6.1.0, another milestone for building scalable, distributed AI pipelines! This major release significantly enhances our capabilities for state-of-the-art multimodal and large language models and universal data ingestion. Upgrade Spark NLP to 6.1.0 to improve both usability and performance across ingestion, inference, and multimodal processing pipelines, all within the native Spark ecosystem.
llama.cpp Integration: We've updated our llama.cpp backend to tag b5932 which supports inference with the latest generation of LLMs.Reader2Doc: Introducing a new annotator that streamlines the process of loading and integrating diverse file formats (PDFs, Word, Excel, PowerPoint, HTML, Text, Email, Markdown) directly into Spark NLP pipelines with a unified and flexible interface.llama.cpp Upgrade: Our llama.cpp backend has been upgraded to version b5932. This update enables native inference for the newest LLMs, such as Gemma 3 and Phi-4, ensuring broader model compatibility and improved performance.
AutoGGUFVisionModel annotator to the latest backend. This means that this annotator will not be available in this version. As a workaround, please use version 6.0.5 of Spark NLP.Reader2Doc Annotator: This new annotator provides a simplified, unified interface for integrating various Spark NLP readers. It supports a wide range of formats, including PDFs, plain text, HTML, Word (.doc/.docx), Excel (.xls/.xlsx), PowerPoint (.ppt/.pptx), email files (.eml, .msg), and Markdown (.md).Let's use a code example to see how easy it is to use:
reader2doc = Reader2Doc() \
.setContentType("application/pdf") \
.setContentPath("./pdf-files") \
.setOutputCol("document")
# other NLP stages in `nlp_stages`
pipeline = Pipeline(stages=[reader2doc] + nlp_stages)
model = pipeline.fit(empty_df)
result_df = model.transform(empty_df)
Check out our full example notebook to see it in action.
pip install spark-nlp==6.1.0
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.1.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.1.0
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.1.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.1.0
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.1.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.1.0</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.1.0</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.1.0</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.5...6.1.0
=======
We're thrilled to announce the release of Spark NLP 6.0.5\! This version introduces a new Markdown Reader, enabling direct processing of Markdown file
We're thrilled to announce the release of Spark NLP 6.0.5! This version introduces a new Markdown Reader, enabling direct processing of Markdown files into structured Spark DataFrames for more diverse NLP workflows. We have also enhanced Microsoft Fabric integration, allowing for seamless model downloads from Lakehouse containers.
MarkdownReader for effortlessly parsing Markdown files into structured Spark DataFrames, paving the way for advanced content analysis and NLP on Markdown content.New MarkdownReader Annotator: Introducing the MarkdownReader, a powerful new feature that allows you to read and parse Markdown files directly into a structured Spark DataFrame. This enables efficient processing and analysis of Markdown content for various NLP applications. We recommend using this reader automatically in our Partition annotator. (Link to notebook)
partitioner = Partition(content_type = "text/markdown"").partition(md_directory)
Microsoft Fabric Integration: Spark NLP now supports downloading models from Microsoft Fabric Lakehouse containers, providing a more integrated and efficient workflow for users leveraging Microsoft Fabric. This enhancement ensures smoother model access and deployment within the Fabric ecosystem. For example, you can define the path to our pretrained models in Spark like so:
from pyspark import SparkConf
conf = SparkConf()
conf.set("spark.jsl.settings.pretrained.cache_folder", "abfss://my_workspace@onelake.dfs.fabric.microsoft.com/lakehouse_folder.Lakehouse/Files/my_models")
We performed crucial maintenance updates to all of our example notebooks, ensuring that they are reproducible and properly displayed in GitHub.
#PyPI
pip install spark-nlp==6.0.5
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.5
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.5
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.5
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.5
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.5
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.5</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.5</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.5</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.5</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.4...6.0.5
=======
We are excited to announce the release of Spark NLP 6.0.4! This version brings advancements in text embeddings with the introduction of the MiniLM fam
We are excited to announce the release of Spark NLP 6.0.4! This version brings advancements in text embeddings with the introduction of the MiniLM family, Spark DataFrame optimizations, and enhanced PDF document parsing. Upgrade to 6.0.4 to leverage these cutting-edge features and expand your NLP capabilities at scale.
Stay updated with our latest examples and tutorials by visiting our Medium - Spark NLP blog!
This release introduces a new family of efficient text embedding models:
MiniLMEmbeddings annotator, enabling the use of MiniLM models for generating highly efficient and effective sentence embeddings. These models are designed to provide strong performance while being significantly smaller and faster than larger alternatives, making them ideal for a wide range of NLP tasks requiring compact and powerful text representations. (Link to notebook)The PDF Reader and PdfToText transformer have been significantly improved for more comprehensive and fault-tolerant document parsing. (Link to notebook)
#PyPI
pip install spark-nlp==6.0.4
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.4
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.4
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.4
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.4
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.4
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.4</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.4</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.4</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.4</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.3...6.0.4
=======
We are excited to announce the release of Spark NLP 6.0.3! This version introduces significant advancements in multimodal capabilities and further ref
We are excited to announce the release of Spark NLP 6.0.3! This version introduces significant advancements in multimodal capabilities and further refines document processing workflows. Upgrade to 6.0.3 to leverage these cutting-edge features and expand your NLP and vision task capabilities at scale.
E5VEmbeddings, enabling universal multimodal embeddings with Multimodal Large Language Models (MLLMs). It can express semantic similarly between texts, images, or a combination of both.Partition and PartitionTransformer annotators with new character and title-based chunking strategies.sparknlp.read().xml() and integrated XML support into the Partition annotator for streamlined XML document processing.This release further boosts Spark NLP's multimodal processing power with the integration of E5-V.
E5VEmbeddings is designed to adapt MLLMs for achieving universal multimodal embeddings. It leverages MLLMs with prompts to effectively bridge the modality gap between different types of inputs, demonstrating strong performance in multimodal embeddings even without fine-tuning. (Link to notebook)The Partition and PartitionTransformer components now include additional chunking strategies and enhancements, which divides content into meaningful units based on the document's structure or number of characters.
maxCharacters): Split documents by number of characters.byTitle): Split documents by titles in the documents. Additional settings:newAfterNChars): Allows for early section breaks before reaching the maxCharacters threshold.overlapAll): Adds trailing context from the previous chunk to the next, improving semantic continuity.pageNumber metadata and starts a new section when a page changes.sparknlp.read().xml(): This method accepts file paths of XML content.Partition annotator by setting content_type = "application/xml".#PyPI
pip install spark-nlp==6.0.3
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.3
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.3
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.3
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.3</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.3</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.3</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.2...6.0.3
=======
We are thrilled to announce the release of Spark NLP 6.0.2! This version introduces powerful new multimodal models and significantly enhances document
We are thrilled to announce the release of Spark NLP 6.0.2! This version introduces powerful new multimodal models and significantly enhances document processing workflows. Upgrade to 6.0.2 to leverage these cutting-edge features and expand your NLP and vision task capabilities at scale.
Stay updated with our latest examples and tutorials by visiting our Medium - Spark NLP blog!
InternVLForMultiModal model, enabling advanced visual question answering with InternVL 2, 2.5, and 3 series models.Florance2Transformer, a sophisticated vision foundation model for diverse prompt-based vision and vision-language tasks like captioning, object detection, and segmentation.Partition and PartitionTransformer annotator for a unified and configurable interface with Spark NLP readers, simplifying unstructured data loading.This release significantly boosts Spark NLP's multimodal processing power with the integration of two new visual language models:
InternVLForMultiModal is a powerful multimodal large language model is specifically designed for visual question answering. This annotator is versatile, supporting the InternVL 2, 2.5, and 3 families of models, allowing users to tackle complex visual-linguistic tasks. (Link to notebook)Florance2Transformer, an advanced vision foundation model. Florence-2 utilizes a prompt-based approach, enabling it to perform a wide array of vision and vision-language tasks. Users can leverage simple text prompts to execute tasks such as image captioning, object detection, and image segmentation with high accuracy. (Link to notebook)Partition and PartitionTransformer annotator.
Partition provides a unified interface for extracting structured content from various document formats into Spark DataFrames. It supports input from files, URLs, in-memory strings, or byte arrays and handles formats such as text, HTML, Word, Excel, PowerPoint, emails, and PDFs. It automatically selects the appropriate reader based on file extension or MIME type and allows customization via parameters. (Link to notebook)PartitionTransformer annotator allows you to use the Partition feature more smoothly within existing Spark NLP workflows, enabling seamless reuse of your pipelines. PartitionTransformer can be used for extracting structured content from various document types using Spark NLP readers. It supports reading from files, URLs, in-memory strings, or byte arrays, and returns parsed output as a structured Spark DataFrame. (Link to notebook)AutoGGUFModel (#14576)#PyPI
pip install spark-nlp==6.0.2
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.2
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.2
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.2
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.2</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.2</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.2</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.1...6.0.2
AutoGGUFModel (#14576)=======
We are pleased to announce the release of Spark NLP 6.0.1, bringing exciting new vision features and continued enhancements. Expand your NLP capabilit
We are pleased to announce the release of Spark NLP 6.0.1, bringing exciting new vision features and continued enhancements. Expand your NLP capabilities at scale for a wide range of tasks by upgrading to 6.0.1 and leverage these powerful new additions and improvements!
We also have been adding blog posts covering various examples for our newest features. Check them out at Medium - Spark NLP!
This release adds support for several cutting-edge VLMs, significantly expanding the range of tasks you can tackle with Spark NLP:
The PDF Reader now includes additional parameters and options, providing users with more flexible and controlled ingestion of PDF documents, improving handling of various PDF structures. (link to notebook)
You can now
splitPage parameter to identify the correct number of pagesonlyPageNum parameter to display only the number of pages of the documenttextStripper parameter used for output layout and formattingsort parameter to enable or disable sorting linesThis release also includes fixes for several issues:
RoBERtaMultipleChoice, preventing these types of annotators to be loaded in Python#PyPI
pip install spark-nlp==6.0.1
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.1
Apple Silicon
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.1
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.1
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.1</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.1</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.1</version>
</dependency>
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/6.0.0...6.0.1
Introducing new Vision Language Models:
Adding additional parameters options to PDF Reader (SPARKNLP-1158)
=======
Nothing published for this version
With Spark NLP 6.0.0, we are setting a new standard for building scalable, distributed AI pipelines. This release transforms Spark NLP from a pure NLP
With Spark NLP 6.0.0, we are setting a new standard for building scalable, distributed AI pipelines. This release transforms Spark NLP from a pure NLP library into the de facto platform for distributed LLM ingestion and multimodal batch processing.
This release introduces native ingestion for enterprise file types including PDFs, Excel spreadsheets, PowerPoint decks, and raw text logs, with automatic structure extraction, semantic segmentation, and metadata preservation — all in scalable, zero-code Spark pipelines.
At the same time, Spark NLP now natively supports Vision-Language Models (VLMs), loading quantized multimodal models like LLAVA, Phi Vision, DeepSeek Janus, and Llama 3.2 Vision directly via Llama.cpp, ONNX, and OpenVINO runtimes with no external inference servers, no API bottlenecks.
With 6.0.0, Spark NLP offers a complete, distributed architecture for universal data ingestion, multimodal understanding, and LLM batch inference at scale — enabling retrieval-augmented generation (RAG), document understanding, compliance audits, enterprise search, and multimodal analytics — all within the native Spark ecosystem.
One unified framework. Text, vision, documents — at Spark scale. Zero boilerplate. Maximum performance.
Spark NLP 6.0.0 introduces the new AutoGGUFVisionModel, enabling native multimodal inference for quantized GGUF models directly within Spark pipelines. Powered by Llama.cpp, this annotator makes it effortless to run Vision-Language Models (VLMs) like LLAVA-1.5-7B Q4_0, Qwen2 VL, and others fully on-premises, at scale, with no external servers or APIs required.
With Spark NLP 6.0.0, Llama.cpp vision models are now first-class citizens inside DataFrames, delivering multimodal inference at scale with native Spark performance.
For the first time, Spark NLP supports pure vision-text workflows, allowing you to pass raw images and captions directly into LLMs that can describe, summarize, or reason over visual inputs.
This unlocks batch multimodal processing across massive datasets with Spark’s native scalability — perfect for product catalogs, compliance audits, document analysis, and more.
ImageAssembler.loadImagesAsBytes to prepare image datasets effortlessly.nCtx), top-k/top-p sampling, temperature, and repeat penalties, allowing fine control over completions.documentAssembler = DocumentAssembler() \
.setInputCol("caption") \
.setOutputCol("caption_document")
imageAssembler = ImageAssembler() \
.setInputCol("image") \
.setOutputCol("image_assembler")
data = ImageAssembler \
.loadImagesAsBytes(spark, "src/test/resources/image/") \
.withColumn("caption", lit("Caption this image."))
model = AutoGGUFVisionModel.pretrained() \
.setInputCols(["caption_document", "image_assembler"]) \
.setOutputCol("completions") \
.setBatchSize(4) \
.setNPredict(40) \
.setTopK(40) \
.setTopP(0.95) \
.setTemperature(0.05)
pipeline = Pipeline().setStages([documentAssembler, imageAssembler, model])
results = pipeline.fit(data).transform(data)
results.selectExpr("reverse(split(image.origin, '/'))[0] as image_name", "completions.result").show(truncate=False)
📚 A full notebook walkthrough is available here.
Font-aware PDF ingestion is now available with automatic page segmentation, encrypted file support, and token-level coordinate extraction, ideal for legal discovery and document Q&A.
Spark NLP can now ingest .xls and .xlsx files directly into Spark DataFrames with automatic schema detection, multiple sheet support, and rich-text extraction for LLM pipelines.
Spark NLP introduces a native reader for .ppt and .pptx files. Capture slides, speaker notes, themes, and alt text at the document level for downstream summarization and retrieval.
New Extractor and Cleaner annotators allow you to pull structured data (emails, IP addresses, dates) from text or clean noisy text artifacts like bullets, dashes, and non-ASCII characters at scale.
A high-performance TextReader is now available to load .txt, .csv, .log and similar files. It automatically detects encoding and line endings for massive ingestion jobs.
Spark NLP now supports vision-language models in GGUF format using the new AutoGGUFVisionModel annotator. Run models like LLAVA-1.5-7B Q4_0 or Qwen2 VL entirely within Spark using Llama.cpp, enabling native multimodal batch inference without servers.
The DeepSeek Janus model, tuned for instruction-following across text and images, is now fully integrated and available via a simple pretrained call.
Support for Alibaba’s Qwen-2 VL series (0.5B to 7B parameters) is now available. Use Qwen-2 checkpoints for OCR, product search, and multimodal retrieval tasks with unified APIs.
The new Phi3Vision annotator brings Microsoft’s Phi-3.5 multimodal model into Spark NLP. Process images and prompts together to generate grounded captions or visual Q&A results, all with a model footprint of less than 1 GB.
Spark NLP now supports LLAVA 1.5 (7B) natively for screenshot Q&A, chart reading, and UI testing tasks. Build fully distributed multimodal inference pipelines without external services or dependencies.
Cohere’s multilingual Command-R models (up to 35B parameters) are now fully integrated. Perform reasoning, RAG, and summarization tasks with no REST API latency and no token limits.
Spark NLP now supports the full OLMo suite of open-weight language models (7B, 1.7B, and more) directly in Scala and Python. OLMo models come with full training transparency, Dolma-sized vocabularies, and reproducible experiment logs, making them ideal for academic research and benchmarking.
New lightweight multiple-choice heads are now available for ALBERT, DistilBERT, RoBERTa, and XLM-RoBERTa models. These are perfect for building auto-grading systems, educational quizzes, and choice ranking pipelines.
AlbertForMultipleChoiceDistilBertForMultipleChoiceRoBertaForMultipleChoiceXlmRoBertaForMultipleChoiceThe Scala API for VisionEncoderDecoder has been fully refactored to expose .generate() parameters like batch size and maximum tokens, aligning it one-to-one with the Python API.
When a GGUF file is missing tensors or uses unsupported quantization, Spark NLP now provides clear and actionable error messages, including guidance on how to fix or convert the model.
A small typo related to the MXBAI integration was corrected to ensure consistency across annotator names and pretrained model references.
The Scala VisionEncoderDecoder wrapper has been updated to fully match the Python API. It now exposes parameters like batch size and maximum tokens, fixing discrepancies that could occur in cross-language pipelines.
Variable naming inconsistencies have been cleaned up throughout the codebase to ensure a more uniform and predictable developer experience.
We have added more than 110,000 new models and pipelines. The complete list of all 88,000+ models & pipelines in 230+ languages is available on our Models Hub.
Python
#PyPI
pip install spark-nlp==6.0.0
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:6.0.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:6.0.0
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:6.0.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:6.0.0
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>6.0.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>6.0.0</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>6.0.0</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>6.0.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-6.0.0.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-6.0.0.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-6.0.0.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-6.0.0.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.5.3...6.0.0
Introducing new large language models:
New MultipleChoice Transformers:
New file format support:
Other improvements:
========
We’re excited to introduce the latest release of Spark NLP 5.5.3, featuring critical enhancements and bug fixes for several of our Text Embeddings ann
We’re excited to introduce the latest release of Spark NLP 5.5.3, featuring critical enhancements and bug fixes for several of our Text Embeddings annotators. These improvements ensure even more reliable and efficient performance for your NLP workflows.
But that’s not all—we’re also celebrating a major milestone: crossing 100,000 truly free and open models on our Models Hub! This achievement underscores our commitment to making state-of-the-art NLP accessible to everyone, forever.
Upgrade today to take advantage of these enhancements, and thank you for being part of the Spark NLP community. Your support and contributions continue to drive innovation forward!
Previously, BGE embeddings used a fixed pooling strategy that didn't match all model variants, resulting in suboptimal performance for some models (cosine similarity around 0.97 compared to the original implementation). Different BGE models are trained with different pooling strategies - some use CLS token pooling while others use attention-based average pooling.
useCLSToken parameter to control embedding pooling strategyval embeddings = BGEEmbeddings.pretrained("bge_small_en_v1.5")
.setUseCLSToken(true) // Use CLS token pooling (default)
.setInputCols("document")
.setOutputCol("embeddings")
Fixed incorrect padding in attention mask calculations for multiple models:
This fix ensures consistent results between native implementations and ONNX versions.
Default Model Change:
Pooling Strategy:
useCLSToken parameter defaults to True// Using new default
val embeddingsNew = BGEEmbeddings.pretrained()
// Using previous default explicitly
val embeddingsOld = BGEEmbeddings.pretrained("bge_base")
// Using CLS token pooling
val embeddingsCLS = BGEEmbeddings.pretrained()
.setUseCLSToken(true)
// Using attention-based average pooling
val embeddingsAvg = BGEEmbeddings.pretrained()
.setUseCLSToken(false)
Python
#PyPI
pip install spark-nlp==5.5.3
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.3
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.3
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.3
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.5.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.5.3</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.5.3</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.5.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.5.3.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.5.3.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.5.3.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.5.3.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.5.2...5.5.3
useCLSToken parameter, which defaults to True. This affects the embedding calculation strategy. Existing users should verify their usage and potentially set this parameter explicitly.MPNet, BGE, E5, Mxbai, Nomic, SnowFlake, and UAE. This resulted in wrong inference results in some cases and not equal to the ONNX version in transformers/sentence-transformers.========
Apache Spark vulnerable Fix by @maziyarpanahi in https://github.com/JohnSnowLabs/spark-nlp/pull/14441
We’re thrilled to introduce the latest enhancements and new features in this release of Spark NLP! These additions bring more powerful model inference capabilities, seamless data ingestion methods, and greater flexibility for scaling your NLP workflows.
Upgrade today to take advantage of these new capabilities and improvements. As always, we look forward to your feedback and contributions, and thank you for being part of the Spark NLP community!
🚀 Major New Features
OpenVINO Support for Transformers (#14408) Many popular transformer-based annotators now leverage OpenVINO for faster inference on Intel hardware. Enjoy speedier pipelines across a wide array of models—such as DeBerta, DistilBert, RoBerta, XlmRoBerta, Albert, and more—enabling efficient, production-grade NLP at scale.
BLIPForQuestionAnswering Transformer (#14422) Introducing BLIPForQuestionAnswering, a new image-based question-answering transformer. Simply provide an image and a question, and BLIP will deliver contextually relevant answers. Perfect for use cases in image analysis, e-commerce, and beyond.
AutoGGUFEmbeddings Annotator (#14433) Seamlessly integrate AutoGGUFModels into your NLP pipeline. The new AutoGGUFEmbeddings annotator provides dense vector embeddings, making it easier than ever to incorporate advanced sentence embeddings into your workflows. We’ve included an end-to-end notebook to help you get started right away.
📜 New Data Ingestions
Parsing HTML to DataFrames (#14449) Need to analyze web content at scale? Use sparknlp.read().html() to parse local or remote HTML files into structured Spark DataFrames. This new feature makes web-scale data analysis and downstream NLP tasks more accessible and scalable.
Email Content to DataFrames (#14455)
Leverage sparknlp.read().email() to transform email content into organized DataFrames. Analyze communications, extract insights, and enrich your NLP pipelines with minimal effort. (Requires #14449 to be merged first.)
Microsoft Word Document Parsing (#14476) Turn .docx and .doc files into structured Spark DataFrames for streamlined integration into your NLP projects. From enterprise documents to reports, this feature simplifies data preparation and analysis at scale.
Microsoft Fabric Integration (#14467) We’ve added support for Microsoft Fabric to store and retrieve word embeddings efficiently. Leverage your existing infrastructure to scale Spark NLP solutions more effectively.
cuDNN Upgrade Instructions for Databricks (#14451) Easily upgrade cuDNN on Databricks to accelerate ONNX model inference on GPU, and take advantage of updated installation instructions for a cleaner setup.
Metadata Preservation in ChunkEmbeddings (#14462) ChunkEmbeddings now retain original metadata, ensuring richer context and more meaningful insights in your downstream tasks.
Default Names and Languages for New Annotators (#14469) We’ve standardized default names and languages in our seq2seq annotators for better clarity, consistency, and ease of use.
Updated:
New Additions for Email and Document Parsing:
Jakarta Mail (jakarta.mail:jakarta.mail-api:2.1.3): Added to support parsing and processing email content.
Angus Mail (org.eclipse.angus:angus-mail:2.0.3): Complementary mail handling library integrated for more robust email parsing capabilities.
Apache POI (org.apache.poi:poi-ooxml:4.1.2 & org.apache.poi:poi-scratchpad:4.1.2): Introduced for parsing Word documents (.docx and .doc) into structured DataFrames, enabling seamless integration of document-based data into Spark NLP workflows.
Python
#PyPI
pip install spark-nlp==5.5.2
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.2
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.2
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.2
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.5.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.5.2</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.5.2</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.5.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.5.2.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.5.2.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.5.2.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.5.2.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.5.1...5.5.2
========
:fire: Enhancements & Bug Fixes
BertForMultipleChoice Transformer Added. Enhanced BERT’s capabilities to handle multiple-choice tasks such as standardized test questions and survey or quiz automation.PromptAssembler Annotator Introduced. Introduced a new annotator that constructs prompts for LLMs using a chat template and a sequence of messages. Accepts an array of tuples with roles (“system”, “user”, “assistant”) and message texts. Utilizes llama.cpp as a backend for template parsing, supporting basic template applications.
Example NotebookpromptAssembler = (
PromptAssembler()
.setInputCol("messages")
.setOutputCol("prompt")
.setChatTemplate(template)
)
DBFS Systems.AutoGGUFModel pipelines on Databricks due to incorrect path handling of gguf files.Python
#PyPI
pip install spark-nlp==5.5.1
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.1
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.1
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.1
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.5.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.5.1</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.5.1</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.5.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.5.1.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.5.1.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.5.1.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.5.1.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.5.0...5.5.1
BertForMultipleChoice Transformer Added. Enhanced BERT’s capabilities to handle multiple-choice tasks such as standardized test questions and survey or quiz automation.PromptAssembler Annotator Introduced. Introduced a new annotator that constructs prompts for LLMs using a chat template and a sequence of messages. Accepts an array of tuples with roles (“system”, “user”, “assistant”) and message texts. Utilizes llama.cpp as a backend for template parsing, supporting basic template applications.========
We're thrilled to announce the release of Spark NLP 5.5.0, a groundbreaking update that pushes the boundaries of natural language processing! This rel
We're thrilled to announce the release of Spark NLP 5.5.0, a groundbreaking update that pushes the boundaries of natural language processing! This release is packed with exciting new features, optimizations, and integrations that will transform your NLP workflows. At the heart of this update is our game-changing integration with Llama.cpp, but that's just the beginning of what's in store!
We're proud to present the centerpiece of Spark NLP 5.5.0: the integration of Llama.cpp! This revolutionary addition brings unparalleled efficiency and performance to large language models within the Spark NLP ecosystem.
AutoGGUFModel annotator.This integration opens up new possibilities for deploying state-of-the-art language models in resource-constrained environments, making advanced NLP capabilities available to a wider range of applications and users.
We extend our heartfelt thanks to all contributors who made this release possible. Your innovative ideas, code contributions, and feedback continue to drive Spark NLP forward. Our Models Hub now contains over 83,000+ free and truly open-source models & pipelines. 🎉
We have added the QWEN2Transformer annotator, supporting the Qwen-2 model architecture known for its efficiency and performance in various NLP tasks like text generation and summarization.
The MiniCPM annotator is now available, providing support for the MiniCPM model designed for efficient language modeling with smaller parameter sizes without compromising performance.
We are excited to include the NLLB annotator, supporting No Language Left Behind models aimed at providing high-quality machine translation capabilities for a wide range of languages, especially low-resource languages.
Introducing support for Nomic Embeddings, which provide robust semantic representations for downstream tasks like clustering and classification.
We have implemented integration with Snowflake, allowing seamless data transfer and processing between Spark NLP and Snowflake data warehouses.
The CamemBertForZeroShotClassification annotator is now available, enabling zero-shot classification capabilities using the CamemBERT model, optimized for French language processing.
We have added support for MxBaiEmbeddings, providing embeddings from the MxBai model designed for multilingual text representation.
We have extended ONNX support to our vision annotators, allowing for optimized and accelerated inference for image-related NLP tasks.
Building upon our commitment to performance optimization, we have added OpenVINO and ONNX support to several additional annotators, ensuring you can leverage hardware acceleration across a broader range of models.
We are excited to introduce the AlbertForZeroShotClassification annotator, bringing zero-shot classification capabilities using the ALBERT model known for its parameter efficiency and strong performance.
We have integrated Phi-3 models into Spark NLP, providing enhanced performance with high-efficiency quantization, supporting INT4 and INT8 quantization for CPUs via OpenVINO.
The StarCoder2 model is now supported for causal language modeling tasks, enabling advanced code generation and understanding capabilities.
Continuing our support for the latest in language modeling, we have introduced support for LLAMA 3, bringing the latest advancements in the LLaMA model series to Spark NLP.
View Pull Requests, View Pull Request
View Pull Requests, View Pull Request, View Pull Request
vimtor/action-zip for creating artifacts to enhance compatibility and performance.Published New OpenVINO Artifacts: Built and published new OpenVINO artifacts for both CPU and GPU to enhance performance and compatibility.
Upgraded ONNX Runtime: Updated onnxruntime to the latest version for improved stability and performance on both CPU and GPU.
We have added more than 50,000 new models and pipelines. The complete list of all 83,000+ models & pipelines in 230+ languages is available on our Models Hub.
Python
#PyPI
pip install spark-nlp==5.5.0
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.5.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.5.0
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.5.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.5.0
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.5.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.5.0</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.5.0</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.5.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.5.0.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.5.0.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.5.0.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.5.0.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.4.2...5.5.0
========
Nothing published for this version
:fire: Enhancements & Bug Fixes
aggressiveMatching parameter to DateMatcher and MultiDateMatcher annotators https://github.com/JohnSnowLabs/spark-nlp/pull/14365aggressiveMatching parameter to DocumentSimilarityRanker annotator https://github.com/JohnSnowLabs/spark-nlp/pull/14370Python
#PyPI
pip install spark-nlp==5.4.2
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.2
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.2
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.2
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.4.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.4.2</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.4.2</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.4.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.4.2.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.4.2.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.4.2.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.4.2.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.4.1...5.4.2
========
:fire: New Features & Enhancements
Python
#PyPI
pip install spark-nlp==5.4.1
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.1
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.1
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.1
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.4.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.4.1</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.4.1</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.4.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.4.1.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.4.1.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.4.1.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.4.1.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.4.0...5.4.1
We're excited to share some amazing updates in the latest Spark NLP release of Spark NLP 🚀 5.4.0! This update is packed with new features and improvem
We're excited to share some amazing updates in the latest Spark NLP release of Spark NLP 🚀 5.4.0! This update is packed with new features and improvements that are set to transform natural language processing. One of the highlights is the integration of OpenVINO Runtime, which significantly boosts performance and efficiency across Intel hardware. You can now enjoy up to a 40% increase in performance compared to TensorFlow, with support for various model formats like ONNX, PaddlePaddle, TensorFlow, and TensorFlow Lite.
We've also added some powerful new annotators: BertEmbeddings, RoBertaEmbeddings, and XlmRoBertaEmbeddings. These are specially fine-tuned to take full advantage of the OpenVINO toolkit, offering better model accuracy and speed.
Another big change is in how we distribute models. We've moved from Broadcast to addFile for model distribution, which makes it easier to scale and manage large language models (LLMs) in cloud environments. This is especially helpful for models with over 7 billion parameters.
In addition, we've introduced the Mistral and Phi-2 architectures, optimized for high-efficiency quantization. There are also practical improvements to core components, like enhanced pooling for BERT-based models and updates to the OpenAIEmbeddings annotator for better performance and integration.
We want to thank our community for their valuable feedback, feature requests, and contributions. Our Models Hub now contains over 37,000+ free and truly open-source models & pipelines. 🎉
NEW Integration: OpenVINO Runtime for Spark NLP 🚀: We're thrilled to announce the integration of OpenVINO Runtime, enhancing Spark NLP with high-performance inference capabilities. OpenVINO Runtime supports direct reading of models in ONNX, PaddlePaddle, TensorFlow, and TensorFlow Lite formats, enabling out-of-the-box optimizations and superior performance on supported Intel hardware.
Enhanced Model Support and Performance Gains: The integration allows Spark NLP to utilize the OpenVINO Runtime API for Java, facilitating the loading and execution of models across various formats including ONNX, PaddlePaddle, TensorFlow, TensorFlow Lite, and OpenVINO IR. Impressively, benchmarks show up to a 40% performance improvement over TensorFlow with no additional tuning required. Additionally, users can harness the full optimization and quantization capabilities of the OpenVINO toolkit via the Model Conversion API.
Enabled Annotators: This update brings OpenVINO compatibility to a range of Spark NLP annotators, including BertEmbeddings, RoBertaEmbeddings, XlmRoBertaEmbeddings, T5Transformer, E5Embeddings, LLAMA2, Mistral, Phi2, and M2M100.
Acknowledgements: This significant enhancement was accomplished during Google Summer of Code 2023. Special thanks to Rajat Krishna (@rajatkrishna) and the entire OpenVINO team for their invaluable support and collaboration. https://github.com/JohnSnowLabs/spark-nlp/pull/14200
<img width="613" alt="bert-large-cased-bs4" src="https://github.com/JohnSnowLabs/spark-nlp/assets/5762953/3595b668-556a-47a7-97b6-a506fb340b89"> <img width="612" alt="roberta-large-bs4" src="https://github.com/JohnSnowLabs/spark-nlp/assets/5762953/1003d687-928a-46e6-96aa-11216a8d8d2d">
Mistral integration, featuring models fine-tuned on the MistralForCasualLM architecture. This addition enhances performance and efficiency by supporting quantization in INT4 and INT8 for CPUs via OpenVINO. https://github.com/JohnSnowLabs/spark-nlp/pull/14318<img width="1622" alt="image" src="https://github.com/JohnSnowLabs/spark-nlp/assets/5762953/d2720f5d-a9a0-4c87-8ce9-3e6caea706c2">
Performance of Mistral 7B and different Llama models on a wide range of benchmarks. For all metrics, all models were re-evaluated with our evaluation pipeline for accurate comparison. Mistral 7B significantly outperforms Llama 2 13B on all metrics, and is on par with Llama 34B (since Llama 2 34B was not released, we report results on Llama 34B). It is also vastly superior in code and reasoning benchmarks. https://mistral.ai/news/announcing-mistral-7b/
Continuing our commitment to user-friendly and scalable solutions, the integration of the Mistral architecture has been designed to be straightforward and easily adoptable, ensuring that users can leverage these enhancements without complexity:
doc_assembler = DocumentAssembler() \
.setInputCol("text") \
.setOutputCol("document")
mistral = MistralTransformer \
.pretrained() \
.setMaxOutputLength(50) \
.setDoSample(False) \
.setInputCols(["document"]) \
.setOutputCol("mistral_generation")
Phi-2, featuring models fine-tuned using the PhiForCausalLM architecture. This update enhances OpenVINO's capabilities, enabling quantization in INT4 and INT8 for CPUs to optimize both performance and efficiency. https://github.com/JohnSnowLabs/spark-nlp/pull/14318Continuing our commitment to user-friendly and scalable solutions, the integration of the Phi architecture has been designed to be straightforward and easily adoptable, ensuring that users can leverage these enhancements without complexity:
doc_assembler = DocumentAssembler() \
.setInputCol("text") \
.setOutputCol("document")
phi2 = Phi2Transformer \
.pretrained() \
.setMaxOutputLength(50) \
.setDoSample(False) \
.setInputCols(["document"]) \
.setOutputCol("phi2_generation")
addFile for deep learning distribution across any cluster. This change addresses the challenges of handling modern LLMs—some boasting over 7 billion parameters—by improving memory management and overcoming serialization limits previously encountered with Java Bytes and Apache Spark's Broadcast method. This update significantly boosts Spark NLP's ability to process LLMs efficiently, underscoring our dedication to delivering scalable NLP solutions.https://github.com/JohnSnowLabs/spark-nlp/pull/14236NEW: MPNetForTokenClassification Annotator: Introducing the MPNetForTokenClassification annotator in Spark NLP 🚀. This annotator efficiently loads MPNet models equipped with a token classification head (a linear layer atop the hidden-states output), ideal for Named-Entity Recognition (NER) tasks. It supports models trained or fine-tuned in ONNX format using MPNetForTokenClassification for PyTorch or TFCamembertForTokenClassification for TensorFlow from HuggingFace 🤗. [View Pull Request](https://github.com/JohnSnowLabs/spark-nlp/pull/14322
Enhanced Pooling for BERT, RoBERTa, and XLM-RoBERTa: We've added support for average pooling in BertSentenceEmbeddings, RoBertaSentenceEmbeddings, and XLMRoBertaEmbeddings annotators. This feature is especially useful when the [CLS] token is not fine-tuned for sentence embeddings via average pooling. View Pull Request
Refined OpenAIEmbeddings: Upgraded to support escape characters to prevent JSON content issues, changed the output annotator type from DOCUMENT to SENTENCE_EMBEDDINGS (note: this affects backward compatibility), enhanced output embeddings with metadata from the document column, introduced a Python unit test class, and added a new submodule for reliable saving/loading of the annotator. View Pull Request
New OpenVINO Notebooks: Released notebooks for exporting HuggingFace models using Optimum Intel and importing into Spark NLP. This update includes notebooks for BertEmbeddings, E5Embeddings, LLAMA2Transformer, RoBertaEmbeddings, XlmRoBertaEmbeddings, and T5Transformer. View Pull Request
Timeout waiting for connection from pool error that occurred when downloading multiple models simultaneously. View Pull Requestkeras.engine by updating the transformers version to 4.34.1. View Pull RequestUnsupported model IR version: 10, max supported IR version: 9 by setting the ONNX version to onnx==1.14.0. View Pull Requestjava.lang.NoSuchMethodError by ensuring compatibility with Spark 3.4 and updating documentation accordingly. View Pull RequestUAEEmbeddings. View Pull Requestonnxruntime to version 1.18.0 for enhanced stability and performance on both CPU and GPU.azure-identity to 1.12.2 and azure-storage-blob to 12.26.0 to improve security and integration with Azure services.Python
#PyPI
pip install spark-nlp==5.4.0
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.4.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.4.0
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.4.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.4.0
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.4.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.4.0</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.4.0</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.4.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.4.0.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.4.0.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.4.0.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.4.0.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.3.3...5.4.0
========
Nothing published for this version
Nothing published for this version
:fire: New Features & Enhancements
UAEEmbeddings for sentence embeddings using Universal AnglE Embedding, aimed at improving semantic textual similarity tasks.UAE is a novel angle-optimized text embedding model, designed to improve semantic textual similarity tasks, which are crucial for Large Language Model (LLM) applications. By introducing angle optimization in a complex space, AnglE effectively mitigates saturation of the cosine similarity function. https://arxiv.org/pdf/2309.12871.pdf
🔥 The universal English sentence embedding WhereIsAI/UAE-Large-V1 achieves SOTA on the MTEB Leaderboard with an average score of 64.64!
metadata.json, enhancing efficiency by avoiding unnecessary downloadsDocumentCharacterTextSplitterDeBertaForZeroShotClassificationBGEEmbeddings and MPNetEmbeddingsMPNetForQuestionAnsweringMPNetForSequenceClassification.onnx_data file, ensuring better reliability in model serialization processesMultilingual_Translation_with_M2M100.ipynb notebook entriesPython
#PyPI
pip install spark-nlp==5.3.3
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.3
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.3
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.3
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.3.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.3.3</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.3.3</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.3.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.3.3.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.3.3.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.3.3.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.3.3.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.3.2...5.3.3
Over 40 new interactive Streamlit demos https://github.com/JohnSnowLabs/spark-nlp/pull/14175
Streamlit demos https://github.com/JohnSnowLabs/spark-nlp/pull/14175XLMRoBertaForQuestionAnswering, XLMRoBertaForTokenClassification, and XLMRoBertaForSequenceClassification: Reverted the change in tfFile naming that was causing exceptions while loading and saving the models https://github.com/JohnSnowLabs/spark-nlp/pull/14204Python
#PyPI
pip install spark-nlp==5.3.2
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.2
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.2
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.2
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.2
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.2
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.3.2</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.3.2</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.3.2</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.3.2</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.3.2.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.3.2.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.3.2.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.3.2.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.3.1...5.3.2
========
Fix M2M100 not working on the second run (closing ONNX session by mistake) https://github.com/JohnSnowLabs/spark-nlp/commit/75d398e5184e91e92e05f6add4
M2M100 not working on the second run (closing ONNX session by mistake) https://github.com/JohnSnowLabs/spark-nlp/commit/75d398e5184e91e92e05f6add4e538cb4ce4ceb3ZeroShotNerClassification issue with NerConverter https://github.com/JohnSnowLabs/spark-nlp/pull/14186Python
#PyPI
pip install spark-nlp==5.3.1
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.1
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.1
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.1
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.1
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.1
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.3.1</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.3.1</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.3.1</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.3.1</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.3.1.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.3.1.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.3.1.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.3.1.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.3.0...5.3.1
We're thrilled to announce the release of Spark NLP 5.3.0, a monumental update that brings cutting-edge advancements and enhancements to the forefront
We're thrilled to announce the release of Spark NLP 5.3.0, a monumental update that brings cutting-edge advancements and enhancements to the forefront of Natural Language Processing (NLP). This release underscores our commitment to providing the NLP community with state-of-the-art tools and models, furthering our mission to democratize NLP technologies.
This release also addresses critical bug fixes, enhancing the stability and reliability of Spark NLP. Fixes include Spark NLP configuration adjustments, score calculation corrections, input validation, notebook improvements, and serialization issues.
We invite the community to explore these new features and enhancements, and we look forward to seeing the innovative applications that Spark NLP 5.3.0 will enable. 🌟
<img width="1990" alt="image" src="https://github.com/JohnSnowLabs/spark-nlp/assets/5762953/660957eb-f153-492b-879c-b0a680a2dbbc">
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our fine-tuned LLMs, called Llama 2-Chat, are optimized for dialogue use cases. Our models outperform open-source chat models on most benchmarks we tested, and based on our human evaluations for helpfulness and safety, may be a suitable substitute for closed-source models. We provide a detailed description of our approach to fine-tuning and safety improvements of Llama 2-Chat in order to enable the community to build on our work and contribute to the responsible development of LLMs. - https://ai.meta.com/research/publications/llama-2-open-foundation-and-fine-tuned-chat-models/
We have made LLAMA2Transformer annotator compatible with ONNX exports and quantizations:
As always, we made this feature super easy and scalable:
doc_assembler = DocumentAssembler() \
.setInputCol("text") \
.setOutputCol("documents")
llama2 = LLAMA2Transformer \
.pretrained() \
.setMaxOutputLength(50) \
.setDoSample(False) \
.setInputCols(["documents"]) \
.setOutputCol("generation")
We will continue improving this annotator and import more models in the future
M2M100 model sets a new benchmark for multilingual translation, supporting direct translation across 9,900 language pairs from 100 languages. This feature represents a significant leap in breaking down language barriers in global communication.Existing work in translation demonstrated the potential of massively multilingual machine translation by training a single model able to translate between any pair of languages. However, much of this work is English-Centric by training only on data which was translated from or to English. While this is supported by large sources of training data, it does not reflect translation needs worldwide. In this work, we create a true Many-to-Many multilingual translation model that can translate directly between any pair of 100 languages. We build and open source a training dataset that covers thousands of language directions with supervised data, created through large-scale mining. Then, we explore how to effectively increase model capacity through a combination of dense scaling and language-specific sparse parameters to create high quality models. Our focus on non-English-Centric models brings gains of more than 10 BLEU when directly translating between non-English directions while performing competitively to the best single systems of WMT. We open-source our scripts so that others may reproduce the data, evaluation, and final M2M-100 model. - https://arxiv.org/pdf/2010.11125.pdf
m2m100 = M2M100Transformer.pretrained() \
.setInputCols(["documents"]) \
.setMaxOutputLength(50) \
.setOutputCol("generation") \
.setSrcLang("zh") \
.setTgtLang("en")
DocumentSimilarity annotator, offering an efficient and scalable solution for ranking documents based on similarity, ideal for retrieval-augmented generation (RAG) applications.query = "Florence in Italy, is among the most beautiful cities in Europe."
doc_similarity_ranker = DocumentSimilarityRankerApproach()\
.setInputCols("sentence_embeddings")\
.setOutputCol("doc_similarity_rankings")\
.setSimilarityMethod("brp")\ # brp for BucketedRandomProjectionLSH and mh for MinHashLSH
.setNumberOfNeighbours(3)\
.setVisibleDistances(True)\
.setIdentityRanking(True)\
.asRetriever(query)
MPNetForSequenceClassification annotator for sequence classification tasks. This annotator is based on the MPNet architecture, enhances our capabilities in sequence classification tasks, offering more precise and context-aware processing.MPNetForQuestionAnswering annotator for question answering tasks. This annotator is based on the MPNet architecture, enhances our capabilities in question answering tasks, offering more precise and context-aware processing.DeBertaForZeroShotClassification annotator, leveraging the DeBERTa architecture, introduces sophisticated zero-shot classification capabilities, enabling the classification of text into predefined classes without direct example training.WordEmbeddingsModel annotator in serverless clusters. We initially introduced the in-memory feature for this annotator for users inside Kubernetes clusters without any HDFS. However, today it runs without any issue locally, on Google Colab, Kaggle, Databricks, AWS EMR, GCP, and AWS Glue.BertForZeroShotClassification annotator14.2, 14.3, 14.2 ML, 14.3 ML, 14.2 GPU, and 14.3 GPU.6.15.0 and 7.0.0.EntityRuler documentation.cluster_tmp_dir on Databricks' DBFS via spark.jsl.settings.storage.cluster_tmp_dir https://github.com/JohnSnowLabs/spark-nlp/issues/14129RoBertaForQuestionAnswering annotator https://github.com/JohnSnowLabs/spark-nlp/pull/14147Python
#PyPI
pip install spark-nlp==5.3.0
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.3.0
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.3.0
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.3.0
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.0
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.3.0
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, and 3.4.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.3.0</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.3.0</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.3.0</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.3.0</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.3.0.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.3.0.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.3.0.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.3.0.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.2.3...5.3.0
MPNetForSequenceClassification annotator for sequence classification tasks. This annotator is based on the MPNet architecture and is designed to classify sequences of text into a set of predefined classes.MPNetForQuestionAnswering annotator for question answering tasks. This annotator is based on the MPNet architecture and is designed to answer questions based on a given context.M2M100 state-of-the-art multilingual translation. M2M100 is a multilingual encoder-decoder (seq-to-seq) model trained for Many-to-Many multilingual translation. The model can directly translate between the 9,900 directions of 100 languages.DeBertaForZeroShotClassification annotator for zero-shot classification tasks. This annotator is based on the DeBERTa architecture and is designed to classify sequences of text into a set of predefined classes.DocumentSimilarityannotator. The new DocumentSimilarity ranker is a powerful tool for ranking documents based on their similarity to a given query document. It is designed to be efficient and scalable, making it ideal for a variety of RAG applications/BertForZeroShotClassification annotator.WordEmbeddingsModel annotator in server-less cluster. We initially introduced in-memory feature for this annotator for users inside Kubernetes cluster without any HDFS, however, today it runs without any issue locally, Google Colab, Kaggle, Databricks, AWS EMR, GCP, and AWS Glue.cluster_tmp_dir on Databricks' DBFS via spark.jsl.settings.storage.cluster_tmp_dir https://github.com/JohnSnowLabs/spark-nlp/issues/14129RoBertaForQuestionAnswering annotator https://github.com/JohnSnowLabs/spark-nlp/pull/14147========
Spark NLP 5.2.3 🚀 comes with an array of exciting features and optimizations. We're thrilled to announce support for ONNX Runtime in XLMRoBertaForToke
Spark NLP 5.2.3 🚀 comes with an array of exciting features and optimizations. We're thrilled to announce support for ONNX Runtime in XLMRoBertaForTokenClassification, XLMRoBertaForSequenceClassification, and XLMRoBertaForQuestionAnswering annotators. This release also showcases a significant refinement in the use of AWS SDK in Spark NLP, shifting from aws-java-sdk-bundle to aws-java-sdk-s3, resulting in a substantial ~320MB reduction in library size and a 20% increase in startup speed, new notebooks to import external models from Hugging Face, over 400+ new LLM models, and more!
We're pleased to announce that our Models Hub now boasts 36,000+ free and truly open-source models & pipelines 🎉. Our deepest gratitude goes out to our community for their invaluable feedback, feature suggestions, and contributions.
XLMRoBertaForTokenClassification annotatorXLMRoBertaForSequenceClassification annotatorXLMRoBertaForQuestionAnswering annotatoraws-java-sdk-bundle to the aws-java-sdk-s3 dependency. This change has resulted in a 318MB reduction in the library's overall size and has enhanced the Spark NLP startup time by 20%. For instance, using sparknlp.start() in Google Colab is now 14 to 20 seconds faster. Special thanks to @c3-avidmych for requesting this feature.DeBertaForQuestionAnswering, DebertaForSequenceClassification, and DeBertaForTokenClassification models from HuggingFaceDocumentTokenSplitter notebookINSTRUCTOR EmbeddingsRoBertaForTokenClassification notebookRoBertaForSequenceClassification notebookOpenAICompletion notebook with new gpt-3.5-turbo-instruct modelBGEEmbeddings not downloading in PythonT4 GPU runtime https://github.com/JohnSnowLabs/spark-nlp/issues/14109Python
#PyPI
pip install spark-nlp==5.2.3
Spark Packages
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x: (Scala 2.12):
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp_2.12:5.2.3
GPU
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-gpu_2.12:5.2.3
Apple Silicon (M1 & M2)
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-silicon_2.12:5.2.3
AArch64
spark-shell --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.2.3
pyspark --packages com.johnsnowlabs.nlp:spark-nlp-aarch64_2.12:5.2.3
Maven
spark-nlp on Apache Spark 3.0.x, 3.1.x, 3.2.x, 3.3.x, 3.4.x, and 3.5.x:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp_2.12</artifactId>
<version>5.2.3</version>
</dependency>
spark-nlp-gpu:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-gpu_2.12</artifactId>
<version>5.2.3</version>
</dependency>
spark-nlp-silicon:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-silicon_2.12</artifactId>
<version>5.2.3</version>
</dependency>
spark-nlp-aarch64:
<dependency>
<groupId>com.johnsnowlabs.nlp</groupId>
<artifactId>spark-nlp-aarch64_2.12</artifactId>
<version>5.2.3</version>
</dependency>
FAT JARs
CPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-assembly-5.2.3.jar
GPU on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-gpu-assembly-5.2.3.jar
M1 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-silicon-assembly-5.2.3.jar
AArch64 on Apache Spark 3.0.x/3.1.x/3.2.x/3.3.x/3.4.x/3.5.x: https://s3.amazonaws.com/auxdata.johnsnowlabs.com/public/jars/spark-nlp-aarch64-assembly-5.2.3.jar
Full Changelog: https://github.com/JohnSnowLabs/spark-nlp/compare/5.2.2...5.2.3
bundle and started to directly using S3 SDK. This will also minimize incompatibilities with other libraries that use AWS SDKsDocumentTokenSplitter notebookRoBertaForTokenClassification notebookRoBertaForSequenceClassification notebookOpenAICompletion notebook with new gpt-3.5-turbo-instruct modelBGEEmbeddings not downloading in Python========
Your coding agent can read these notes before it upgrades. Set up the MCP server →