NewYour coding agent can read the release notes before it upgrades.Set up the MCP server →
PyPI · #4587 most downloaded on PyPI
A Python NLP Library for Many Human Languages, by the Stanford NLP Group
Last release 3 days ago
01 Oct 2026
Release timing varies
gaps range from 2 weeks to 10 months
Nearly every release is documented
notes for 31 of 33 stable releases
Nothing withdrawn
no release was ever pulled
9 years old
34 releases · first in 2017
All tokenizer, MWT, POS, lemmatizer, and dependency parser models have been rebuilt using UD 2.18 datasets. The combined English, Spanish, and French
All tokenizer, MWT, POS, lemmatizer, and dependency parser models have been rebuilt using UD 2.18 datasets. The combined English, Spanish, and French packages have also been refreshed from the most recent development-branch snapshots to reflect recent improvements.
orjson. See GHSA-gh9r-c94j-cmp5. #1649Fix multi-word token IDs becoming lists instead of tuples after a JSON round-trip through to_serialized() / from_serialized(), which produced invalid CoNLL-U output (e.g. [3, 4] instead of 3-4) and broke downstream dictionary operations. Introduced in v1.14.0; addresses #1662. Thanks @Arthur031221! #1663
Fix empty words (enhanced UD) being misidentified as multi-word tokens when reconstructing a Document from dicts, causing incorrect token structure after a round-trip through to_dict(). Thanks @Arthur031221! #1664
Fix start/end character offsets not being assigned to words and tokens when reading a CoNLL-U Document. The fix populates offsets either by aligning tokens against the sentence text or by reading SpaceAfter annotations already present. #1656
Fix a KeyError when loading the English MIMIC no-CharLM lemmatizer package: charlm_forward_file and charlm_backward_file are now treated as optional keys rather than required ones. Addresses #1651. #1668
Embed external segmentation dictionaries (used by Thai, Japanese, and Chinese tokenizers) directly in the tokenizer model files, so they are not lost if models are rebuilt without them. The dictionaries are stored compressed. Adds unit tests to verify that models which should include a dictionary do. #1688
Allow en-dashes and em-dashes to function as standalone tokens without suppressing comma→dash augmentation, improving the tokenizer's ability to learn dash-connected word patterns from diverse training data. #1687
Add a warning when the constraint-repair loop reaches its maximum iteration count without fully resolving all violations, making incomplete repairs visible during debugging. #1643
Refactor the dependency parser DataLoader to separate PyTorch and non-PyTorch code paths, enabling dynamic augmentation of individual training examples on-the-fly rather than at initialization time. This gives more balanced augmentation across training and extends coverage to silver-annotated datasets. #1650
Generalize mixed_odia_dataset.py into mixed_indic_dataset.py, which can now build combined training datasets for any low-resource Indic target language. Originally developed for Sindhi, this release also uses it for Bhojpuri. The interface replaces five language-specific flags with a single --donors parameter; a DONOR_CONFIGS dictionary at the top of the script makes adding new donor languages a one-line change. #1655
Add a script for building tokenizer training sets from a mixture of multiple languages, useful for training tokenizers on related-language groups such as the Indic family. #1672
Move initial-punctuation stripping from data preparation into the DataLoader for tokenizer, POS tagger, and dependency parser, so it is applied dynamically during training rather than baked into data files. #1653
Integrate speaker information from ingested UDCoref documents into the coreference training data pipeline. #1645
Add a conversion script for the IIT (BHU) Bhojpuri POS corpus, transforming it from its mixed flat/SSF-bracket format into a standardized one-token-per-line layout. Thanks @abhiprd2000! #1675
Add multi-column xpos tagging support, allowing the tagger to train across multiple datasets with differing xpos schemes simultaneously. Applied to Bhojpuri (BHTB + IIT corpus), where it yields substantial improvements. Experiments on English (ParTUT and LinES) showed no benefit — English xpos accuracy is already saturated around 97.4%, which makes it a poor test case for a technique aimed at low-resource settings with small main treebanks. #1680
Upgrade the CoreNLP semgrex communication protocol to support enhanced queries. #1685
Update the CoreNLP installation script to report what was installed, be more conservative about which version to download, and fix a bug in the DEFAULT_CORENLP_URL constant. #1686
One column per quarter.
Fix a potential zip slip vulnerability when extracting downloaded model archives. While low-risk given that Stanza controls the resources being downlo…
Fix a potential zip slip vulnerability when extracting downloaded model archives. While low-risk given that Stanza controls the resources being downloaded, extraction now validates that no file paths escape the target directory. See GHSA-2fwf-f686-7p34. #1621
Restrict the unpickler used when deserializing annotated Documents, and add a deprecation warning: in a future release, Document serialization will move to JSON entirely, removing the pickle dependency. See GHSA-487q-m798-cp85. #1626
Remove shell subprocess calls from make_lm_data.py, addressing GHSA-c9h2-qmqw-qf6h. As a side benefit, the charlm data preparation script is now fully portable to Windows. #1623
ner= MISC key instead of coref_chains= when serializing a Document to CoNLL-U, causing collision with real NER labels on the same token. Thank you @devteamaegis! #1628sl_combined package mixes the SSJ and SST treebanks (reported by Kaja Dobrovoljc to be highly compatible), augments lemma and POS training with SUK 1.1 data, builds a lemma dictionary from Sloleks 3.1, and adds contextual lemma classifiers for the ambiguous pairs del/delo and rok/roka. #1625Reorganize the lemmatizer dictionary to use a pos → word → lemma layout and store it gzip-compressed. This dramatically reduces load time for large models — Slovenian drops from 30+ seconds to under 5 seconds — and shrinks model sizes considerably. A conversion script for updating locally trained 1.13.0 models is included.
Note: lemmatizer models from v1.13.0 are not compatible with v1.14.0. Please re-download or convert existing models. #1627
Reduce the hidden dimension of the contextual lemma classifier, making models smaller and faster without hurting accuracy. #1629
nsubj/csubj and obj relations: if the graph parser produces a node with multiple subjects or direct objects, the parser now reruns Chu-Liu-Edmonds iteratively (reusing the original neural scores) to find the best-scoring repair. This is on by default in the Pipeline. Addresses #1340. #1638Move comma-transposition augmentations from the data preparation script into the DataLoader, so that augmentation is applied on-the-fly during training rather than being baked in once at preprocessing time. This produces more balanced training and avoids accidentally affecting other annotators' data files. #1624
Add new structural feature functions to the tokenizer to help distinguish address-line formatting from normal running text, laying groundwork for fixing sentence-splitting errors on non-prose inputs. Addresses #1640. #1642
Add a stanza.utils.list_installed script that lists all locally cached Stanza models and their versions, without modifying anything on disk. Addresses #1542. #1632
Add a tokenize_with_speakers() convenience function for processing transcript-style text where each line begins with a speaker label, automatically assigning speaker metadata to sentences before passing them to the coref annotator. #1631
For researchers building character language models for new languages, this release includes updated tooling for collecting and deduplicating training data from OSCAR. The previous OSCAR 2023 source is no longer accessible to new users and is broken with datasets >= 4.0; the new scripts target the OSCAR Community Crawl instead. Addresses #1622.
Add a download script for the OSCAR Community Crawl that bypasses load_dataset (which has a known bug with OSCAR), along with an inventory script to inspect the language breakdown of downloaded chunks. Also adds OSCAR language codes to constant.py and fixes a bug where extra language name aliases were being silently clobbered. #1633
Switch the near-deduplication strategy from TLSH to MinHash LSH. MinHash is faster, retains more content, and still achieves satisfactory deduplication rates as verified by the diagnostic script included in this PR. #1639
Switch model downloads from raw requests calls to the huggingface_hub library when downloading from Hugging Face. This should be more reliable, take a
requests calls to the huggingface_hub library when downloading from Hugging Face. This should be more reliable, take advantage of XET when available, and benefit from the HF local cache. Addresses #1619, a report of intermittent download failures. #1614Add a Slovenian NER model trained on the UNER dataset, along with the scripts needed to process UNER data into Stanza's internal NER format. #1615
Update the default word vectors for Erzya (MYV) to use rootroo embeddings, and rebuild all MYV models accordingly. The rootroo vectors show clear improvements on POS (dev UPOS 90.81 vs 90.21 with mokha vectors) with a small gain on depparse as well. #1606
Fix a longstanding bug in the constituency parser output layer: the nonlinearity was missing between the last two linear layers. The buggy forward pass made those two layers mathematically fusable, so existing models have been condensed to 2 output layers with no loss in accuracy. #1610
Set the default number of output layers to 2. Experiments showed that 2- and 3-layer configurations perform equivalently (once the nonlinearity bug above is corrected), so models will now train with the smaller default. #1611
At the end of training, automatically condense output layer rows that weight decay has trained toward zero, shrinking the saved model and improving inference speed slightly. #1613
Several efficiency improvements to the parser state representation, improving throughput by roughly 20%. Changes include using type() instead of isinstance() for type checks (with appropriate guards), switching TreeStack from a namedtuple to __slots__, and storing transition scheme information as attributes on transition objects to avoid repeated accessor calls. #1603 #1612
Add a script for visualizing constituency parser model weights: outputs heatmaps of linear layer weights and time-series plots of LSTM gate statistics, useful for analyzing training behavior. #1605
Add support for controlling forget gate initialization and applying a separate weight decay to LSTM biases, based on Jozefowicz et al. Experiments showed a small improvement; this is now the default going forward. New models will be trained using this configuration. Future work: apply these LSTM training changes to other models, especially depparse #1609
Add a --gradient_checkpointing flag to the dependency parser training script, allowing fine-tuning of larger transformers under tighter memory constraints. #1592
Add a freeze → warmup → plateau learning rate scheduler (WarmupThenPlateauScheduler) for use in the dependency parser. This gives finer control over the stages of transformer fine-tuning. #1589
Detach transformer embeddings from the computation graph whenever the transformer is not actively being trained (e.g. during the frozen stage when bert_finetune=True). This reduces memory usage and speeds up the stages of training that don't update the transformer. #1590
Update language codes and treebank names throughout the codebase to align with UD 2.18. #1599
Add a --additional_files flag for building combined training datasets, and add the ability to construct a combined dataset for any language that has train/dev/test splits. #1598
Refactor bert_layer_mix into a trainable parameter that is passed directly into the bert_embeddings function, removing the need for each model to process the returned embeddings separately. #1597
Unify word embedding storage across models: models that were storing word vectors directly now use PretrainedWordVocab from the shared Pretrain object instead. This removes redundant storage and makes embedding handling consistent across annotators. #1600
This is the same as v1.12.1, with the following important update: all model now enforce weights_only=True when loading. This may obsolete some old mod
This is the same as v1.12.1, with the following important update: all model now enforce weights_only=True when loading. This may obsolete some old models (all models distributed with Stanza are already patched). If that happens, please load and then resave your model with an older version of Stanza, such as v1.12.1.
Enforce weights_only=True when loading the lemma classifier, addressing part of the security advisory GHSA-v5jw-96jm-7h2c. This should already be the default in later versions of PyTorch, but is now explicitly enforced. #1584
All models now have the code path which allows for weights_only=False removed. Instead, attempting to load a legacy model will throw an exception. If that happens, please load your model with an older version (such as 1.12.1) and resave it before proceeding. #1587
Add control characters to the set of characters treated as whitespace when tokenizing, fixing a bug where certain Unicode control characters (such as "region end" markers) were incorrectly attached to words. #1573 Addresses #1257
Add tokenizer augmentation that occasionally replaces commas with en-dashes or em-dashes, so that models trained on datasets that lack those characters learn to treat them similarly to commas. #1573
Add regression tests for Spanish tokenization errors reported in #1257 and tests for the whitespace/control-character handling and tokenizer augmentations. #1573
Enforce weights_only=True when loading the lemma classifier, avoiding a possible security risk. #1584
The lemma classifier for ja_gsd is now also attached to ja_combined. #1584
Train and attach two lemma classifiers to en_combined — both 's and her can be reliably classified from the available data. #1584
Add end-to-end unit tests for run_lemma.py, including training a lemmatizer and attaching multiple lemma classifiers. #1586
Add a silver dataset covering como_VERB in Spanish to the combined Spanish training data, addressing #1440. Also adds a utility to print a confusion matrix of tagging results filtered by a word regex (e.g. --upos_word_regex "^(?i:como)$"), making it easier to isolate the effects of annotation changes. #1579 stanfordnlp/handparsed-treebank@d0c29a3
Add silver training sentences covering unknown Spanish VERB lemmas to the combined Spanish lemmatizer, addressing #1255. Also includes a script to check lemmatizer results for a batch of word/POS combinations. #1580 stanfordnlp/handparsed-treebank@11327ef
Rewrite stanza-parseviewer.js to use a proper constituency parse visualizer instead of a repurposed dependency parse visualizer, fixing the broken vertical striping. Also adds a table of morphological features to the visualization. #1581 Addresses #1358
Various small improvements to the web demo: route all responses to /; templatize stanza-brat.html so the version number is sourced from _version.py; move the logo to the demo directory for easier serving; add favicon support to the pipeline demo; guard against empty POST requests. #1582
Enforce weights_only=True when loading the lemma classifier, addressing part of the security advisory GHSA-v5jw-96jm-7h2c. This should already be the
weights_only=True when loading the lemma classifier, addressing part of the security advisory GHSA-v5jw-96jm-7h2c. This should already be the default in later versions of PyTorch, but is now explicitly enforced. #1584Add control characters to the set of characters treated as whitespace when tokenizing, fixing a bug where certain Unicode control characters (such as "region end" markers) were incorrectly attached to words. #1573 Addresses #1257
Add tokenizer augmentation that occasionally replaces commas with en-dashes or em-dashes, so that models trained on datasets that lack those characters learn to treat them similarly to commas. #1573
Add regression tests for Spanish tokenization errors reported in #1257 and tests for the whitespace/control-character handling and tokenizer augmentations. #1573
Enforce weights_only=True when loading the lemma classifier, avoiding a possible security risk. #1584
The lemma classifier for ja_gsd is now also attached to ja_combined. #1584
Train and attach two lemma classifiers to en_combined — both 's and her can be reliably classified from the available data. #1584
Add end-to-end unit tests for run_lemma.py, including training a lemmatizer and attaching multiple lemma classifiers. #1586
Add a silver dataset covering como_VERB in Spanish to the combined Spanish training data, addressing #1440. Also adds a utility to print a confusion matrix of tagging results filtered by a word regex (e.g. --upos_word_regex "^(?i:como)$"), making it easier to isolate the effects of annotation changes. #1579 stanfordnlp/handparsed-treebank@d0c29a3
Add silver training sentences covering unknown Spanish VERB lemmas to the combined Spanish lemmatizer, addressing #1255. Also includes a script to check lemmatizer results for a batch of word/POS combinations. #1580 stanfordnlp/handparsed-treebank@11327ef
Rewrite stanza-parseviewer.js to use a proper constituency parse visualizer instead of a repurposed dependency parse visualizer, fixing the broken vertical striping. Also adds a table of morphological features to the visualization. #1581 Addresses #1358
Various small improvements to the web demo: route all responses to /; templatize stanza-brat.html so the version number is sourced from _version.py; move the logo to the demo directory for easier serving; add favicon support to the pipeline demo; guard against empty POST requests. #1582
New transformer packages, POS and depparse, for multiple languages. More can be added if a language you want to use does not already have a default_ac
New transformer packages, POS and depparse, for multiple languages. More can be added if a language you want to use does not already have a default_accurate package! Please file an issue on our github for that.
| Lang Code | Language | Transformer Model |
|---|---|---|
| bg | Bulgarian | rmihaylov/bert-base-bg |
| el | Greek | nlpaueb/bert-base-greek-uncased-v1 |
| hr | Croatian | classla/bcms-bertic |
| ka | Georgian | xlm-roberta-large |
| mt | Maltese | MaCoCu/MaltBERTa |
| nl | Dutch | DTAI-KULeuven/robbert-2023-dutch-large |
| ru | Russian | DeepPavlov/rubert-base-cased |
| sl | Slovenian | EMBEDDIA/crosloengual-bert |
| sr | Serbian | classla/bcms-bertic |
| sv | Swedish | KBLab/bert-base-swedish-cased |
Depparse bug fix: incorporate the bias in the biaffine model. Also, properly transpose the inputs. This actually did not change scores on average, weirdly enough. https://github.com/stanfordnlp/stanza/issues/308
Depparse can train with silver dataset: https://github.com/stanfordnlp/stanza/commit/1f2828df148287b5da9f230217ed7011c8fec479 This can also be used to train with two different datasets in equal weights or weighted via the --silver_weight flag
Depparse training option: can finetune only the last N layers of a transformer https://github.com/stanfordnlp/stanza/commit/e7245e3e5639e5fc81d950b5fdad069d013ee89e
Improved depparse optimizer scheduling - https://github.com/stanfordnlp/stanza/commit/0cf8654fbd09f4d52964a8c043e6c0a8ebb9fe99 The code was previously released in v1.11.1, but now all of the models are retrained with the new training scheme. A small sample of the results tested on a few transformer based depparse models, either with less strict stopping threshold or the two stage optimization, shows clear improvement (similar improvement in test scores):
| 5 model dev avg LAS | 1 stage | 1 stage 2k | 1 stage 4k | 2 stage |
|---|---|---|---|---|
| de_gsd | 89.03 | 89.50 | 89.71 | 89.83 |
| en_ewt | 93.47 | 93.69 | 93.74 | 93.89 |
| fi_tdt | 92.16 | 92.56 | 92.69 | 93.15 |
| it_vit | 90.12 | 90.37 | 90.44 | 90.60 |
| ta_ttb | 71.26 | 71.39 | 71.45 | 72.19 |
| zh-hans_gsdsimp | 85.47 | 85.69 | 85.76 | 85.89 |
"Smooth" MWT training by including a small fraction of non-MWT words in the training. https://github.com/stanfordnlp/stanza/pull/1568 Solves the problem of Finnish MWT having "t" at the end, but not at the start or middle, so natural words with "t" at the start would lead to the seq2seq model going haywire. https://github.com/stanfordnlp/stanza/issues/1562
Include in the default Finnish models a small snippet of sentences with non-MWT tokenization for certain non-MWT words. Addresses that some words such as tollei were treated as MWT https://github.com/stanfordnlp/stanza/commit/380aecd837229a7c8e9bca8bac111357a4756f25
The lemmatizer trains with silver tags (lower scores on gold, but better performance against real world text) https://github.com/stanfordnlp/stanza/issues/1567 https://github.com/stanfordnlp/stanza/commit/829f22ed901546bfd803f4feb7a0603639af82ae https://github.com/stanfordnlp/stanza/commit/14a97397a0bcdd97849b8ad81efd6a38ebc91e0b This update will be used for retraining against UD 2.18 when it is available.
Bugfix for training new MWT models from scratch - UD to internal format converter was not working https://github.com/stanfordnlp/stanza/commit/31df8e3f2e0183014bb9c6f2a833afa30ea0a7ab
Make it so download etc. don't automatically reset the logging level. That only happens if the user specifically sets the logging level in the function call https://github.com/stanfordnlp/stanza/pull/1551 https://github.com/stanfordnlp/stanza/issues/1418 https://github.com/stanfordnlp/stanza/pull/1569 Thank you @haoyu-haoyu
Update usage of morphseg: the latest version has a cleaner interface to the underlying model https://github.com/stanfordnlp/stanza/pull/1550 Thank you @TheWelcomer
Multi-doc wrapper to bulk_process Thank you @Rakshitha-Ireddi https://github.com/stanfordnlp/stanza/pull/1570
Utils for processing the coref output format Thank you @Rakshitha-Ireddi https://github.com/stanfordnlp/stanza/pull/1571
Add a human-readable coref output: https://github.com/stanfordnlp/stanza/issues/1560 https://github.com/stanfordnlp/stanza/commit/19c2b07043307cb752ca7c5ea3dbaf89de14dac5
Add a connection to the Morpheme Segmentation processor: 450ca74 #1527 https://github.com/TheWelcomer/MorphSeg Thank you @TheWelcomer !
platformdirs to put the downloaded models in the system cache directory by default. Thank you @McSinyx ! #1541 Note that this means if you have not set a default path, your existing ~/stanza_resources will be obsolete and you will now have models (with a version number) in .cache/stanzaAdd Abkhaz models from the fasttextwiki word vectors and the abnc UD dataset. This involved making the tagger & depparse train finetuned word vectors with a lower cutoff for small pretrains, as the fasttextwiki vectors were quite small for Abkhaz. #485 49f97a4 76f3335 Can add more test-only UD datasets on request, but the results seem low enough that we aren't doing it by default
ANG NER model downloaded from here: https://github.com/dmetola/Old_English-OEDT/tree/main 714072d
<PAD> as a relation type. A more principled fix would be to rebuild all the models, but this will work for now 284e9b4When training smaller POS datasets, finetune more words if the embedding is small. Makes it more likely that a small embedding is useful, since we can cover everything in the training set. 76f3335
Process !!! and ??? the same as ! and ? in the pos and depparse, addressing the downstream errors caused by unknown strings of punct. 5fd1d50 #1532
Train the tokenizer to recognize non-ascii variants of ! and ? with augmentation, addressing the tokenizer errors found for punct that doesn't exist in the training data #1532 d69c33f
Modify the depparse model to scale scores so that only one root is ever chosen. See https://aclanthology.org/2020.emnlp-main.390.pdf https://aclanthology.org/2021.emnlp-main.823/ 88c0cf6 c50fa5c
Fix random typos 4af05f6 Thanks @thecaptain789
Add an early termination for coref option, as requested in #1531 1d30e90
Update the semgrex client to allow for results to come back in non-sentence order, allowing for future addition of sort operators 46eb340
Fix a minor memory waste f8d62fe
Use the UD udtools package instead of having our own copy of the scoring script b20cd3a
As requested in #1523, allow for speaker information passing to the coref annotator: c4201b9 dc50998 1df3f8b
Add a convenience method to retag a conllu file in a Pipeline. call pipe.process_conllu(text) where text is a conllu file. 74fbdc4
cover up a jieba warning - package has not been updated in many years, not likely to be updated to fix deprecation errors any time soon. 0afdb61 thank…
Tokenizer can support the pretrained charlm now. This significantly improves the MWT performance on Hebrew, for example. #1511
Building tokenizers with pretrained charlm exposed a possible issue with the tokenizer including spaces when tokenizing when an MWT is split across two words. The effect occurred in Hebrew, but an English example would be wo n't tokenized as a single token with embedded space. Augmenting the training to enforce word splits across those spaces fixed the issue. 52cea78
use PackedSequence for the tokenizer - is slower, but results are stable when using inputs of different lengths: 4433e83 #1472
If a Tokenizer training set consistently has spaces between the ends of words and punctuations, the resulting trained model may not properly recognize the same text with periods at the end of the word. For example, this is a test . vs this is a test. Reported in #1504 Fixed for VI by 6878d8e
Coref now includes a zeros predictor - this predicts when a mention for certain datasets (such as Spanish) is a pro-drop mention. This behavior occurs by adding an empty node to the sentence. It can be disabled with the coref_use_zeros=False flag to the Pipeline. #1502
Sindhi pipeline based on the ISRA UD dataset, published at SyntaxFest 2025, with annotation support from MLtwist: https://aclanthology.org/2025.udw-1.11/
Tamil coreference model from KBC
update English lemmatizer with more verbs and ADJ from Prof. Lapalme
also, French lemmatizer changes with corrections from Prof. Lapalme
create a German lemmatizer using GSD data and a set of ADJ from Wiktionary
add GRC models mixed with a copy of the data with the diacritics stripped. because those work worse on GRC with diacritics, the originals are still the default: 5beca58
add a Thai TUD dataset from https://github.com/nlp-chula/TUD (not yet included in UD): bca078c
NER model for ANG: 68a56aa https://github.com/dmetola/Old_English-OEDT/tree/main
NER models for Hindi, Telegu, and Urdu: #1469, model built from https://github.com/ltrc/IL-NER, added in a4902df
improve efficiency of reading conllu documents: f15f0bc
sort CoNLLU features when outputting a doc, as is standard: aa20fbb
semgrex interface improvements: search all files, only output failed matches, process all documents at once
allow for combined depparse models with multiple training files in a zip file (easier to mix training data): be94ac6
lemmatizer can skip blank lemmas (useful when training using partially complete lemma data): 7c34714
if using pretokenized text in the NER, try to use the token text to extract the text (previously would crash): ab249f6
remove stray test output files: 2e4735a thanks to @otakutyrant
relative attention layer, similar to that used in https://aclanthology.org/2023.findings-emnlp.25/ #1474
output some basic analysis of errors: 5503c4c
current best conparser published at SyntaxFest 2025: https://aclanthology.org/2025.iwpt-1.4/
remove verbose from ReduceLROnPlateau: 1015b6b thanks to @otakutyrant
update usage of xml.etree.ElementTree to match updated python interface: 7ca8750 thanks to @otakutyrant
cover up a jieba warning - package has not been updated in many years, not likely to be updated to fix deprecation errors any time soon. 0afdb61 thanks to @otakutyrant
drop support for Python 3.8: 6420c3d thanks to @otakutyrant
update tomli version requirement, #1444 thanks to @BLKSerene
In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish.
In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish. We also add an Albanian model composed of the two available UD treebanks and an Old English model based on a prototype dataset not yet published in UD.
Other notable changes:
's -> be or have in the default_accurate package. Also built is a HI model. Others potentially to follow. Now with fewer bugs at startup. https://github.com/stanfordnlp/stanza/pull/1422weights_only=True when loading models. https://github.com/stanfordnlp/stanza/pull/1430 https://github.com/stanfordnlp/stanza/issues/1429' characters, including " used in "s - https://github.com/stanfordnlp/stanza/pull/1437 https://github.com/stanfordnlp/stanza/issues/1436CorrectForm annotations in the UD treebanks https://github.com/stanfordnlp/stanza/commit/dbdf429aff4175fec33856501e6899e96b390e86Bugfixes:
raise_for_status earlier when failing to download something, so that the proper error gets displayed.
Thank you @pattersam https://github.com/stanfordnlp/stanza/pull/1432In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish. We also add an Albanian model composed of the two available UD treebanks and an Old English model based on a prototype dataset not yet published in UD.
Other notable changes:
's -> be or have in the default_accurate package. Also built is a HI model. Others potentially to follow. Now with fewer bugs at startup. #1422weights_only=True when loading models. #1430 #1429' characters, including " used in "s - #1437 #1436CorrectForm annotations in the UD treebanks dbdf429Bugfixes:
raise_for_status earlier when failing to download something, so that the proper error gets displayed.In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish.
In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish. We also add an Albanian model composed of the two available UD treebanks and an Old English model based on a prototype dataset not yet published in UD.
Other notable changes:
's -> be or have in the default_accurate package. Also built is a HI model. Others potentially to follow. https://github.com/stanfordnlp/stanza/pull/1422weights_only=True when loading models. https://github.com/stanfordnlp/stanza/pull/1430 https://github.com/stanfordnlp/stanza/issues/1429' characters, including " used in "s - https://github.com/stanfordnlp/stanza/pull/1437 https://github.com/stanfordnlp/stanza/issues/1436CorrectForm annotations in the UD treebanks https://github.com/stanfordnlp/stanza/commit/dbdf429aff4175fec33856501e6899e96b390e86Bugfixes:
raise_for_status earlier when failing to download something, so that the proper error gets displayed.
Thank you @pattersam https://github.com/stanfordnlp/stanza/pull/1432In this release, we rebuild all of the models with UD 2.15, allowing for new languages such as Georgian, Komi Zyrian, Low Saxon, and Ottoman Turkish. We also add an Albanian model composed of the two available UD treebanks and an Old English model based on a prototype dataset not yet published in UD.
Other notable changes:
's -> be or have in the default_accurate package. Also built is a HI model. Others potentially to follow. #1422weights_only=True when loading models. #1430 #1429' characters, including " used in "s - #1437 #1436CorrectForm annotations in the UD treebanks dbdf429Bugfixes:
raise_for_status earlier when failing to download something, so that the proper error gets displayed.Added models which cover several different languages: one for combined Germanic and Romance languages, one for the Slavic languages available in UDCor
download_method=None now turns off HF downloads as well, for use in instances with no access to internet https://github.com/stanfordnlp/stanza/pull/1408 https://github.com/stanfordnlp/stanza/issues/1399Added models which cover several different languages: one for combined Germanic and Romantic languages, one for the Slavic languages available in UDCo
download_method=None now turns off HF downloads as well, for use in instances with no access to internet https://github.com/stanfordnlp/stanza/pull/1408 https://github.com/stanfordnlp/stanza/issues/1399Added models which cover several different languages: one for combined Germanic and Romantic languages, one for the Slavic languages available in UDCo
download_method=None now turns off HF downloads as well, for use in instances with no access to internet https://github.com/stanfordnlp/stanza/pull/1408 https://github.com/stanfordnlp/stanza/issues/1399Add an Old English pipeline, improve the handling of MWT for cases that should be easy, and improve the memory management of our usage of transformers
Add an Old English pipeline, improve the handling of MWT for cases that should be easy, and improve the memory management of our usage of transformers with adapters.
Fix words ending with -nna split into MWT https://github.com/stanfordnlp/handparsed-treebank/commit/2c48d4093daddc790bf89d7b35c47ee4d7d272d1 https://github.com/stanfordnlp/stanza/issues/1366
Fix MWT for English splitting into weird words by enforcing that the pieces add up to the whole (which is always the case in the English treebanks) https://github.com/stanfordnlp/stanza/issues/1371 https://github.com/stanfordnlp/stanza/pull/1378
Mark start_char and end_char on an MWT if it is composed of exactly its subwords https://github.com/stanfordnlp/stanza/commit/23840891c37d54a5cf491ea58b0702987dd4a6d7 https://github.com/stanfordnlp/stanza/issues/1361
Fix crash when trying to load previously unknown language https://github.com/stanfordnlp/stanza/issues/1360 https://github.com/stanfordnlp/stanza/commit/381736f8fb9b60a929002cc750bd0df3d7dad03a
Check that sys.stderr has isatty before manipulating it with tqdm, in case sys.stderr was monkeypatched: https://github.com/stanfordnlp/stanza/commit/d180ae02b278dd09dff53bc910e7aa43656e944d https://github.com/stanfordnlp/stanza/issues/1367
Try to avoid OOM in the POS in the Pipeline by reducing its max batch length https://github.com/stanfordnlp/stanza/commit/42718135e2ab4b145bbb5861d55bb9424ca3549f
Fix usage of gradient checkpointing & a weird interaction with Peft (thanks to @Jemoka) https://github.com/stanfordnlp/stanza/commit/597d48f1ead89fa9a0cca86cf9f0b530ed249792
Add * to the list of functional tags to drop in the constituency parser, helping Icelandic annotation https://github.com/stanfordnlp/stanza/commit/57bfa8bbd8d3d42d4ee29d4a406640b126ce0f46 https://github.com/stanfordnlp/stanza/issues/1356#issuecomment-1981216912
Can train depparse without using any of the POS columns, especially useful if training a cross-lingual parser: https://github.com/stanfordnlp/stanza/commit/4048caed1b89030082d23b8f71d23bae6c9c54f1 https://github.com/stanfordnlp/stanza/commit/15b136bb30dda272d318a61a5f602e7fc81e7a31
Add a constituency model for German https://github.com/stanfordnlp/stanza/commit/7a4f48c738f0db8923aa5da88d0a9743eaee4c6a https://github.com/stanfordnlp/stanza/commit/86ddaab31c73a7d0a389d0557f3696c29d441657 https://github.com/stanfordnlp/stanza/issues/1368
Fix deprecation warnings for escape sequences: https://github.com/stanfordnlp/stanza/pull/1321 https://github.com/stanfordnlp/stanza/issues/1293 Thank…
We integrate PEFT into our training pipeline for several different models. This greatly reduces the size of models with finetuned transformers, letting us make the finetuned versions of those models the default_accurate model.
The biggest gains observed are with the constituency parser and the sentiment classifier.
Previously, the default_accurate package used transformers where the head was trained but the transformer itself was not finetuned.
download_resources_json was broken: https://github.com/stanfordnlp/stanza/pull/1318 https://github.com/stanfordnlp/stanza/issues/1317 Thank you @ider-zh.get() https://github.com/stanfordnlp/stanza/commit/13ee3d5cbc2c9174c3e0c67ca75b580e4fe683b1 https://github.com/stanfordnlp/stanza/issues/1357device arg in MultilingualPipeline would crash if device was passed for an individual Pipeline: https://github.com/stanfordnlp/stanza/commit/44058a0ec296c6da5997bfaf8911a26d425d2cecFix deprecation warnings for escape sequences: https://github.com/stanfordnlp/stanza/pull/1321 https://github.com/stanfordnlp/stanza/issues/1293 Thank…
We integrate PEFT into our training pipeline for several different models. This greatly reduces the size of models with finetuned transformers, letting us make the finetuned versions of those models the default_accurate model.
The biggest gains observed are with the constituency parser and the sentiment classifier.
Previously, the default_accurate package used transformers where the head was trained but the transformer itself was not finetuned.
download_resources_json was broken: https://github.com/stanfordnlp/stanza/pull/1318 https://github.com/stanfordnlp/stanza/issues/1317 Thank you @ider-zhRemove deprecated output methods such as conll_as_string and doc2conll_text. Use "{:C}".format(doc) instead https://github.com/stanfordnlp/stanza/comm…
Conjunction-Aware Word-Level Coreference Resolution https://arxiv.org/abs/2310.06165 original implementation: https://github.com/KarelDO/wl-coref/tree/master
Updated form of Word-Level Coreference Resolution https://aclanthology.org/2021.emnlp-main.605/ original implementation: https://github.com/vdobrovolskii/wl-coref
If you use Stanza's coref module in your work, please be sure to cite both of the above papers.
Special thanks to vdobrovolskii, who graciously agreed to allow for integration of his work into Stanza, to @KarelDO for his support of his training enhancement, and to @Jemoka for the LoRA PEFT integration, which makes the finetuning of the transformer based coref annotator much less expensive.
Currently there is one model provided, a transformer based English model trained from OntoNotes. The provided model is currently based on Electra-Large, as that is more harmonious with the rest of our transformer architecture. When we have LoRA integration with POS, depparse, and the other processors, we will revisit the question of which transformer is most appropriate for English.
Future work includes ZH and AR models from OntoNotes, additional language support from UD-Coref, and lower cost non-transformer models
https://github.com/stanfordnlp/stanza/pull/1309
English now has an MWT model by default. Text such as won't is now marked as a single token, split into two words, will and not. Previously it was expected to be tokenized into two pieces, but the Sentence object containing that text would not have a single Token object connecting the two pieces. See https://stanfordnlp.github.io/stanza/mwt.html and https://stanfordnlp.github.io/stanza/data_objects.html#token for more information.
Code that used to operate with for word in sentence.words will continue to work as before, but for token in sentence.tokens will now produce one object for MWT such as won't, cannot, Stanza's, etc.
Pipeline creation will not change, as MWT is automatically (but not silently) added at Pipeline creation time if the language and package includes MWT.
https://github.com/stanfordnlp/stanza/pull/1314/commits/f22dceb93275fc724536b03b31c08a94617880ca https://github.com/stanfordnlp/stanza/pull/1314/commits/27983aefe191f6abd93dd49915d2515d7c3973d1
conll_as_string and doc2conll_text. Use "{:C}".format(doc) instead https://github.com/stanfordnlp/stanza/commit/e01650f9c56382495082a9a24fa0310414c46651doc_id field if the document they are created from has a doc_id. https://github.com/stanfordnlp/stanza/pull/1314/commits/8e2201f42cb99a5a3d8358ce59501c1d88f2585epeft module used for finetuning the transformer used in the coref processor does not support those versions.peft as an optional dependency to transformer based installationsnetworkx as a dependency for reading enhanced dependencies. Added toml as a dependency for reading the coref config.V1.6.1 is a patch of a bug in the Arabic POS tagger.
V1.6.1 is a patch of a bug in the Arabic POS tagger.
We also mark Python 3.11 as supported in the setup.py classifiers. This will be the last release that supports Python 3.6
The package parameter for building the Pipeline now has three default settings:
default, the same as before, where POS, depparse, and NER use the charlm, but lemma does notdefault-fast, where POS and depparse are built without the charlm, making them substantially faster on CPU. Some languages currently have non-charlm NER as welldefault-accurate, where the lemmatizer also uses the charlm, and other models use transformers if we have one for that language. Suggestions for more transformers to use are welcomeFurthermore, package dictionaries are now provided for each UD dataset which encompass the default versions of models for that dataset, although we do not further break that down into -fast and -accurate versions for each UD dataset.
PR: https://github.com/stanfordnlp/stanza/pull/1287
addresses https://github.com/stanfordnlp/stanza/issues/1259 and https://github.com/stanfordnlp/stanza/issues/1284
The NER models now can learn multiple output layers at once.
https://github.com/stanfordnlp/stanza/pull/1289
Theoretically this could be used to save a bit of time on the encoder while tagging multiple classes at once, but the main use case was to crosstrain the OntoNotes model on the WorldWide English newswire data we collected. The effect is that the model learns to incorporate some named entities from outside the standard OntoNotes vocabulary into the main 18 class tagset, even though the WorldWide training data is only 8 classes.
Results of running the OntoNotes model, with charlm but not transformer, on the OntoNotes and WorldWide test sets:
original ontonotes on worldwide: 88.71 69.29
simplify-separate 88.24 75.75
simplify-connected 88.32 75.47
We also produced combined models for nocharlm and with Electra as the input encoding. The new English NER models are the packages ontonotes-combined_nocharlm, ontonotes-combined_charlm, and ontonotes-combined_electra-large.
Future plans include using multiple NER datasets for other models as well.
Postprocessing of proposed tokenization possible with dependency injection on the Pipeline (ty @Jemoka). When creating a Pipeline, you can now provide a callable via the tokenize_postprocessor parameter, and it can adjust the candidate list of tokens to change the tokenization used by the rest of the Pipeline https://github.com/stanfordnlp/stanza/pull/1290
Finetuning for transformers in the NER models: have not yet found helpful settings, though https://github.com/stanfordnlp/stanza/commit/45ef5445f44222df862ed48c1b3743dc09f3d3fd
SE and SME should both represent Northern Sami, a weird case where UD didn't use the standard 2 letter code https://github.com/stanfordnlp/stanza/issues/1279 https://github.com/stanfordnlp/stanza/commit/88cd0df5da94664cb04453536212812dc97339bb
charlm for PT (improves accuracy on non-transformer models): https://github.com/stanfordnlp/stanza/commit/c10763d0218ce87f8f257114a201cc608dbd7b3a
build models with transformers for a few additional languages: MR, AR, PT, JA https://github.com/stanfordnlp/stanza/commit/45b387531c67bafa9bc41ee4d37ba0948daa9742 https://github.com/stanfordnlp/stanza/commit/0f3761ee63c57f66630a8e94ba6276900c190a74 https://github.com/stanfordnlp/stanza/commit/c55472acbd32aa0e55d923612589d6c45dc569cc https://github.com/stanfordnlp/stanza/commit/c10763d0218ce87f8f257114a201cc608dbd7b3a
V1.6.1 fixes a bug in the Arabic POS model which was an unfortunate side effect of the NER change to allow multiple tag sets at once: https://github.com/stanfordnlp/stanza/commit/b56f442d4d179c07411a44a342c224408eb6a6a9
Scenegraph CoreNLP connection needed to be checked before sending messages: https://github.com/stanfordnlp/CoreNLP/issues/1346#issuecomment-1713267522 https://github.com/stanfordnlp/stanza/commit/c71bf3fdac8b782a61454c090763e8885d0e3824
run_ete.py was not correctly processing the charlm, meaning the whole thing wouldn't actually run https://github.com/stanfordnlp/stanza/commit/16f29f3dcf160f0d10a47fec501ab717adf0d4d7
Chinese NER model was pointing to the wrong pretrain https://github.com/stanfordnlp/stanza/issues/1285 https://github.com/stanfordnlp/stanza/commit/82a02151da17630eb515792a508a967ef70a6cef
The package parameter for building the Pipeline now has three default settings:
The package parameter for building the Pipeline now has three default settings:
default, the same as before, where POS, depparse, and NER use the charlm, but lemma does notdefault-fast, where POS and depparse are built without the charlm, making them substantially faster on CPU. Some languages currently have non-charlm NER as welldefault-accurate, where the lemmatizer also uses the charlm, and other models use transformers if we have one for that language. Suggestions for more transformers to use are welcomeFurthermore, package dictionaries are now provided for each UD dataset which encompass the default versions of models for that dataset, although we do not further break that down into -fast and -accurate versions for each UD dataset.
PR: https://github.com/stanfordnlp/stanza/pull/1287
addresses https://github.com/stanfordnlp/stanza/issues/1259 and https://github.com/stanfordnlp/stanza/issues/1284
The NER models now can learn multiple output layers at once.
https://github.com/stanfordnlp/stanza/pull/1289
Theoretically this could be used to save a bit of time on the encoder while tagging multiple classes at once, but the main use case was to crosstrain the OntoNotes model on the WorldWide English newswire data we collected. The effect is that the model learns to incorporate some named entities from outside the standard OntoNotes vocabulary into the main 18 class tagset, even though the WorldWide training data is only 8 classes.
Results of running the OntoNotes model, with charlm but not transformer, on the OntoNotes and WorldWide test sets:
original ontonotes on worldwide: 88.71 69.29
simplify-separate 88.24 75.75
simplify-connected 88.32 75.47
We also produced combined models for nocharlm and with Electra as the input encoding. The new English NER models are the packages ontonotes-combined_nocharlm, ontonotes-combined_charlm, and ontonotes-combined_electra-large.
Future plans include using multiple NER datasets for other models as well.
Postprocessing of proposed tokenization possible with dependency injection on the Pipeline (ty @Jemoka). When creating a Pipeline, you can now provide a callable via the tokenize_postprocessor parameter, and it can adjust the candidate list of tokens to change the tokenization used by the rest of the Pipeline https://github.com/stanfordnlp/stanza/pull/1290
Finetuning for transformers in the NER models: have not yet found helpful settings, though https://github.com/stanfordnlp/stanza/commit/45ef5445f44222df862ed48c1b3743dc09f3d3fd
SE and SME should both represent Northern Sami, a weird case where UD didn't use the standard 2 letter code https://github.com/stanfordnlp/stanza/issues/1279 https://github.com/stanfordnlp/stanza/commit/88cd0df5da94664cb04453536212812dc97339bb
charlm for PT (improves accuracy on non-transformer models): https://github.com/stanfordnlp/stanza/commit/c10763d0218ce87f8f257114a201cc608dbd7b3a
build models with transformers for a few additional languages: MR, AR, PT, JA https://github.com/stanfordnlp/stanza/commit/45b387531c67bafa9bc41ee4d37ba0948daa9742 https://github.com/stanfordnlp/stanza/commit/0f3761ee63c57f66630a8e94ba6276900c190a74 https://github.com/stanfordnlp/stanza/commit/c55472acbd32aa0e55d923612589d6c45dc569cc https://github.com/stanfordnlp/stanza/commit/c10763d0218ce87f8f257114a201cc608dbd7b3a
Scenegraph CoreNLP connection needed to be checked before sending messages: https://github.com/stanfordnlp/CoreNLP/issues/1346#issuecomment-1713267522 https://github.com/stanfordnlp/stanza/commit/c71bf3fdac8b782a61454c090763e8885d0e3824
run_ete.py was not correctly processing the charlm, meaning the whole thing wouldn't actually run https://github.com/stanfordnlp/stanza/commit/16f29f3dcf160f0d10a47fec501ab717adf0d4d7
Chinese NER model was pointing to the wrong pretrain https://github.com/stanfordnlp/stanza/issues/1285 https://github.com/stanfordnlp/stanza/commit/82a02151da17630eb515792a508a967ef70a6cef
depparse can have transformer as an embedding https://github.com/stanfordnlp/stanza/pull/1282/commits/ee171cd167900fbaac16ff4b1f2fbd1a6e97de0a
depparse can have transformer as an embedding https://github.com/stanfordnlp/stanza/pull/1282/commits/ee171cd167900fbaac16ff4b1f2fbd1a6e97de0a
Lemmatizer can remember word,pos it has seen before with a flag https://github.com/stanfordnlp/stanza/issues/1263 https://github.com/stanfordnlp/stanza/commit/a87ffd0a4f43262457cf7eecf5555a621c6dc24e
Scoring scripts for Flair and spAcy NER models (requires the appropriate packages, of course) https://github.com/stanfordnlp/stanza/pull/1282/commits/63dc212b467cd549039392743a0be493cc9bc9d8 https://github.com/stanfordnlp/stanza/pull/1282/commits/c42aed569f9d376e71708b28b0fe5b478697ba05 https://github.com/stanfordnlp/stanza/pull/1282/commits/eab062341480e055f93787d490ff31d923a68398
SceneGraph connection for the CoreNLP client https://github.com/stanfordnlp/stanza/pull/1282/commits/d21a95cc90443ec4737de6d7ba68a106d12fb285
Update constituency parser to reduce the learning rate on plateau. Fiddling with the learning rates significantly improves performance https://github.com/stanfordnlp/stanza/pull/1282/commits/f753a4f35b7c2cf7e8e6b01da3a60f73493178e1
Tokenize [] based on () rules if the original dataset doesn't have [] in it https://github.com/stanfordnlp/stanza/pull/1282/commits/063b4ba3c6ce2075655a70e54c434af4ce7ac3a9
Attempt to finetune the charlm when building models (have not found effective settings for this yet) https://github.com/stanfordnlp/stanza/pull/1282/commits/048fdc9c9947a154d4426007301d63d920e60db0
Add the charlm to the lemmatizer - this will not be the default, since it is slower, but it is more accurate https://github.com/stanfordnlp/stanza/pull/1282/commits/e811f52b4cf88d985e7dbbd499fe30dbf2e76d8d https://github.com/stanfordnlp/stanza/pull/1282/commits/66add6d519deb54ca9be5fe3148023a5d7d815e4 https://github.com/stanfordnlp/stanza/pull/1282/commits/f086de2359cce16ef2718c0e6e3b5deef1345c74
Forgot to include the lemmatizer in CoreNLP 4.5.3, now in 4.5.4 https://github.com/stanfordnlp/stanza/commit/4dda14bd585893044708c70e30c1c3efec509863 https://github.com/bjascob/LemmInflect/issues/14#issuecomment-1470954013
prepare_ner_dataset was always creating an Armenian pipeline, even for non-Armenian langauges https://github.com/stanfordnlp/stanza/commit/78ff85ce7eed596ad195a3f26474065717ad63b3
Fix an empty bulk_process throwing an exception https://github.com/stanfordnlp/stanza/pull/1282/commits/5e2d15d1aa59e4a1fee8bba1de60c09ba21bf53e https://github.com/stanfordnlp/stanza/issues/1278
Unroll the recursion in the Tarjan part of the Chuliu-Edmonds algorithm - should remove stack overflow errors https://github.com/stanfordnlp/stanza/pull/1282/commits/e0917b0967ba9752fdf489b86f9bfd19186c38eb
Put NER and POS scores on one line to make it easier to grep for: https://github.com/stanfordnlp/stanza/commit/da2ae33e8ef9e48842685dfed88896b646dba8c4 https://github.com/stanfordnlp/stanza/commit/8c4cb04d38c1101318755270f3aa75c54236e3fe
Switch all pretrains to use a name which indicates their source, rather than the dataset they are used for: https://github.com/stanfordnlp/stanza/pull/1282/commits/d1c68ed01276b3cf1455d497057fbc0b82da49e5 and many others
Pipeline uses torch.no_grad() for a slight speed boost https://github.com/stanfordnlp/stanza/pull/1282/commits/36ab82edfc574d46698c5352e07d2fcb0d68d3b3
Generalize save names, which eventually allows for putting transformer, charlm or nocharlm in the save name - this lets us distinguish different complexities of model https://github.com/stanfordnlp/stanza/pull/1282/commits/cc0845826973576d8d8ed279274e6509250c9ad5 for constituency, and others for the other models
Add the model's flags to the --help for the run scripts, such as https://github.com/stanfordnlp/stanza/pull/1282/commits/83c0901c6ca2827224e156477e42e403d330a16e https://github.com/stanfordnlp/stanza/pull/1282/commits/7c171dd8d066c6973a8ee18a016b65f62376ea4c https://github.com/stanfordnlp/stanza/pull/1282/commits/8e1d112bee42f2211f5153fcc89083b97e3d2600
Remove the dependency on six https://github.com/stanfordnlp/stanza/pull/1282/commits/6daf97142ebc94cca7114a8cda5a20bf66f7f707 (thank you @BLKSerene )
VLSP constituency https://github.com/stanfordnlp/stanza/commit/500435d3ec1b484b0f1152a613716565022257f2
VLSP constituency -> tagging https://github.com/stanfordnlp/stanza/commit/cb0f22d7be25af0b3b2790e3ce1b9dbc277c13a7
CTB 5.1 constituency https://github.com/stanfordnlp/stanza/pull/1282/commits/f2ef62b96c79fcaf0b8aa70e4662d33b26dadf31
Add support for CTB 9.0, although those models are not distributed yet https://github.com/stanfordnlp/stanza/pull/1282/commits/1e3ea8a10b2e485bc7c79c6ab41d1f1dd8c2022f
Added an Indonesian charlm
Indonesian constituency from ICON treebank https://github.com/stanfordnlp/stanza/pull/1218
All languages with pretrained charlms now have an option to use that charlm for dependency parsing
French combined models out of GSD, ParisStories, Rhapsodie, and Sequoia https://github.com/stanfordnlp/stanza/pull/1282/commits/ba64d37d3bf21af34373152e92c9f01241e27d8b
UD 2.12 support https://github.com/stanfordnlp/stanza/pull/1282/commits/4f987d2cd708ce4ca27935d347bb5b5d28a78058
Headlining this release is the initial release of Ssurgeon, a rule-based dependency graph editing tool. Along with the existing Semgrex integration wi
Headlining this release is the initial release of Ssurgeon, a rule-based dependency graph editing tool. Along with the existing Semgrex integration with CoreNLP, Ssurgeon allows for rewriting of dependencies such as in the UD datasets. More information is in the GURT 2023 paper, https://aclanthology.org/2023.tlt-1.7/
In addition to this addition, there are two other CoreNLP integrations, a long list of bugfixes, a few other minor features, and a long list of constituency parser experiments which were somewhere between "ineffective" and "small improvements" and are available for people to experiment with.
detach().cpu() speeds things up significantly in some cases https://github.com/stanfordnlp/stanza/commit/ccfbc56b3b312fdde1350104a0d0d5645c9c80cc"{:C}" for document objects which prints out documents as CoNLL: https://github.com/stanfordnlp/stanza/pull/1169en_combined_bert model, others to come https://github.com/stanfordnlp/stanza/pull/1132Pipeline cache in Multilingual is a single OrderedDict https://github.com/stanfordnlp/stanza/issues/1115#issuecomment-1239759362 https://github.com/st
Pipeline cache in Multilingual is a single OrderedDict https://github.com/stanfordnlp/stanza/issues/1115#issuecomment-1239759362 https://github.com/stanfordnlp/stanza/commit/ba3f64d5f571b1dc70121551364fc89d103ca1cd
Don't require pytest for all installations unless needed for testing
https://github.com/stanfordnlp/stanza/issues/1120
https://github.com/stanfordnlp/stanza/commit/8c1d9d80e2e12729f60f05b81e88e113fbdd3482
hide SiLU and Minh imports if the version of torch installed doesn't have those nonlinearities https://github.com/stanfordnlp/stanza/issues/1120 https://github.com/stanfordnlp/stanza/commit/6a90ad4bacf923c88438da53219c48355b847ed3
Reorder & normalize installations in setup.py https://github.com/stanfordnlp/stanza/pull/1124
We improve the quality of the POS, constituency, and sentiment models, add an integration to displaCy, and add new models for a variety of languages.
We improve the quality of the POS, constituency, and sentiment models, add an integration to displaCy, and add new models for a variety of languages.
New Polish NER model based on NKJP from Karol Saputa and ryszardtuora https://github.com/stanfordnlp/stanza/issues/1070 https://github.com/stanfordnlp/stanza/pull/1110
Make GermEval2014 the default German NER model, including an optional Bert version https://github.com/stanfordnlp/stanza/issues/1018 https://github.com/stanfordnlp/stanza/pull/1022
Japanese conversion of GSD by Megagon https://github.com/stanfordnlp/stanza/pull/1038
Marathi NER dataset from L3Cube. Includes a Sentiment model as well https://github.com/stanfordnlp/stanza/pull/1043
Thai conversion of LST20 https://github.com/stanfordnlp/stanza/commit/555fc0342decad70f36f501a7ea1e29fa0c5b317
Kazakh conversion of KazNERD https://github.com/stanfordnlp/stanza/pull/1091/commits/de6cd25c2e5b936bc4ad2764b7b67751d0b862d7
Sentiment conversion of Tass2020 for Spanish https://github.com/stanfordnlp/stanza/pull/1104
VIT constituency dataset for Italian https://github.com/stanfordnlp/stanza/pull/1091/commits/149f1440dc32d47fbabcc498cfcd316e53aca0c6 ... and many subsequent updates
Combined UD models for Hebrew https://github.com/stanfordnlp/stanza/issues/1109 https://github.com/stanfordnlp/stanza/commit/e4fcf003feb984f535371fb91c9e380dd187fd12
For UD models with small train dataset & larger test dataset, flip the datasets UD_Buryat-BDT UD_Kazakh-KTB UD_Kurmanji-MG UD_Ligurian-GLT UD_Upper_Sorbian-UFAL https://github.com/stanfordnlp/stanza/issues/1030 https://github.com/stanfordnlp/stanza/commit/9618d60d63c49ec1bfff7416e3f1ad87300c7073
Spanish conparse model from multiple sources - AnCora, LDC-NW, LDC-DF https://github.com/stanfordnlp/stanza/commit/47740c6252a6717f12ef1fde875cf19fa1cd67cc
Pretrained charlm integrated into POS. Gives a small to decent gain for most languages without much additional cost https://github.com/stanfordnlp/stanza/pull/1086
Pretrained charlm integrated into Sentiment. Improves English, others not so much https://github.com/stanfordnlp/stanza/pull/1025
LSTM, 2d maxpool as optional items in the Sentiment
from the paper Text Classification Improved by Integrating Bidirectional LSTM with Two-dimensional Max Pooling
https://github.com/stanfordnlp/stanza/pull/1098
First learn with AdaDelta, then with another optimizer in conparse training. Very helpful https://github.com/stanfordnlp/stanza/commit/b1d10d3bdd892c7f68d2da7f4ba68a6ae3087f52
Grad clipping in conparse training https://github.com/stanfordnlp/stanza/commit/365066add019096332bcba0da4a626f68b70d303
GPU memory savings: charlm reused between different processors in the same pipeline https://github.com/stanfordnlp/stanza/pull/1028
Word vectors not saved in the NER models. Saves bandwidth & disk space https://github.com/stanfordnlp/stanza/pull/1033
Functions to return tagsets for NER and conparse models https://github.com/stanfordnlp/stanza/issues/1066 https://github.com/stanfordnlp/stanza/pull/1073 https://github.com/stanfordnlp/stanza/commit/36b84db71f19e37b36119e2ec63f89d1e509acb0 https://github.com/stanfordnlp/stanza/commit/2db43c834bc8adbb8b096cf135f0fab8b8d886cb
displaCy integration with NER and dependency trees https://github.com/stanfordnlp/stanza/commit/20714137d81e5e63d2bcee420b22c4fd2a871306
Fix that it takes forever to tokenize a single long token (catastrophic backtracking in regex) TY to Sk Adnan Hassan (VT) and Zainab Aamir (Stony Brook) https://github.com/stanfordnlp/stanza/pull/1056
Starting a new corenlp client w/o server shouldn't wait for the server to be available TY to Mariano Crosetti https://github.com/stanfordnlp/stanza/issues/1059 https://github.com/stanfordnlp/stanza/pull/1061
Read raw glove word vectors (they have no header information) https://github.com/stanfordnlp/stanza/pull/1074
Ensure that illegal languages are not chosen by the LangID model https://github.com/stanfordnlp/stanza/issues/1076 https://github.com/stanfordnlp/stanza/pull/1077
Fix cache in Multilingual pipeline https://github.com/stanfordnlp/stanza/issues/1115 https://github.com/stanfordnlp/stanza/commit/cdf18d8b19c92b0cfbbf987e82b0080ea7b4db32
Fix loading of previously unseen languages in Multilingual pipeline https://github.com/stanfordnlp/stanza/issues/1101 https://github.com/stanfordnlp/stanza/commit/e551ebe60a4d818bc5ba8880dda741cc8bd1aed7
Fix that conparse would occasionally train to NaN early in the training https://github.com/stanfordnlp/stanza/commit/c4d785729e42ac90f298e0ef4ab487d14fa35591
W&B integration for all models: can be activated with --wandb flag in the training scripts https://github.com/stanfordnlp/stanza/pull/1040
New webpages for building charlm, NER, and Sentiment https://stanfordnlp.github.io/stanza/new_language_charlm.html https://stanfordnlp.github.io/stanza/new_language_ner.html https://stanfordnlp.github.io/stanza/new_language_sentiment.html
Script to download Oscar 2019 data for charlm from HF (requires datasets module)
https://github.com/stanfordnlp/stanza/pull/1014
Unify sentiment training into a Python script, replacing the old shell script https://github.com/stanfordnlp/stanza/pull/1021 https://github.com/stanfordnlp/stanza/pull/1023
Convert sentiment to use .json inputs. In particular, this helps with languages with spaces in words such as Vietnamese https://github.com/stanfordnlp/stanza/pull/1024
Slightly faster charlm training https://github.com/stanfordnlp/stanza/pull/1026
Data conversion of WikiNER generalized for retraining / add new WikiNER models https://github.com/stanfordnlp/stanza/pull/1039
XPOS factory now determined at start of POS training. Makes addition of new languages easier https://github.com/stanfordnlp/stanza/pull/1082
Checkpointing and continued training for charlm, conparse, sentiment https://github.com/stanfordnlp/stanza/pull/1090 https://github.com/stanfordnlp/stanza/commit/0e6de808eacf14cd64622415eeaeeac2d60faab2 https://github.com/stanfordnlp/stanza/commit/e5793c9dd5359f7e8f4fe82bf318a2f8fd190f54
Option to write the results of a NER model to a file https://github.com/stanfordnlp/stanza/pull/1108
Add fake dependencies to a conllu formatted dataset for better integration with evaluation tools https://github.com/stanfordnlp/stanza/commit/6544ef3fa5e4f1b7f06dbcc5521fbf9b1264197a
Convert an AMT NER result to Stanza .json https://github.com/stanfordnlp/stanza/commit/cfa7e496ca7c7662478e03c5565e1b2b2c026fad
Add a ton of language codes, including 3 letter codes for languages we generally treat as 2 letters https://github.com/stanfordnlp/stanza/commit/5a5e9187f81bd76fcd84ad713b51215b64234986 https://github.com/stanfordnlp/stanza/commit/b32a98e477e9972737ad64deea0bda8d6cebb4ec and others
As part of the new Stanza release, we integrate transformer inputs to the NER and conparse modules. In addition, we now support several additional lan
As part of the new Stanza release, we integrate transformer inputs to the NER and conparse modules. In addition, we now support several additional languages for NER and conparse.
Download resources.json and models into temp dirs first to avoid race conditions between multiple processors https://github.com/stanfordnlp/stanza/issues/213 https://github.com/stanfordnlp/stanza/pull/1001
Download models for Pipelines automatically, without needing to call stanza.download(...)
https://github.com/stanfordnlp/stanza/issues/486
https://github.com/stanfordnlp/stanza/pull/943
Add ability to turn off downloads https://github.com/stanfordnlp/stanza/commit/68455d895986357a2c1f496e52c4e59ee0feb165
Add a new interface where both processors and package can be set https://github.com/stanfordnlp/stanza/issues/917 https://github.com/stanfordnlp/stanza/commit/f37042924b7665bbaf006b02dcbf8904d71931a1
When using pretokenized tokens, get character offsets from text if available https://github.com/stanfordnlp/stanza/issues/967 https://github.com/stanfordnlp/stanza/pull/975
If Bert or other transformers are used, cache the models rather than loading multiple times https://github.com/stanfordnlp/stanza/pull/980
Allow for disabling processors on individual runs of a pipeline https://github.com/stanfordnlp/stanza/issues/945 https://github.com/stanfordnlp/stanza/pull/947
Add # text and # sent_id to conll output https://github.com/stanfordnlp/stanza/discussions/918 https://github.com/stanfordnlp/stanza/pull/983 https://github.com/stanfordnlp/stanza/pull/995
Add ner to the token conll output https://github.com/stanfordnlp/stanza/discussions/993 https://github.com/stanfordnlp/stanza/pull/996
Fix missing Slovak MWT model https://github.com/stanfordnlp/stanza/issues/971 https://github.com/stanfordnlp/stanza/commit/5aa19ec2e6bc610576bc12d226d6f247a21dbd75
Upgrades to EN, IT, and Indonesian models https://github.com/stanfordnlp/stanza/issues/1003 https://github.com/stanfordnlp/stanza/pull/1008 IT improvements with the help of @attardi and @msimi
Fix improper tokenization of Chinese text with leading whitespace https://github.com/stanfordnlp/stanza/issues/920 https://github.com/stanfordnlp/stanza/pull/924
Check if a CoreNLP model exists before downloading it (thank you @interNULL) https://github.com/stanfordnlp/stanza/pull/965
Convert the run_charlm script to python https://github.com/stanfordnlp/stanza/pull/942
Typing and lint fixes (thank you @asears) https://github.com/stanfordnlp/stanza/pull/833 https://github.com/stanfordnlp/stanza/pull/856
stanza-train examples now compatible with the python training scripts https://github.com/stanfordnlp/stanza/issues/896
Bert integration (not by default, thank you @vythaihn) https://github.com/stanfordnlp/stanza/pull/976
Swedish model (thank you @EmilStenstrom) https://github.com/stanfordnlp/stanza/issues/912 https://github.com/stanfordnlp/stanza/pull/857
Persian model https://github.com/stanfordnlp/stanza/issues/797
Danish model https://github.com/stanfordnlp/stanza/pull/910/commits/3783cc494ee8c6b6d062c4d652a428a04a4ee839
Norwegian model (both NB and NN) https://github.com/stanfordnlp/stanza/pull/910/commits/31fa23e5239b10edca8ecea46e2114f9cc7b031d
Use updated Ukrainian data (thank you @gawy) https://github.com/stanfordnlp/stanza/pull/873
Myanmar model (thank you UCSY) https://github.com/stanfordnlp/stanza/pull/845
Training improvements for finetuning models https://github.com/stanfordnlp/stanza/issues/788 https://github.com/stanfordnlp/stanza/pull/791
Fix inconsistencies in B/S/I/E tags https://github.com/stanfordnlp/stanza/issues/928#issuecomment-1027987531 https://github.com/stanfordnlp/stanza/pull/961
Add an option for multiple NER models at the same time, merging the results together https://github.com/stanfordnlp/stanza/issues/928 https://github.com/stanfordnlp/stanza/pull/955
Dynamic oracle (improves accuracy a bit) https://github.com/stanfordnlp/stanza/pull/866
Missing tags now okay in the parser https://github.com/stanfordnlp/stanza/issues/862 https://github.com/stanfordnlp/stanza/commit/04dbf4f65e417a2ceb19897ab62c4cf293187c0b
bugfix of () not being escaped when output in a tree https://github.com/stanfordnlp/stanza/commit/eaf134ca699aca158dc6e706878037a20bc8cbd4
charlm integration by default https://github.com/stanfordnlp/stanza/pull/799
Bert integration (not the default model) (thank you @vythaihn and @hungbui0411) https://github.com/stanfordnlp/stanza/commit/05a0b04ee6dd701ca1c7c60197be62d4c13b17b6 https://github.com/stanfordnlp/stanza/commit/0bbe8d10f895560a2bf16f542d2e3586d5d45b7e
Preemptive bugfix for incompatible devices from @zhaochaocs https://github.com/stanfordnlp/stanza/issues/989 https://github.com/stanfordnlp/stanza/pull/1002
New models: DA, based on Arboretum IT, based on the Turin treebank JA, based on ALT PT, based on Cintil TR, based on Starlang ZH, based on CTB7
Stanza 1.3.0 introduces a language id model, a constituency parser, a dictionary in the tokenizer, and some additional features and bugfixes.
Stanza 1.3.0 introduces a language id model, a constituency parser, a dictionary in the tokenizer, and some additional features and bugfixes.
Langid model and multilingual pipeline Based on "A reproduction of Apple's bi-directional LSTM models for language identification in short strings." by Toftrup et al 2021 (https://github.com/stanfordnlp/stanza/commit/154b0e8e59d3276744ae0c8ea56dc226f777fba8)
Constituency parser
Based on "In-Order Transition-based Constituent Parsing" by Jiangming Liu and Yue Zhang. Currently an en_wsj model available, with more to come.
(https://github.com/stanfordnlp/stanza/commit/90318023432d584c62986123ef414a1fa93683ca)
Evalb interface to CoreNLP Useful for evaluating the parser - requires CoreNLP 4.3.0 or later
Dictonary tokenizer feature Noticeably improved performance for ZH, VI, TH (https://github.com/stanfordnlp/stanza/pull/776)
HuggingFace integration No more git issues complaining about unavailable models! (Hopefully) (https://github.com/stanfordnlp/stanza/commit/f7af5049568f81a716106fee5403d339ca246f38)
Sentiment processor crashes on certain inputs (issue https://github.com/stanfordnlp/stanza/issues/804, fixed by https://github.com/stanfordnlp/stanza/commit/e232f67f3850a32a1b4f3a99e9eb4f5c5580c019)
In anticipation of a larger release with some new features, we make a small update to fix some existing bugs and add two more NER models.
In anticipation of a larger release with some new features, we make a small update to fix some existing bugs and add two more NER models.
Sentiment models would crash on no text (issue https://github.com/stanfordnlp/stanza/issues/769, fixed by https://github.com/stanfordnlp/stanza/pull/781/commits/47889e3043c27f9c5abd9913016929f1857de7bf)
Java processes as a context were not properly closed (https://github.com/stanfordnlp/stanza/pull/781/commits/a39d2ff6801a23aa73add1f710d809a9c0a793b1)
Downloading tokenize now downloads mwt for languages which require it (issue https://github.com/stanfordnlp/stanza/issues/774, fixed by https://github.com/stanfordnlp/stanza/pull/777, from davidrft)
NER model can finetune and save to/from different filenames (https://github.com/stanfordnlp/stanza/pull/781/commits/0714a0134f0af6ef486b49ce934f894536e31d43)
NER model now displays a confusion matrix at the end of training (https://github.com/stanfordnlp/stanza/pull/781/commits/9bbd3f712f97cb2702a0852e1c353d4d54b4b33b)
Afrikaans, trained in NCHLT (https://github.com/stanfordnlp/stanza/pull/781/commits/6f1f04b6d674691cf9932d780da436063ebd3381)
Italian, trained on a model from FBK (https://github.com/stanfordnlp/stanza/pull/781/commits/d9a361fd7f13105b68569fddeab650ea9bd04b7f)
A regression in NER results occurred in 1.2.1 when fixing a bug in VI models based around spaces.
A regression in NER results occurred in 1.2.1 when fixing a bug in VI models based around spaces.
Fix Sentiment not loading correctly on Windows because of pickling issue (https://github.com/stanfordnlp/stanza/pull/742) (thanks to @BramVanroy)
Fix NER bulk process not filling out data structures as expected (https://github.com/stanfordnlp/stanza/issues/721) (https://github.com/stanfordnlp/stanza/pull/722)
Fix NER space issue causing a performance regression (https://github.com/stanfordnlp/stanza/issues/739) (https://github.com/stanfordnlp/stanza/pull/732)
All models other than NER and Sentiment were retrained with the new UD 2.8 release. All of the updates include the data augmentation fixes applied in
All models other than NER and Sentiment were retrained with the new UD 2.8 release. All of the updates include the data augmentation fixes applied in 1.2.0, along with new augmentations tokenization issues and end-of-sentence issues. This release also features various enhancements, bug fixes, and performance improvements, along with 4 new NER models.
Add Bulgarian, Finnish, Hungarian, Vietnamese NER models
Use new word vectors for Armenian, including better coverage for the new Western Armenian dataset(https://github.com/stanfordnlp/stanza/pull/718/commits/d9e8301addc93450dc880b06cb665ad10d869242)
Add copy mechanism in the seq2seq model. This fixes some unusual Spanish multi-word token expansion errors and potentially improves lemmatization performance. (https://github.com/stanfordnlp/stanza/pull/692 https://github.com/stanfordnlp/stanza/issues/684)
Fix Spanish POS and depparse mishandling a leading ¿ missing (https://github.com/stanfordnlp/stanza/pull/699 https://github.com/stanfordnlp/stanza/issues/698)
Fix tokenization breaking when a newline splits a Chinese token(https://github.com/stanfordnlp/stanza/pull/632 https://github.com/stanfordnlp/stanza/issues/531)
Fix tokenization of parentheses in Chinese(https://github.com/stanfordnlp/stanza/commit/452d842ed596bb7807e604eeb2295fd4742b7e89)
Fix various issues with characters not present in UD training data such as ellipses characters or unicode apostrophe (https://github.com/stanfordnlp/stanza/pull/719/commits/db0555253f0a68c76cf50209387dd2ff37794197 https://github.com/stanfordnlp/stanza/pull/719/commits/f01a1420755e3e0d9f4d7c2895e0261e581f7413 https://github.com/stanfordnlp/stanza/pull/719/commits/85898c50f14daed75b96eed9cd6e9d6f86e2d197)
Fix a variety of issues with Vietnamese tokenization - remove language specific model improvement which got roughly 1% F1 but caused numerous hard-to-track issues (https://github.com/stanfordnlp/stanza/pull/719/commits/3ccb132e03ce28a9061ec17d2c0ae84cc2000548)
Fix spaces in the Vietnamese words not being found in the embedding used for POS and depparse(https://github.com/stanfordnlp/stanza/pull/719/commits/197212269bc33b66759855a5addb99d1f465e4f4)
Include UD_English-GUMReddit in the GUM models(https://github.com/stanfordnlp/stanza/pull/719/commits/9e6367cb9bdd635d579fd8d389cb4d5fa121c413)
Add Pronouns & PUD to the mixed English models (various data improvements made this more appealing)(https://github.com/stanfordnlp/stanza/pull/719/commits/f74bef7b2ed171bf9c027ae4dfd3a10272040a46)
Add ability to pass a Document to the pipeline in pretokenized mode(https://github.com/stanfordnlp/stanza/commit/f88cd8c2f84aedeaec34a11b4bc27573657a66e2 https://github.com/stanfordnlp/stanza/issues/696)
Track comments when reading and writing conll files (https://github.com/stanfordnlp/stanza/pull/676 originally from @danielhers in https://github.com/stanfordnlp/stanza/pull/155)
Add a proxy parameter for downloads to pass through to the requests module (https://github.com/stanfordnlp/stanza/pull/638)
add sent_idx to tokens (https://github.com/stanfordnlp/stanza/commit/ee6135c538e24ff37d08b86f34668ccb223c49e1)
Fix Windows encoding issues when reading conll documents from @yanirmr (b40379eaf229e7ffc7580def57ee1fad46080261 https://github.com/stanfordnlp/stanza/pull/695)
Fix tokenization breaking when second batch is exactly eval_length(https://github.com/stanfordnlp/stanza/commit/726368644d7b1019825f915fabcfe1e4528e068e https://github.com/stanfordnlp/stanza/issues/634 https://github.com/stanfordnlp/stanza/issues/631)
Bulk process for tokenization - greatly speeds up the use case of many small docs (https://github.com/stanfordnlp/stanza/pull/719/commits/5d2d39ec822c65cb5f60d547357ad8b821683e3c)
Optimize MWT usage in pipeline & fix MWT bulk_process (https://github.com/stanfordnlp/stanza/pull/642 https://github.com/stanfordnlp/stanza/pull/643 https://github.com/stanfordnlp/stanza/pull/644)
Add a UD Enhancer tool which interfaces with CoreNLP's generic enhancer (https://github.com/stanfordnlp/stanza/pull/675)
Add an interface to CoreNLP tokensregex using stanza tokenization (https://github.com/stanfordnlp/stanza/pull/659)
Nothing published for this version
This release features support for extending the capability of the Stanza pipeline with customized processors, a new sentiment analysis tool, improveme
This release features support for extending the capability of the Stanza pipeline with customized processors, a new sentiment analysis tool, improvements to the CoreNLPClient functionality, new models for a few languages (including Thai, which is supported for the first time in Stanza), new biomedical and clinical English packages, alternative servers for downloading resource files, and various improvements and bugfixes.
New Sentiment Analysis Models for English, German, Chinese: The default Stanza pipelines for English, German and Chinese now include sentiment analysis models. The released models are based on a convolutional neural network architecture, and predict three-way sentiment labels (negative/neutral/positive). For more information and details on the datasets used to train these models and their performance, please visit the Stanza website.
New Biomedical and Clinical English Model Packages: Stanza now features syntactic analysis and named entity recognition functionality for English biomedical literature text and clinical notes. These newly introduced packages include: 2 individual biomedical syntactic analysis pipelines, 8 biomedical NER models, 1 clinical syntactic pipelines and 2 clinical NER models. For detailed information on how to download and use these pipelines, please visit Stanza's biomedical models page.
Support for Adding User Customized Processors via Python Decorators: Stanza now supports adding customized processors or processor variants (i.e., an alternative of existing processors) into existing pipelines. The name and implementation of the added customized processors or processor variants can be specified via @register_processor or @register_processor_variant decorators. See Stanza website for more information and examples (see custom Processors and Processor variants). (PR https://github.com/stanfordnlp/stanza/pull/322)
Support for Editable Properties For Data Objects: We have made it easier to extend the functionality of the Stanza neural pipeline by adding new annotations to Stanza's data objects (e.g., Document, Sentence, Token, etc). Aside from the annotation they already support, additional annotation can be easily attached through data_object.add_property(). See our documentation for more information and examples. (PR https://github.com/stanfordnlp/stanza/pull/323)
Support for Automated CoreNLP Installation and CoreNLP Model Download: CoreNLP can now be easily downloaded in Stanza with stanza.install_corenlp(dir='path/to/corenlp/installation'); CoreNLP models can now be downloaded with stanza.download_corenlp_models(model='english', version='4.1.0', dir='path/to/corenlp/installation'). For more details please see the Stanza website. (PR https://github.com/stanfordnlp/stanza/pull/363)
Japanese Pipeline Supports SudachiPy as External Tokenizer: You can now use the SudachiPy library as tokenizer in a Stanza Japanese pipeline. Turn on this when building a pipeline with nlp = stanza.Pipeline('ja', processors={'tokenize': 'sudachipy'}. Note that this will require a separate installation of the SudachiPy library via pip. (PR https://github.com/stanfordnlp/stanza/pull/365)
New Alternative Server for Stable Download of Resource Files: Users in certain areas of the world that do not have stable access to GitHub servers can now download models from alternative Stanford server by specifying a new resources_url argument. For example, stanza.download(lang='en', resources_url='stanford') will now download the resource file and English pipeline from Stanford servers. (Issue https://github.com/stanfordnlp/stanza/issues/331, PR https://github.com/stanfordnlp/stanza/pull/356)
CoreNLPClient Supports New Multiprocessing-friendly Mechanism to Start the CoreNLP Server: The CoreNLPClient now supports a new Enum values with better semantics for its start_server argument for finer-grained control over how the server is launched, including a new option called StartServer.TRY_START that launches the CoreNLP Server if one isn't running already, but doesn't fail if one has already been launched. This option makes it easier for CoreNLPClient to be used in a multiprocessing environment. Boolean values are still supported for backward compatibility, but we recommend StartServer.FORCE_START and StartSerer.DONT_START for better readability. (PR https://github.com/stanfordnlp/stanza/pull/302)
New Semgrex Interface in CoreNLP Client for Dependency Parses of Arbitrary Languages: Stanford CoreNLP has a module which allows searches over dependency graphs using a regex-like language. Previously, this was only usable for languages which CoreNLP already supported dependency trees. This release expands it to dependency graphs for any language. (Issue https://github.com/stanfordnlp/stanza/issues/399, PR https://github.com/stanfordnlp/stanza/pull/392)
New Tokenizer for Thai Language: The available UD data for Thai is quite small. The authors of pythainlp helped provide us two tokenization datasets, Orchid and Inter-BEST. Future work will include POS, NER, and Sentiment. (Issue https://github.com/stanfordnlp/stanza/issues/148)
Support for Serialization of Document Objects: Now you can serialize and deserialize the entire document by running serialized_string = doc.to_serialized() and doc = Document.from_serialized(serialized_string). The serialized string can be decoded into Python objects by running objs = pickle.loads(serialized_string). (Issue https://github.com/stanfordnlp/stanza/issues/361, PR https://github.com/stanfordnlp/stanza/pull/366)
Improved Tokenization Speed: Previously, the tokenizer was the slowest member of the neural pipeline, several times slower than any of the other processors. This release brings it in line with the others. The speedup is from improving the text processing before the data is passed to the GPU. (Relevant commits: https://github.com/stanfordnlp/stanza/commit/546ed13563c3530b414d64b5a815c0919ab0513a, https://github.com/stanfordnlp/stanza/commit/8e2076c6a0bc8890a54d9ed6931817b1536ae33c, https://github.com/stanfordnlp/stanza/commit/7f5be823a587c6d1bec63d47cd22818c838901e7, etc.)
User provided Ukrainian NER model: We now have a model built from the lang-uk NER dataset, provided by a user for redistribution.
Token.id is Tuple and Word.id is Integer: The id attribute for a token will now return a tuple of integers to represent the indices of the token (or a singleton tuple in the case of a single-word token), and the id for a word will now return an integer to represent the word index. Previously both attributes are encoded as strings and requires manual conversion for downstream processing. This change brings more convenient handling of these attributes. (Issue: https://github.com/stanfordnlp/stanza/issues/211, PR: https://github.com/stanfordnlp/stanza/pull/357)
Changed Default Pipeline Packages for Several Languages for Improved Robustness: Languages that have changed default packages include: Polish (default is now PDB model, from previous LFG, https://github.com/stanfordnlp/stanza/issues/220), Korean (default is now GSD, from previous Kaist, https://github.com/stanfordnlp/stanza/issues/276), Lithuanian (default is now ALKSNIS, from previous HSE, https://github.com/stanfordnlp/stanza/issues/415).
CoreNLP 4.1.0 is required: CoreNLPClient requires CoreNLP 4.1.0 or a later version. The client expects recent modifications that were made to the CoreNLP server.
Properties Cache removed from CoreNLP client: The properties_cache has been removed from CoreNLPClient and the CoreNLPClient's annotate() method no longer has a properties_key argument. Python dictionaries with custom request properties should be directly supplied to annotate() via the properties argument.
Fixed Logging Behavior: This is mainly for fixing the issue that Stanza will override the global logging setting in Python and influence downstream logging behaviors. (Issue https://github.com/stanfordnlp/stanza/issues/278, PR https://github.com/stanfordnlp/stanza/pull/290)
Compatibility Fix for PyTorch v1.6.0: We've updated several processors to adapt to new API changes in PyTorch v1.6.0. (Issues https://github.com/stanfordnlp/stanza/issues/412 https://github.com/stanfordnlp/stanza/issues/417, PR https://github.com/stanfordnlp/stanza/pull/406)
Improved Batching for Long Sentences in Dependency Parser: This is mainly for fixing an issue where long sentences will cause an out of GPU memory issue in the dependency parser. (Issue https://github.com/stanfordnlp/stanza/issues/387)
Improved neural tokenizer robustness to whitespaces: the neural tokenizer is now more robust to the presence of multiple consecutive whitespace characters (PR https://github.com/stanfordnlp/stanza/pull/380)
Resolved properties issue when switching languages with requests to CoreNLP server: An issue with default properties has been resolved. Users can now switch between CoreNLP supported languages with and get expected properties for each language by default.
This is a maintenance release of Stanza. It features new support for jieba as Chinese tokenizer, faster lemmatizer implementation, improved compatibil
This is a maintenance release of Stanza. It features new support for jieba as Chinese tokenizer, faster lemmatizer implementation, improved compatibility with CoreNLP v4.0.0, and many more!
Supporting jieba library as Chinese tokenizer. The Stanza (simplified and traditional) Chinese pipelines now support using the jieba Chinese word segmentation library as tokenizer. Turn on this feature in a pipeline with: nlp = stanza.Pipeline('zh', processors={'tokenize': 'jieba'}, or by specifying argument tokenize_with_jieba=True.
Setting resource directory with environment variable. You can now override the default model location $HOME/stanza_resources by setting an environmental variable STANZA_RESOURCES_DIR (https://github.com/stanfordnlp/stanza/issues/227). The new directory will then be used to store and look up model files. Thanks to @dhpollack for implementing this feature.
Faster lemmatizer implementation. The lemmatizer implementation has been improved to be about 3x faster on CPU and 5x faster on GPU (https://github.com/stanfordnlp/stanza/issues/249). Thanks to @mahdiman for identifying the original issue.
Improved compatibility with CoreNLP 4.0.0. The client is now fully compatible with the latest v4.0.0 release of the CoreNLP package.
Correct character offsets in NER outputs from pre-tokenized text. We fixed an issue where the NER outputs from pre-tokenized text may be off-by-one (https://github.com/stanfordnlp/stanza/issues/229). Thanks to @RyanElliott10 for reporting the issue.
Correct Vietnamese tokenization on sentences beginning with punctuation. We fixed an issue where the Vietnamese tokenizer may throw an AssertionError on sentences that begin with a punctuation (https://github.com/stanfordnlp/stanza/issues/217). Thanks to @aryamccarthy for reporting this issue.
Correct pytorch version requirement. Stanza is now asking for pytorch>=1.3.0 to avoid a runtime error raised by pytorch ((https://github.com/stanfordnlp/stanza/issues/231)). Thanks to @Vodkazy for reporting this.
Default Korean Kaist tokenizer failing on punctuation. The default Korean Kaist model is reported to have issues with separating punctuations during tokenization (https://github.com/stanfordnlp/stanza/issues/276). Switching to the Korean GSD model may solve this issue.
Default Polish LFG POS tagger incorrectly labeling last word in sentence as PUNCT. The default Polish model trained on the LFG treebank may incorrectly tag the last word in a sentence as PUNCT (https://github.com/stanfordnlp/stanza/issues/220). This issue may be solved by switching to the Polish PDB model.
Note that if your code was developed on a previous version of the package, there are potentially many breaking changes in this release. The most notab…
This is the first major release of Stanza (previously known as StanfordNLP), a software package to process many human languages. The main features of this release are
conda install -c stanfordnlp stanza.CoreNLPClient to access the Java CoreNLP software from Python code. It is also forward compatible with the next major release of CoreNLP.This release also contains many enhancements and bugfixes:
sentence.text. (#80)logging for all procedual logging, which can be controlled globally either through logging_level or a verbose shortcut. See this page for more information. (#81)tokenize_no_ssplit to True at pipeline instantiation. (#108)depparse_pretagged to True at pipeline instantiation. (#141) Thanks @mrapacz for the contribution!CoreNLPClient. (#154)Note that if your code was developed on a previous version of the package, there are potentially many breaking changes in this release. The most notable changes are in the Document objects, which contain all the annotations for the raw text or document fed into the Stanza pipeline. The underlying implementation of Document and all related data objects have broken away from using the CoNLL-U format as its internal representation for more flexibility and efficiency accessing their attributes, although it is still compatible with CoNLL-U to maintain ease of conversion between the two. Moreover, many properties have been renamed for clarity and sometimes aliased for ease of access. Please see our documentation page about these data objects for more information.
Nothing published for this version
Nothing published for this version
Your coding agent can read these notes before it upgrades. Set up the MCP server →