| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| README.md | 2026-10-01 | 5.9 kB | |
| Stanza v1.15.0 Release Notes - Updated for UD 2.18 source code.tar.gz | 2026-10-01 | 1.9 MB | |
| Stanza v1.15.0 Release Notes - Updated for UD 2.18 source code.zip | 2026-10-01 | 2.2 MB | |
| Totals: 3 Items | 4.1 MB | 0 | |
Stanza v1.15.0 Release Notes
Updated Models
All tokenizer, MWT, POS, lemmatizer, and dependency parser models have been rebuilt using UD 2.18 datasets. The combined English, Spanish, and French packages have also been refreshed from the most recent development-branch snapshots to reflect recent improvements.
- Add an NER model for Shahmukhi Punjabi (the Persian-script variety of Punjabi). #1666
Security Fixes
- Replace pickle-based lemmatizer model serialization with
orjson. See GHSA-gh9r-c94j-cmp5. #1649
Bugfixes
-
Fix multi-word token IDs becoming lists instead of tuples after a JSON round-trip through
to_serialized()/from_serialized(), which produced invalid CoNLL-U output (e.g.[3, 4]instead of3-4) and broke downstream dictionary operations. Introduced in v1.14.0; addresses #1662. Thanks @Arthur031221! #1663 -
Fix empty words (enhanced UD) being misidentified as multi-word tokens when reconstructing a
Documentfrom dicts, causing incorrect token structure after a round-trip throughto_dict(). Thanks @Arthur031221! #1664 -
Fix start/end character offsets not being assigned to words and tokens when reading a CoNLL-U
Document. The fix populates offsets either by aligning tokens against the sentence text or by readingSpaceAfterannotations already present. #1656 -
Fix a
KeyErrorwhen loading the English MIMIC no-CharLM lemmatizer package:charlm_forward_fileandcharlm_backward_fileare now treated as optional keys rather than required ones. Addresses #1651. #1668
Tokenizer
-
Embed external segmentation dictionaries (used by Thai, Japanese, and Chinese tokenizers) directly in the tokenizer model files, so they are not lost if models are rebuilt without them. The dictionaries are stored compressed. Adds unit tests to verify that models which should include a dictionary do. #1688
-
Allow en-dashes and em-dashes to function as standalone tokens without suppressing comma→dash augmentation, improving the tokenizer's ability to learn dash-connected word patterns from diverse training data. #1687
Dependency Parser
-
Add a warning when the constraint-repair loop reaches its maximum iteration count without fully resolving all violations, making incomplete repairs visible during debugging. #1643
-
Refactor the dependency parser
DataLoaderto separate PyTorch and non-PyTorch code paths, enabling dynamic augmentation of individual training examples on-the-fly rather than at initialization time. This gives more balanced augmentation across training and extends coverage to silver-annotated datasets. #1650
Training Infrastructure
-
Generalize
mixed_odia_dataset.pyintomixed_indic_dataset.py, which can now build combined training datasets for any low-resource Indic target language. Originally developed for Sindhi, this release also uses it for Bhojpuri. The interface replaces five language-specific flags with a single--donorsparameter; aDONOR_CONFIGSdictionary at the top of the script makes adding new donor languages a one-line change. #1655 -
Add a script for building tokenizer training sets from a mixture of multiple languages, useful for training tokenizers on related-language groups such as the Indic family. #1672
-
Move initial-punctuation stripping from data preparation into the
DataLoaderfor tokenizer, POS tagger, and dependency parser, so it is applied dynamically during training rather than baked into data files. #1653 -
Integrate speaker information from ingested UDCoref documents into the coreference training data pipeline. #1645
-
Add a conversion script for the IIT (BHU) Bhojpuri POS corpus, transforming it from its mixed flat/SSF-bracket format into a standardized one-token-per-line layout. Thanks @abhiprd2000! #1675
-
Add multi-column xpos tagging support, allowing the tagger to train across multiple datasets with differing xpos schemes simultaneously. Applied to Bhojpuri (BHTB + IIT corpus), where it yields substantial improvements. Experiments on English (ParTUT and LinES) showed no benefit — English xpos accuracy is already saturated around 97.4%, which makes it a poor test case for a technique aimed at low-resource settings with small main treebanks. #1680
CoreNLP Integration
-
Upgrade the CoreNLP semgrex communication protocol to support enhanced queries. #1685
-
Update the CoreNLP installation script to report what was installed, be more conservative about which version to download, and fix a bug in the
DEFAULT_CORENLP_URLconstant. #1686