Download Latest Version Stanza v1.15.0 Release Notes - Updated for UD 2.18 source code.zip (2.2 MB) Google Add to Preferred Sources
Home / v1.15.0
Name Modified Size InfoDownloads / Week
Parent folder
README.md 2026-10-01 5.9 kB
Stanza v1.15.0 Release Notes - Updated for UD 2.18 source code.tar.gz 2026-10-01 1.9 MB
Stanza v1.15.0 Release Notes - Updated for UD 2.18 source code.zip 2026-10-01 2.2 MB
Totals: 3 Items   4.1 MB 0

Stanza v1.15.0 Release Notes

Updated Models

All tokenizer, MWT, POS, lemmatizer, and dependency parser models have been rebuilt using UD 2.18 datasets. The combined English, Spanish, and French packages have also been refreshed from the most recent development-branch snapshots to reflect recent improvements.

  • Add an NER model for Shahmukhi Punjabi (the Persian-script variety of Punjabi). #1666

Security Fixes

Bugfixes

  • Fix multi-word token IDs becoming lists instead of tuples after a JSON round-trip through to_serialized() / from_serialized(), which produced invalid CoNLL-U output (e.g. [3, 4] instead of 3-4) and broke downstream dictionary operations. Introduced in v1.14.0; addresses #1662. Thanks @Arthur031221! #1663

  • Fix empty words (enhanced UD) being misidentified as multi-word tokens when reconstructing a Document from dicts, causing incorrect token structure after a round-trip through to_dict(). Thanks @Arthur031221! #1664

  • Fix start/end character offsets not being assigned to words and tokens when reading a CoNLL-U Document. The fix populates offsets either by aligning tokens against the sentence text or by reading SpaceAfter annotations already present. #1656

  • Fix a KeyError when loading the English MIMIC no-CharLM lemmatizer package: charlm_forward_file and charlm_backward_file are now treated as optional keys rather than required ones. Addresses #1651. #1668

Tokenizer

  • Embed external segmentation dictionaries (used by Thai, Japanese, and Chinese tokenizers) directly in the tokenizer model files, so they are not lost if models are rebuilt without them. The dictionaries are stored compressed. Adds unit tests to verify that models which should include a dictionary do. #1688

  • Allow en-dashes and em-dashes to function as standalone tokens without suppressing comma→dash augmentation, improving the tokenizer's ability to learn dash-connected word patterns from diverse training data. #1687

Dependency Parser

  • Add a warning when the constraint-repair loop reaches its maximum iteration count without fully resolving all violations, making incomplete repairs visible during debugging. #1643

  • Refactor the dependency parser DataLoader to separate PyTorch and non-PyTorch code paths, enabling dynamic augmentation of individual training examples on-the-fly rather than at initialization time. This gives more balanced augmentation across training and extends coverage to silver-annotated datasets. #1650

Training Infrastructure

  • Generalize mixed_odia_dataset.py into mixed_indic_dataset.py, which can now build combined training datasets for any low-resource Indic target language. Originally developed for Sindhi, this release also uses it for Bhojpuri. The interface replaces five language-specific flags with a single --donors parameter; a DONOR_CONFIGS dictionary at the top of the script makes adding new donor languages a one-line change. #1655

  • Add a script for building tokenizer training sets from a mixture of multiple languages, useful for training tokenizers on related-language groups such as the Indic family. #1672

  • Move initial-punctuation stripping from data preparation into the DataLoader for tokenizer, POS tagger, and dependency parser, so it is applied dynamically during training rather than baked into data files. #1653

  • Integrate speaker information from ingested UDCoref documents into the coreference training data pipeline. #1645

  • Add a conversion script for the IIT (BHU) Bhojpuri POS corpus, transforming it from its mixed flat/SSF-bracket format into a standardized one-token-per-line layout. Thanks @abhiprd2000! #1675

  • Add multi-column xpos tagging support, allowing the tagger to train across multiple datasets with differing xpos schemes simultaneously. Applied to Bhojpuri (BHTB + IIT corpus), where it yields substantial improvements. Experiments on English (ParTUT and LinES) showed no benefit — English xpos accuracy is already saturated around 97.4%, which makes it a poor test case for a technique aimed at low-resource settings with small main treebanks. #1680

CoreNLP Integration

  • Upgrade the CoreNLP semgrex communication protocol to support enhanced queries. #1685

  • Update the CoreNLP installation script to report what was installed, be more conservative about which version to download, and fix a bug in the DEFAULT_CORENLP_URL constant. #1686

Contributors

Source: README.md, updated 2026-10-01