| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| grobid-trainer-0.9.1.jar | 2026-08-04 | 16.4 MB | |
| grobid-service-0.9.1.jar | 2026-08-04 | 3.3 MB | |
| grobid-core-0.9.1.jar | 2026-08-04 | 17.2 MB | |
| 0.9.1 source code.tar.gz | 2026-08-04 | 393.2 MB | |
| 0.9.1 source code.zip | 2026-08-04 | 401.7 MB | |
| README.md | 2026-08-04 | 7.7 kB | |
| Totals: 6 Items | 831.7 MB | 4 | |
What's Changed
Added
- Training web API to list ongoing trainings and interrupt a running one, releasing the per-model lock [#1483]
- Development/model debug API exposing the first-level models for inspecting intermediate results [#1439]
- Push-based export of JVM/process runtime metrics over OTLP, complementing the Prometheus scrape endpoint [#1479]
/metrics/prometheusendpoint now serves the Prometheus exposition format with JVM/process metrics, plus Prometheus/Grafana setup docs [#1473]- Linking-aware
affiliation_linkedmetric in the end-to-end HEADER evaluation [#1493] [#1467] - Rootless Docker images (Kubernetes/OpenShift friendly) [#1442]
- Fail fast at startup when the grobid-home path contains spaces and Wapiti is the configured engine [#1481]
- Apache 2.0 licence headers in source files [#1485]
- Documentation: community page [#1447], PDF-TEI Editor reference [#1448], and instructions for adding new model flavors [#1465]
- CI: CodeQL analysis workflow [#1475], dependabot updates for GitHub Actions [#1494], and runs on forks without publishing credentials [#1451]
Changed
- Updated pdfalto to 0.6.2. Notable for GROBID: deterministic output (an uninitialised read and a use-after-free made line/block grouping depend on heap contents, so the same PDF could yield different
<TextLine>/<TextBlock>structure across runs), bounded peak memory on very large or vector-heavy documents, superscript citation/affiliation callouts separated from adjacent words (token boundary now driven by pdfalto's own superscript detection), and several text-extraction fixes. The command-line options GROBID passes are unchanged. - pdfalto is now given GROBID's configured temp directory (
grobid.temp, by defaultgrobid-home/tmp) throughTMPDIR. Since 0.6.1 pdfalto streams large page DOMs to a scratch file to bound peak memory and picks its location fromTMPDIR, falling back to/tmp— which is tmpfs in most containers, so spilling there leaves peak memory unchanged and risks ENOSPC. - Reworked author–affiliation linking into a dedicated, unit-tested
AuthorAffiliationAssignerwith a priority-based strategy [#1467] - Preserve extracted affiliations when header consolidation rewrites authors, via staged reconciliation [#1488]
- Parallelized the end-to-end evaluation scoring phase (~20 min → ~2 min), with upfront setup verification in the Python script [#1487]
- Rewrote sentence-segmentation re-alignment as a drift-free forward two-pointer alignment and made language detection thread-safe [#1457]
- Language-sensitive tokenization in
PDFALTOSaxHandlerinstead of always using the default analyzer [#1480] - Only one training per model at a time; a concurrent request for the same model returns 409 Conflict [#1477]
- Unified training-data file generation between the web API and batch mode [#1508]
- Lexicon: added
Lexicon.builder()for optional eager gazetteer pre-loading (lazy stays the default);getInstance()is now deprecated [#1440] - Lexicon: added 4 missing ISO 3166-1 country codes (BQ, CW, SS, SX) and migrated AN to its ISO 3166-3 form ANHH
- Documentation: expanded the end-to-end evaluation and configuration guides and recorded the
grobid-evaluationdataset DOI [#1419] [#1430] [#1501] - Dependency updates: OpenNLP 1.9.4 → 2.5, Jetty, jackson 2.21.4, DeLFT 0.4.6; dropped unused dependencies [#1449] [#1469] [#1423]
- Replaced Powermock with Mockito [#1458]
- Adopted Spotless for code formatting [#1384], rewrote overloaded methods [#1401], and added codespell spell-checking [#1365] [#1450] [#1478]
Fixed
- Handle the exit codes pdfalto 0.6.1 introduced. Code 5 reports that the ALTO was written correctly but page streaming was disabled mid-run, so peak memory was no longer bounded; both execution paths rejected any non-zero exit and would have discarded these successful conversions. It is now a logged warning. Codes 4 (ALTO write failed) and 98 (allocation failure) are reported as
PDFALTO_CONVERSION_FAILUREinstead of falling through toBAD_INPUT_DATA, which blamed the PDF for a pdfalto-side failure. - JVM shutdown deadlock when closing JEP Python interpreters: close now runs on the owning worker thread and is idempotent [#1506]
- HTTP 500 from
processFulltextDocumentwhen two footnotes share the same superscript marker; such callouts now fall back to plain text [#1472] - TEI paragraph boundaries no longer collapse when a paragraph starts right after a trailing reference marker [#1482]
- Invalid/unbalanced XML in generated training data across models [#1470], and unclosed
<bibl>before<other>in reference-segmenter training data [#1466] xml:idvalues in generated training files are now valid NCNames [#1508]- NPE in reference-segmenter training-data generation after segmentation retraining [#1490]
- Document language is detected for segmentation training data instead of hardcoding
xml:lang="en"[#1460] - Textual year/month/day fields stay in sync with the normalized publication date for library callers and non-TEI output paths [#1463]
TextUtilities.clean()no longer folds the letters æ/Æ and œ/Œ to ASCII "ae"/"oe"; typographic ligature expansion (fi/fl/ff) is unchanged [#1461]- Guard against out-of-range page index in citation annotation, which surfaced as HTTP 500 on malformed PDFs [#1459]
- Model selection with mixed DeLFT/Wapiti engines and flavor selections, with clearer logging when a flavor falls back to the base model [#1455]
- Citations consisting only of non-breaking spaces now return 204 No Content instead of HTTP 500 [#1407]
- biblio-glutton health probe in the evaluation configuration check now targets an existing endpoint (
/service/data) instead of always reporting a healthy glutton as unreachable [#1492] - Block/segmentation desync warning now includes the page number and a text excerpt so occurrences can be located and reproduced [#1471]
./gradlew install[#1427], git revision information [#1433], and Docker image summary [#1429]
Security
- Prevent command injection through crafted PDF file names in the non-server
pdfaltopath: the command is no longer interpolated into abash -cstring but passed as positional parameters and exec'd via"$@"(GHSA-mgxf-7mg7-qpmf) [#1477] - Stop leaking a JVM thread per request on the
/api/modelTrainingendpoint by shutting down the per-request executor (GHSA-g2r5-4c8r-c84f) [#1477] - Remove the vulnerable JLine telnet server module from the classpath by depending on
jline-terminalonly instead of theorg.jline:jlineuber-jar pulled in transitively byprogressbar(GHSA-47qp-hqvx-6r3f, GHSA-2r2c-cx56-8933) [#1469] - Upgrade jackson (core, databind, afterburner, dataformat-yaml) to 2.21.4 to address CVE-2026-54513 (array-element type allowlist bypass in polymorphic type validation) [#1469]
- Upgrade Apache OpenNLP to 2.5 (arbitrary class loading via model manifest) and Jetty (HTTP request smuggling via chunked extension quoted-string parsing) [#1449]
- Harden
ZipUtilsagainst zip-slip: each entry's canonical output path is validated against the target directory before any write (flagged by CodeQL) [#1486]
New Contributors
- @flrjrf made their first contribution in https://github.com/grobidOrg/grobid/pull/1427
- @luismmontilla made their first contribution in https://github.com/grobidOrg/grobid/pull/1430
- @yarikoptic made their first contribution in https://github.com/grobidOrg/grobid/pull/1365
- @lfoppiano with @Copilot made their first contribution in https://github.com/grobidOrg/grobid/pull/1480
Full Changelog: https://github.com/grobidOrg/grobid/compare/0.9.0...0.9.1