Download Latest Version base-memory.engukr.zip (24.8 MB)
Email in envelope

Get an email when there's a new version of OPolyglot

Home / data / tessdata
Name Modified Size InfoDownloads / Week
Parent folder
LICENSE 2025-11-12 11.4 kB
README.md 2025-11-12 1.4 kB
fra.traineddata.zip 2025-11-11 6.2 MB
eng.traineddata.zip 2025-11-11 10.9 MB
ita.traineddata.zip 2025-11-10 6.9 MB
por.traineddata.zip 2025-11-09 6.7 MB
spa.traineddata.zip 2025-11-08 8.3 MB
deu.traineddata.zip 2025-11-08 7.1 MB
fin.traineddata.zip 2025-11-07 9.0 MB
rus.traineddata.zip 2025-11-01 8.6 MB
ukr.traineddata.zip 2025-11-01 5.2 MB
pol.traineddata.zip 2025-11-01 8.3 MB
Totals: 12 Items   77.3 MB 418

tessdata

These language data files only work with Tesseract 4.0.0 and newer versions. They are based on the sources in tesseract-ocr/langdata on GitHub. (still to be updated for 4.0.0 - 20180322)

These have models for legacy tesseract engine (--oem 0) as well as the new LSTM neural net based engine (--oem 1).

The LSTM models (--oem 1) in these files have been updated to the integerized versions of tessdata_best on GitHub. So, they should be faster but probably a little less accurate than tessdata_best.

tessdata_fast on GitHub provides an alternate set of integerized LSTM models which have been built with a smaller network. tessdata_fast files are the ones packaged for Debian and Ubuntu.

The legacy tesseract models (--oem 0) have been removed for Indic and Arabic script language files.

tessdata for 3.04 or 3.05

Get language data files for Tesseract 3.04 or 3.05 from the 3.04 tree.

More information and a complete list of all languages is available in the Tesseract wiki.

All data in the repository are licensed under the Apache-2.0 License, see file LICENSE.

Source: README.md, updated 2025-11-12