Download Latest Version docubrowser-foss-1.5.2-1-windows.zip (822.9 kB) Google Add to Preferred Sources
Home / v1.4.0
Name Modified Size InfoDownloads / Week
Parent folder
docubrowser-foss-1.4.0-1.tar.gz 2026-09-10 663.6 kB
DocuBrowse v1.4.0 -- Multi-Language Support (Japanese First) source code.tar.gz 2026-09-10 2.3 MB
DocuBrowse v1.4.0 -- Multi-Language Support (Japanese First) source code.zip 2026-09-10 2.4 MB
README.md 2026-09-10 4.1 kB
Totals: 4 Items   5.4 MB 0

DocuBrowse v1.4.0 — Multi-Language Support (Japanese First)

DocuBrowse can now run entirely in a second language. This release adds a per-install internationalization framework and ships Japanese as the first fully-supported language alongside English — localized UI, language-appropriate AI models, and CJK-aware keyword search — plus two fixes surfaced during Japanese testing.

🌐 Internationalization (i18n)

  • Per-install language. One install serves one language — the interface, embedding model, synopsis model, and FTS tokenizer are all selected together. English (en, default) and Japanese (ja) are supported today.
  • Data-driven, not code-driven. Adding a language is a LANG_MODELS row in lang_models.py plus a locales/<code>.json file — no new code paths.
  • Localized UI. All interface strings live in per-language locale files (locales/en.json, locales/ja.json) and are resolved client-side; the active locale ships with GET /api/config.
  • Language-appropriate AI. Japanese uses the bge-m3 multilingual embedding model and a Japanese synopsis model; models are pulled from Ollama on demand at first run and on language switch — nothing extra is bundled.
  • First-run language prompt. A fresh install asks English or Japanese (answer non-interactively with DOCUBROWSE_LANG); upgrades never re-prompt and always preserve an existing lang.
  • Settings language switcher. Switch language from the Settings gear (POST /api/language); it provisions the new language's models and warns that a rescan is needed to rebuild embeddings/index.
  • Per-language search shaping. Stopword stripping is language-aware (English articles/conjunctions; Japanese no-op), and the A–Z/0–9 index bar is hidden for CJK languages.
  • Reliable 2+ character CJK matching. Japanese keyword search uses app-side character-bigram segmentation over the unicode61 tokenizer (not trigram), so common 2-character kanji compounds (熟語 such as 栽培 / 品種) match instead of returning nothing. Single-character queries surface via Both/semantic mode.
  • Zero dependency, uniform for CJK. No MeCab/jieba/konlpy — the same path applies to Chinese and Korean when they land.

🐛 Bug fixes

  • Japanese keyword search returned nothing. The configured language was not threaded into the database connection, so a Japanese install built its full-text index with the English tokenizer. get_db() now resolves the install language, fixing keyword search end-to-end.
  • Japanese AI summaries came back blank. Hybrid reasoning synopsis models (e.g. the Japanese nemotron model) spent their whole budget "thinking" and returned nothing; reasoning is now disabled for synopsis generation, and the synopsis timeout is raised for large cold models.

📚 Documentation & quality

  • Japanese README (README-ja.md) — a full translation, kept structurally parallel to README.md, with a language switcher on both.
  • New Languages section in the README documenting per-install language, model selection, the CJK bigram approach, and what is not yet supported (mixed-language corpora, kana/pinyin index, Japanese "My Number" PII).
  • pylint 10/10 maintained on the new and touched source files; the i18n code was run through an over-engineering pass.

⚠️ Notes & not-yet-supported

  • One install = one language; mixed-language or per-document corpora are not supported. Switching an instance's language does not translate existing documents and requires a rescan to rebuild the index for the new language.
  • PII detection remains US-pattern only (Japanese "My Number" not yet implemented).
  • Languages beyond English/Japanese, and a kana/reading (or pinyin) index bar for CJK, are future work.

Tarball only release

  • Once Korean and Chinese are in place will do a full set of packages.
Source: README.md, updated 2026-09-10