| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| docubrowser-foss-1.4.0-1.tar.gz | 2026-09-10 | 663.6 kB | |
| DocuBrowse v1.4.0 -- Multi-Language Support (Japanese First) source code.tar.gz | 2026-09-10 | 2.3 MB | |
| DocuBrowse v1.4.0 -- Multi-Language Support (Japanese First) source code.zip | 2026-09-10 | 2.4 MB | |
| README.md | 2026-09-10 | 4.1 kB | |
| Totals: 4 Items | 5.4 MB | 0 | |
DocuBrowse v1.4.0 — Multi-Language Support (Japanese First)
DocuBrowse can now run entirely in a second language. This release adds a per-install internationalization framework and ships Japanese as the first fully-supported language alongside English — localized UI, language-appropriate AI models, and CJK-aware keyword search — plus two fixes surfaced during Japanese testing.
🌐 Internationalization (i18n)
- Per-install language. One install serves one language — the interface,
embedding model, synopsis model, and FTS tokenizer are all selected together.
English (
en, default) and Japanese (ja) are supported today. - Data-driven, not code-driven. Adding a language is a
LANG_MODELSrow inlang_models.pyplus alocales/<code>.jsonfile — no new code paths. - Localized UI. All interface strings live in per-language locale files
(
locales/en.json,locales/ja.json) and are resolved client-side; the active locale ships withGET /api/config. - Language-appropriate AI. Japanese uses the
bge-m3multilingual embedding model and a Japanese synopsis model; models are pulled from Ollama on demand at first run and on language switch — nothing extra is bundled. - First-run language prompt. A fresh install asks English or Japanese
(answer non-interactively with
DOCUBROWSE_LANG); upgrades never re-prompt and always preserve an existinglang. - Settings language switcher. Switch language from the Settings gear
(
POST /api/language); it provisions the new language's models and warns that a rescan is needed to rebuild embeddings/index. - Per-language search shaping. Stopword stripping is language-aware (English articles/conjunctions; Japanese no-op), and the A–Z/0–9 index bar is hidden for CJK languages.
🈶 CJK keyword search
- Reliable 2+ character CJK matching. Japanese keyword search uses app-side
character-bigram segmentation over the
unicode61tokenizer (nottrigram), so common 2-character kanji compounds (熟語 such as 栽培 / 品種) match instead of returning nothing. Single-character queries surface via Both/semantic mode. - Zero dependency, uniform for CJK. No MeCab/jieba/konlpy — the same path applies to Chinese and Korean when they land.
🐛 Bug fixes
- Japanese keyword search returned nothing. The configured language was not
threaded into the database connection, so a Japanese install built its
full-text index with the English tokenizer.
get_db()now resolves the install language, fixing keyword search end-to-end. - Japanese AI summaries came back blank. Hybrid reasoning synopsis models (e.g. the Japanese nemotron model) spent their whole budget "thinking" and returned nothing; reasoning is now disabled for synopsis generation, and the synopsis timeout is raised for large cold models.
📚 Documentation & quality
- Japanese README (
README-ja.md) — a full translation, kept structurally parallel toREADME.md, with a language switcher on both. - New Languages section in the README documenting per-install language, model selection, the CJK bigram approach, and what is not yet supported (mixed-language corpora, kana/pinyin index, Japanese "My Number" PII).
- pylint 10/10 maintained on the new and touched source files; the i18n code was run through an over-engineering pass.
⚠️ Notes & not-yet-supported
- One install = one language; mixed-language or per-document corpora are not supported. Switching an instance's language does not translate existing documents and requires a rescan to rebuild the index for the new language.
- PII detection remains US-pattern only (Japanese "My Number" not yet implemented).
- Languages beyond English/Japanese, and a kana/reading (or pinyin) index bar for CJK, are future work.
Tarball only release
- Once Korean and Chinese are in place will do a full set of packages.