Originally created by: sachin-detrax
Closes [#27].
The local model list was a fixed array of nine models compiled into models-catalog.ts. Anything else on Hugging Face was unreachable without running llama-server by hand or waiting for a maintainer to ship a new release. Models are now open-ended.
A user-added entry is a full LocalModelDef with a custom- prefixed id, persisted in localModels.customModels (config v33 → v34; older files inherit []). loadConfig registers them in a catalog registry, so getLocalModelDef / isKnownLocalModelId resolve curated and custom entries through one lookup — daemon start, installer, TUI rows and CLI all picked them up without changes.
+ Add a model from Hugging Face..., pinned under Local text models. Enter opens a prompt taking either a reference (resolve and add) or free text (search, then pick by digit).models add <ref> and models search <query>./models add and /models search./resolve/ and /blob/ file URLshf://owner/repo[@rev]/file.ggufhf download ... command (one- or two-argument, trailing flags dropped)owner/nameNaming only a repo picks a 4-bit quant and any mmproj projector, so vision repos come out vision-capable. HF_TOKEN is honoured for gated repos, including on the weights download.
hf:// is parsed by hand rather than with new URL: that puts the owner in the host slot and lowercases it, and HF owners are case-sensitive, so hf://Qwen/... would silently 404.
Both surfaced while testing this, and both could shred the panel on any daemon failure:
toStatusLine() truncates at every error-message write and at the render point, with wrap="truncate-end" on the Ink Text.LlmPanel frame budget did not account for overlay modal / banner rows, so opening a modal during an error overran the frame. estimateOverlayRows() now covers them.Regression tests pin the frame height for both.
extractLoadFailure pulls the diagnostic reason out of llama-server logs so a failed load reports why rather than a generic message.
35 files changed, +2104/−52, including new suites for the HF reference parser, panel rendering, frame height, and the reducer.
npx tsc -p tsconfig.json --noEmit # clean
npx vitest run src/local-llm/ src/tui/ src/cli/
# before: 6 failed | 797 passed (803)
# after: 6 failed | 835 passed (841)
The 6 failures are identical by name on main and on this branch (ChatLog/SplashBanner/TuiApp smoke, llm-panel selectors, persistEmbeddingHybridRecall) — pre-existing, untouched by this change. 38 tests added, none broken.
🤖 Generated with Claude Code
Originally posted by: sosidudku1
Strong PR, thank you. The input parser is the best part of it: repo URLs,
hf://refs, direct file links and even a pastedhf downloadcommand all resolve correctly, and falling back to a download-sorted live search is the right default. A few things are worth aligning before merge.Two blocking items, both small:
-00001-of-NNNNN.ggufshard, only the first part is downloaded and the model will not start, and the user finds out only at llama-server launch. Please either reject sharded picks with a clear error ("sharded models are not supported yet") or put an explicit warning into the "added" message.MTP/folder ormtp-/-mtppatterns closes this; it is a few lines inpickDefaultGgufFile.Nice to have, your call whether in this PR or as follow-ups:
listHuggingFaceGgufFilesalready returns every file with sizes, and the numbered-pick modal from search results could be reused as a second step, with Enter keeping the recommended quant. Seeing "Q4_K_M 4.2 GB vs Q8_0 8.1 GB" before committing to a download is a real UX win.Two things we will split into separate issues rather than grow this PR: memory-requirement estimates read from the GGUF header instead of the file-size heuristic (you left a note about this yourself), and hardware-aware quant selection instead of the constant Q4 preference. This PR does not need to carry them.
With 1 and 2 in, this is good to merge from my side.
Originally posted by: sosidudku1
We merged [#41], [#48] and [#50] today, which is why this branch now shows conflicts, they touch the same files. Could you rebase onto main? The two blocking items from the review (sharded repos, MTP filter) still apply after the rebase. Everything else is ready on our side.
Related
Tickets:
#41Tickets:
#48Tickets:
#50Originally posted by: sosidudku1
Rebased onto main per the author's request (config bump lands as v36, after stopOnExit and progressIndicator took 34 and 35) and picked up the two blocking review items: sharded picks are now rejected outright with a clear error (first part included, plus a dedicated message for sharded-only repos), and MTP/NextN companion files are excluded from the default-quant fallback. All 62 tests from the original branch pass unchanged, 4 new ones cover the guards. Ready for review.