| Name | Modified | Size | Downloads / Week |
|---|---|---|---|
| Parent folder | |||
| kev-0.8b.tar.gz | 2026-10-01 | 45.8 MB | |
| kev-4b.tar.gz | 2026-10-01 | 130.8 MB | |
| kev-9b.tar.gz | 2026-10-01 | 172.1 MB | |
| SHA256SUMS.txt | 2026-10-01 | 242 Bytes | |
| Kev 1.0 source code.tar.gz | 2026-10-01 | 140.5 MB | |
| Kev 1.0 source code.zip | 2026-10-01 | 141.8 MB | |
| README.md | 2026-10-01 | 11.3 kB | |
| Totals: 7 Items | 630.9 MB | 0 | |
Kev 1.0 is the first versioned release of the whole Kev family: four decision models that read a document and a set of typed questions and return calibrated probabilities over the options, in one forward pass, behind TypeSafe's System One API. Nothing in it is newly trained. It pins the checkpoints, model cards, evaluation suites and serving code that the next generation of Kev is measured against, with the same tag (v1.0) on every Hub repo.
What's in 1.0
| Model | Hub repo | Weights revision | Form | Base | Temperature | Validated context |
|---|---|---|---|---|---|---|
| Kev-0.8B | jaredpalmer/kev-0.8b |
9a45d25e |
LoRA adapter + head | Qwen3.5-0.8B-Base (Apache-2.0) | 2.35 | 8,192 tokens |
| Kev-4B | jaredpalmer/kev-4b |
139fdd94 |
LoRA adapter + head | Qwen3.5-4B-Base (Apache-2.0) | 2.41 | 8,192 tokens |
| Kev-9B (v2) | jaredpalmer/kev-9b |
b5d8c18e |
LoRA adapter + head | Qwen3.5-9B-Base (Apache-2.0) | 2.19 | 8,192 tokens |
| Kev-27B (v2) | jaredpalmer/kev-27b |
28be62e9 |
full bf16 weights (51 GB) + head | Qwen3.8-27B, post-trained (Apache-2.0) | 1.32 | 65,536 tokens |
Headline numbers (fp32 evaluation path, each model at its shipped temperature; the transfer-v4 test is locked and was read once per model):
| Kev-0.8B | Kev-4B | Kev-9B | Kev-27B | Jev | |
|---|---|---|---|---|---|
| Held-out datasets: breadth-v1 test, chance-corrected index | 23.3 | 38.0 | 41.0 | 52.3 | 54.0 |
| Out-of-domain: transfer-v4 development accuracy | 0.648 | 0.817 | 0.820 | 0.851 | 0.857 |
| Out-of-domain: transfer-v4 locked test accuracy / Brier | 0.697 / 0.397 | 0.838 / 0.224 | 0.852 / 0.199 | 0.889 / 0.154 | – |
| Skills: hard-v1 test | 0.665 | 0.803 | 0.834 | 0.918 | – |
| Developer tooling: devtools-v1 test, all sources | 0.637 | 0.756 | 0.791 | 0.790 | – |
| Real documents: documents-v1 test | 0.851 | 0.903 | 0.900 | 0.908 | – |
| MMLU-Pro (transfer-v9 development) | 0.230 | 0.565 | 0.590 | 0.675 | 0.840 |
hard-v1, devtools-v1 and documents-v1 have training splits that every Kev trained on: those rows measure held-out items of trained families, not transfer. The breadth-v1 and transfer-v4 rows are datasets no Kev trained on. Jev was read on the development partitions and on breadth-v1 test only. Every number traces to a committed report through docs/claims.json; the model cards (docs/model-cards/) have the rest, with intervals.
What changed since the last family release
Measured from the GitHub release kev-family as first assembled for the current family on 2026-09-24 (Kev-27B v1, Kev-9B v1, and the same Kev-4B and Kev-0.8B as here). Its 2026-09-30 updates (Kev-27B v2, Kev-9B v2) are listed here too, since 1.0 is where they become part of a versioned release.
- Kev-27B v2: full weights. Every weight of Qwen3.8-27B fine-tuned for one epoch on a 145,840-record corpus, then averaged 0.85 / 0.15 with v1. Against v1 on test: held-out datasets +1.2 pp [+0.3, +2.2], held-out task families +5.3 [+3.7, +6.8], skills, tooling and documents +8.9 [+7.5, +10.3]; locked out-of-domain test 0.889 against 0.896, Brier 0.154 against 0.160. Worse and overconfident on long contracts (CUAD ECE 0.053 against 0.007). v1 is at
jaredpalmer/kev-27b@v1-lora. - Kev-9B v2. v1 plus one epoch on the documents and skills data Kev-4B and Kev-0.8B already had. Against v1 on test: hard-v1 + devtools-v1 +18.7 pp [+16.7, +20.8], documents-v1 +7.1 [+4.7, +9.2]; level on the locked out-of-domain test (0.852 both) with Brier 0.199 against 0.224. v1 is at
jaredpalmer/kev-9b@v1. - No silent truncation. The server used to cut a state longer than its limit without saying so. It now refuses a state over 65,536 tokens with a 422 that names the token count and the limit;
KEV_TRUNCATE_STATES=1opts back into truncation, and every response from such a server then saystruncated. The deploy and fine-tune skills pinKEV_REFto a commit that has this fix and the long-document and MLX changes below (71d4829), and the Space was republished after the fix. - Long documents on every size. The evaluation path kept fp32 attention for long rows on the math kernel, so Kev-0.8B, 4B and 9B ran out of GPU memory on 32k–64k-token states. Long rows now run the memory-efficient kernel in fp32: Kev-4B reads a 61k-token state in 17.2 s with 17.2 GiB above the weights on an H100, and shorter rows keep their logits bit for bit. This is what makes the validated context lengths above measurable.
- Apple Silicon. The MLX backend loads full-weight checkpoints as saved, with no merge, which gives Kev-27B a Mac path (expected to need about 51 GB plus working memory; not yet run at that size). Long states are prefilled 1,024 tokens at a time and the cache evicts before a pass, so Kev-4B serves a 65,000-token state on a 32 GB M5 at a 13.0 GB peak (84.5 s new, 716 ms cached).
- Kernel provenance. Every evaluation report and trial now records the kernel set its logits depend on (package versions, GPU, dtype, attention and DeltaNet implementations), after a kernel change in the evaluation image was found to move Kev-27B v1's reads by 0.03–0.06 in probability with no change to Kev's own code.
- Evaluation audit. Three suites were removed as unsound for selecting models: scienthoon (templated tickets, one question the text cannot answer), WANLI-v2 / WANLI-v1 (a quarter of gold labels are one of two disagreeing annotators) and TypeSafe's public evals (gold from two closed models, too few questions). Headline panels exclude items the audit found unanswerable or unlabelled. Past releases' figures on those suites are kept in their records, not on the 1.0 cards.
- Calibration. Kev-4B and Kev-0.8B ship temperatures fitted on held-out items of their training data. A registered refit on held-out datasets was evaluated for both and adopted for neither: it did not improve Kev-4B (Brier difference −0.0001 [−0.0005, +0.0003]), and it made Kev-0.8B worse calibrated on its document and skill families by more than the registered tolerance. Kev-9B and Kev-27B already ship held-out-dataset temperatures.
- Training data published. The documents-v1 and hard-v1 training partitions are in the
jaredpalmer/kev-suitesdataset, so the small models' training data can be fetched and hash-checked. - Validated context length. Each card now states the longest state at which accuracy on CUAD contracts stays within 3 pp (95 % lower bound) of the same model at 8k tokens. Kev-27B holds to 65,536 tokens, the serving limit (its 64k lower bound is −2.4 pp). Kev-0.8B, 4B and 9B validate only their trained 8,192: each already misses the tolerance at 16k (lower bounds −8.5, −3.4 and −3.7 pp), so past 8k tokens their answers on long documents are not covered by the measurement.
- Formal model cards. All four cards follow one structure: summary, details, intended and out-of-scope uses, how to use, training data and procedure, evaluation, limitations, risks, compute, provenance.
Known limitations
- In-distribution gains. The large gains of the last year are on suites whose training splits are in the training data. On datasets no Kev trained on, Kev-27B is 1.7 index points below Jev on breadth-v1 test, and the smaller sizes are 13–31 points below.
- Untrained lengths. Kev-0.8B, 4B and 9B trained on states of at most 7,552 tokens and Kev-27B on at most 32,768; the server accepts 65,536. Use the validated context length, not the serving limit.
- Kev-27B on long contracts is less accurate than v1 and overconfident (CUAD test ECE 0.053 against 0.007); refit the temperature on your own documents or use
@v1-lorafor contract review. - Kev-0.8B and tool routing. Its When2Call accuracy fell below chance (0.133 on test) after its documents-and-skills stage; do not use it for tool-call routing.
- Date arithmetic is the weakest family at every size below 27B (
deadlinepolicy accuracy 0.35 / 0.65 / 0.725 against Jev's 0.95);KEV_DATE_FACTS=1helps. - Knowledge is set by the base (MMLU-Pro 0.230–0.675 against Jev's 0.840).
- Kev-9B on a Mac has not been measured, and Kev-27B on a Mac is expected to fit 96–128 GB but has not been run.
- Selection. Kev-27B v2 and Kev-9B v2 were re-selected under the audited rule with earlier development reads known; their test margins are optimistic.
- One temperature per model cannot reorder confidences, so at a 5 % error budget the models automate fewer decisions than Jev out of domain.
How to run
:::bash
git clone https://github.com/jaredpalmer/kev.git && cd kev && uv sync --extra serve
uv run --extra serve python -m kev.serve --run jaredpalmer/kev-4b@v1.0 --port 8009 # CUDA, or MLX on Apple Silicon
From a release tarball:
:::bash
shasum -a 256 -c SHA256SUMS.txt
tar -xzf kev-4b.tar.gz
uv run --extra serve python -m kev.serve --run kev-4b --port 8009
The TypeSafe SDK works unchanged: TypeSafeClient(api_key="local", base_url="http://127.0.0.1:8009", model="kev-latest"). Kev-27B needs one B200, H200 or H100 80 GB: --run jaredpalmer/kev-27b@v1.0. To deploy an HTTPS endpoint on Modal, see skills/kev-deploy.
Assets
Each tarball holds one checkpoint as it is on the Hub at its weights revision (LoRA adapter, head.pt with the temperature, tokenizer files, the training trial's result.json, provenance.json, training_config.json, training_metrics.json and log), the Kev 1.0 model card as README.md, and the locked transfer-v4 read as locked_test.json. They are built by scripts/build_release_assets.py from docs/releases/kev-1.0-assets.json, and rebuilding gives the same bytes. Every file in them is dated 2026-10-01 00:00 UTC, which kev.serve reports as the release date of an unpacked checkpoint. The tarballs first attached on 2026-10-01 dated their files 1970-01-01, so kev.serve reported 1969-12-31; they were replaced the same day by these. The weights and every other file are byte-identical; only the tarball hashes changed.
| File | SHA-256 | Checkpoint | adapter / head SHA-256 |
|---|---|---|---|
kev-0.8b.tar.gz (46 MB) |
0ae144c7675f0c3f333be0bb878a0f202ab9e6fa84169fb7cb16efe6176c1ef1 |
jaredpalmer/kev-0.8b@9a45d25e |
9b908623… / f400bd12… |
kev-4b.tar.gz (131 MB) |
2e707e2ebd08980dc7881222b7024cea5606401441c1a086afb170ae7784201c |
jaredpalmer/kev-4b@139fdd94 |
90e81735… / dd633435… |
kev-9b.tar.gz (172 MB) |
acd13320b7d1b052ce989f19ca9d1d9ba5219b8beced0ee67337908aef1deb3f |
jaredpalmer/kev-9b@b5d8c18e |
2b2a70cf… / 8e1dab2c… |
Kev-27B is not attached, because its 51 GB of weights are over GitHub's 2 GB limit per asset. Download it from the Hub: jaredpalmer/kev-27b@v1.0 (weights commit 28be62e9, head.pt 7968f17b…).
On every Hub repo, the v1.0 tag points to the commit that uploaded the Kev 1.0 card. That commit changed only README.md, so its weights are the same bytes as the weights revision in the first table: kev-0.8b bf75a6a8, kev-4b 6cfce5c2, kev-9b db029f08, kev-27b af0e6d55.