Skip to main content

OpenJEV — One Refusal Test, an 80 MB File That Is Real, and the Option That Did Not Help

· 19 min read
Baris Bayrak
Software Engineer

On 15 September 2026 TypeSafe released Jev, a model that takes a state and a set of typed questions and returns decisions with probabilities rather than text. Within eleven days there were at least ten open reimplementations of it, in five architecturally different ways, and one of them was trending on GitHub. That speed is itself the finding: "don't generate tokens, return probabilities" is an interface, and an interface takes an afternoon to copy. What does not copy is the part nobody puts in the launch video — whether the model can say nothing here applies.

We ran three of them against a real tool surface, 48 requests, one machine. Choosing the correct tool was close: 14 of 24 for a 121M on-device model, 15 of 24 for a 151M encoder-based one, 10 of 24 for a 144M multilingual one. Refusing a request that no tool could serve was not close at all: 3 of 24, 16 of 24, and 0 of 24. And when we handed the third model an explicit "none of these" option, it ignored it on every single one of the 24. That last result cost us a claim we had already written down, which is the most useful thing in this post.

The Question the Model Card Never Answers​

In an agent, a decision model sits in front of the tool list and answers "which of these applies?". The interesting case is not the request that matches a tool. It is the request that matches none of them, where the only correct output is an empty result.

That framing is inherited, and it is worth naming where from. A tool from an earlier project on this site audits technical claims against the hard ceiling of their own physics, and the single rule that made it work was keeping a control set of real products the tool must never flag. Without it, three separate domains shipped a confidently wrong bound. The same discipline applies here with the sign flipped: the control set is 24 requests that no tool in the set can serve, and the correct behaviour on every one of them is silence. A decision model that always returns an answer is not a router. It is a guessing machine with a probability attached.

So: 11 tools, 24 requests each way. The tools are the real ones registered in our own Rust agent library, read from the source rather than invented for the test:

In scope — 24 requests, correct answer is a named tool"show me the unstaged changes in this repo", "run cargo test", "search for TODO in the rust sources", "remember that the store is SQLite with FTS5", …
Off-topic — 24 requests, correct answer is nothing"what is the capital of France?", "give me a recipe for lentil soup", "play some jazz music", "write a haiku about the sea", …

What We Ran, and One Control First​

Needle 3 (Cactus Compute, 121M parameters, Apache-2.0) is the only one of the nine that ships a callable artifact: a static C engine plus a 2-bit container. We linked libneedle.a into a Rust binary, loaded the 35,335,380-byte weight file, and drove the C API directly. The whole runtime — weights, engine, tokenizer, config — is 36.7 MB. No GPU, no Python, no network.

Verdict / OpenJev (151M on a ModernBERT encoder, Apache-2.0) is a fine-tune of a GLiClass zero-shot classifier with an explicit abstention candidate in its schema. It exports to ONNX, so it is reachable from Rust, though we ran it through PyTorch. Its weights are 605.5 MB in fp32 — 16× the whole Needle runtime, which is the price of the encoder design.

The first thing we did with Verdict was not run our test. It was to run theirs. An earlier pass of our own — 12 labels described as bare imperatives — produced a near-uniform distribution that looked exactly like a broken model, and it was actually a mis-shaped instrument: the engine renders every option as the literal sentence It is {description}, so "Run a program or shell command" was never a sentence fragment that template could use. Their README example caught it in one run: 2 of 2 correct on its home domain, and it abstains cleanly on an off-topic request. Only then was our test fair. We discarded that first pass rather than publish its number.

Julia-1 (Supersonic Labs, 2026-09-26) is the most carefully documented release in this set and the reason is worth stating plainly: it publishes its own failure. A 144.3M-parameter model built on jhu-clsp/mmBERT-small, Apache-2.0, trained to choose among 2–20 supplied answers, running on a CPU. It ships its evaluation as machine-readable metrics, a provenance.json, an inference-policy.json, and a reproduction script pinned to a dataset revision — and we verified the weights before using them: model.safetensors is 577,189,056 bytes and its SHA-256 is df853bf7fe42…, matching the hash in both their policy file and their reproduction script.

Two things in that paperwork made it the obvious third engine. It reports coverage alongside accuracy, and its Banking77 result lists {"ok": 99, "abstained": 1} — so this is a release that tracks abstention as a first-class outcome. And its inference-policy.json sets "calibration": null, meaning it explicitly does not claim calibrated probabilities. Both are rare. Its card states the weights occupy 550.5 MiB, and they do: 577,189,056 bytes is 550.5 MiB exactly. Running it cost us an afternoon: 577 MB, no GPU, and the same PyTorch stack they used.

The Result​

24 off-topic inputs where no tool appliesNeedle 3 (121M, on-device C)Verdict (151M, encoder + abstain route)Julia-1 (144M, mmBERT-small)
Refused3 of 24 — 12.5%16 of 24 — 67%0 of 24 — 0%
Picked the right tool14 of 2415 of 2410 of 24
Mean confidence, wrong call0.6220.2830.648
Mean confidence, right call0.8600.4970.839
On disk36.7 MB605.5 MB577.2 MB
Per-turn, this machine57–460 ms~1.4 s39–63 ms

The consequence for an agent is not symmetric. A wrong tool call is a bad turn. A router that answers 21 of 24 unusable requests is a system that cannot tell the user "I can't do that here" — and this is not a small model doing its best under a hard constraint, because the same model refuses correctly some of the time, which we will come to.

Verdict's failures are the mirror image, and it publishes them: four of its misses were false abstentions — it declined requests that did have a tool, including "build the project with cargo build" and "remember that the store is SQLite with FTS5". Its own documentation puts abstention recall at 97.5% on held-out in-domain data with precision down to 89.45% and 23 false abstentions. That is a calibration trade, and it is the right trade for a router: a false "I can't" costs one clarification, a confident wrong tool call costs the turn and possibly the session.

Confidence Is Not a Refusal Signal​

This is the part we did not expect, and it is the most useful thing in the whole exercise.

Needle's three refusals arrived at 0.986, 0.990 and 0.990 confidence. Its most confident wrong call was 0.891. Whatever the refusal path is measuring, it is not the model's uncertainty: it declines when it is certain there is no tool, and calls when it is uncertain but has a plausible match. So confidence is not merely useless as a scope gate — in the tail it points the wrong way.

Concretely: a gate set at 0.7 would have acted on 9 of 24 requests that no tool could serve. Confidence does separate right from wrong on average — 0.860 for correct calls against 0.622 for wrong ones — and that separation reproduced exactly across two sample sizes. It is simply a different quantity from "is this request in scope", and the two get conflated constantly because both are reported as a number between 0 and 1.

Eleven Releases, Five Families​

The releases are not variations on one design. They are five different bets:

FamilyWhoWhat it actually is
Read the answer scores off a stock LLMSemIf (MIT), Simple Jev (featherless-ai)No new weights at all — intercept the model before it generates and read its option logits
Train a dedicated non-generative modelLaya, Von (395M), Verdict (151M), Nimble, Julia-1 (144M)Purpose-built encoders with decision heads; the honest projects publish calibration curves and failure slices
Retrofit an LLM with a decision headKev (Qwen3.5 0.8B–8B)Keep the LLM, add the head, share one reading of the state across many questions
Diffusion fill-inOpenJev on DiffusionGemma-26BBlank the answer slots and fill them all in one shot, like inpainting
Freeze the LLM, train contrastive headsCLMThe encoder is never touched; what you download is a head, not a model
Not a reimplementationCUA-S1-FORMS (trycua)Research code, no weights distributed — the one release here that cannot be run at all

Only one of these families scores 16 of 24 on the refusal test, and it is the one whose schema carries an explicit insufficient evidence route. That correlation is real; the explanation we first gave for it was wrong, and the third engine is what showed us. We had written that a schema of five options makes "none of them" unrepresentable, "no matter how good the weights are". Julia-1 does not have that limitation — and it still refused nothing. Representation is necessary and nowhere near sufficient: a model refuses only if the option is there and it was trained to use it. Verdict's route is trained (its card names the objective, RLCD); the others have never seen an abstain option, so handing them one puts it out of distribution.

The most valuable sentence in any of the eleven releases is Verdict measuring itself against the thing it reimplements: on TypeSafe's public benchmark OpenJev scores 48.07%, against 90.80% for Jev and 88.43% for a 26B diffusion model. A small model honestly reporting that it is half as accurate as the model it copies — alongside a cardinality study showing its own top-1 accuracy falling from 97% at three candidates to 72% at 25 — is worth more than any launch chart. Julia-1's equivalent is a Banking77 pilot at 64 of 100 against Jev's 87%, published next to the three tasks where it wins.

The Option That Did Not Help​

This is the experiment we would run first, if we were starting again, because it converts a correlation into a cause. On the same 24 off-topic requests and the same 11 tools, we ran Julia-1 twice: once with the 11 tool options exactly as every other engine saw them, and once with a twelfth option described as "a request that none of these tools can serve".

Julia-1optionsOff-topic refusalsTool correctmean p, wrong / right
A — tools only110 of 2410 of 240.648 / 0.839
B — tools + an offered abstain option120 of 248 of 240.636 / 0.992

Offering the option changed nothing. On all 24 unusable requests it picked a tool anyway — 16 times execute_command, plus three web_fetch, three read_file and two web_search — and on "what is the capital of France?" it reached for a shell command. Adding the option did not merely fail to help; it also cost two points of tool accuracy, which is what you would expect if the extra label is noise to a model that has no idea what to do with it.

The conclusion is narrower than the one we published first, and more useful. A reimplementation can copy the interface in a day, but the interface is not what makes a decision model safe to put in front of a tool list. What makes it safe is a route that is present and trained, and that is a data and objective decision made long before anyone writes a launch post.

The 80 MB File Called "8B"​

CLM arrived with a claim that does not survive its own file listing: an "8B" model whose only weights file is 75.6 MB. Eight billion parameters at fp16 is about 16 GB, so those numbers cannot describe the same object. They don't, and the model card says so plainly: CLM is "two small projection heads on top of a frozen Qwen3-8B encoder".

The arithmetic confirms it exactly. Each head is a projection 4096 → 1536 → 1536 → 512, which is 4096×1536 + 1536 + 1536×1536 + 1536 + 2×1536 + 512×1536 + 512 = 9,443,840 parameters, so the pair is 18,887,680 in fp32 = 75,550,720 bytes against a file of 75,557,149. The remaining 6,429 bytes — 0.0085% — are torch archive overhead, not weights.

So the file is real, correctly described, and not the good news either. The base it requires, Qwen3-8B, is 8,190,726,144 parameters at bf16 = 16,381,452,288 bytes, which is 216.8× the size of the artifact you download. Deploying CLM means serving an 8B encoder on a GPU with vLLM; the 19M of heads are a rounding error on top. It is the opposite of the on-device proposition, and it will never be small.

Two more caveats are in the card rather than the announcement. The headline agentic results — DeepSWE 81.6%, Terminal-Bench 2.1 87.6% — come from fine-tuned heads in a separate repository, not from this checkpoint, which is the starting point for that fine-tune. And its "up to 9×" is the card's own "13× faster with ~1k candidates", which is a caching result.

And the structural point that matters here: CLM cannot abstain at all. Searching its schema, server, engine, client and fine-tuning guide for any notion of abstention or refusal returns nothing. Its yes/no type is strictly false/true, and its choice type is a softmax over whatever candidates you supply — the card's own words are "no generation: CLM only scores the candidates you give it, and its probabilities are relative to that set". On our 24 controls it would return an answer every time, 0 of 24 by construction, and adding a "none of the above" option is a modelling decision the user has to make because the API does not make it.

The Idea Worth Stealing​

CLM's speed claim is not about its architecture being clever. It is that states and actions are encoded separately, so the action side is computed once and reused while the state changes. For a fixed action set that is exactly the right decomposition — and it describes our problem precisely: an agent's tool list is fixed, and what changes per turn is the request. Encode the 11 tools once at startup, encode only the state per turn, score. That works with any scorer in the table above, needs no new model, and is the cheapest real improvement in this entire body of work.

Candidate Count Matters, and Not Monotonically​

One more measurement, because the candidate menu turns out to be a variable rather than a constant. Same 24 inputs at three menu sizes:

Verdict surfacelabelsOff-topic refusalsTool correct
first 4 tools519 of 24 (79%)10 of 24
first 8 tools98 of 24 (33%)14 of 24
all 11 tools1216 of 24 (67%)15 of 24

Refusal falls hard from 5 labels to 9, then rises at 12. Three points at n = 24 cannot explain that, and we are not going to pretend otherwise — but it is enough to say the effect is real and not simply "more options, worse", which is consistent with Verdict's own note that its abstention "depends heavily on the seen prompt structure". Their published curve, on 100 samples per slice, is the version to trust: 97% → 96% → 91% → 78% → 72% accuracy as the menu grows from 3 to 25. An 11-tool router is a 12-candidate problem, which is where their own data already says to be careful.

What It Cannot Do​

  • It is 24 hand-written requests per side, not a benchmark. The differences are large enough to act on; they are not precision measurements, and the sweep above shows how sharp the sensitivity to framing is.
  • One tool surface, one language, one working domain. These models are classifiers over text. Every number here is out-of-domain for the encoders involved.
  • Nothing was fine-tuned. All three engines ran zero-shot. The largest advertised gains in this field come from per-product fine-tuning, which we did not do.
  • No engine was tested inside a live agent loop. This measures whether the decision is right, not whether a turn improves.
  • CLM was read, not run. Reaching it means vLLM plus 16 GB of Qwen3-8B, which is not worth an afternoon to confirm something its own schema already settles.
  • Julia-1's own abstentions do not come from the model. Its metrics report coverage and an abstained status on Banking77, but the word does not appear in any code file it ships, and its reproduction script always returns a choice. Whatever produced those abstentions lives in a Banking77-specific harness that is not in the repository.
  • Julia-1's probabilities are declared uncalibrated and are then cosmetically sharpened. Its policy file sets "calibration": null, and a presentation step rounds a winner above 0.95 with all others under 0.045 to a one-hot 1.0. That is a display rule, disclosed in the docstring, and it is why confident answers read as exactly 1.000 in our table. Do not read those ones as measured certainty.
  • A refusal is not a safety property. "Nothing applies" is a scope judgement. It is not a check that a request is safe, and nothing here should be put in a security path in place of a deterministic one.

What We Learned​

What worked: three things, all cheap. The first is the control set — 24 off-topic requests that no tool can serve — which is the only reason "3 of 24" is a number at all rather than a story about a model that seemed fine. The second is running the vendor's own example before judging it on ours; that single control separated "broken model" from "mis-shaped test" in one run, and the first number we would otherwise have published was wrong. And a third, quieter one: sometimes the answer is in the source rather than the weights. CLM's inability to refuse is not a measurement, it is a grep, and it took a minute. Julia-1 made the same point from the other direction — its published SHA-256 let us verify 577 MB of weights before running anything, and its own policy file told us not to trust its calibration before we had formed an opinion about it.

What didn't — seven of our own instruments, all caught before publication. A C API whose return value is not a byte length, which silently truncated every response into what looked like a model that refuses everything. An off-by-one in our own validator that marked every call as ungrounded and turned a real signal into noise. A 12-label test that was mis-shaped for the model's label template. A summary line that reported "4 refusals" in a run where zero occurred, and whose first fix reported 3 — both wrong, in a log that a reader would have believed instead of the rows. A sample of 12 that let us write "it never refuses" when 24 makes it refuse three times. A sweep we declared unreproducible and marked do-not-cite because we had read its log file while the job was still running; the file was complete, and we had condemned valid evidence on a stale read. And a version pin we ignored: Julia-1 declares transformers>=5.0,<5.1, our shared environment had 5.17, and their encoder patch broke in a way that looked like a model failure until we built the environment they actually specified. Two of those generalise. A background job's output is not evidence until the job has exited. And when a release pins a dependency range, that range is part of the artifact.

The uncomfortable part: we published an explanation and the next release falsified it. Our claim was that abstention is a property of the schema — put a "none of these" option in the contract and the model has somewhere to put the answer. The first half is true and the second half is not: Julia-1 was offered exactly that option and used it zero times in 24, then lost accuracy for having been given it. What we got wrong was treating representability as sufficient, when it is only necessary — Verdict's route works because it was trained with an objective that rewards honest abstention, not because the label existed. That is a narrower claim than the one we first made, and it is the one the evidence supports. It is also the reason to run a control you did not design the model for: the 24 requests could only ever have shown us that the engines differ, and it took adding one option to a third engine to show us why.


Every number above is re-read from committed run artifacts by scripts/check-openjev-post.py — the Needle counts are parsed out of the harness logs, the Verdict and Julia-1 counts out of the per-request JSON, and the CLM arithmetic is re-derived from its config. It fails if any figure drifts, or if a figure appears in this post with no derivation behind it. scripts/mutation-test-openjev-checker.py corrupts the post ten ways and requires every corruption to be caught, so that the checker is not merely a thing that passes. --live re-fetches the pinned vendor claims, which are the parts that can change under a published post.

Updating this post. This is meant to be re-run, not rewritten: the harness, the evidence rows for all three engines, and the checker are all committed. To add a model, run it on the same 24 + 24 requests against the same 11 tools, add its rows, add one line to the family table, and let the checker decide whether the prose still holds.

Updated 2026-09-26: Julia-1 added and measured, which required correcting the explanation this post first gave for why abstention separates the engines. The correction is in place above rather than in a footnote, because we would rather the post be wrong in public for a day than quietly right later.