Research
Decision models compared: Jev, Clef, d1 and Perplexity
Every answer on Knowviq is checked before you see it. The checking is done by a decision model. TypeSafe Jev, the model Knowviq uses today, now has company: Cloudflare released Clef and Clef-flash on 1 October 2026, and Liquid AI and Perplexity both offer decision models through their APIs. We ran all five on the same 938 checks from our own traffic. This post explains what decision models are, what we measured, and why the main site is staying on Jev for now. You can also try the Clef version yourself.
What a decision model is
A decision model reads some input (the "state") and answers questions about it with a typed value and probabilities. It does not write text. TypeSafe's API has three question types (TypeSafe docs):
- noul: a yes/no question, answered as the probability of yes (0 to 1).
- choice: pick one option from a list, with a probability for every option and a confidence.
- score: rate on an ordered scale of 2 to 10 levels, returning a possibly fractional score, the probability of each level and a confidence.
That sits between two familiar tools. A large language model can answer any question, but it answers in prose, which code then has to parse, and its stated certainty is just more text. A classic classifier gives clean probabilities, but only for the one task it was trained on. A decision model takes questions written in plain language, like an LLM, and returns a probability your code can compare against a threshold, like a classifier. Many questions can be asked about the same state in one call.
The probabilities are meant to be calibrated: across many answers, things given 0.8 should be true about 80% of the time. TypeSafe is careful to say calibration is measured across groups of predictions and does not guarantee any single answer (System One concepts).
TypeSafe Jev
Jev is TypeSafe's flagship model, which it calls the first "System One" model, after the fast, intuitive System 1 thinking Daniel Kahneman described. It takes text only (strings, JSON objects and arrays), not images (TypeSafe). Knowviq has used Jev since launch.
Cloudflare Clef and Clef-flash
Cloudflare announced Clef and Clef-flash on 1 October 2026 (Cloudflare blog). They run on Workers AI and accept the same request shape as Jev. What Cloudflare says about them:
- Clef is built on a frozen Qwen3.8-27B backbone and Clef-flash on Qwen3.5-9B. The weights are open, under Apache 2.0, on Hugging Face.
- They accept images as well as text, with a 64k context.
- Median latency of 209.3 ms for Clef and 38.8 ms for Clef-flash, against 524.1 ms for Jev (p95: 238.6, 122.4 and 536.0 ms).
- On TypeSafe's own evaluation suite, the Clef models beat Jev in 3 of 4 workflow areas (Jev led on agent trace observability).
- Workers AI lists Clef at US$0.24 and Clef-flash at US$0.09 per million input tokens (model pages). What that meant for us per call is in the results below.
These benchmarks are Cloudflare's own and self-reported. Latency in particular depends on where you call from. That is why we ran our own test.
Liquid AI d1
d1 is Liquid AI's first decision model
(Liquid docs).
It uses the same System One request and response format as Jev, with
the same noul, choice and score questions, at an endpoint with the
same name (/decisions/v1/systemone). Liquid's docs show
TypeSafe's own SDKs calling it by pointing them at Liquid's address.
We tested the d1:free tier. We could not find a paid
price on Liquid's own pages; Vercel's AI Gateway lists d1 at US$0.04
per million input tokens
(Vercel).
Perplexity Decisions API
Perplexity's Decisions API serves one model,
pplx-decider-v1-27b
(Perplexity docs).
It takes the same three question types and accepts text, JSON or
images as the state. It costs US$0.04 per million input tokens, output
is free, and each organisation can send 10 requests a second.
How Knowviq uses a decision model
For each search, three kinds of decision decide what you see:
- Result scoring. Each result is scored on how directly it answers the search. Only verified results are shown, with a percentage.
- Sentence support. The draft answer is split into sentences, and each one is checked against evidence excerpts. A sentence is shown only if the model says the evidence supports it. Unsupported sentences are never shown.
- The answerable gate. A yes/no check on whether the results actually answer the question. If not, we don't show an answer.
Our head-to-head
Method
We rebuilt 938 real Knowviq checks with our own production code: 670 result-scoring requests and 268 sentence-check requests, about 9,300 individual decisions. 100 of the result requests pair a search with another search's results, as known negatives. The identical request went to every model, and the answers were run through Knowviq's own rules. Jev and the two Clef models were tested first; d1 and Perplexity were added a few hours later the same day with the same requests, labels and scoring. Every model answered every request in the end, with no malformed answers; d1 and Perplexity needed some polite retries when they asked us to slow down. All tests ran on 2 October 2026 (AEST).
Results
| Measure (items) | Jev | Clef | Clef-flash | d1 | Perplexity |
|---|---|---|---|---|---|
| Result relevance, AUC (2,228) | 0.899 | 0.894 | 0.873 | 0.905 | 0.907 |
| Answerable gate, AUC (647) | 0.983 | 0.982 | 0.979 | 0.984 | 0.982 |
| Good answers hidden by the gate at 0.75 | 22.2% | 7.3% | 6.2% | 10.7% | 8.0% |
| Best result picked correctly (240) | 77.9% | 80.0% | 78.7% | 78.7% | 78.7% |
| Sentence support, AUC (random 50) | 0.978 | 0.786 | 0.729 | 0.935 | 0.970 |
| Supported sentences hidden (random 50) | 4.9% | 34% | 41% | 0% | 4.9% |
| Weakly supported sentences shown (of 15) | 7 | 2 | 2 | 10 | 7 |
| Median time from our test box, results / sentences | 98 / 122 ms | 1.29 / 1.28 s* | 0.73 / 0.76 s* | 0.78 / 1.23 s | 283 / 445 ms |
| Cost per 1,000 calls, our mix | $0.21 | $0.93 | $0.35 | free tier† | $1.28 |
* Through a local development proxy, not from inside a Worker; see the speed caveats below. † No paid price published by Liquid; at Vercel's listed price our mix would cost about $1.30 per 1,000 calls. None of the five models showed a sentence we labelled unsupported.
In plain terms:
- Ranking results: all close. Perplexity and d1 rank results slightly better than Jev and Clef; Clef-flash is a little behind.
- The answerable gate: Jev is the outlier. All five rank answerable questions equally well, but at our 0.75 threshold Jev hides 22.2% of answers we judged good. The others hide 6–11%. Clef, Clef-flash and Perplexity let through slightly more bad ones (4.5% of negatives against 2.7% for Jev and d1).
- Sentence support: Jev and Perplexity lead, d1 is close, Clef trails. On the random 50, Jev and Perplexity are statistically tied. d1 hid no supported sentences but let more weakly supported ones through. Clef often marks sentences unsupported even when the evidence says the same thing word for word.
- Our badge rule doesn't fit every model. Our rule (score ≥ 2.5 and confidence ≥ 0.7) hides 41% of relevant results under Jev, 33% under Perplexity, 59% under d1, and 82% and 87% under Clef and Clef-flash, whose confidence values are much less peaked.
- Cost depends on token counting. d1 and Perplexity charge little per token, but for the same requests they reported about 6.4 times as many input tokens as Jev. So in our workload Perplexity cost about 6 times as much as Jev per call, Clef 4.4 times and Clef-flash 1.6 times. d1 is free for now.
Caveats on the labels
- Answerable labels are our own judgements of past answers. Only 12 were "no", plus the 100 mismatched-results negatives, so false-show rates rest on few items.
- Relevance labels come from matching answer keys in result text. That is noisy: it misses paraphrases and can match out of context.
- We labelled 150 sentence/evidence pairs ourselves. A second check by another model agreed 91.3% of the time. Only 50 of the 150 are a random sample; the rest were picked where Jev and the Clef models disagreed or scored low, before d1 and Perplexity were added. So we lead with the random 50, where small gaps (Jev 0.978 against Perplexity 0.970) are within noise.
Caveats on speed
From our test machine in US West, Jev answered in a median of 98 to 122 ms, Perplexity in 283 to 445 ms and d1 in 0.78 to 1.23 s. We first reached Clef through a local Wrangler development proxy, where it took about 1.28 s, and even a one-question call took a median of 654 ms. Once /clef was live, we measured Clef from inside our own Worker: about 1.1 to 1.4 s per result-scoring call, with a median of about 1.4 s. That is much slower than Cloudflare's published 209 ms median, but our requests are large (thousands of tokens and many questions each), so it isn't a like-for-like comparison.
Clef-flash stalled during a slow, one-at-a-time run between 5:15 and 5:30 AEST: a one-question call took a median of 13.7 s, and one took 108 s. It was stable under our batch run. d1's free tier sometimes asked us to slow down even when we sent one request at a time, so it would need a paid plan before carrying real traffic.
Why the main site stays on Jev
Showing only supported sentences is the core of Knowviq. Jev and Perplexity are the best at it and tied in our test, with d1 close behind. Jev is also the fastest from where we tested and by far the cheapest per check in our workload. Perplexity ranks results slightly better and is better calibrated on the answerable gate, which makes it the strongest alternative and a sensible backup, but it costs about 6 times as much per check. d1 is promising and free today, but it is slower, rate-limited on the free tier, and has no published paid price. Clef is better calibrated on the gate but well behind on sentence support. None of the four clearly beats Jev for what we do.
The other models did teach us something about Jev. All five rank answerable questions equally well, so Jev's high false-hide rate comes from our 0.75 threshold, not from Jev's judgement. That threshold looks too strict, and we plan to test a lower one, such as 0.6, once we have more judged "no" answers.
Knowviq with Clef
knowviq.com/clef runs the same
search and the same pipeline, but every decision is made by Clef
(@cf/cloudflare/clef). We chose Clef over Clef-flash
because it was more accurate in our test and didn't stall. Copying
Jev's thresholds would have made Clef absurdly strict, so we set
Clef's own from the same data:
- Result badge: score ≥ 2.5 and confidence ≥ 0.3 (Jev: 0.7). On our labelled results this shows wrong results 9.4% of the time and hides 41.4% of relevant ones, about the same as Jev's rule on Jev (9.1% and 41.1%). The percentage shown is the probability that the result is relevant.
- Answerable gate: 0.75, unchanged. Clef is not strict here.
- Best result: confidence ≥ 0.5 (Jev: 0.7). At 0.5 Clef's pick is used 49 times and right 69% of the time, close to Jev at 0.7 (71 times, 68%).
- Sentences: shown only when Clef's probability of support is at least 0.5 and the sentence cites a source (Jev: 0.7). In the random 50, that shows 31 sentences (Jev shows 41), with 2 of the 9 weak or unsupported ones getting through for both models and no unsupported ones. Anything Clef rates below 0.4 counts as unsupported and is never shown.
If Clef errors or times out, /clef says so. It never quietly falls back to Jev. The footer shows recent Clef decision times measured inside our Worker, which will help with the latency question. Follow-ups and photo search are off on /clef.
What we'll retest
- Clef again in 4 to 8 weeks, including latency from inside a Worker.
- Sentence support with one sentence per request, to see whether bundling about ten questions per call hurts Clef.
- Whether Clef-flash still stalls under light traffic.
- d1 once Liquid publishes paid pricing and limits, and Perplexity as a backup provider for Jev.
- A lower answerable threshold for Jev.
- Larger label sets: 300 randomly sampled sentences with a second human pass, and judged relevance in place of answer-key matching.
Try it
Search the same thing on Knowviq and on Knowviq with Clef and compare. All of these models are calibrated tools, not proofs, and all of them make mistakes.
Sources
- Cloudflare: Clef decision models (1 October 2026)
- Workers AI model page: clef and clef-flash
- TypeSafe docs: Introduction and System One
- Liquid AI docs: Decision Models and Vercel AI Gateway: d1
- Perplexity docs: Decisions API
- Knowviq's own evaluation, run on 2 October 2026 (numbers above).