Blog

Research

Decision models compared: Jev, Clef, d1 and Perplexity

Every answer on Knowviq is checked before you see it. The checking is done by a decision model. TypeSafe Jev, the model Knowviq uses today, now has company: Cloudflare released Clef and Clef-flash on 1 October 2026, and Liquid AI and Perplexity both offer decision models through their APIs. We ran all five on the same 938 checks from our own traffic. This post explains what decision models are, what we measured, and why the main site is staying on Jev for now. You can also try the Clef version yourself.

What a decision model is

A decision model reads some input (the "state") and answers questions about it with a typed value and probabilities. It does not write text. TypeSafe's API has three question types (TypeSafe docs):

That sits between two familiar tools. A large language model can answer any question, but it answers in prose, which code then has to parse, and its stated certainty is just more text. A classic classifier gives clean probabilities, but only for the one task it was trained on. A decision model takes questions written in plain language, like an LLM, and returns a probability your code can compare against a threshold, like a classifier. Many questions can be asked about the same state in one call.

The probabilities are meant to be calibrated: across many answers, things given 0.8 should be true about 80% of the time. TypeSafe is careful to say calibration is measured across groups of predictions and does not guarantee any single answer (System One concepts).

TypeSafe Jev

Jev is TypeSafe's flagship model, which it calls the first "System One" model, after the fast, intuitive System 1 thinking Daniel Kahneman described. It takes text only (strings, JSON objects and arrays), not images (TypeSafe). Knowviq has used Jev since launch.

Cloudflare Clef and Clef-flash

Cloudflare announced Clef and Clef-flash on 1 October 2026 (Cloudflare blog). They run on Workers AI and accept the same request shape as Jev. What Cloudflare says about them:

These benchmarks are Cloudflare's own and self-reported. Latency in particular depends on where you call from. That is why we ran our own test.

Liquid AI d1

d1 is Liquid AI's first decision model (Liquid docs). It uses the same System One request and response format as Jev, with the same noul, choice and score questions, at an endpoint with the same name (/decisions/v1/systemone). Liquid's docs show TypeSafe's own SDKs calling it by pointing them at Liquid's address. We tested the d1:free tier. We could not find a paid price on Liquid's own pages; Vercel's AI Gateway lists d1 at US$0.04 per million input tokens (Vercel).

Perplexity Decisions API

Perplexity's Decisions API serves one model, pplx-decider-v1-27b (Perplexity docs). It takes the same three question types and accepts text, JSON or images as the state. It costs US$0.04 per million input tokens, output is free, and each organisation can send 10 requests a second.

How Knowviq uses a decision model

For each search, three kinds of decision decide what you see:

  1. Result scoring. Each result is scored on how directly it answers the search. Only verified results are shown, with a percentage.
  2. Sentence support. The draft answer is split into sentences, and each one is checked against evidence excerpts. A sentence is shown only if the model says the evidence supports it. Unsupported sentences are never shown.
  3. The answerable gate. A yes/no check on whether the results actually answer the question. If not, we don't show an answer.

Our head-to-head

Method

We rebuilt 938 real Knowviq checks with our own production code: 670 result-scoring requests and 268 sentence-check requests, about 9,300 individual decisions. 100 of the result requests pair a search with another search's results, as known negatives. The identical request went to every model, and the answers were run through Knowviq's own rules. Jev and the two Clef models were tested first; d1 and Perplexity were added a few hours later the same day with the same requests, labels and scoring. Every model answered every request in the end, with no malformed answers; d1 and Perplexity needed some polite retries when they asked us to slow down. All tests ran on 2 October 2026 (AEST).

Results

Measure (items)JevClefClef-flashd1Perplexity
Result relevance, AUC (2,228)0.8990.8940.8730.9050.907
Answerable gate, AUC (647)0.9830.9820.9790.9840.982
Good answers hidden by the gate at 0.7522.2%7.3%6.2%10.7%8.0%
Best result picked correctly (240)77.9%80.0%78.7%78.7%78.7%
Sentence support, AUC (random 50)0.9780.7860.7290.9350.970
Supported sentences hidden (random 50)4.9%34%41%0%4.9%
Weakly supported sentences shown (of 15)722107
Median time from our test box, results / sentences98 / 122 ms1.29 / 1.28 s*0.73 / 0.76 s*0.78 / 1.23 s283 / 445 ms
Cost per 1,000 calls, our mix$0.21$0.93$0.35free tier†$1.28

* Through a local development proxy, not from inside a Worker; see the speed caveats below. † No paid price published by Liquid; at Vercel's listed price our mix would cost about $1.30 per 1,000 calls. None of the five models showed a sentence we labelled unsupported.

In plain terms:

Caveats on the labels

Caveats on speed

From our test machine in US West, Jev answered in a median of 98 to 122 ms, Perplexity in 283 to 445 ms and d1 in 0.78 to 1.23 s. We first reached Clef through a local Wrangler development proxy, where it took about 1.28 s, and even a one-question call took a median of 654 ms. Once /clef was live, we measured Clef from inside our own Worker: about 1.1 to 1.4 s per result-scoring call, with a median of about 1.4 s. That is much slower than Cloudflare's published 209 ms median, but our requests are large (thousands of tokens and many questions each), so it isn't a like-for-like comparison.

Clef-flash stalled during a slow, one-at-a-time run between 5:15 and 5:30 AEST: a one-question call took a median of 13.7 s, and one took 108 s. It was stable under our batch run. d1's free tier sometimes asked us to slow down even when we sent one request at a time, so it would need a paid plan before carrying real traffic.

Why the main site stays on Jev

Showing only supported sentences is the core of Knowviq. Jev and Perplexity are the best at it and tied in our test, with d1 close behind. Jev is also the fastest from where we tested and by far the cheapest per check in our workload. Perplexity ranks results slightly better and is better calibrated on the answerable gate, which makes it the strongest alternative and a sensible backup, but it costs about 6 times as much per check. d1 is promising and free today, but it is slower, rate-limited on the free tier, and has no published paid price. Clef is better calibrated on the gate but well behind on sentence support. None of the four clearly beats Jev for what we do.

The other models did teach us something about Jev. All five rank answerable questions equally well, so Jev's high false-hide rate comes from our 0.75 threshold, not from Jev's judgement. That threshold looks too strict, and we plan to test a lower one, such as 0.6, once we have more judged "no" answers.

Knowviq with Clef

knowviq.com/clef runs the same search and the same pipeline, but every decision is made by Clef (@cf/cloudflare/clef). We chose Clef over Clef-flash because it was more accurate in our test and didn't stall. Copying Jev's thresholds would have made Clef absurdly strict, so we set Clef's own from the same data:

If Clef errors or times out, /clef says so. It never quietly falls back to Jev. The footer shows recent Clef decision times measured inside our Worker, which will help with the latency question. Follow-ups and photo search are off on /clef.

What we'll retest

Try it

Search the same thing on Knowviq and on Knowviq with Clef and compare. All of these models are calibrated tools, not proofs, and all of them make mistakes.

Sources