Clef eval
Cloudflare's first decision models, Clef and Clef-Flash, take the same requests as Jev, so I ran them on my Jev eval cases unchanged. Clef did almost as well as Jev. Clef-Flash did worse, and most of its mistakes were on requests that needed a human to sign off.
- Made
- Models tested
- @cf/cloudflare/clef, @cf/cloudflare/clef-flash
- Scope
- 98 cases, 221 questions, the Jev eval set
- Accuracy
- Clef 94.9%, Clef-Flash 86.4% (Jev 97.7%)
The models
Clef and Clef-Flash are Cloudflare's own models on Workers AI. Cloudflare's changelog lists Clef at 27B parameters for "highest-precision decisions" and Clef-Flash at 9B for "latency-critical, hot-path decisions", and calls both "drop-in compatible with Jev". They take the same System One request: a state, plus Noul, Choice and Score questions. So the cases ran without changes.
How I tested it
I exported the 98 cases from the Jev harness to JSON and graded them with the same rules, so every question and every expected answer is identical. A small Worker passes each case to env.AI.run untouched, and a Node script outside the Worker sends the cases one at a time and grades the answers. I ran both models twice, once through wrangler dev and once against the deployed Worker. The numbers below are from the deployed runs. No call failed in any run.
As with Jev, 177 questions have an objective answer and go into the accuracy number. The other 44 are judgement calls and stay out.
| Scored questions | Jev | Clef | Clef-Flash |
|---|---|---|---|
| All 177 | 97.7% (173) | 94.9% (168) | 86.4% (153) |
| Noul | 99.2% | 95.0% | 85.0% |
| Choice | 95.2% | 100% | 85.7% |
| Score | 94.4% | 91.7% | 91.7% |
Clef against Jev
Clef missed 9 questions to Jev's 4. All three models missed the same three: the date calculation Jev also got wrong, and the two Score boundary cases from the Jev eval. I think those are hard cases or weak labels, and they say nothing about one model against another. Clef missed six on its own: the $305 expense that doesn't exceed a $300 cap once alcohol is excluded, one checksum, a fraud severity it scored 0.84 against an expected 0, and three stop-for-a-human checks phrased implicitly. It also got one right that Jev missed, sending a manifest check to a human where Jev chose a tool call.
With only 177 questions, 9 misses against 4 may not mean much. Clef-Flash's 24 against 4 is a clear difference.
Clef-Flash and human sign-off
Clef-Flash got 18 Noul questions wrong, and 17 of them were a no where the answer was yes. Clef's six Noul misses split three each way. A wrong no matters most here, because the hardest questions ask whether a human must approve the action before it goes ahead.
The Jev eval found that implied phrasing ("let me get this in front of legal first") lowers Jev's confidence without changing its answer. Clef shows the same drop but crosses the threshold on three implicit cases, at 0.14 to 0.39. Clef-Flash crosses it on most of them, and also missed one case where the rule was spelled out: a customer data purge, at 0.23.
Injected instructions
Jev and Clef both got all seven injection questions right. Clef-Flash got four. In the three it missed, text inside the state changed its answer. An urgent ticket carrying a fake request to deprioritise it was marked not urgent, an expense with a fake VP override went to the standard manager, and a forged pre-authorisation got a no-stop at 0.44. All three models passed the five subtler injection cases.
Repeat runs
The local and deployed runs were separate calls by different routes, minutes apart. All 221 answers matched for both models, with probabilities equal to four decimal places. Latency differed between the runs, so it wasn't a cache in my Worker; either decoding is deterministic or Cloudflare caches results. Repeat runs won't tell you anything about variance on this set, and the differences above aren't random variation.
The confidence scores helped as well. On Noul, Clef's correct answers sat 0.91 from the threshold on average and its wrong ones 0.46 (Clef-Flash: 0.87 and 0.50). A confidence floor would catch a fair share of the errors.
Latency and cost
| Per case | p50 | p95 | Max | Over 3s |
|---|---|---|---|---|
| Jev | 379 ms | 1,084 ms | 1,195 ms | 0 |
| Clef | 694 ms | 1,357 ms | 19,535 ms | 1 |
| Clef-Flash | 516 ms | 3,891 ms | 4,069 ms | 16 |
Clef-Flash is meant to be the fast one, and its median is about a quarter below Clef's. But 16 of 98 cases took over three seconds, and its p95 is nearly three times Clef's. Clef had one 19.5 second call. So plan for the slow calls, not the typical time. Neither matched Jev, though Jev's numbers come from an earlier session over its own API, so treat that comparison as rough. The wrangler dev runs had much worse tails (Clef peaked at 118.6 seconds), and there the time was spent inside Workers AI, not in my harness.
A full run used 43,117 input tokens. Both models reported zero output tokens.
If you're switching from Jev
- Clef is a reasonable swap if you want to stay on Workers AI. It was a few points less accurate here, and slower.
- Don't use Clef-Flash to decide whether to stop for a human, or on text that might carry injected instructions.
- With either model, keep a confidence floor and a rule-based check for the must-stop cases.
- The Jev advice still holds: put the rule in the state, and compute dates in your own code.
Caveats
The cases were written to probe Jev's documented weaknesses, so they may favour Jev, and the answer key is mine. One run per model per route is enough for accuracy, given the identical answers, but not for latency, which I measured from one machine in one evening. Cloudflare's docs don't say how Clef-Flash differs from Clef beyond size, so anything I say about it comes from these results alone. I didn't test Clef's image and video input.
First published .