JevBench results, explained
JevBench is the independent leaderboard for Jev-style decision models. In JevBench v1.4.2.2 three open 4B models finish above Jev on the composite score, led by Imajev-4B at 67.4 against Jev's 63.3. Jev still has the highest intelligence score in the top ten. Here is the board, what the score measures, and why those two statements are both true.
- 95systems listed · 91 of them ranked
- 842decisions per system · 534 public + 308 sealed
- 29.3%chance on the sealed set · best in top ten is 38.0%; Jev gets 36.7%
- 4equally weighted axes · intelligence, calibration, speed, cost
The top ten
Composite first place: Imajev-4B. Highest intelligence: Jev 1.13.0. Best on the hard tier: Plumb-4B. Fastest: Plumb-4B. No one system leads everything — which is why the composite alone is a poor guide.
| Rank | System | Score | Intelligence | Calibration | Speed | Hard tier | Sealed | Runs as |
|---|---|---|---|---|---|---|---|---|
| #1 | Imajev-4B | 67.4 | 52.2 | 80.4 | 90.6 | — | 37% | Self-hosted GPU |
| #2 | Plumb-4B | 65.8 | 53 | 75.5 | 93.5 | 77.7% | 38% | Self-hosted GPU |
| #3 | decider-4b v2 | 64.1 | 49.4 | 75 | 92.9 | 67.3% | 34.7% | Self-hosted, RTX PRO 6000 |
| #4 | Jev 1.13.0 | 63.3 | 53.1 | 76.3 | 83.3 | 74.1% | 36.7% | Hosted API |
| #5 | JevK5 v0.2.0 | 62 | 48.9 | 74.5 | 91.1 | 70% | 33.1% | Self-hosted, RunPod GPU |
| #6 | Cygnet | 61.8 | 49.5 | 74.9 | 90.7 | 75.5% | 33.8% | Self-hosted, RTX PRO 6000 |
| #7 | Hopper | 59.4 | 48 | 79.1 | 86.8 | 65% | 34.1% | Self-hosted, RTX A6000 |
| #8 | Winnow-12B Q8 | 55.6 | 48.3 | 64.8 | 82.3 | 70.9% | 33.1% | Self-hosted, 16 GB GPU |
| #9 | reflex 4B | 54 | 47.5 | 70.4 | 68 | 63.2% | 28.2% | Self-hosted GPU |
| #10 | djev | 52.2 | 47 | 55.4 | 91.4 | 69.5% | 29.9% | Hosted API, preview |
Score, intelligence, calibration and speed are 0–100, higher is better. The hard tier is 220 decisions; sealed is 308 private ones. For Imajev-4B the benchmark publishes the hard tier's two halves — 72.1% public, 75.2% held out — but no combined figure. The fourth axis, cost, is part of the score but not shown; the full table is at the source.
Same numbers, different weighting
The composite weighs speed and cost as heavily as getting the answer right. JevBench also publishes other views of the same measurements — change the weighting and the top of the board reorders.
| System | Official composite | Accuracy weighted 60% | Intelligence only |
|---|---|---|---|
| Imajev-4B | #1 | #1 | #3 |
| Plumb-4B | #2 | #2 | #2 |
| decider-4b v2 | #3 | #4 | #7 |
| Jev 1.13.0 | #4 | #3 | #1 |
| JevK5 v0.2.0 | #5 | #6 | #8 |
If your pipeline is high-volume and mostly routine, the composite is a fair guide. If the decisions that cost you money are the ambiguous ones, read the accuracy columns instead — that is where Jev's post-training shows up.
How to read the score
- Composite — harmonic mean of intelligence, calibration, speed and cost, equally weighted. One weak axis pulls it down hard; a cheap fast 4B can outrank a smarter hosted model.
- Intelligence — how often the system picks the right answer. Sealed decisions make up 20% of this axis.
- Calibration — whether a stated 0.8 really means about 80%. Critical if you set thresholds like “auto-approve above 0.9”.
- Tiers — 72 easy, 96 standard, 146 judge-style, 220 hard public cases. Easy and standard separate almost nothing; the hard tier is where systems disagree.
- Sealed — 308 private cases introduced in v1.4, chance 29.3%. The integrity check on everything above.
Head-to-head pages
Each comparison page breaks the same board down to one pair — key differences, full tier table, and when to pick which.
FAQ
What is JevBench?
JevBench is an independent leaderboard for Jev-style System One decision models, published by Benchmark Heaven. Each system answers the same 842 decisions — 534 public across easy, standard, judge and hard tiers, plus 308 sealed cases whose text stays private.
Is Jev the best model on JevBench?
Not on the official composite — Jev ranks #4 at 63.3, behind three open 4B models led by Imajev-4B at 67.4. On the intelligence-only view Jev ranks #1 in the top ten. The composite weighs speed and cost as heavily as accuracy, so which is “best” depends on what you weight.
Why do small open models rank above Jev?
The composite is a harmonic mean of four equally weighted axes: intelligence, calibration, speed and cost. A self-hosted 4B scores very well on speed (no network) and cost (no API bill), which can outweigh a few points of accuracy — the harmonic mean punishes any weak axis hard.
What are sealed decisions?
308 of the 842 decisions per system are sealed: their text is never published, so no system could have been tuned on them. Chance is 29.3%. They are the best check on whether public-tier accuracy generalises.
How current is this table?
This page reflects JevBench v1.4.2.2, published 27 September 2026: 95 systems listed, 91 ranked. The live board is at benchmarkheaven.com/jev-models; we link it below and update this page when a new version ships.