jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Leaderboard guide · JevBench v1.4.2.2 · updated September 28, 2026

JevBench results, explained

JevBench is the independent leaderboard for Jev-style decision models. In JevBench v1.4.2.2 three open 4B models finish above Jev on the composite score, led by Imajev-4B at 67.4 against Jev's 63.3. Jev still has the highest intelligence score in the top ten. Here is the board, what the score measures, and why those two statements are both true.

01

The top ten

Composite first place: Imajev-4B. Highest intelligence: Jev 1.13.0. Best on the hard tier: Plumb-4B. Fastest: Plumb-4B. No one system leads everything — which is why the composite alone is a poor guide.

RankSystemScoreIntelligenceCalibrationSpeedHard tierSealedRuns as
#1 Imajev-4BMohit Garg · Qwen3.5-4B LoRA 67.4 52.2 80.4 90.6 — 37% Self-hosted GPU
#2 Plumb-4Bcrh225 · JevK5 v0.2 + LoRA 65.8 53 75.5 93.5 77.7% 38% Self-hosted GPU
#3 decider-4b v2Mapika · Qwen3.5-4B 64.1 49.4 75 92.9 67.3% 34.7% Self-hosted, RTX PRO 6000
#4 Jev 1.13.0TypeSafe 63.3 53.1 76.3 83.3 74.1% 36.7% Hosted API
#5 JevK5 v0.2.0allebee · Qwen3.5-4B 62 48.9 74.5 91.1 70% 33.1% Self-hosted, RunPod GPU
#6 Cygnetblockbrain · frozen Gemma 4 12B 61.8 49.5 74.9 90.7 75.5% 33.8% Self-hosted, RTX PRO 6000
#7 HopperHopitAI · Qwen3.5-4B LoRA 59.4 48 79.1 86.8 65% 34.1% Self-hosted, RTX A6000
#8 Winnow-12B Q8EldanRing · Gemma 4 12B 55.6 48.3 64.8 82.3 70.9% 33.1% Self-hosted, 16 GB GPU
#9 reflex 4Bkshetrajna12 54 47.5 70.4 68 63.2% 28.2% Self-hosted GPU
#10 djevMaisa · DiffusionGemma 52.2 47 55.4 91.4 69.5% 29.9% Hosted API, preview

Score, intelligence, calibration and speed are 0–100, higher is better. The hard tier is 220 decisions; sealed is 308 private ones. For Imajev-4B the benchmark publishes the hard tier's two halves — 72.1% public, 75.2% held out — but no combined figure. The fourth axis, cost, is part of the score but not shown; the full table is at the source.

02

Same numbers, different weighting

The composite weighs speed and cost as heavily as getting the answer right. JevBench also publishes other views of the same measurements — change the weighting and the top of the board reorders.

SystemOfficial compositeAccuracy weighted 60%Intelligence only
Imajev-4B#1#1#3
Plumb-4B#2#2#2
decider-4b v2#3#4#7
Jev 1.13.0#4#3#1
JevK5 v0.2.0#5#6#8

If your pipeline is high-volume and mostly routine, the composite is a fair guide. If the decisions that cost you money are the ambiguous ones, read the accuracy columns instead — that is where Jev's post-training shows up.

03

How to read the score

  • Composite — harmonic mean of intelligence, calibration, speed and cost, equally weighted. One weak axis pulls it down hard; a cheap fast 4B can outrank a smarter hosted model.
  • Intelligence — how often the system picks the right answer. Sealed decisions make up 20% of this axis.
  • Calibration — whether a stated 0.8 really means about 80%. Critical if you set thresholds like “auto-approve above 0.9”.
  • Tiers — 72 easy, 96 standard, 146 judge-style, 220 hard public cases. Easy and standard separate almost nothing; the hard tier is where systems disagree.
  • Sealed — 308 private cases introduced in v1.4, chance 29.3%. The integrity check on everything above.
04

Head-to-head pages

Each comparison page breaks the same board down to one pair — key differences, full tier table, and when to pick which.

FAQ

What is JevBench?

JevBench is an independent leaderboard for Jev-style System One decision models, published by Benchmark Heaven. Each system answers the same 842 decisions — 534 public across easy, standard, judge and hard tiers, plus 308 sealed cases whose text stays private.

Is Jev the best model on JevBench?

Not on the official composite — Jev ranks #4 at 63.3, behind three open 4B models led by Imajev-4B at 67.4. On the intelligence-only view Jev ranks #1 in the top ten. The composite weighs speed and cost as heavily as accuracy, so which is “best” depends on what you weight.

Why do small open models rank above Jev?

The composite is a harmonic mean of four equally weighted axes: intelligence, calibration, speed and cost. A self-hosted 4B scores very well on speed (no network) and cost (no API bill), which can outweigh a few points of accuracy — the harmonic mean punishes any weak axis hard.

What are sealed decisions?

308 of the 842 decisions per system are sealed: their text is never published, so no system could have been tuned on them. Chance is 29.3%. They are the best check on whether public-tier accuracy generalises.

How current is this table?

This page reflects JevBench v1.4.2.2, published 27 September 2026: 95 systems listed, 91 ranked. The live board is at benchmarkheaven.com/jev-models; we link it below and update this page when a new version ships.

Sources