jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs Winnow

Gemma 4 12B decision fine-tune — decisions, chat and vision from one local model.

On this site

Jev 1.13.0

TypeSafe · System One model

63.3JevBench rank #4

Highest intelligence in the top ten: 53.1.

Alternative

Winnow

EldanRing · Gemma 4 12B decision fine-tune — decisions, chat and vision from one local model

55.6JevBench rank #8

Eighth of 91: one model for decisions, chat and images on a 16 GB GPU.

01

Key differences

  1. Where it runs
    JevHosted API, generally available
    WinnowSelf-hosted GGUF; Q8 tested on a 16 GB GPU
  2. What it does
    JevTyped decisions only
    WinnowTyped decisions, chat and image input from one server
  3. Evaluation-style cases
    Jev94.5% correct
    Winnow91.1% correct
  4. Hardest test cases
    Jev74.1% correct
    Winnow70.9% correct
  5. Sealed decisions
    Jev36.7% correct
    Winnow33.1% correct
  6. Probabilities
    JevCalibrated — 76.3 on JevBench
    WinnowNormalised logits, no fitted calibration — 64.8
02

JevBench v1.4.2.2, measured the same way

One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗

Composite score

  1. #1 Imajev-4B 67.4
  2. #2 Plumb-4B 65.8
  3. #3 decider-4b v2 64.1
  4. #4 Jev 1.13.0 63.3
  5. #5 JevK5 v0.2.0 62
  6. #6 Cygnet 61.8
  7. #7 Hopper 59.4
  8. #8 Winnow-12B Q8 55.6
  9. #9 reflex 4B 54
  10. #10 djev 52.2
  11. #13 SemIf 47.7
  12. #29 OpenJev 36.9

The top four finish within 4.1 points of each other, then the board falls away sharply.

Capability, higher is better

IntelligenceHow often it picks the right answer
Jev53.1
Winnow48.3
CalibrationWhether 0.8 really means about 80%
Jev76.3
Winnow64.8
SpeedMeasured response time
Jev83.3
Winnow82.3

Accuracy by how hard the decision is

Easy72 straightforward cases
Jev 100% Winnow 100%
Standard96 everyday cases
Jev 99% Winnow 96.9%
Judge146 evaluation-style calls
Jev 94.5% Winnow 91.1%
Hard220 genuinely ambiguous cases
Jev 74.1% Winnow 70.9%
Sealed308 private cases · chance is 29.3%
Jev 36.7% Winnow 33.1%

Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.

03

What Winnow is

Winnow is EldanRing’s Gemma 4 12B fine-tune that keeps the base model’s chat and vision abilities and adds the System One decision interface. Run as a GGUF (the Q8 quant was measured on a 16 GB GPU), it gives you one local model that answers typed questions, talks, and reads images.

On JevBench v1.4.2.2 it ranks #8 at 55.6. The gap to Jev is mainly calibration — 64.8 vs 76.3 — because Winnow emits normalised logits rather than a fitted probability; you would calibrate it on your own labels. Accuracy tiers are respectable: 91.1% judge, 70.9% hard, 33.1% sealed.

Where Winnow falls short

  • No fitted calibration: probabilities are normalised logits until you fit them on your data.
  • A 12B model is bigger and slower to serve than the 4B entries above it.
  • Judge-tier accuracy trails Jev by ~3.4 points; sealed by ~3.6.
04

When to use which

Choose Jev if

  • You need calibrated probabilities out of the box — 76.3 vs 64.8 on JevBench.
  • Decisions are your workload, not chat — Jev does one thing and is versioned for it.
  • You want a hosted endpoint instead of serving a 12B GGUF.
Try Jev free

Choose Winnow if

  • One model for decisions, chat and images on a single 16 GB GPU.
  • You are willing to fit calibration on your own labels.
  • Everything stays local — no per-token API cost.
Get Winnow

FAQ

Is Winnow better than Jev?

On JevBench, no — #8 at 55.6 vs Jev’s #4 at 63.3. Its pitch is different: one local model that does decisions, chat and vision, at the cost of calibration (64.8 vs 76.3) and judge-tier accuracy (91.1% vs 94.5%).

Can Winnow really do chat and decisions?

Yes — it is a Gemma 4 12B fine-tune that keeps the base chat and vision abilities while adding the typed-decision interface. That breadth is the trade: it is not a dedicated decision model.

What hardware does Winnow need?

A 16 GB GPU for the Q8 GGUF build — far friendlier than most 12B options, heavier than the 4B entries.

Are Winnow’s probabilities reliable?

They are normalised logits with no fitted calibration — 64.8 on JevBench’s calibration axis vs Jev’s 76.3. Fit a temperature or isotonic map on your own labels before trusting thresholds.

Sources