Jev vs Winnow
Gemma 4 12B decision fine-tune — decisions, chat and vision from one local model.
Jev 1.13.0
TypeSafe · System One model
63.3JevBench rank #4
Highest intelligence in the top ten: 53.1.
Winnow
EldanRing · Gemma 4 12B decision fine-tune — decisions, chat and vision from one local model
55.6JevBench rank #8
Eighth of 91: one model for decisions, chat and images on a 16 GB GPU.
Key differences
- Where it runs JevHosted API, generally availableWinnowSelf-hosted GGUF; Q8 tested on a 16 GB GPU
- What it does JevTyped decisions onlyWinnowTyped decisions, chat and image input from one server
- Evaluation-style cases Jev94.5% correctWinnow91.1% correct
- Hardest test cases Jev74.1% correctWinnow70.9% correct
- Sealed decisions Jev36.7% correctWinnow33.1% correct
- Probabilities JevCalibrated — 76.3 on JevBenchWinnowNormalised logits, no fitted calibration — 64.8
JevBench v1.4.2.2, measured the same way
One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗
Composite score
The top four finish within 4.1 points of each other, then the board falls away sharply.
Capability, higher is better
Accuracy by how hard the decision is
Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.
What Winnow is
Winnow is EldanRing’s Gemma 4 12B fine-tune that keeps the base model’s chat and vision abilities and adds the System One decision interface. Run as a GGUF (the Q8 quant was measured on a 16 GB GPU), it gives you one local model that answers typed questions, talks, and reads images.
On JevBench v1.4.2.2 it ranks #8 at 55.6. The gap to Jev is mainly calibration — 64.8 vs 76.3 — because Winnow emits normalised logits rather than a fitted probability; you would calibrate it on your own labels. Accuracy tiers are respectable: 91.1% judge, 70.9% hard, 33.1% sealed.
Where Winnow falls short
- No fitted calibration: probabilities are normalised logits until you fit them on your data.
- A 12B model is bigger and slower to serve than the 4B entries above it.
- Judge-tier accuracy trails Jev by ~3.4 points; sealed by ~3.6.
When to use which
Choose Jev if
- You need calibrated probabilities out of the box — 76.3 vs 64.8 on JevBench.
- Decisions are your workload, not chat — Jev does one thing and is versioned for it.
- You want a hosted endpoint instead of serving a 12B GGUF.
Choose Winnow if
- One model for decisions, chat and images on a single 16 GB GPU.
- You are willing to fit calibration on your own labels.
- Everything stays local — no per-token API cost.
FAQ
Is Winnow better than Jev?
On JevBench, no — #8 at 55.6 vs Jev’s #4 at 63.3. Its pitch is different: one local model that does decisions, chat and vision, at the cost of calibration (64.8 vs 76.3) and judge-tier accuracy (91.1% vs 94.5%).
Can Winnow really do chat and decisions?
Yes — it is a Gemma 4 12B fine-tune that keeps the base chat and vision abilities while adding the typed-decision interface. That breadth is the trade: it is not a dedicated decision model.
What hardware does Winnow need?
A 16 GB GPU for the Q8 GGUF build — far friendlier than most 12B options, heavier than the 4B entries.
Are Winnow’s probabilities reliable?
They are normalised logits with no fitted calibration — 64.8 on JevBench’s calibration axis vs Jev’s 76.3. Fit a temperature or isotonic map on your own labels before trusting thresholds.