Jev vs Cygnet
A frozen Gemma 4 12B with a one-token readout — no training at all.
Jev 1.13.0
TypeSafe · System One model
63.3JevBench rank #4
Highest intelligence in the top ten: 53.1.
Cygnet
blockbrain · A frozen Gemma 4 12B with a one-token readout — no training at all
61.8JevBench rank #6
Sixth of 91 with zero fine-tuning, and ahead of Jev on the hardest public cases.
Key differences
- Where it runs JevHosted API, generally availableCygnetSelf-hosted on vLLM; measured on 48–96 GB GPUs
- Training JevPost-trained decision modelCygnetNone — stock Gemma 4 12B plus one temperature
- Evaluation-style cases Jev94.5% correctCygnet95.2% correct
- Hardest test cases Jev74.1% correctCygnet75.5% correct
- Sealed decisions Jev36.7% correctCygnet33.8% correct
- Options and input size JevUp to 255 options, 64k tokensCygnetUp to 26 options, 16,384 tokens
JevBench v1.4.2.2, measured the same way
One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗
Composite score
The top four finish within 4.1 points of each other, then the board falls away sharply.
Capability, higher is better
Accuracy by how hard the decision is
Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.
What Cygnet is
Cygnet is blockbrain’s proof that you do not need to train anything: a frozen Gemma 4 12B with a one-token readout — the model’s next-token distribution over the answer labels, read once — plus a single temperature. There are no weights to download beyond the stock model, just a recipe and a calibration step.
On JevBench v1.4.2.2 it ranks #6 at 61.8, and it is the only entry in the top ten that beats Jev on the hardest public tier (75.5% vs 74.1%) — plus a small edge on judge-style cases (95.2% vs 94.5%). The sealed set tells the other half: 33.8% vs Jev’s 36.7%, where cases nobody could tune for live.
Where Cygnet falls short
- Needs a 12B model on 48–96 GB of GPU under vLLM — heavier than every 4B entry above it.
- Gemma’s licence applies, not Apache — check the terms for your use.
- One-token readout caps choice questions at 26 options and context at 16,384 tokens.
- Sealed accuracy trails Jev by ~3 points: 33.8% vs 36.7%.
When to use which
Choose Jev if
- You want sealed-set robustness and calibrated probabilities as delivered.
- Your choice questions exceed 26 options or states exceed 16k tokens.
- You want a hosted endpoint instead of operating vLLM on a 12B model.
Choose Cygnet if
- You want zero training: stock Gemma 4 12B plus a readout recipe.
- Your hardest public cases dominate — it edges Jev there (75.5% vs 74.1%).
- You already run a vLLM fleet with 48 GB+ GPUs.
FAQ
Is Cygnet better than Jev?
On the composite, no — #6 at 61.8 vs Jev’s #4 at 63.3. But it is the only top-ten system ahead of Jev on the hardest public tier (75.5% vs 74.1%) and on judge-style cases (95.2% vs 94.5%). Jev answers with the sealed set: 36.7% vs 33.8%.
Does Cygnet require fine-tuning?
No — that is the point. It is a frozen Gemma 4 12B with a one-token readout recipe and one temperature. No weights beyond the stock model.
What licence is Cygnet under?
The readout recipe is blockbrain’s; the model itself is Google’s Gemma 4 12B under the Gemma licence — not Apache-2.0. Check the Gemma terms for your deployment.
What hardware does Cygnet need?
vLLM serving Gemma 4 12B — the benchmark measured it on 48 GB and 96 GB GPUs. Much heavier than the 4B entries it sits behind on composite.