jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs decider-4b

Open Qwen3.5-4B System One rebuild, 8.4 GB of Apache-2.0 weights.

On this site

Jev 1.13.0

TypeSafe · System One model

63.3JevBench rank #4

Highest intelligence in the top ten: 53.1.

Alternative

decider-4b

Mapika · Open Qwen3.5-4B System One rebuild, 8.4 GB of Apache-2.0 weights

64.1JevBench rank #3

Third of 91 on the composite — ahead of Jev on the score, behind it on the answers.

01

Key differences

  1. Where it runs
    JevHosted API, generally available
    decider-4bSelf-hosted on one GPU, about 8.4 GB in bf16
  2. Licence
    JevProprietary, hosted
    decider-4bApache-2.0 weights and package
  3. Speed
    Jev83.3 · 0.65s median over the network
    decider-4b92.9 · 17ms raw on the benchmark’s GPU
  4. Evaluation-style cases
    Jev94.5% correct
    decider-4b87.7% correct
  5. Hardest test cases
    Jev74.1% correct
    decider-4b67.3% correct
  6. Sealed decisions
    Jev36.7% correct
    decider-4b34.7% correct
  7. Context
    Jev64k tokens per request
    decider-4b32k tokens
02

JevBench v1.4.2.2, measured the same way

One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗

Composite score

  1. #1 Imajev-4B 67.4
  2. #2 Plumb-4B 65.8
  3. #3 decider-4b v2 64.1
  4. #4 Jev 1.13.0 63.3
  5. #5 JevK5 v0.2.0 62
  6. #6 Cygnet 61.8
  7. #7 Hopper 59.4
  8. #8 Winnow-12B Q8 55.6
  9. #9 reflex 4B 54
  10. #10 djev 52.2
  11. #13 SemIf 47.7
  12. #29 OpenJev 36.9

The top four finish within 4.1 points of each other, then the board falls away sharply.

Capability, higher is better

IntelligenceHow often it picks the right answer
Jev53.1
decider-4b49.4
CalibrationWhether 0.8 really means about 80%
Jev76.3
decider-4b75
SpeedMeasured response time
Jev83.3
decider-4b92.9

Accuracy by how hard the decision is

Easy72 straightforward cases
Jev 100% decider-4b 100%
Standard96 everyday cases
Jev 99% decider-4b 96.9%
Judge146 evaluation-style calls
Jev 94.5% decider-4b 87.7%
Hard220 genuinely ambiguous cases
Jev 74.1% decider-4b 67.3%
Sealed308 private cases · chance is 29.3%
Jev 36.7% decider-4b 34.7%

Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.

03

Which decider?

Mapika publishes several decider models and releases come quickly. JevBench has ranked three of them. The ranked one is decider-4b v2: Qwen3.5-4B-Base, 4.2B parameters, about 8.4 GB in bf16, released 24 September — the version kept under the Hugging Face tag v2.

decider-4b v2.1 is the current default, released the same day: v2 plus a further LoRA stage. Its author reports it slightly weaker on JevBench’s public hard items (0.649 vs 0.676) and less well calibrated on hard items, so it is not the version the board measured.

decider-35b-a3b (Qwen3.5-35B-A3B-Base, 34.7B parameters with 3B active, ~65 GB bf16 or 19.6 GB NVFP4) ranks #21 at 41.2, held back by cost. decider-2b (Qwen3.5-2B-Base, 1.9B) ranks #41 at 30.7, held back by calibration. decider-0.8b and decider-2b-vision are not on the board.

Where decider-4b falls short

  • The ranked checkpoint is v2 — the newer v2.1 default is weaker on hard items per its own author.
  • 32k-token context is half of Jev’s 64k request budget.
  • Judge-tier accuracy trails Jev by nearly 7 points: 87.7% vs 94.5%.
04

When to use which

Choose Jev if

  • Evaluation-style prompts are your workload: judge-tier accuracy is 94.5% vs 87.7%.
  • Your ambiguous cases matter — hard tier 74.1% vs 67.3%, sealed 36.7% vs 34.7%.
  • You want a hosted, versioned endpoint instead of GPU operations.
Try Jev free

Choose decider-4b if

  • You need Apache-2.0 weights on your own GPU — about 8.4 GB in bf16.
  • Raw latency matters: 17ms measured raw on the benchmark GPU vs a networked call.
  • You want to fine-tune the decision head on your own labels.
Get decider-4b

FAQ

Is decider-4b better than Jev?

On the JevBench composite, yes — #3 at 64.1 vs Jev’s #4 at 63.3, mostly through speed and cost. On the accuracy tiers Jev leads everywhere it counts: judge 94.5% vs 87.7%, hard 74.1% vs 67.3%, sealed 36.7% vs 34.7%, and intelligence 53.1 vs 49.4.

Which decider-4b version should I run?

JevBench ranked v2 (the Hugging Face tag v2). The newer v2.1 default adds a LoRA stage its author reports is slightly weaker on hard items and less calibrated — check the repo’s notes before defaulting to it.

Is decider-4b open source?

Yes — Apache-2.0 weights and package from Mapika, a Qwen3.5-4B-Base rebuild of the System One request shape. Independent of TypeSafe.

Can decider-4b replace a Jev API call?

It speaks the same state-plus-questions contract at 32k tokens. For judge-style and ambiguous cases, test your own prompts first — Jev leads those tiers by several points.

Sources