jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs JevK5

Open Qwen3.5-4B with a distilled LoRA — the base of the #2 model on the board.

On this site

Jev 1.13.0

TypeSafe · System One model

63.3JevBench rank #4

Highest intelligence in the top ten: 53.1.

Alternative

JevK5

allebee · Open Qwen3.5-4B with a distilled LoRA — the base of the #2 model on the board

62JevBench rank #5

Fifth of 91 ranked systems, and the weights are yours to run.

01

Key differences

  1. Where it runs
    JevHosted API, generally available
    JevK5Self-hosted on your GPU, about 9 GB in bf16
  2. Licence
    JevProprietary, hosted
    JevK5Apache-2.0 weights and code
  3. Evaluation-style cases
    Jev94.5% correct
    JevK594.5% correct — a tie
  4. Hardest test cases
    Jev74.1% correct
    JevK570.0% correct
  5. Sealed decisions
    Jev36.7% correct
    JevK533.1% correct
  6. Options and input size
    JevUp to 255 options, 64k tokens
    JevK5Up to 16 options, 16,384 tokens
02

JevBench v1.4.2.2, measured the same way

One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗

Composite score

  1. #1 Imajev-4B 67.4
  2. #2 Plumb-4B 65.8
  3. #3 decider-4b v2 64.1
  4. #4 Jev 1.13.0 63.3
  5. #5 JevK5 v0.2.0 62
  6. #6 Cygnet 61.8
  7. #7 Hopper 59.4
  8. #8 Winnow-12B Q8 55.6
  9. #9 reflex 4B 54
  10. #10 djev 52.2
  11. #13 SemIf 47.7
  12. #29 OpenJev 36.9

The top four finish within 4.1 points of each other, then the board falls away sharply.

Capability, higher is better

IntelligenceHow often it picks the right answer
Jev53.1
JevK548.9
CalibrationWhether 0.8 really means about 80%
Jev76.3
JevK574.5
SpeedMeasured response time
Jev83.3
JevK591.1

Accuracy by how hard the decision is

Easy72 straightforward cases
Jev 100% JevK5 100%
Standard96 everyday cases
Jev 99% JevK5 95.8%
Judge146 evaluation-style calls
Jev 94.5% JevK5 94.5%
Hard220 genuinely ambiguous cases
Jev 74.1% JevK5 70%
Sealed308 private cases · chance is 29.3%
Jev 36.7% JevK5 33.1%

Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.

03

What JevK5 is

JevK5 is allebee’s open Qwen3.5-4B System One rebuild: base weights plus a distilled LoRA, Apache-2.0 code and weights, about 9 GB in bf16. It is also the substrate for Plumb-4B — crh225’s LoRA fine-tune that sits at #2 on the same board.

On JevBench v1.4.2.2 it ranks #5 at 62.0, 1.3 points behind Jev. It is the only system in the top five that ties Jev on the judge tier (94.5%). Where it gives ground is the hard tier (70.0% vs 74.1%), the sealed set (33.1% vs 36.7%), and interface limits: 16 options per choice question and 16,384 tokens of context, versus Jev’s 255 options and 64k tokens.

Where JevK5 falls short

  • Choice questions cap at 16 options — Jev supports up to 255.
  • Context is 16,384 tokens, a quarter of Jev’s 64k-token request budget.
  • JevBench did not re-measure its latency in v1.4 — check the source table before planning around the speed figure.
04

When to use which

Choose Jev if

  • Your choice questions have more than 16 options, or states past 16k tokens.
  • You want calibration and sealed-set robustness without fitting anything.
  • You want a hosted, versioned endpoint instead of GPU operations.
Try Jev free

Choose JevK5 if

  • You want Apache-2.0 weights — the strongest judge-tier parity with Jev on the board.
  • You plan to fine-tune: Plumb-4B (#2) shows what a LoRA on this base reaches.
  • Decisions must stay inside your own network on a ~9 GB model.
Get JevK5

FAQ

Is JevK5 better than Jev?

Not on this board: #5 at 62.0 vs Jev’s #4 at 63.3. It ties Jev on judge-style cases (94.5%) but trails on hard (70.0% vs 74.1%) and sealed (33.1% vs 36.7%) tiers, and caps at 16 options / 16k tokens vs 255 / 64k.

How does JevK5 relate to Plumb-4B?

Plumb-4B is crh225’s LoRA fine-tune on top of JevK5 v0.2 — it sits at #2 on the same board at 65.8. JevK5 is the base you run or extend.

Is JevK5 open source?

Yes — Apache-2.0 weights and code by allebee, a Qwen3.5-4B rebuild of the System One interface. Independent of TypeSafe.

Can I serve JevK5 behind the Jev API shape?

It implements the same state-plus-questions contract. Watch the limits: 16 options per choice and 16,384-token context, both well under Jev’s.

Sources