jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs Imajev-4B

Open Qwen3.5-4B decision model that also reads photos.

On this site

Jev 1.13.0

TypeSafe · System One model

63.3JevBench rank #4

Highest intelligence in the top ten: 53.1.

Alternative

Imajev-4B

Mohit Garg · Open Qwen3.5-4B decision model that also reads photos

67.4JevBench rank #1

First of 91 on JevBench, with the best calibration in the top ten.

01

Key differences

  1. Where it runs
    JevHosted API, generally available
    Imajev-4BSelf-hosted on a Mac or one GPU
  2. Inputs
    JevText only
    Imajev-4BText, plus up to two photos per request
  3. “Can’t tell”
    JevNot documented as a separate output
    Imajev-4BA trained unknown probability on every answer
  4. Evaluation-style cases
    Jev94.5% correct
    Imajev-4B89.0% correct
  5. Sealed decisions
    Jev36.7% correct
    Imajev-4B37.0% correct
  6. Calibration (JevBench)
    Jev76.3
    Imajev-4B80.4
  7. Input size
    Jev64k tokens per request
    Imajev-4BState up to 32 KB, about 8k tokens
02

JevBench v1.4.2.2, measured the same way

One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗

Composite score

  1. #1 Imajev-4B 67.4
  2. #2 Plumb-4B 65.8
  3. #3 decider-4b v2 64.1
  4. #4 Jev 1.13.0 63.3
  5. #5 JevK5 v0.2.0 62
  6. #6 Cygnet 61.8
  7. #7 Hopper 59.4
  8. #8 Winnow-12B Q8 55.6
  9. #9 reflex 4B 54
  10. #10 djev 52.2
  11. #13 SemIf 47.7
  12. #29 OpenJev 36.9

The top four finish within 4.1 points of each other, then the board falls away sharply.

Capability, higher is better

IntelligenceHow often it picks the right answer
Jev53.1
Imajev-4B52.2
CalibrationWhether 0.8 really means about 80%
Jev76.3
Imajev-4B80.4
SpeedMeasured response time
Jev83.3
Imajev-4B90.6

Accuracy by how hard the decision is

Easy72 straightforward cases
Jev 100% Imajev-4B 100%
Standard96 everyday cases
Jev 99% Imajev-4B —
Judge146 evaluation-style calls
Jev 94.5% Imajev-4B 89%
Hard220 genuinely ambiguous cases
Jev 74.1% Imajev-4B —
Sealed308 private cases · chance is 29.3%
Jev 36.7% Imajev-4B 37%

Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.

03

What Imajev-4B is

Imajev-4B is an independent Apache-2.0 project by Mohit Garg: a Qwen3.5-4B LoRA that mirrors Jev’s /v1/systemone request contract and adds two things Jev does not have — image inputs (up to two photos per call) and a trained unknown_probability field so the model can abstain. Its author states no Jev outputs were used in training.

On JevBench v1.4.2.2 it entered at #1 with 67.4, ahead of Jev’s 63.3 on the composite — driven mainly by the best calibration in the top ten (80.4) and strong measured speed (90.6). Jev keeps the higher intelligence score (53.1 vs 52.2), a clear lead on judge-style cases (94.5% vs 89.0%), and the pair are level on the sealed set.

Where Imajev-4B falls short

  • Without its calibration file the model is over-confident on hard items.
  • An empty field can be read as “no” instead of unknown, and English is the only supported language.
  • Its own image benchmark, ImajevBench, uses AI-generated images that have not yet had a human audit.
  • State is limited to 32 KB (~8k tokens) — a sixth of Jev’s 64k-token request budget.
04

When to use which

Choose Jev if

  • Your inputs are long text: documents, traces or tickets past about 8k tokens.
  • Your hard cases turn on probabilities, reworded questions or multi-step lookups.
  • You want a hosted, versioned model with nothing to serve.
Try Jev free

Choose Imajev-4B if

  • Your decisions depend on a photo: a listing against its picture, a return against what was shipped, a part against a known-good one.
  • You want the model to say it cannot tell, and send those cases to a person.
  • You need open weights that run on a Mac or one GPU inside your own network.
Get Imajev-4B

FAQ

Is Imajev-4B better than Jev?

On JevBench’s composite, yes: #1 at 67.4 against Jev’s #4 at 63.3, with better calibration and speed. Jev keeps a slightly higher intelligence score, 53.1 against 52.2, and a clear lead on the judge tier, 94.5% against 89.0%. On the sealed set they are level.

Can Imajev read images?

Yes. It takes up to two images per request — for example a reference and a target — alongside the state and questions. JevBench, which ranks it first, is a text-only benchmark and does not test that.

Is Imajev made by TypeSafe?

No. It is an independent Apache-2.0 project by Mohit Garg that mirrors Jev’s request contract. Its author states that no Jev outputs were used in training.

Do Jev requests work with Imajev?

Its README says text-only requests written for Jev’s /v1/systemone work unchanged, and the response adds unknown_probability and abstained fields. Check your state size against its 32 KB limit first.

Sources