jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev comparison · JevBench v1.4.2.2 · updated September 28, 2026

Jev vs SemIf

MIT-licensed logit reader on Qwen3.5-4B — a fast open rebuild on your own GPU.

On this site

Jev 1.13.0

TypeSafe · System One model

63.3JevBench rank #4

Highest intelligence in the top ten: 53.1.

Alternative

SemIf

TheoLeeCJ · MIT-licensed logit reader on Qwen3.5-4B — a fast open rebuild on your own GPU

47.7JevBench rank #13

Thirteenth of 91 — edges Jev on judge-style cases, falls short where it gets hard.

01

Key differences

  1. Where it runs
    JevHosted API — nothing for you to operate
    SemIfSelf-hosted on a GPU you own or rent
  2. Licence
    JevProprietary, hosted
    SemIfMIT code over open weights
  3. Evaluation-style cases
    Jev94.5% correct
    SemIf95.2% correct
  4. Hardest test cases
    Jev74.1% correct
    SemIf59.5% correct
  5. Probabilities
    JevCalibrated as delivered — 76.3 / 100
    SemIfNeeds fitting on your own workload — 66.8 / 100
  6. Where your data goes
    JevSent to TypeSafe’s API
    SemIfNever leaves your network
02

JevBench v1.4.2.2, measured the same way

One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗

Composite score

  1. #1 Imajev-4B 67.4
  2. #2 Plumb-4B 65.8
  3. #3 decider-4b v2 64.1
  4. #4 Jev 1.13.0 63.3
  5. #5 JevK5 v0.2.0 62
  6. #6 Cygnet 61.8
  7. #7 Hopper 59.4
  8. #8 Winnow-12B Q8 55.6
  9. #9 reflex 4B 54
  10. #10 djev 52.2
  11. #13 SemIf 47.7
  12. #29 OpenJev 36.9

The top four finish within 4.1 points of each other, then the board falls away sharply.

Capability, higher is better

IntelligenceHow often it picks the right answer
Jev53.1
SemIf44.4
CalibrationWhether 0.8 really means about 80%
Jev76.3
SemIf66.8
SpeedMeasured response time
Jev83.3
SemIf83.7

Accuracy by how hard the decision is

Easy72 straightforward cases
Jev 100% SemIf 100%
Standard96 everyday cases
Jev 99% SemIf 97.9%
Judge146 evaluation-style calls
Jev 94.5% SemIf 95.2%
Hard220 genuinely ambiguous cases
Jev 74.1% SemIf 59.5%
Sealed308 private cases · chance is 29.3%
Jev 36.7% SemIf 26.3%

Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.

03

What SemIf is

SemIf is TheoLeeCJ’s MIT-licensed logit reader over open Qwen3.5-4B weights: it reads the model’s label logits directly rather than training a decision head. The appeal is deployment — a familiar LLM on a GPU you already own or rent, with your data never leaving your network.

On JevBench v1.4.2.2 it ranks #13 at 47.7. The surprise: it edges Jev on judge-style cases, 95.2% vs 94.5%. The shortfall shows up exactly where decisions get expensive — hard tier 59.5% vs 74.1%, sealed 26.3% vs 36.7% (chance on the sealed set is 29.3%), and calibration at 66.8 that needs fitting per workload.

Where SemIf falls short

  • Sealed-set accuracy is 26.3% — below the 29.3% chance line on the sealed set.
  • Hard-tier accuracy trails Jev by ~15 points: 59.5% vs 74.1%.
  • Probabilities are unfitted — calibrate on your own workload before using thresholds.
04

When to use which

Choose Jev if

  • Ambiguous or adversarial cases are where your cost lives — hard 74.1% vs 59.5%, sealed 36.7% vs 26.3%.
  • You need calibrated probabilities without fitting anything.
  • You want a hosted endpoint instead of GPU operations.
Try Jev free

Choose SemIf if

  • Data cannot leave your network — everything runs on your GPU.
  • Your workload is judge-style cases, where it matches or edges Jev.
  • MIT code over open weights you can read end to end.
Get SemIf

FAQ

Is SemIf better than Jev?

Only on judge-style cases: 95.2% vs 94.5%. Everywhere else Jev is well ahead — composite #4 vs #13, hard tier 74.1% vs 59.5%, sealed 36.7% vs 26.3%, calibration 76.3 vs 66.8.

Is SemIf open source?

Yes — MIT-licensed code by TheoLeeCJ reading logits from open Qwen3.5-4B weights. Independent of TypeSafe.

Why is SemIf’s sealed score so low?

26.3% on 308 sealed cases — below the 29.3% chance line. A logit reader without a trained decision head has little to anchor on for cases nobody could tune for. Jev’s post-training shows up exactly there.

Who should pick SemIf?

Teams whose data cannot leave their network and whose workload is mostly judge-style or routine calls — with calibration fitted on their own labels before trusting any threshold.

Sources