Jev vs SemIf
MIT-licensed logit reader on Qwen3.5-4B — a fast open rebuild on your own GPU.
Jev 1.13.0
TypeSafe · System One model
63.3JevBench rank #4
Highest intelligence in the top ten: 53.1.
SemIf
TheoLeeCJ · MIT-licensed logit reader on Qwen3.5-4B — a fast open rebuild on your own GPU
47.7JevBench rank #13
Thirteenth of 91 — edges Jev on judge-style cases, falls short where it gets hard.
Key differences
- Where it runs JevHosted API — nothing for you to operateSemIfSelf-hosted on a GPU you own or rent
- Licence JevProprietary, hostedSemIfMIT code over open weights
- Evaluation-style cases Jev94.5% correctSemIf95.2% correct
- Hardest test cases Jev74.1% correctSemIf59.5% correct
- Probabilities JevCalibrated as delivered — 76.3 / 100SemIfNeeds fitting on your own workload — 66.8 / 100
- Where your data goes JevSent to TypeSafe’s APISemIfNever leaves your network
JevBench v1.4.2.2, measured the same way
One benchmark measured 95 systems under one method and ranked 91, so these figures are comparable in a way vendor-published numbers are not. The composite weighs four axes — intelligence, calibration, speed and cost — and hides where systems actually differ, so the charts below break it apart. Full method at the source ↗
Composite score
The top four finish within 4.1 points of each other, then the board falls away sharply.
Capability, higher is better
Accuracy by how hard the decision is
Easy and standard decisions separate almost nothing. The hard tier is where these systems stop agreeing, and the sealed tier shows how much of that holds on questions nobody could have tuned for. A dash means the benchmark published no combined figure. The composite also weighs a fourth axis — cost — which we do not reproduce; see the source table.
What SemIf is
SemIf is TheoLeeCJ’s MIT-licensed logit reader over open Qwen3.5-4B weights: it reads the model’s label logits directly rather than training a decision head. The appeal is deployment — a familiar LLM on a GPU you already own or rent, with your data never leaving your network.
On JevBench v1.4.2.2 it ranks #13 at 47.7. The surprise: it edges Jev on judge-style cases, 95.2% vs 94.5%. The shortfall shows up exactly where decisions get expensive — hard tier 59.5% vs 74.1%, sealed 26.3% vs 36.7% (chance on the sealed set is 29.3%), and calibration at 66.8 that needs fitting per workload.
Where SemIf falls short
- Sealed-set accuracy is 26.3% — below the 29.3% chance line on the sealed set.
- Hard-tier accuracy trails Jev by ~15 points: 59.5% vs 74.1%.
- Probabilities are unfitted — calibrate on your own workload before using thresholds.
When to use which
Choose Jev if
- Ambiguous or adversarial cases are where your cost lives — hard 74.1% vs 59.5%, sealed 36.7% vs 26.3%.
- You need calibrated probabilities without fitting anything.
- You want a hosted endpoint instead of GPU operations.
Choose SemIf if
- Data cannot leave your network — everything runs on your GPU.
- Your workload is judge-style cases, where it matches or edges Jev.
- MIT code over open weights you can read end to end.
FAQ
Is SemIf better than Jev?
Only on judge-style cases: 95.2% vs 94.5%. Everywhere else Jev is well ahead — composite #4 vs #13, hard tier 74.1% vs 59.5%, sealed 36.7% vs 26.3%, calibration 76.3 vs 66.8.
Is SemIf open source?
Yes — MIT-licensed code by TheoLeeCJ reading logits from open Qwen3.5-4B weights. Independent of TypeSafe.
Why is SemIf’s sealed score so low?
26.3% on 308 sealed cases — below the 29.3% chance line. A logit reader without a trained decision head has little to anchor on for cases nobody could tune for. Jev’s post-training shows up exactly there.
Who should pick SemIf?
Teams whose data cannot leave their network and whose workload is mostly judge-style or routine calls — with calibration fitted on their own labels before trusting any threshold.