Try it with your own rules
Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.
What Jev returns
Illustrative output for the preset below — run it live to get real values for your input.
How long do I have to return an unused lamp, and who pays return shipping?
Unused lamps can be returned within 30 days of delivery. The buyer pays return shipping unless the item arrived damaged.
You can return an unused lamp within 30 days of delivery. Return shipping is free for every order.
Using only reference, assess every factual claim in candidate_answer.
Does candidate_answer directly address the user_question?
P(yes) = 0.97 — relevance is judged separately from correctness
How much of user_question does candidate_answer address?
From one example to a reusable workflow
Define the reference
Include the exact source passages and the user request. The judge can only check the evidence you provide; it does not browse for missing facts.
Separate the criteria
Ask source support, relevance and coverage as separate questions. Edit the answer labels and score levels to match your application.
Evaluate and reuse
Run one example in the playground, then apply the same question set to every row of your evaluation set with the API.
Keep the criteria separate
| Check | What it measures | How to use it |
|---|---|---|
| Groundedness · Choice | Does the supplied reference support each factual claim? | Distinguish supported, contradicted and insufficient evidence — never collapse them into one score. |
| Relevance · Yes/No | Does the answer address the actual question? | Use a noul probability independently of the factual verdict; a correct answer to the wrong question still fails. |
| Completeness · Score | Does the answer cover every requested part? | Use descriptive score levels; do not treat the score as factual accuracy. |
Why a typed judge beats a chat verdict
A general LLM asked to "rate this answer" returns prose you still have to parse. Jev returns a verdict your code can act on directly — a choice label for the factual verdict, a probability for relevance, a score for coverage — each independently logged and thresholded.
One state, every criterion
Ask all rubric questions about the same state in a single request. They share the state input cost and run together, so a three-question judge costs and lands like one call.
Calibration over vibes
Because answers are typed, you can measure the judge itself: agreement rate with human labels, false-contradiction rate, and score distributions over a held-out set. A judge you cannot measure is just another prompt.
FAQ
Can Jev judge answers in languages other than English?
The judge checks the state you pass — question, reference and answer can be any language. Rubric instructions stay in English or the language you define them in.
How many questions can one judge run?
Up to 8 questions per request on jevmodel.org. Groundedness, relevance and coverage fit comfortably in one call.
Is Jev judging deterministic?
No — it returns probabilities, not guarantees. Treat the verdict like a sensor reading: threshold it, log it, and spot-check against human labels.