jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev Model · Evaluation & quality

LLM as a Judge

Give every AI answer a clear quality check.

Use Jev for the LLM-as-a-judge workflow: supply the question, reference and candidate answer, then ask narrow questions about its quality. Get typed verdicts and scores you can inspect, compare and use in your application.

One statefocused questions
  • User questionWhat the candidate answer should satisfy
  • Reference evidenceThe source the answer must stay grounded in
  • Candidate answerThe generated text under evaluation
Structured answersfor your application
  • Source-support verdictsupported · contradicted · insufficient_evidence
  • Relevance probabilitynoul: does it address the user question
  • Completeness scorescore on your own ordered rubric

Try it with your own rules

Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.

ChoiceUsing only reference, assess every factual claim in candidate_answer.
Yes / NoDoes candidate_answer directly address the user_question?
ScoreHow much of user_question does candidate_answer address?
Open in playground

What Jev returns

Illustrative output for the preset below — run it live to get real values for your input.

user_question

How long do I have to return an unused lamp, and who pays return shipping?

reference

Unused lamps can be returned within 30 days of delivery. The buyer pays return shipping unless the item arrived damaged.

candidate_answer

You can return an unused lamp within 30 days of delivery. Return shipping is free for every order.

From one example to a reusable workflow

01

Define the reference

Include the exact source passages and the user request. The judge can only check the evidence you provide; it does not browse for missing facts.

02

Separate the criteria

Ask source support, relevance and coverage as separate questions. Edit the answer labels and score levels to match your application.

03

Evaluate and reuse

Run one example in the playground, then apply the same question set to every row of your evaluation set with the API.

Keep the criteria separate

CheckWhat it measuresHow to use it
Groundedness · ChoiceDoes the supplied reference support each factual claim?Distinguish supported, contradicted and insufficient evidence — never collapse them into one score.
Relevance · Yes/NoDoes the answer address the actual question?Use a noul probability independently of the factual verdict; a correct answer to the wrong question still fails.
Completeness · ScoreDoes the answer cover every requested part?Use descriptive score levels; do not treat the score as factual accuracy.
01

Why a typed judge beats a chat verdict

A general LLM asked to "rate this answer" returns prose you still have to parse. Jev returns a verdict your code can act on directly — a choice label for the factual verdict, a probability for relevance, a score for coverage — each independently logged and thresholded.

02

One state, every criterion

Ask all rubric questions about the same state in a single request. They share the state input cost and run together, so a three-question judge costs and lands like one call.

03

Calibration over vibes

Because answers are typed, you can measure the judge itself: agreement rate with human labels, false-contradiction rate, and score distributions over a held-out set. A judge you cannot measure is just another prompt.

FAQ

Can Jev judge answers in languages other than English?

The judge checks the state you pass — question, reference and answer can be any language. Rubric instructions stay in English or the language you define them in.

How many questions can one judge run?

Up to 8 questions per request on jevmodel.org. Groundedness, relevance and coverage fit comfortably in one call.

Is Jev judging deterministic?

No — it returns probabilities, not guarantees. Treat the verdict like a sensor reading: threshold it, log it, and spot-check against human labels.

API keys

Call Jev with your key.

Send state plus typed questions to the local Jev endpoint. Successful requests charge input tokens. Add an Idempotency-Key header when retrying.

+ New key