jevmodel.org account

Sign in for 100,000 free input tokens

About 500 Jev requests in the playground or from your own code with an API key. Packs start at $9.90 when you need more. No card required.

By signing in you agree to our terms. We only use your email for your account.
Jev Model · Evaluation & quality

AI Agent Evaluation

Check agent claims against what the tools actually did.

Agents claim "done" whether or not the tool calls actually accomplished the task. Give Jev the task, the tool trace and the claimed result, and get a typed verdict on completion, compliance and quality.

One statefocused questions
  • TaskWhat the agent was assigned to do
  • Tool traceThe actions it executed and their results
  • Claimed resultThe outcome the agent reported
Structured answersfor your application
  • Completion verdictsuccess · partial · failed
  • In-policy probabilitynoul: stayed within allowed actions
  • Quality scorescore on your execution rubric

Try it with your own rules

Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.

ChoiceDoes the tool trace prove the claimed result was achieved?
Yes / NoDid every tool call stay within the allowed actions?
ScoreHow well was the task executed end to end?
Open in playground

What Jev returns

Illustrative output for the preset below — run it live to get real values for your input.

task

Find the 3 cheapest direct flights SFO→BOS next Friday and email the list to ops.

tool_trace

search_flights returned 5 options → sent email with 3 lowest to ops@…

claimed_result

Emailed the 3 cheapest direct flights to ops.

From one example to a reusable workflow

01

Capture the trace

Log the tool calls with their results — the judge needs evidence, not the agent's narration.

02

Ask about the claim, not the transcript

Question the gap between claim and evidence: did the trace actually produce the reported outcome?

03

Score for regression

Run the same judge over every run in CI or production sampling; watch the score distribution move as you change the agent.

Keep the criteria separate

CheckWhat it measuresHow to use it
Completion · ChoiceDoes the tool evidence support the claimed outcome?Split success / partial / failed — partial runs are where agent evals earn their keep.
Compliance · Yes/NoDid the agent stay inside its allowed actions?A noul gate catches out-of-scope tool use before it becomes an incident.
Execution quality · ScoreHow well was the task done, not just whether?Track the score distribution over time — it drifts before completions break.
01

Claims are not outcomes

An agent's final message is self-reported. The trace is what happened. Jev judges the evidence: whether the search results support the summary, whether the file was actually written, whether the email went to the right list.

02

Typed verdicts scale to regression suites

Run the same question set against every recorded trace. Because answers are choice labels, probabilities and scores — not prose — a CI gate can assert on them directly.

03

Pair evaluation with guardrails

Evaluation checks what an agent did; a tool-call guardrail checks what it is about to do. Use both: gate before execution, evaluate after, and feed disagreements back into the prompt or policy.

FAQ

Does Jev replay the tool calls?

No — Jev judges the state you pass: task, recorded trace and claimed result. Your harness owns execution and logging.

How is this different from a tool-call guardrail?

A guardrail runs before an action executes; evaluation runs after the run completes. They share the same typed-question pattern.

Can I evaluate multi-step traces?

Yes — serialize the trace into the state (up to 8,000 characters). Ask one question about overall completion and narrower ones about specific steps.

API keys

Call Jev with your key.

Send state plus typed questions to the local Jev endpoint. Successful requests charge input tokens. Add an Idempotency-Key header when retrying.

+ New key