Try it with your own rules
Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.
What Jev returns
Illustrative output for the preset below — run it live to get real values for your input.
Find the 3 cheapest direct flights SFO→BOS next Friday and email the list to ops.
search_flights returned 5 options → sent email with 3 lowest to ops@…
Emailed the 3 cheapest direct flights to ops.
Does the tool trace prove the claimed result was achieved?
Did every tool call stay within the allowed actions?
P(yes) = 0.99 — only approved search and email tools were used
How well was the task executed end to end?
From one example to a reusable workflow
Capture the trace
Log the tool calls with their results — the judge needs evidence, not the agent's narration.
Ask about the claim, not the transcript
Question the gap between claim and evidence: did the trace actually produce the reported outcome?
Score for regression
Run the same judge over every run in CI or production sampling; watch the score distribution move as you change the agent.
Keep the criteria separate
| Check | What it measures | How to use it |
|---|---|---|
| Completion · Choice | Does the tool evidence support the claimed outcome? | Split success / partial / failed — partial runs are where agent evals earn their keep. |
| Compliance · Yes/No | Did the agent stay inside its allowed actions? | A noul gate catches out-of-scope tool use before it becomes an incident. |
| Execution quality · Score | How well was the task done, not just whether? | Track the score distribution over time — it drifts before completions break. |
Claims are not outcomes
An agent's final message is self-reported. The trace is what happened. Jev judges the evidence: whether the search results support the summary, whether the file was actually written, whether the email went to the right list.
Typed verdicts scale to regression suites
Run the same question set against every recorded trace. Because answers are choice labels, probabilities and scores — not prose — a CI gate can assert on them directly.
Pair evaluation with guardrails
Evaluation checks what an agent did; a tool-call guardrail checks what it is about to do. Use both: gate before execution, evaluate after, and feed disagreements back into the prompt or policy.
FAQ
Does Jev replay the tool calls?
No — Jev judges the state you pass: task, recorded trace and claimed result. Your harness owns execution and logging.
How is this different from a tool-call guardrail?
A guardrail runs before an action executes; evaluation runs after the run completes. They share the same typed-question pattern.
Can I evaluate multi-step traces?
Yes — serialize the trace into the state (up to 8,000 characters). Ask one question about overall completion and narrower ones about specific steps.