Try it with your own rules
Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.
What Jev returns
Illustrative output for the preset below — run it live to get real values for your input.
The EU AI Act applies to open-source models
eur-lex.europa.eu: "…does not apply to AI systems released under free and open-source licences…" (2024-07-12) · openai.com/blog: "GPAI provisions do cover open weights above thresholds" (2025-08-01)
Do the supplied sources support the claim as written?
Is the claim true as literally stated?
P(yes) = 0.38 — applies only above capability thresholds, not categorically
How strong and current is the supplied evidence?
From one example to a reusable workflow
Fetch first, decide second
Your code owns retrieval — pull snippets from your search provider and serialize them with dates and sources into the state.
Separate evidence from answer
Ask whether the sources support the claim (choice) and whether the claim is true (noul) as different questions — strong evidence for a false claim is a real outcome.
Score the source quality
An evidence-quality score tells you whether a surprising verdict is grounded in a primary source or a stray blog post.
Keep the criteria separate
| Check | What it measures | How to use it |
|---|---|---|
| Verdict · Choice | Do the supplied sources support the claim? | supported / contradicted / partial keeps "the evidence is mixed" distinct from "no evidence". |
| Answer · Yes/No | Is the claim true given this evidence? | The noul probability is your fact-check signal — threshold it per how costly a wrong yes is. |
| Evidence quality · Score | Are the sources primary, recent and independent? | Down-weight verdicts built on weak evidence instead of treating all snippets equally. |
Grounding is a decision problem
The hard part of live context is not fetching — it is deciding whether what you fetched actually answers the question. That judgment is exactly what a typed decision returns.
Dates are part of the evidence
Include source dates in the state and say in the instructions how to weigh them. A 2024 regulation page and a 2026 changelog disagree because the world changed — let the instructions say which wins.
Compose with RAG evaluation
Web context checks claims against evidence; RAG evaluation gates which evidence enters the context at all. Use both when your pipeline retrieves from the open web.
FAQ
Does Jev browse the web itself?
No — on jevmodel.org you supply the snippets. Fetch with your search provider, pass results as state, and Jev weighs them.
How much evidence fits?
State is limited to 8,000 characters serialized — roughly 10–20 focused snippets with sources and dates.
Is the verdict reproducible?
The same state gives similar probabilities, but evidence changes. Log inputs with timestamps so a verdict stays explainable later.