Try it with your own rules
Start from the preset below, adapt the questions, and inspect the typed answers. Run it live in the playground — it is the same request shape your application will send.
What Jev returns
Illustrative output for the preset below — run it live to get real values for your input.
Nordwind GmbH, Köln — VAT DE812345678 — nordwind.de
Nordwind GmbH & Co. KG, Cologne — VAT DE812345678 — nordwind-cologne.de
Are record_a and record_b the same legal entity?
Is this pair safe to merge without human review?
P(yes) = 0.91 — shared VAT ID outweighs the name suffix difference
How costly would a wrong merge be for this pair?
From one example to a reusable workflow
Define the policy in the instructions
Say which fields are strong evidence (VAT, domain), which are weak (city spelling), and what counts as a conflict.
Split the decision from the action
Use the choice verdict for same/different/review, and the noul only for "safe to auto-merge" — never let confidence alone trigger an irreversible merge.
Route the review bucket
Pairs marked review go to a human queue with the full record pair attached — that is where recall lives.
Keep the criteria separate
| Check | What it measures | How to use it |
|---|---|---|
| Verdict · Choice | Are the two records the same entity? | Three labels: same, different, review — review is a first-class outcome, not an error. |
| Auto-merge · Yes/No | Is this pair safe to merge unattended? | Threshold the noul per field quality; 0.9+ might auto-merge, 0.6–0.9 reviews. |
| Blast radius · Score | How expensive is a wrong merge here? | High-value accounts get a stricter threshold than marketing contacts. |
Why rules alone fail
Exact-match joins miss suffixes and translations; fuzzy thresholds on edit distance merge things that share a name. A typed decision lets you state the policy in words — "shared tax ID is decisive" — and apply it consistently.
Keep the gray zone explicit
The difference between a good and a bad matcher is what happens between obvious same and obvious different. A review label routes those pairs to humans instead of silently merging or splitting.
Idempotent and auditable
Each decision returns the full probability distribution. Log it: when a merge is disputed you can show exactly which evidence weighed in and how confident the verdict was.
FAQ
How big can the records be?
Both records plus your match policy must serialize under 8,000 characters of state. Company and product records fit easily.
Does this replace my candidate-generation step?
No — use blocking or embeddings to find candidate pairs, then Jev judges each pair. It is a decision layer, not an index.
Can I tune for precision over recall?
Yes — that is what thresholds are for. Raise the auto-merge noul bar and let the review bucket absorb the rest.