🔍 Read the full analysis: 24 Practical Paths To AI Decision Modeling With Jev on ThorstenMeyerAI.com
Get smart everyday buys delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer has published an analysis mapping 24 practical uses for Jev, a tool that returns calibrated answers to typed questions for AI decision modeling. Three uses are live in his publishing operation covering roughly 90,000 decisions, 12 are strong fits, 7 need measurement first, and 2 are rejected as poor fits.
Publisher Thorsten Meyer has published an analysis mapping 24 concrete use cases for Jev, a decision-modeling tool that returns calibrated, machine-readable answers to typed questions rather than prose. According to his September 29, 2026 write-up, three uses are already live in his own publishing operation — about 90,000 decisions processed so far — while 12 more meet his criteria as strong fits, 7 require measurement first, and 2 are rejected outright. The analysis includes a four-condition test for evaluating where narrow AI decision models apply, along with cost and accuracy figures from live production use.
Meyer describes Jev’s mechanics simply: a caller sends a state — text or JSON — plus a set of typed questions, and receives calibrated answers the code can branch on, with no prose to parse. A single call carrying the state and all questions takes roughly 0.3 to 0.9 seconds and costs about $0.04 per million input tokens, according to his figures. The tool supports three answer types: noul (a probability of yes from 0 to 1, used for gates, flags and filters), choice (one selected option with per-option probabilities and confidence, for routing and classification), and score (a position on ordered levels plus confidence, for quality or severity judgments).
According to Meyer, the central feature is confidence calibration. In his own measurement on a 31-topic classification task, Jev agreed with a frontier LLM 97 to 99% of the time when its confidence was 0.8 or higher — but only 42% of the time below 0.5. He reports that this asymmetry informs the design pattern behind the use cases: acting on the clear cases and routing the gray zone to a person or a stronger model. Each use case in the article pairs a question with an explicit rule, such as auto-approving only at confidence 0.9 or higher.
The three live uses demonstrate the economics. A language check that scanned 78,889 articles in one night for $2.01 flagged 1,576 non-English pieces and fixed 1,553 of them. A relevance gate judged about 10,000 story-site pairings over three days, finding only 22% clearly on-topic. A classifier fallback, used when the primary LLM errors, agreed with a frontier model 89% overall and 97 to 99% at high confidence. Meyer reports that 88% of the news items he processes start from a bare headline — the finding behind a proposed “thin-source detector” that would flag whether a source contains enough verifiable facts to write from without invention.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Reported Cost Structure and Rejections
The report documents the cost structure of automated judgment calls in Meyer’s operation. He reports that when a check costs a fraction of a cent, running it over an entire content archive rather than a sample becomes feasible, citing the full-site language audit that ran overnight for two dollars. The mapped use cases identify where, per Meyer, that cost structure applies: high volume, narrow questions, and cheap errors.
The report also documents rejections. Meyer marks a same-event deduplication use case as a poor fit — not because the question was hard, but because his canary test found zero duplicates to fix, meaning there was no measured problem for the tool to solve. His framing: “If your canary finds nothing to fix, there is no problem for Jev to solve yet.” That measurement-before-wiring approach is reflected throughout the analysis: 7 of the 24 candidates are tagged as unproven pending measurement, and disclosure-detection misses are routed to human review rather than auto-published because, per the report, a miss carries compliance risk.
The Four-Condition Fit Test Behind the Map
Meyer’s screening method asks four conditions before any wiring: high volume (thousands of small calls, not a handful of big ones), a narrow question with no multi-step reasoning, cheap errors (a wrong answer costs little, or unsure cases escalate), and a visibly failing heuristic — measured, not assumed. If an existing keyword rule works, he says, keep it.
Deployment follows a fixed sequence: replay 300 to 500 real past decisions, compare results overall and per confidence band, read 20 disagreements to decide who was right, and wire in only where the high-confidence band reaches 95%. Each integration gets its own feature flag, off by default, canaried on 5 to 10 units before rollout. Every use case carries one of four tags: live, strong fit, measure first, or poor fit. The live relevance gate illustrates the safe-switching design: it combines three answers into one decision and only acts when Jev is confident, so the uncertain middle keeps its existing path.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, publisher
What Meyer Has Not Yet Proven
Seven of the 24 use cases remain unproven because the fourth fit condition — a visibly failing heuristic — has not been measured, including the thin-source detector, product-fit checks for buying guides, and headline-quality scoring. The accuracy figures come from Meyer’s own single measurement on one 31-topic classification task, not an independent benchmark, and the 97 to 99% agreement figure compares Jev to a frontier LLM rather than to ground truth. Cost figures reflect his workload and token mix. The published source excerpt cuts off partway through the commerce section, so details of the commerce and home use cases, and the full reasoning behind both poor-fit rejections, are not fully visible in the available material.
Measurements Needed Before Rollout
According to Meyer’s own framework, the next step for the seven “measure first” candidates is a shadow replay of 300 to 500 past decisions to establish whether current heuristics actually fail. The thin-source detector requires validating that a confidence threshold can reliably separate dense wire items from thin teasers; the product-fit check needs a measured error rate for the existing matcher; and headline quality would launch as a pre-publish nudge, never a sole gate. The same-event dedupe case is set aside unless a duplicate problem is measured in the future. The four-condition test and shadow-replay protocol are documented in the report and can be applied to other workflows independently of Jev, according to Meyer.
Key Questions
What is Jev, in practical terms?
According to Meyer, Jev is a decision-modeling service: you send a state (text or JSON) and typed questions, and it returns calibrated answers — probabilities, choices with confidence, or scores — that code can branch on directly. It does not write, summarise or extract text.
How accurate and how expensive is it?
In Meyer’s measurement, Jev agreed with a frontier LLM 97 to 99% of the time at confidence 0.8 or higher, and 42% below 0.5. He reports costs of about $0.04 per million input tokens, with calls taking 0.3 to 0.9 seconds; one overnight scan of 78,889 articles cost $2.01.
When should you not use this kind of tool?
Meyer’s test rejects use cases that fail any of four conditions: low volume, questions needing multi-step reasoning, expensive errors, or no measured evidence that a simpler rule already fails. He rejected same-event deduplication because a canary test found zero duplicates to fix.
How were the live use cases rolled out safely?
Each was validated by replaying real past decisions, comparing per confidence band, and wiring in only where the high-confidence band reached 95%. Live integrations run behind feature flags that default to off, canaried on 5 to 10 units before rollout, and uncertain answers keep their old processing path.
Are the 24 use cases all production-ready?
No. Per Meyer’s breakdown, 3 are live, 12 are strong fits meeting all four conditions, 7 need a measurement first because the failing-heuristic condition is unproven, and 2 are poor fits with the reasons stated.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
