🔍 Read the full analysis: 24 Examples Of Jev At Work In AI Decision-Making on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Thorsten Meyer’s Sept. 29 article maps 24 potential uses for Jev, a tool that returns typed answers to narrow questions so software can route decisions. Meyer says three applications are live, 12 meet his fit test, seven need measurement and two are poor fits; those figures and performance results come from his own operation and have not been independently verified.
Thorsten Meyer published a 24-use-case assessment of Jev on September 29, reporting that three applications are already running in his publishing operation and identifying 12 further strong fits under his four-condition test. The article matters to teams considering automated classification and checks because it sets out when the tool may be useful, as well as cases where Meyer says evidence is still missing or the fit is poor.
Meyer describes Jev as a system that takes text or JSON state plus typed questions and returns answers software can use for branching. The response types include a yes-or-no probability, a choice with probabilities and confidence, or a score on ordered levels. Jev does not write or summarize, according to the article; the surrounding code decides what to do with each answer. A single call carries the state and questions, takes about 0.3 to 0.9 seconds, and is priced by Meyer at about $0.04 per million input tokens.
Three applications are described as live: a relevance gate for stories and sites, a language check, and a fallback classifier for a 31-topic taxonomy. Meyer says the language check scanned 78,889 articles for $2.01, found 1,576 non-English articles, and fixed 1,553. He reports 89% agreement with a frontier large language model for the fallback classifier, rising to 97% to 99% when Jev’s confidence was at least 0.8. These are results reported by Meyer; the source provides no independent evaluation methodology or external validation.
The broader list covers publishing, commerce, software, business operations and home uses. In publishing, examples include disclosure checks and comment moderation, while other ideas such as detecting thin sources or judging headline quality are marked for measurement first. Meyer categorizes 12 use cases as strong fits, seven as needing measurement and two as poor fits, in addition to the three already live. The article excerpt gives detailed examples from publishing but does not include the full list of 24 cases.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Automated Checks May Help
The proposal is aimed at repeated, narrow decisions where leaving every item to a larger model or a person may be costly, but where an unchecked error could still matter. Meyer’s design uses confidence to act on clear cases and route uncertain ones. For example, a high-confidence language result could trigger a rewrite, while a low-confidence disclosure check could be sent to a reviewer.
That boundary matters because a fast or inexpensive answer is useful only if it addresses a real failure. Meyer calls out the duplicate-story detector as a poor fit after his canary found zero duplicates. His point is that low-cost automation does not justify itself by existing; operators first need evidence that the current rule or workflow is failing often enough to warrant a change.
For readers assessing Jev, the reported metrics are an initial account from one operator, not a general guarantee of accuracy, savings or suitability. Results will depend on the questions, data, confidence thresholds and cost of mistakes in each deployment. The article’s recommendations are therefore best understood as a framework for testing candidate applications, rather than proof that all 24 uses will work elsewhere.
Meyer’s Four Conditions for Jev
Meyer says a use case should meet four conditions before Jev is wired into a workflow: high volume, a narrow question that needs no multistep reasoning, cheap errors or a route for uncertain cases, and a heuristic that has been shown to fail. He advises replaying 300 to 500 past decisions, comparing results overall and by confidence band, and reviewing 20 disagreements to determine which answer was right.
His suggested rollout depends on that evidence: wire Jev in only where the high-confidence band reaches 95%, use a separate feature flag that is off by default, test on 5% to 10% of units, and then expand. The article says the live relevance gate judged about 10,000 story-site pairings in three days, with 22% clearly on-topic. Meyer reports that the live language check fixed most of the non-English articles it found, while leaving 23 unresolved in the figures provided.
For the classifier, Meyer reports 97% to 99% agreement with a frontier LLM at confidence of 0.8 or higher, compared with 42% below 0.5 in a 31-topic measurement. The source does not identify the evaluation set, sample size, or how agreement was calculated. Those omissions limit what can be concluded from the figures beyond Meyer’s own reported test.
“Jev is the right tool wherever a system needs thousands of small judgements and can hand the unclear ones to something smarter.”
— Thorsten Meyer, article author
Evidence Still Needed for Wider Use
The source is a first-person account from Meyer, who says the live uses run in his own publishing operation. It does not provide independent validation, full datasets, detailed evaluation methods or results from other organizations. The reported agreement rates and cost figures should not be assumed to transfer to other tasks or deployments.
The article excerpt also ends partway through its commerce and customer operations section. Although it says the full assessment contains 24 examples, it does not show all 24 in the supplied material. The basis for classifying each omitted case, and the particular two cases labeled poor fits beyond the detailed duplicate example, cannot be fully assessed from this excerpt.
Several proposed uses remain explicitly unproven. Meyer says a thin-source detector needs testing against a character-count rule, a product-match scorer needs a measured error rate for the existing matcher, and a headline-quality score should be a pre-publication prompt rather than the sole gate. Whether these checks reduce mistakes without introducing new ones remains unclear.
Testing Before Wider Rollout
Meyer recommends shadow testing candidate workflows against 300 to 500 historical decisions before enabling them. Teams would compare Jev’s results by confidence level and manually review disagreements, then limit deployment to cases where high-confidence accuracy reaches his stated 95% threshold. His proposed feature flag and 5% to 10% canary offer a way to observe performance on a small share before a wider rollout.
The next evidence needed is case-specific: whether the existing heuristic fails, how often Jev is right at each confidence level, and what happens when it is wrong. The article does not announce an external launch, independent audit or follow-up evaluation date. Any broader claims about Jev’s effectiveness will depend on further results from those tests and on whether other operators reproduce Meyer’s findings.
Key Questions
What is Jev, according to Meyer?
Meyer describes Jev as a tool that receives text or JSON state and typed questions, then returns structured answers such as probabilities, category choices or scores. Software uses those answers to make workflow decisions.
How many of the 24 uses are already running?
Meyer says three are live in his publishing operation. He classifies 12 as strong fits, seven as needing measurement and two as poor fits.
Are the reported accuracy figures independently verified?
No independent verification is provided in the source material. The agreement rates are results Meyer reports from his own measurement, and the article excerpt does not give the full methodology or evaluation data.
What does Meyer recommend before deploying Jev?
He recommends replaying 300 to 500 past decisions, reviewing disagreements, checking accuracy by confidence band, and limiting rollout with a feature flag and a small canary. He says high-confidence performance should reach 95% before wiring a use case into the workflow.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
