🔍 Read the full analysis: Can Jev Fit Your AI Workflow? 24 Decision-Model Uses on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
Thorsten Meyer has published a map of 24 possible uses for Jev, a tool that returns typed decision answers for software workflows. He says three uses are already running in his publishing operation, while other cases range from strong fits to ideas that need measurement or fail his test.
Thorsten Meyer published a guide on September 29 mapping 24 potential uses for Jev, a tool that supplies structured answers to questions posed by software. Meyer says three uses are live in his publishing operation and have processed about 90,000 decisions; the guide also classifies 12 ideas as strong fits, seven as needing measurement and two as poor fits.
Meyer describes Jev as a way to send a state, such as text or JSON, alongside typed questions and receive answers that code can act on. The tool does not write or summarize content, according to his account. Its answer types include a yes-or-no probability, a choice among options with probabilities and confidence, and a score on ordered levels. He reports that a call takes about 0.3 to 0.9 seconds and costs about $0.04 per million input tokens.
The three live examples are a relevance gate for matching stories to a site, a language check and a fallback topic classifier. For the language check, Meyer says a scan of 78,889 articles cost $2.01, found 1,576 non-English items and fixed 1,553. He reports 89% agreement with a frontier large language model for the classifier, rising to 97% to 99% when Jev’s confidence was at least 0.8. These are figures from Meyer’s own operation and measurements.
His proposed fit test has four parts: high volume, a narrow question, low-cost errors or a route for uncertain cases, and evidence that an existing heuristic fails. Meyer recommends replaying 300 to 500 past decisions, comparing results across confidence bands and reviewing 20 disagreements. He says a workflow should be wired in only where the high-confidence band reaches 95%, with a separate flag initially off and a small canary rollout.
24 use cases for Jev at a glance
Every use case, coloured by how well it fits
Proven in production
1Relevance gate: story and site2Language check3Classifier fallbackPublishing and content
4Thin-source detector5Same-event dedupe6Product fits the roundup7Disclosure present8Headline quality9Comment moderationCommerce and support
10Support-ticket routing11Return-reason coding12Review to feature complaints13Catalogue taxonomy14Order-fraud pre-triageSoftware and AI systems
15LLM guardrail16RAG passage filter17Citation check18Tool and intent routing19Log-line triage20PR risk triageBusiness ops and home
21Inbox triage22Expense categorisation23Lead qualification24Smart-home intent15 of 24 are ready to build or already running
Where Cheap Checks Can Help
The guide’s practical point is that low-cost decisions can be applied at scale when software can act on clear answers and pass uncertain cases to a person or a more capable system. That may let publishers check more items than they could review manually, while retaining an existing path for ambiguous cases. Meyer’s examples focus on filtering, classification and routing rather than asking Jev to create finished content.
The proposed test also sets a limit on adoption: a task should not be automated simply because it can be phrased as a short question. A tool needs a measured problem to solve, and mistakes must be affordable or caught by escalation. Meyer marks same-event deduplication as a poor fit after his canary found no duplicates. That result illustrates why he treats a visible failure in the current method as a condition for use.
From Publishing Checks to Wider Uses
Meyer’s map spans publishing, commerce, software, business operations and home workflows, though the supplied account gives detail only on the first six publishing cases and introduces commerce before ending. It is therefore not possible from this material to describe all 24 examples. Within publishing, the cases include detecting thin source material, checking whether a product belongs in a roundup, spotting required disclosures, assessing headlines and moderating comments.
The fit labels distinguish uses that Meyer says are running from those he considers ready to build, those that need an initial measurement and those he considers poor fits. In publishing, he calls disclosure checks and comment moderation strong fits. He says the disclosure check could catch paraphrased statements about free products or affiliate links that a regular expression might miss; he recommends human review for misses, rather than automatic publication. For comment moderation, his suggested rule auto-approves clearly acceptable comments or hides clearly identified spam at confidence of 0.9 or higher, while queuing other cases.
Other publishing proposals remain conditional. A thin-source detector would assess whether a source contains enough verifiable facts to support a report without invention. Meyer says 88% of the news items he processes start from a bare headline, but labels the use case “measure first.” A product-fit score and headline-quality check also need evidence of current error rates. The headline check is described as a pre-publication nudge, not a sole gate.
“Jev does not write, summarise or extract.”
— Thorsten Meyer, describing Jev’s role
Which Workflows Still Need Proof
The figures are Meyer’s own reported results; the supplied material does not provide independent testing, detailed evaluation methods or a comparison with alternatives. It also does not include the remaining commerce, software, operations and home examples, despite describing the guide as a 24-use map. Readers cannot assess those proposals from the available account.
Several cases lack a measured baseline. Meyer says the thin-source detector, product-fit matcher and headline-quality score need further measurement; his deduplication canary found no cases to fix. The account does not state how the reported 90,000 decisions were divided among the three live uses, or provide error rates for each one. It also does not specify the model or service behind every comparison with a frontier LLM.
Measure Before Wider Rollout
Meyer’s next step for unproven cases is to replay 300 to 500 real past decisions, compare outcomes by confidence band and inspect disagreements. Only workflows meeting his stated 95% threshold in the high-confidence band would move toward integration. He recommends a separate feature flag, a 5% to 10% canary and then a wider rollout if the results support it.
The guide does not give a date for further results or identify a broader release schedule. For now, the reported live uses are confined to Meyer’s publishing operation, while the other proposed applications remain at different stages of evaluation.
Key Questions
What does Jev do?
It returns structured answers to typed questions about text or JSON, which software can use to make decisions. Meyer says it does not write, summarize or extract content.
How many Jev uses does Meyer say are live?
Meyer reports three live uses in his publishing operation: story relevance, language checking and fallback topic classification. He says they have processed about 90,000 decisions in total.
What makes a workflow a strong fit?
Meyer’s test calls for high volume, a narrow question, low-cost errors or an escalation path, and evidence that the current heuristic fails visibly.
Are all 24 proposed uses proven?
No. Meyer labels 12 strong fits, seven measure-first cases and two poor fits, alongside the three uses he says are already live. The available source does not detail every case or provide independent validation.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
