🔍 Read the full analysis: My September 2026 AI Routine: Opus Builds, Sol Digs, Jev Decides on ThorstenMeyerAI.com
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
TL;DR
On 29 September 2026, the day GPT-6.1 Sol launched, Thorsten Meyer published his updated AI working routine: Claude Opus 5.5 as main builder, GPT-6.1 Sol for detail work and review, and Jev, a decision-only model, for high-volume routing. The rationale is economic: six top models now score within about 20 index points while task costs differ by roughly 100x.
Thorsten Meyer published his September 2026 AI routine on 29 September, arguing that the frontier model market has shifted from a capability race to a price curve: six leading models now sit within roughly 20 index points of each other on the Artificial Analysis Intelligence Index while their cost per task differs by about 100x. His working answer pairs Claude Opus 5.5 as the main builder with GPT-6.1 Sol, released the same day, as a low-cost second reviewer — plus Jev, a decision model that cannot write prose, for high-volume yes/no and routing judgements.
The routine assigns each model a role based on score versus cost per task. Opus 5.5, released 22 September, tops the index at 58 on its max setting at $5.98 per task, and serves as the main model for features, APIs and multi-file work. GPT-6.1 Sol, launched 29 September at $2/$10 per million tokens, scores 51 at xhigh for $0.39 per task, making a review pass cheap enough to run routinely. GPT-6 Luna ($0.07 per task, 1,429 tasks per $100) handles classification and extraction.
Meyer reports three findings from the index data. First, Opus 5.5 outscores its more expensive sibling Claude Fable 5.1 by 5 points while costing less per task. Second, Sonnet 5.5 at max effort costs more per task than Opus at max while scoring 2 points lower. Third, Sol xhigh costs about one-eighth of GPT-6 Astra and roughly one-twentieth of Fable per task for a score only 1 to 2 points lower.
The effort setting, Meyer argues, is the bigger cost lever than model choice: on Opus 5.5, moving from xhigh to max adds 2 index points for 73% more cost per task, and medium-to-max multiplies cost 4.46x for 7 points. He runs Opus at high (54 points, $1.82) or xhigh (56 points, $3.46) for hard problems, reserving max. Sol’s trade-offs are noted: 57 to 69 seconds to first token at high and xhigh settings, making it unsuitable for interactive use at those levels, with low and max settings not yet published by Artificial Analysis.
Opus builds. Sol reviews. Jev decides.
One price tape, six models
Score against cost, at every effort setting
The effort dial moves the bill more than the model
Claude Opus 5.5
Claude Sonnet 5.5
GPT-6.1 Sol: near-Astra scores at a fraction of the price
Three published settings
| Setting | Index | Cost per task | Output tokens | First token |
|---|---|---|---|---|
| medium | 48 | $0.21 | 15M | 5.3 s |
| high | 50 | $0.32 | 25M | 57 s |
| xhigh | 51 | $0.39 | 36M | 69 s |
Same score band, very different bill
My stack: who builds, who reviews
Cheaper tokens are not cheaper work
Read the numbers with four warnings
Part 2: Jev, the model that decides instead of writing
One call in, typed answers out
Three question types
Confidence is the superpower
Three uses running in my publishing operation
The fit test, then the shadow test
- Replay 300 to 500 past decisions
- Compare overall and per confidence band
- Read 20 disagreements, decide who was right
- High band at 95% or better?
- Own flag, off by default
- Canary on 5 to 10 units
- Roll out in the confident band only
24 use cases, sorted by how well they fit
Proven in production
- 1Relevance gate
- 2Language check
- 3Classifier fallback
Publishing and content
- 4Thin-source detector
- 5Same-event dedupe
- 6Product fits roundup
- 7Disclosure present
- 8Headline quality
- 9Comment moderation
Commerce and support
- 10Support-ticket routing
- 11Return-reason coding
- 12Review to feature complaints
- 13Catalogue taxonomy
- 14Order-fraud pre-triage
Software and AI systems
- 15LLM guardrail
- 16RAG passage filter
- 17Citation check
- 18Tool and intent routing
- 19Log-line triage
- 20PR risk triage
Business ops and home
- 21Inbox triage
- 22Expense categorisation
- 23Lead qualification
- 24Smart-home intent
Limits, cost and one hard rule
Why Model Routing Now Beats Model Loyalty
The piece documents a broader shift in how practitioners choose AI models: when capability gaps between leaders shrink to single index points while price per task differs by two orders of magnitude, the practical question becomes which model clears a quality bar at the lowest cost, not which model is smartest. Meyer’s review-seat logic is the core of the argument — a different model family reviewing Opus’s output is a stronger check than Opus reviewing itself, and at $0.39 per task, an independent review pass on every meaningful change becomes affordable.
Meyer also cautions against over-reading cheap tokens. Halving model price saves only 12.5% of real cost by his illustrative example, and a single extra minute of human review erases the saving — a point he flags as illustrative rather than measured. His four operating rules include: effort is not capability, a different model reading the same flawed spec is not an independent review, and passing tests are not approval to ship.
The September 2026 Release Sprint
The routine covers a crowded month: Claude Fable 5.1 (1 September), GPT-6 Astra (3 September), Opus 5.5 and GPT-6 Luna (22 September), Claude Sonnet 5.5 (28 September) and GPT-6.1 Sol (29 September). Sol launched at the same token pricing as its week-old predecessor, with its medium setting already matching the earlier GPT-6 Sol’s score of 48 at one-fifth the cost per task ($0.21 versus $1.06). Sol is also unusually concise: its high setting used 25M output tokens on the index against a median of 82M for comparable models, and Sonnet 5.5 at max wrote about 193k output tokens per task — the most Artificial Analysis has measured. All scores derive from Artificial Analysis Intelligence Index v4.3.x.
“In four weeks, the AI frontier stopped being a leaderboard and became a price curve.”
— Thorsten Meyer
What the Index Does Not Yet Show
Several gaps remain. Artificial Analysis has not yet published low or max settings for GPT-6.1 Sol, so its full cost-performance range is unknown. Meyer notes that one index point is inside measurement noise, meaning the 1-2 point gaps separating Sol from Astra and Fable should not be treated as decisive. The index itself measures general capability, not any specific workload. The cost-of-work example — that halving model price saves 12.5% of real cost — is explicitly illustrative, not measured. Little detail is given about Jev beyond its role in high-volume decisions, and readers’ own task mixes may favor different routing choices.
Watching Sol’s Missing Settings
The immediate data gap is Artificial Analysis’s pending publication of Sol’s low and max effort settings, which will complete its cost-performance picture. Meyer’s stated practice is to keep shadow-testing models against his own workload before switching defaults, and to revise the routine as new releases land — a cadence that September’s near-weekly launches suggests will continue. Readers considering a similar stack are advised to validate the routing logic against their own task mix rather than adopting the index rankings directly.
Key Questions
What is the core claim of this September 2026 routine?
That with six top models within about 20 index points but roughly 100x apart in cost per task, the practical choice is routing work to the cheapest model that meets a quality bar — Opus 5.5 for building, GPT-6.1 Sol for review, Luna for bulk tasks.
Why use GPT-6.1 Sol for review instead of a higher-scoring model?
Per Meyer, a different model family provides a better check than Opus reviewing itself, and Sol’s $0.32–$0.39 cost per task makes running that review on every meaningful change affordable.
What are GPT-6.1 Sol’s main drawbacks?
High and xhigh settings take 57 to 69 seconds to produce a first token, ruling out interactive use, and Opus 5.5 still leads it by 5 index points at xhigh. Low and max settings have not yet been benchmarked.
Is a higher effort setting always better?
No. On Opus 5.5, going from xhigh to max adds 2 points for 73% more cost per task, and medium-to-max costs 4.46x for 7 points. Meyer’s rule: effort is not capability.
Do these index scores apply to everyone’s workload?
Meyer explicitly cautions they do not — the Artificial Analysis Intelligence Index is a map of general capability, and he recommends shadow-testing any model on your own tasks before switching.
Source: ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
