AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Why AI Needs A Better System For Checking Its Work on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get office and shipping supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

Examples in mathematics, software and contract work point to a widening gap between the speed of AI-generated output and the capacity to check it. The figures come from sources cited by ThorstenMeyerAI.com, including vendors with commercial interests, so they warrant caution; the broader question of accountability and reviewer capacity remains unresolved.

A new analysis of AI-generated mathematics, software and contract work argues that production is scaling faster than verification, leaving human reviewers to decide which results can be trusted and used. Its examples include 722 mathematical manuscripts attributed to OpenAI and software-team data showing more code changes alongside longer review times, though several figures come from companies that sell review tools.

The analysis says OpenAI posed about 4,000 mathematical problems and produced 722 manuscripts grouped into 372 families. Some results were formally checked using Lean, a proof-assistant system. OpenAI cautioned that some manuscripts without formal verification could contain issues. The account contrasts that output with the extensive expert review of an earlier result from the programme: a counterexample to an Erdős conjecture was checked by five leading mathematicians. The comparison illustrates the central claim, but does not establish that every manuscript requires the same review process or effort.

In software, the analysis cites a Faros AI comparison of teams during periods of high and low AI adoption. It reports that teams merged 98% more pull requests while review time rose 91%. LinearB, which the source says analysed 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin. It also reported acceptance rates of 32.7% for AI-written changes and 84.4% for human-written ones. These are findings attributed to the companies, not independently established sector-wide rates.

The source also cites a peer-reviewed 2026 study in which 61% of AI-agent pull requests received no human review before being merged or closed. It says Faros reported a 31.3% rise in merges with no review during high-adoption periods. For professional services, the article describes an OpenAI partnership with contract-software company Ironclad and says GPT-6 Astra met 55% of evaluation criteria on average across 11 tasks. The source does not provide the full evaluation method or a detailed account of the remaining criteria.

At a glance
analysisWhen: Source analysis published this week; so…
The developmentA source analysis of AI-generated mathematics, software and contract work argues that verification capacity is emerging as a constraint on using AI output.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Limits AI Adoption

If generation becomes cheaper while checking remains labor-intensive, organisations may struggle to turn more AI output into usable work. The constraint is not just the number of reviewers: it is whether they can assess the right question, identify consequential errors and take responsibility for a decision. In a mathematical proof, that can mean checking whether the argument establishes a meaningful claim. In software, it can mean determining whether tests cover the intended behavior. In a contract, it can mean spotting a missed approval rule or unsuitable clause.

The source analysis describes three possible responses when review capacity falls short: rubber-stamping, delaying or deprioritising AI work, and relying more heavily on the producer’s own selection of what to publish. Each carries a different risk. Work may pass without adequate scrutiny, useful changes may wait because they are machine-generated, or the party that created the output may also become its main gatekeeper. The cited numbers suggest these patterns deserve attention, but they do not show how common each response is across all industries.

There is also a training issue. Experienced reviewers generally build judgment through years of doing the underlying work. If junior staff increasingly oversee AI drafts rather than learning to write code, prepare contracts or develop proofs themselves, organisations could weaken the future supply of skilled reviewers. That is a risk raised by the analysis, not a measured outcome in the figures it cites. It points to a practical question for employers: how to gain productivity from AI while preserving work that teaches people to spot errors.

Amazon

AI review and verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Examples Across Three Fields

The analysis links three areas that use different standards of evidence. Mathematics can use formal systems such as Lean to verify that a proof follows from its stated assumptions. Software teams combine automated tests with human code review. Contract work depends on applying language and rules to particular transactions. In each case, a check can confirm some properties without resolving every judgment that matters: a proof checker does not decide whether a theorem is useful, tests cannot cover requirements they omit, and a model’s score on a task set does not by itself establish that its contract work is ready for use.

The source refers to this mismatch as “verification abundance, adjudication scarcity.” That phrase captures a distinction between checking a defined property and deciding whether the work is appropriate, complete and reliable in context. The distinction helps explain why automated verification can reduce some review labor without removing the need for people who understand the task and can answer for the final decision.

The article cites company analyses alongside a peer-reviewed study, but cautions that several data providers sell code-review products. That commercial interest does not by itself invalidate their findings, yet it is a reason to examine methods, definitions and comparison periods before treating the reported percentages as representative of every team. The supplied material does not give enough detail to independently reproduce the calculations.

Amazon

code review software for AI-generated code

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

How Broad Is the Review Gap?

The examples do not establish a single, comparable measure of verification costs across mathematics, software and legal work. The source gives figures from different methods and periods, and does not provide the underlying datasets or complete study details for each claim. It also notes that some software-data providers sell review tools, making independent scrutiny of their methods especially relevant.

It remains unclear how much of the reported delay or lower acceptance rate is caused by AI authorship itself, differences in the complexity of changes, team practices or other factors. The source also does not establish that every AI-generated result needs human review at the same level, or that automated checks cannot improve. The longer-term effect on junior workers’ training and the supply of experienced reviewers is presented as a concern, not a demonstrated trend.

Amazon

mathematical proof assistant software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Better Evidence and Review Practices

The next useful step is clearer reporting: studies that define review time, acceptance and “no review” consistently, disclose evaluation methods and compare similar work across teams. Independent research could help determine whether AI-generated changes create distinct review burdens or whether the reported differences reflect how teams select and route that work.

For organisations using AI, the immediate issue is to match output volume with a review process suited to the consequences of error. That may include formal checks for properties that can be specified, human review for contextual judgment and clear responsibility for approvals. Employers will also need to decide how junior staff can gain the experience needed to become reliable reviewers. The source analysis does not identify a settled solution or a timetable; whether review tools and training practices can keep pace with AI output remains an open question.

Amazon

AI project review management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main finding of the analysis?

It argues that AI can generate work faster than people can verify it. Examples from mathematics, software and contract work illustrate the gap, but the cited figures come from different sources and should not be treated as one universal measure.

Were all 722 mathematical manuscripts formally checked?

No. The source says some results were checked in Lean and quotes OpenAI warning that some unformalized results could have issues. It does not specify in the supplied material how many manuscripts received formal verification.

What did the software figures measure?

The source attributes figures to Faros AI and LinearB, including changes in pull-request volume and review time, review-start delays and acceptance rates. These are company-reported findings; the source notes that several cited providers sell code-review tools.

Can AI check AI-generated work?

Automated systems can check defined properties, such as whether a proof meets formal rules or code passes specified tests. Those checks do not necessarily establish that the original claim, test or work product addresses the right problem. The analysis argues that human judgment and accountability remain necessary in many settings.

What remains unknown?

The evidence presented does not show how representative the reported figures are across industries, how much AI authorship causes the review differences, or how future training and reviewer supply will change. Independent, comparable studies would help clarify those questions.

Source: ThorstenMeyerAI.com

This content is for general information only and is not financial, tax or legal advice. Consult a qualified professional for decisions about your money.
HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What’s Changed For Anthropic Model Inference On Amazon Bedrock?

Anthropic says its models on Amazon Bedrock are available for in-region inference in Seoul and Singapore; model and data-handling details remain unspecified.

Is Walmart Open On 4Th Of July

Walmart is open on July 4th, with store hours varying by location. Find out what to expect for Independence Day shopping.

Comcast to split into two companies, spin off NBCUniversal and Sky

Comcast will divide into two firms, spinning off NBCUniversal and Sky, in a strategic move aimed at unlocking value and focusing on core businesses.

Top Links 1165 Broken Confidence. Patient Capital In China. Forgotten Feminist. On The Banks Of The Dordogne.

Examining recent issues of broken confidence, patient capital challenges, and forgotten feminists in China, with context and future outlook.