🔍 Read the full analysis: AI’s Cheap Creation Comes With A Costly Checking Problem on ThorstenMeyerAI.com
Get school and study supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
A source essay points to a widening gap between AI-generated work and the human capacity to check it, citing OpenAI mathematics output, software-review data and a contract-workflow evaluation. The figures suggest review can become a bottleneck, though some cited data comes from vendors and the article does not establish a single, comparable measure across fields.
A source essay argues that AI is making work faster to produce than to verify, citing OpenAI’s publication of 722 mathematical manuscripts and software data showing longer review waits or limited human checks. The examples point to a growing verification bottleneck, though the figures come from different sources and do not establish one common measure of the problem.
According to the source, OpenAI posed about 4,000 mathematical problems to a model and published 722 resulting manuscripts, grouped into 372 families. The source says the average result took about three hours of compute. Some results were formally checked using Lean, while OpenAI warned that some unformalized results could have issues. The source contrasts that volume with the careful review by five leading mathematicians of an earlier result from the same programme: a proposed counterexample to an Erdős conjecture.
In software, the essay cites figures from two industry analytics firms. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB, analyzing 8.1 million pull requests across 4,800 organizations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A 2026 peer-reviewed study cited by the essay found 61% of AI-agent pull requests had no human review before they were merged or closed.
The source also describes an OpenAI partnership with contract-software company Ironclad. In an evaluation across 11 tasks, OpenAI’s GPT-6 Astra met 55% of evaluation criteria on average, according to the essay. That indicates performance against the stated criteria, not that the system independently completed legally usable work. A human would still need to identify errors and determine whether the output fits the relevant workflow.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Sets the Pace
If organizations produce more AI-assisted work than reviewers can assess, the constraint shifts from drafting to judgment and accountability. That can delay useful work, let unchecked errors pass or push reviewers to reject machine-generated material by default. The essay cites LinearB data saying 38% of reviewers deliberately deprioritize AI-generated changes; that figure describes reported reviewer behavior, not a universal practice.
The concern extends beyond immediate quality checks. The source argues that review skill grows through doing the underlying work: writing code, drafting contracts and proving results. If AI tools take over much of that junior work, organizations may have fewer people gaining the experience needed to assess later output. This is a risk raised by the essay, not a demonstrated outcome of the cited figures.
The practical implication is that the value of AI depends not only on how much it can generate but on whether an organization has enough qualified people and procedures to validate what it uses. Senior engineers, specialist lawyers, auditors and scientific reviewers may become limits on adoption when their time is scarce. The source calls this a “referee premium”; it is an economic interpretation, not a measured wage trend.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Verification Gap
The examples in mathematics, software and contracting differ, but each separates producing an answer from establishing that it is correct and fit for purpose. In mathematics, software can check whether a proof follows from its stated assumptions. That does not establish that the theorem is the right one or that the result matters. In software, tests can check specified behavior but cannot prove that the tests capture every real requirement.
The source essay describes this distinction as “verification abundance, adjudication scarcity”: automated tools can check narrow properties, while people still decide whether the work addresses the right question and what should be trusted. It also notes that several software-data providers cited sell code-review products. That commercial connection does not invalidate their figures, but it is relevant when interpreting them; the source gives no independent, unified dataset covering all three fields.
Responsibility is another part of the gap. A person signs a contract, an engineer may approve a design and named authors answer for research. Automated checks can support these processes, but they do not by themselves assign professional or legal accountability. The extent to which organizations can change those arrangements is not addressed by the examples.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Enough?
The cited figures do not show whether the increase in review time is caused by AI adoption alone. The source does not provide study methods, definitions of review time, or comparable before-and-after windows for every data point. Faros AI and LinearB sell products in the software-development market, and their results may not represent every organization. The 2026 peer-reviewed study is identified, but its title, sample details and methods are not supplied in the source material.
It is also unclear how many of the 722 mathematical manuscripts were correct, useful or later accepted by independent experts. The source says some were formally checked and notes OpenAI’s warning about unformalized results, but it does not give a complete error rate. Likewise, the 55% contract evaluation result does not specify the consequences of missed criteria or how the evaluation maps to real legal work.
More broadly, the examples support a concern about review capacity, but they do not prove that the same bottleneck is affecting every industry or that AI necessarily makes checking more expensive in every setting. The scale of any effect will depend on task complexity, safeguards, review practices and the costs of errors.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Evidence to Watch in Practice
The next useful evidence would show whether organizations can increase review capacity without lowering standards: for example, whether AI-assisted checks catch errors reliably, whether human review times change over longer periods, and how often unchecked output produces problems. Independent studies with clear comparison periods and shared definitions would help distinguish a broad pattern from results tied to particular tools or teams.
For the mathematics example, subsequent formalization and independent assessment could clarify how many generated results withstand scrutiny and what kinds of mistakes remain. For software and contract work, organizations’ own reporting on review coverage, correction rates and outcomes would help show whether faster generation creates a net gain after verification costs are included.
Until such evidence is available, the source’s central point remains an argument supported by several examples rather than a settled cross-industry conclusion: AI can expand output quickly, while trusted use still depends on human expertise, effective checks and clear accountability.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The source essay argues that AI-generated work is increasing faster than the capacity to check it. It supports that argument with examples from mathematical research, software development and contract workflows.
Did OpenAI formally verify all 722 mathematical manuscripts?
No. The source says some results were checked in Lean and reports OpenAI’s warning that some unformalized results could have issues. It does not provide a full breakdown of which manuscripts were checked or their accuracy.
What do the software-review figures show?
The cited reports describe longer review waits, lower acceptance rates for AI-generated changes in one dataset, and a peer-reviewed study in which 61% of AI-agent pull requests received no human review before merging or closing. The measures come from different sources and should not be treated as one combined estimate.
Does the evidence prove AI makes all review more expensive?
No. The examples raise a concern about review bottlenecks, but the source does not establish a universal effect. Some figures come from companies that sell code-review products, and methods and comparison periods are not fully described.
What should readers watch for next?
Look for independent, longer-term evidence on review time, error rates, human-review coverage and real-world outcomes. Those measures could show whether organizations are gaining usable productivity after accounting for verification.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
