🔍 Read the full analysis: What AI’s Low-Cost Production Means For Quality Control on ThorstenMeyerAI.com
Get the latest gadgets delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
AI tools can produce mathematical manuscripts, software changes and contract drafts at growing volume, while checking them still takes human time. The supplied report points to longer review queues and gaps in oversight, but some figures come from companies selling review tools and need careful interpretation.
AI tools are producing more work at lower cost in mathematics, software and contract drafting, while human review remains slower and limited, according to a report drawing on OpenAI announcements, industry datasets and research. The gap matters because organizations can only safely use output they can check, and the evidence supplied points to longer review queues and instances of work moving ahead with little or no human scrutiny.
OpenAI said its model was given about 4,000 mathematical problems and produced 722 manuscripts, grouped into 372 families. The source says the average result used about three hours of compute. Some results were checked in Lean, a proof-assistant system; OpenAI cautioned that some results without formal verification could have issues. The account contrasts that volume with the careful review by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.
In software, the report cites separate measurements from Faros AI and LinearB. Faros reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organizations found AI-generated changes waited 4.6 times longer for review to start and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A 2026 peer-reviewed study cited in the source found 61% of AI-agent pull requests received no human review before being merged or closed.
The source also describes an OpenAI partnership with contract-software company Ironclad. In an evaluation of 11 contracting tasks, GPT-6 Astra met an average of 55% of the criteria, which the source characterizes as an improvement over the prior model. That result indicates progress on the evaluated tasks, but also leaves criteria unmet; the source does not provide the full evaluation design or task-by-task scores.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Limits AI Output
The immediate consequence is operational: more generated work does not automatically mean more usable work. If expert review cannot keep pace, a company may face backlogs, accept work with inadequate checks, or delay machine-generated work because reviewers distrust it. The report cites LinearB data that 38% of reviewers deliberately deprioritize AI-generated changes, a response that may protect against errors but can also hold up sound work.
The issue extends beyond efficiency. In software, a missed defect can affect users or systems; in contracts, a missed approval requirement or unsuitable clause can have legal consequences. In mathematics, formal verification can establish that a proof follows from stated premises, but human experts still assess whether the claim is relevant, correctly framed and meaningful. These examples show that verification includes more than checking whether an output meets a narrow test.
The report argues that expertise itself may become a constraint. Senior reviewers typically build judgment through years of doing the underlying work. If junior staff mainly supervise AI-generated drafts instead of learning to write code, proofs or contracts, organizations may weaken the pipeline that develops future reviewers. That is a concern raised by the report, not a measured outcome established by the figures cited.
As an affiliate, we earn on qualifying purchases.
Three Fields, One Review Gap
The source frames the evidence as a shared pattern across mathematics, software and professional workflows: generation can scale quickly, while deciding whether an output is correct and fit for purpose continues to require people, formal systems or both. Its phrase “verification abundance, adjudication scarcity” captures the distinction between checking a defined answer and deciding whether the right question was asked.
Those are different tasks. A proof checker can test a proof against a theorem as written; it cannot by itself determine whether the theorem captures the intended claim. Software tests check the cases they cover, not every real-world requirement. Likewise, an automated review may flag a contractual issue without taking responsibility for the final agreement. The source says an earlier mathematical counterexample was disputed, but gives no detailed account of the dispute, so it does not establish precisely how or why the result failed.
The software figures also come with a qualification: Faros AI and LinearB sell code-review tools, according to the source, so their measurements should be read with that commercial context in mind. The metrics use different samples and definitions and should not be treated as directly comparable. The cited peer-reviewed study offers another measure, but its specific methods and scope are not included in the material provided.
software code review automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Limits of the Available Evidence
The figures do not establish a single industry-wide rate for AI errors or review delays. The source does not specify the dates and methods behind every dataset, nor does it explain how high- and low-adoption periods were defined by Faros. LinearB’s acceptance figures also do not, on their own, show why changes were rejected or whether the comparison groups were otherwise similar.
For mathematics, the supplied account does not list how many of the 722 manuscripts were formally checked, what standards were used to select them, or how independent the review was. For Ironclad’s evaluation, the full criteria and baseline model score are not given. It is also unclear how often the cited unreviewed software changes caused defects, and whether organizations are adding review staff or changing their processes in response.
The broad claim that AI makes trust more expensive is the report’s interpretation. The evidence describes specific workflows and measurements, but does not show that every use of AI raises review costs or that human review is the only effective safeguard.
As an affiliate, we earn on qualifying purchases.
Track Review and Training Measures
The next useful evidence will show whether review capacity is expanding alongside AI output. Organizations and researchers can track review wait times, the share of changes receiving human scrutiny, rejection reasons and defects found after release, while explaining how each measure is defined. Those details would help distinguish a temporary adjustment period from a persistent bottleneck.
For AI-assisted mathematics and contracting, fuller evaluation disclosures would clarify which tasks are formally verified, which criteria systems miss, and how performance compares with earlier tools under the same tests. The source material does not identify a scheduled follow-up release or a specific policy change. For now, the central question is whether workplaces can scale not only generation, but also the expertise and training needed to judge the results.
mathematical proof verification software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is the main development described?
The report describes AI-generated work increasing across several fields while human review remains time-consuming and limited. It uses mathematical manuscripts, software pull requests and contract tasks as examples.
Did OpenAI’s mathematical manuscripts all receive formal verification?
No. The source says some results were checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state how many manuscripts received formal checks.
What do the software figures show?
The cited datasets report increased pull-request volume alongside longer review waits, lower acceptance for AI-generated changes in one analysis, and substantial shares of changes receiving no human review in another study. The measures have different samples and methods and should not be combined into one rate.
Does this prove AI output is unreliable?
No. The figures do not establish that all AI-generated work is unreliable. They describe review patterns and evaluation results; they do not provide a universal error rate or show that every output requires the same kind of checking.
What remains unknown?
The source does not provide full methods for several datasets, complete details of the contract evaluation, or the proportion of mathematical manuscripts formally verified. It also does not show whether review backlogs are causing more defects or how quickly organizations are building review capacity.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
