AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: What AI’s Low-Cost Production Means For Quality Control on ThorstenMeyerAI.com

Before you orderOffer from Amazon

Get the latest gadgets delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

AI tools can produce mathematical manuscripts, software changes and contract drafts at growing volume, while checking them still takes human time. The supplied report points to longer review queues and gaps in oversight, but some figures come from companies selling review tools and need careful interpretation.

AI tools are producing more work at lower cost in mathematics, software and contract drafting, while human review remains slower and limited, according to a report drawing on OpenAI announcements, industry datasets and research. The gap matters because organizations can only safely use output they can check, and the evidence supplied points to longer review queues and instances of work moving ahead with little or no human scrutiny.

OpenAI said its model was given about 4,000 mathematical problems and produced 722 manuscripts, grouped into 372 families. The source says the average result used about three hours of compute. Some results were checked in Lean, a proof-assistant system; OpenAI cautioned that some results without formal verification could have issues. The account contrasts that volume with the careful review by five leading mathematicians of an earlier result from the same programme, described as a counterexample to an old Erdős conjecture.

In software, the report cites separate measurements from Faros AI and LinearB. Faros reported that teams merged 98% more pull requests during high-AI-adoption periods, while review time rose 91%. LinearB said its analysis of 8.1 million pull requests across 4,800 organizations found AI-generated changes waited 4.6 times longer for review to start and were accepted 32.7% of the time, compared with 84.4% for human-written changes. A 2026 peer-reviewed study cited in the source found 61% of AI-agent pull requests received no human review before being merged or closed.

The source also describes an OpenAI partnership with contract-software company Ironclad. In an evaluation of 11 contracting tasks, GPT-6 Astra met an average of 55% of the criteria, which the source characterizes as an improvement over the prior model. That result indicates progress on the evaluated tasks, but also leaves criteria unmet; the source does not provide the full evaluation design or task-by-task scores.

At a glance
reportWhen: Source material describes developments…
The developmentA report on AI-assisted work across mathematics, software and contracts describes a growing mismatch between the low cost of generating output and the limited capacity to verify it.
The Referee Shortage — Post-Labor
AI Dispatch · Post-Labor · 7 October 2026

The referee shortage: AI made doing cheap and checking expensive

OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.

One pattern, three fields
Mathematics
722
manuscripts, ~3h compute each

Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.

Software
+98% / +91%
more PRs merged / longer review

Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.

Professional workflows
55%
of criteria met — Astra on Ironclad

Real progress. Someone still has to find the other 45% before the work can be used.

Generation collapsed. Verification didn’t. (conceptual, not to scale)
Cost to produce a resultdown
Cost to check a resultnot down
No author intent

Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.

Checks the answer, not the question

A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.

Someone must be accountable

Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.

Illustrative: $1 of model time + 4 minutes of review at $45/hour. Halving the model price saves 12.5%; one extra review minute erases it. In that example, review is three-quarters of the bill.
What happens when referees run out — already visible
Rubber-stamping
61%

of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).

Triage by suspicion
38%

of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.

Producer as filter
~4,000 → 372

OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.

The apprenticeship paradox: reviewers are made by doing the work. The work AI absorbs — writing code, drafting contracts, proving lemmas — is exactly what trained the reviewers. Demand for judgement rises as its supply line shrinks.
What to do
Price verification

Budget review hours next to model spend.

Formalise checks

Provers, types, tests, policy engines.

Tier the review

Experts only where consequences are high.

Fund the referees

Who profits from generation pays for checking.

Protect apprenticeship

Keep some production human for learners.

The take

The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.

Sources: OpenAI maths release & Erdős verification as covered here; arXiv:2608.28997; OpenAI × Ironclad (6 Oct 2026); Faros AI; LinearB 2026 (8.1M PRs); Duma et al., EASE 2026 — via secondary reporting. Several code-review sources sell review tools. Review-cost example illustrative. Analysis is the author’s.
thorstenmeyerai.com

Review Capacity Limits AI Output

The immediate consequence is operational: more generated work does not automatically mean more usable work. If expert review cannot keep pace, a company may face backlogs, accept work with inadequate checks, or delay machine-generated work because reviewers distrust it. The report cites LinearB data that 38% of reviewers deliberately deprioritize AI-generated changes, a response that may protect against errors but can also hold up sound work.

The issue extends beyond efficiency. In software, a missed defect can affect users or systems; in contracts, a missed approval requirement or unsuitable clause can have legal consequences. In mathematics, formal verification can establish that a proof follows from stated premises, but human experts still assess whether the claim is relevant, correctly framed and meaningful. These examples show that verification includes more than checking whether an output meets a narrow test.

The report argues that expertise itself may become a constraint. Senior reviewers typically build judgment through years of doing the underlying work. If junior staff mainly supervise AI-generated drafts instead of learning to write code, proofs or contracts, organizations may weaken the pipeline that develops future reviewers. That is a concern raised by the report, not a measured outcome established by the figures cited.

Amazon

AI quality control review tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Three Fields, One Review Gap

The source frames the evidence as a shared pattern across mathematics, software and professional workflows: generation can scale quickly, while deciding whether an output is correct and fit for purpose continues to require people, formal systems or both. Its phrase “verification abundance, adjudication scarcity” captures the distinction between checking a defined answer and deciding whether the right question was asked.

Those are different tasks. A proof checker can test a proof against a theorem as written; it cannot by itself determine whether the theorem captures the intended claim. Software tests check the cases they cover, not every real-world requirement. Likewise, an automated review may flag a contractual issue without taking responsibility for the final agreement. The source says an earlier mathematical counterexample was disputed, but gives no detailed account of the dispute, so it does not establish precisely how or why the result failed.

The software figures also come with a qualification: Faros AI and LinearB sell code-review tools, according to the source, so their measurements should be read with that commercial context in mind. The metrics use different samples and definitions and should not be treated as directly comparable. The cited peer-reviewed study offers another measure, but its specific methods and scope are not included in the material provided.

Amazon

software code review automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Limits of the Available Evidence

The figures do not establish a single industry-wide rate for AI errors or review delays. The source does not specify the dates and methods behind every dataset, nor does it explain how high- and low-adoption periods were defined by Faros. LinearB’s acceptance figures also do not, on their own, show why changes were rejected or whether the comparison groups were otherwise similar.

For mathematics, the supplied account does not list how many of the 722 manuscripts were formally checked, what standards were used to select them, or how independent the review was. For Ironclad’s evaluation, the full criteria and baseline model score are not given. It is also unclear how often the cited unreviewed software changes caused defects, and whether organizations are adding review staff or changing their processes in response.

The broad claim that AI makes trust more expensive is the report’s interpretation. The evidence describes specific workflows and measurements, but does not show that every use of AI raises review costs or that human review is the only effective safeguard.

Amazon

contract drafting AI software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Track Review and Training Measures

The next useful evidence will show whether review capacity is expanding alongside AI output. Organizations and researchers can track review wait times, the share of changes receiving human scrutiny, rejection reasons and defects found after release, while explaining how each measure is defined. Those details would help distinguish a temporary adjustment period from a persistent bottleneck.

For AI-assisted mathematics and contracting, fuller evaluation disclosures would clarify which tasks are formally verified, which criteria systems miss, and how performance compares with earlier tools under the same tests. The source material does not identify a scheduled follow-up release or a specific policy change. For now, the central question is whether workplaces can scale not only generation, but also the expertise and training needed to judge the results.

Amazon

mathematical proof verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What is the main development described?

The report describes AI-generated work increasing across several fields while human review remains time-consuming and limited. It uses mathematical manuscripts, software pull requests and contract tasks as examples.

Did OpenAI’s mathematical manuscripts all receive formal verification?

No. The source says some results were checked in Lean and quotes OpenAI warning that unformalized results could have issues. It does not state how many manuscripts received formal checks.

What do the software figures show?

The cited datasets report increased pull-request volume alongside longer review waits, lower acceptance for AI-generated changes in one analysis, and substantial shares of changes receiving no human review in another study. The measures have different samples and methods and should not be combined into one rate.

Does this prove AI output is unreliable?

No. The figures do not establish that all AI-generated work is unreliable. They describe review patterns and evaluation results; they do not provide a universal error rate or show that every output requires the same kind of checking.

What remains unknown?

The source does not provide full methods for several datasets, complete details of the contract evaluation, or the proportion of mathematical manuscripts formally verified. It also does not show whether review backlogs are causing more defects or how quickly organizations are building review capacity.

Source: ThorstenMeyerAI.com

HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Nanotech Gender Gap: Where Are the Women in Nano?

AIThis post was created with the assistance of artificial intelligence (AI).Women remain…

AmenGate: The Moment Before the Scroll

AmenGate introduces a prayer lock for iPhone, replacing mindless scrolling with meaningful prayer, built on system-level frameworks for trust and longevity.

The Truth About August 2 And The State Of AI Today

Key developments on August 2, 2026, reveal the EU AI Act’s compliance deadlines shifted, but critical transparency rules remain in effect. What this means for AI compliance.

Rogue One: The Andor Cut — On Fan Editing as Tonal Reverse-Engineering

A fan editor releases a reimagined version of Rogue One, blending tonal elements from Andor to explore a different narrative feel, raising questions about fan editing and storytelling.