🔍 Read the full analysis: AI Can Produce More For Less. What About Reviewing It? on ThorstenMeyerAI.com
Get hardware and tech essentials delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
TL;DR
OpenAI published 722 mathematical manuscripts produced from roughly 4,000 problems, while a prior result from the programme received careful scrutiny from five leading mathematicians. Software studies also report longer review waits and, in one study, many AI-agent pull requests received no human review. The figures point to a growing review bottleneck, though some software data comes from vendors and the long-term effects on jobs and training remain uncertain.
OpenAI published 722 mathematical manuscripts this week, generated from roughly 4,000 problems, while a counterexample to an earlier Erdős conjecture from the same programme received careful verification from five leading mathematicians. The contrast highlights a developing problem across research and professional work: AI can generate material quickly, but checking whether it is correct, relevant and safe to use still depends heavily on limited human expertise.
OpenAI’s mathematical work produced manuscripts across 372 problem families, with an average result taking about three hours of compute, according to the source material. Some results were formally checked using Lean, a proof assistant. OpenAI cautioned that some results that were not formalized could have issues. Formal verification can establish that a proof follows from a stated theorem, but it does not by itself establish that the theorem is the right one or that the result is meaningful.
Software studies describe a similar tension. Faros AI reported that teams merged 98% more pull requests during high-AI-adoption periods while review time increased 91%. LinearB, analysing 8.1 million pull requests across 4,800 organisations, reported that AI-generated changes waited 4.6 times longer for review to begin and were accepted 32.7% of the time, compared with 84.4% for human-written changes. These are findings from separate datasets and should not be treated as one universal measure.
A peer-reviewed 2026 study found that 61% of AI-agent pull requests received no human review before they were merged or closed. Faros also reported a 31.3% rise in merges with zero review during high-adoption periods. Several cited software sources sell code-review products, a commercial interest worth keeping in mind when interpreting their results; the studies do not establish that every team or codebase is experiencing the same pattern.
The referee shortage: AI made doing cheap and checking expensive
OpenAI’s model produced a maths result in about three hours of compute. Verifying one earlier result took five of the world’s leading mathematicians. That ratio is the next decade of work: producing is cheap and abundant; trusting is slow, human and scarce.
Some Lean-checked; OpenAI warns unformalized ones “could have issues.” Verification abundance, adjudication scarcity.
Faros AI. LinearB (8.1M PRs): AI changes wait 4.6× longer, accepted 32.7% vs 84.4%.
Real progress. Someone still has to find the other 45% before the work can be used.
Machine output arrives without reasoning a reviewer can interrogate. It looks locally clean and gives no clue where it’s wrong.
A prover confirms the proof proves its statement; tests confirm what tests check. Neither confirms it’s what was needed.
Contracts are signed, designs stamped, papers defended. Responsibility is institutional — you can’t hold a model to it.
of AI-agent pull requests got no human review at all (EASE 2026). Zero-review merges up 31.3% (Faros).
of reviewers deliberately deprioritise AI changes (LinearB). Good machine work waits behind bad.
OpenAI chose which maths families were significant. When referees can’t keep up, the producer’s filter becomes the review.
Budget review hours next to model spend.
Provers, types, tests, policy engines.
Experts only where consequences are high.
Who profits from generation pays for checking.
Keep some production human for learners.
The first automation question was which jobs AI would do. The better one is which jobs AI makes more necessary: the ones that check, adjudicate and take responsibility. Expect a referee premium — senior engineers, auditors, specialist lawyers, reviewing scientists become the binding constraint on how much AI output anyone can use.Accountability — standing behind a result — may be the most durable form of human work there is.
Review Capacity Shapes AI Use
When systems generate more work than specialists can examine, the practical limit on AI use may shift from production to review capacity. Organisations may have code, research or draft contracts ready faster, yet still be unable to rely on them until someone checks the substance and accepts responsibility.
The consequences are not limited to slower approvals. If reviewers are overwhelmed, work may be waved through, delayed because it came from AI, or screened mainly by the people who produced it. Each response carries a different risk: unchecked errors, wasted time, or dependence on a producer’s own judgment without independent scrutiny. The evidence supplied points to these pressures, but does not quantify their overall cost across the economy.
There is also a workforce question. Experienced reviewers typically develop judgment through years of doing the underlying work. If AI reduces opportunities for junior staff to write code, draft contracts or build research skills, organisations could weaken the future supply of people qualified to review machine output. That is a plausible concern raised by the source, not an established long-term outcome.
As an affiliate, we earn on qualifying purchases.
A Pattern Across Three Fields
The source frames mathematics, software and contract work as examples of the same mismatch. In mathematics, proof assistants can check formal steps, while specialists still judge whether a result addresses the intended question and deserves attention. The earlier Erdős-conjecture counterexample prompted scrutiny by five leading mathematicians, illustrating that producing a result and establishing confidence in it are distinct tasks.
In professional workflows, OpenAI partnered with contract-software company Ironclad to train GPT-6 Astra on real contracting tasks. The source says the model met an average of 55% of evaluation criteria across 11 tasks, an improvement over its predecessor. That evaluation indicates progress on the measured tasks; it does not show that the system can independently complete legal work or that every missed criterion has equal importance.
AI can assist with checking, including by running tests or formal proof tools. But those tools assess what they have been asked to assess. They cannot automatically determine whether requirements, tests or assumptions reflect the real need. Contracts also require accountable signatories, and engineering designs or research claims are subject to professional and institutional responsibilities.
As an affiliate, we earn on qualifying purchases.
How Much Review Is Missing?
The available figures do not establish a single, comparable rate of review failure across industries. The software studies use different datasets, time periods and measures, and some sources sell code-review tools. The supplied material does not include study methods in enough detail to independently assess how representative each sample is.
It is also unclear how often unreviewed AI output causes consequential errors, how much review is later performed through testing or other controls, and whether teams will add reviewers or develop reliable automated checks. The 55% contract-task score does not disclose which criteria were missed or how the tasks were weighted. The longer-term effect of AI on junior workers’ training and the future supply of expert reviewers remains unknown.
peer review software for developers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Tracking Review and Training
The next useful evidence will show whether organisations change review staffing and procedures as AI-generated work increases. For software, that means tracking review delays, defects, overrides and the share of changes merged without human examination, with methods and comparison periods clearly stated. For mathematical work, readers will need to distinguish formal verification from expert assessment of a result’s assumptions and significance.
Further evaluations of AI contract tools should identify which criteria systems miss and whether human reviewers catch those gaps before documents are used. The source material does not give a timetable for such follow-up or announce a new review standard. For now, the central question is whether human expertise and training can keep pace with the growing supply of AI-produced work.
As an affiliate, we earn on qualifying purchases.
Key Questions
What did OpenAI publish this week?
OpenAI published 722 mathematical manuscripts produced from roughly 4,000 problems and grouped into 372 problem families, according to the source material. Some results were formally checked in Lean; OpenAI said some non-formalized results could have issues.
Does formal verification prove an AI result is useful?
No. Formal verification can check whether a proof establishes its stated theorem. It does not, by itself, show that the theorem captures the right question, that the result matters, or that its assumptions fit a real-world task.
What do the software studies report?
Separate studies report longer review waits for AI-generated changes and a high share of AI-agent pull requests receiving no human review in one peer-reviewed study. The datasets and measures differ, and some sources sell review tools, so the numbers should not be treated as a single industry-wide rate.
Is AI replacing human reviewers?
The supplied evidence does not establish that reviewers are being broadly replaced. It instead raises questions about whether human review capacity is keeping up with output, and whether reduced opportunities to do entry-level work could affect how future experts gain experience.
Source: ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
