📊 Full opportunity report: Engineering Is Automated. Research Is the Residual. on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
AI systems are now capable of automating most engineering tasks in AI research, reaching near-saturation levels on key benchmarks. Research remains less automated, but evidence suggests the gap may close faster than expected, potentially transforming AI development timelines.
Recent advances in AI capability demonstrate that systems can now automate the core engineering tasks involved in AI research, with benchmarks approaching saturation levels. This development shifts the focus from engineering to research, raising questions about the future pace of AI innovation and the remaining human role.
Multiple independent benchmarks—CORE-Bench, MLE-Bench, and kernel design research—show AI systems have achieved near-complete automation of key engineering skills relevant to AI research. For example, CORE-Bench, which measures the ability to reproduce computational research papers, reached 95.5% success in December 2025, with the benchmark’s author declaring it ‘solved.’ Similarly, MLE-Bench, assessing performance in Kaggle competitions, hit 64.4% in February 2026, reaching a level comparable to mid-tier human practitioners. These patterns suggest that the bottleneck in AI research is shifting from engineering execution to the more cognitively complex research phase, which remains less automated. Experts like Thorsten Meyer interpret this as evidence that the residual research challenge may close faster than the engineering milestone, especially if research itself becomes a form of engineering at scale.Engineering is automated.
Research is the residual.
Six skill benchmarks. Edison’s framing. The question Clark leaves open is whether research is just engineering at scale.
Jack Clark’s Import AI #455 catalogs six benchmarks measuring AI capability on AI R&D tasks and concludes “AI can today automate vast swatches, perhaps the entirety, of AI engineering.” The residual question is research. The structural read on the residual: it may not be a permanent moat.
Six skills. One trajectory.
Clark catalogs six benchmarks measuring AI capability on AI R&D-relevant tasks. Each individual benchmark could be noise. Six benchmarks moving together is a curve. The pattern is the cascade observed across the broader Clark series — visible here in the specific R&D-skill domain.

The No-BS Guide to AI for Trading & Market Research: How to Use ChatGPT, Claude & AI Tools for Market Analysis, Stock Research & Data-Driven Trading … … Required (The No-BS AI Playbooks Book 3)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Three data points. Mixed signal.
Clark provides three data points on the creative-spark question. Yes-evidence: Erdős-1051, centaur math discovery, sporadic Move-37-style moments. No-evidence: low yield, framing dependence, absence of acceleration. The mixed signal is the honest read.
The data supports two readings. Pessimistic: rare moments suggest creative insight is qualitatively distinct from engineering work. Optimistic: rare moments are an artifact of low-volume exploration; more shots on goal yields more discoveries. Both readings are consistent with Clark’s “vast swatches, perhaps the entirety” claim. They differ on the residual.

AI Simply Explained by a Software Engineer
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five dimensions Clark gestures at but leaves underdeveloped.
Clark’s section is rigorous on the empirical evidence. Five strategic dimensions matter for the institutional response that the Clark series synthesis argues is structurally inadequate.

Evals for AI Engineers: Systematically Measuring and Improving AI Applications
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Two readings. Different equilibria.
The structural question Clark leaves open: is research a permanent moat that bounds automated AI R&D, or is it engineering at scale that dissolves with more shots on goal? Both readings are consistent with the current data. They differ by orders of magnitude in consequences.
Productivity multiplier years
Recursive loop operational

AI Robotic Arm Kit Hiwonder SO-ARM101 Embodied Imitation Learning Open Source 6-Axis Robot Arm 12 High-Torque Bus Servo Motors AI Vision Recognition (Advanced Kit, Included 3D Printed Part, Assembled)
【End-to-End Imitation Learning】Hiwonder SO-ARM101 robot arm is an embodied intelligent hardware platform compatible with the Lerobot open-source framework….
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Five audiences. Asymmetric cost of being wrong.
The institutional response should not bet on inspiration being a permanent moat. If the distinction holds, capacity built is still useful. If it closes, capacity is necessary. Asymmetric cost-of-being-wrong points toward building now.
IN INDUSTRY
IN ACADEMIA
POLICYMAKERS
INVESTORS
EVERYONE ELSE
Engineering is automated. The residual is the question. The institutional response should not bet on inspiration being a permanent moat.
Implications for AI Development and Human Roles
The automation of core engineering tasks means AI systems can now handle much of the technical execution involved in research. This could accelerate AI development timelines significantly, reducing the need for human intervention in routine engineering work. However, the remaining research phase—conceptualization, hypothesis generation, and innovation—may still require human insight. If research also becomes automatable, the pace of AI progress could accelerate further, potentially leading to a new era of rapid innovation and raising questions about the future of human involvement in AI research and development.Progress Patterns in AI Capabilities and Benchmark Saturation
Historically, AI development has been constrained by the difficulty of automating core engineering and research tasks. Recent data from 2024-2026 shows a rapid progression across multiple benchmarks. CORE-Bench, measuring research reproduction, improved from 21.5% in September 2024 to 95.5% in December 2025. MLE-Bench, evaluating Kaggle competition performance, rose from 16.9% in October 2024 to 64.4% in February 2026. The pattern across these benchmarks indicates a nearing saturation point, suggesting that current AI systems are approaching the limits of automation in engineering tasks. Meanwhile, research-focused capabilities are still developing but show promising signs of rapid progress, with continuous publication of new kernel design techniques and automation methods throughout 2025 and 2026.“The pattern across multiple benchmarks shows AI nearing saturation in core engineering skills, shifting the residual challenge to research, which may close faster than expected.”
— Thorsten Meyer
Unresolved Questions About Research Automation Pace
While engineering tasks are nearing full automation, it remains unclear how quickly research itself can be fully automated. The structural question is whether research can be reduced to engineering at scale, which could accelerate progress further. The timeline for this transition is still uncertain, and the technological breakthroughs needed to automate higher-level research tasks are not yet confirmed.
Monitoring Progress in AI Research Automation
Researchers and industry observers will focus on the continued development of AI systems capable of automating research tasks, including hypothesis generation, experimental design, and innovation. Key milestones include further improvements in benchmarks, publication of new automation techniques, and potential breakthroughs in automating research-level cognition. The next 12-24 months will be critical for assessing whether research automation can match engineering progress, potentially reshaping the timeline of AI development.
Key Questions
What are the main benchmarks showing AI automation progress?
CORE-Bench measures research reproduction, MLE-Bench assesses Kaggle competition performance, and kernel design research evaluates automation in low-level infrastructure tasks. All three show rapid progress approaching saturation.
Does automation mean AI can now fully replace human researchers?
While core engineering tasks are highly automated, research involves complex cognition that remains less automated. The timeline for full automation of research is still uncertain.
What are the implications for AI development timelines?
If research automation accelerates as engineering has, the pace of AI progress could increase significantly, potentially leading to faster breakthroughs and shifts in industry leadership.
Are there risks associated with automating research?
Automating research could lead to rapid progress but also raises concerns about oversight, safety, and the potential for unintended consequences. These issues require careful attention as automation advances.
Source: ThorstenMeyerAI.com