0 cumulative citations
View corpus contextLeading AI researchers warn automating AI research could spark recursive self‑improvement and constitute one of the field’s gravest risks, though they diverge on timing and policy; frontier lab staff are more likely than academics to expect explosive growth, and most predict advanced R&D systems will be kept private.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
2 cumulative citations
View corpus contextMany leading AI researchers expect AI development to exceed the transformative impact of all previous technological revolutions. This belief is based on the idea that AI will be able to automate the process of AI research itself, leading to a positive feedback loop. In August and September of 2025, we interviewed 25 leading researchers from frontier AI labs and academia, including participants from Google DeepMind, OpenAI, Anthropic, Meta, UC Berkeley, Princeton, and Stanford to understand researcher perspectives on these scenarios. Though AI systems have not yet been able to recursively improve, 20 of the 25 researchers interviewed identified automating AI research as one of the most severe and urgent AI risks. Participants converged on predictions that AI agents will become more capable at coding, math and eventually AI development, gradually transitioning from `assistants' or `tools' to `autonomous AI developers,' after which point, predictions diverge. While researchers agreed upon the possibility of recursive improvement, they disagreed on basic questions of timelines or appropriate governance mechanisms. For example, an epistemic divide emerged between frontier lab researchers and academic researchers, the latter of which expressed more skepticism about explosive growth scenarios. Additionally, 17/25 participants expected AI systems with advanced coding or R&D capabilities to be increasingly reserved for internal use at AI companies or governments, unseen by the public. Participants were split as to whether setting regulatory ``red lines" was a good idea, though almost all favored transparency-based mitigations.
Summary
Main Finding
Leading AI researchers (25 interviews) converge on a plausible, staged pathway by which AI systems could meaningfully automate AI R&D (ASARA)—from coding assistants to autonomous AI developers—and view that prospect as a major risk. While most interviewees accept the possibility of recursive improvement, they sharply disagree on timelines, the likelihood of an “intelligence explosion,” and appropriate governance. A majority expect the most capable ASARA systems to be kept internal by frontier AI labs, raising concerns about concentrated power and unobserved, AI-augmented progress.
Key Points
- Core consensus trajectory (three stages):
- Research speedup: tools that multiply human researchers’ productivity (e.g., coding/engineering acceleration).
- Collaboration: AI handles substantive research sub‑tasks while humans keep high‑level control.
- Full loop automation: AI executes entire research cycles autonomously (human bottleneck).
- Terminology: The study uses “ASARA” = AI Systems for AI R&D Automation; most participants were comfortable discussing “intelligence explosion” / “takeoff” scenarios even when disagreeing about timing or magnitude.
- Risk salience: 20/25 participants identified automating AI research as one of the most severe and urgent AI risks.
- Deployment expectations: 17/25 participants expected ASARA-capable systems (or their full-capability versions) to be kept internal by frontier labs rather than publicly released.
- Reasons labs might keep ASARA internal:
- Preserve competitive advantage.
- Limited compute / opportunity cost of serving the public versus using compute for internal R&D.
- Avoid capability diffusion, distillation, or misuse.
- Reasons labs might deploy publicly:
- Financial pressures and investor expectations (need for revenue / demonstrable products).
- Regulatory interventions or government mandatory access.
- Business model, corporate culture, or the economics of specialized external deployment (selling distilled / restricted versions).
- Organizational and epistemic divide: Frontier lab researchers tended toward less skepticism about fast takeoffs than many academic participants.
- Governance views: Participants were split on prescriptive “red lines” (e.g., bans on self‑improvement), but almost all favored transparency-based mitigations (audits, disclosure, third‑party evaluation); some argued that some mitigations could be worse than the risks they address.
- Metrics & benchmarks: Several participants emphasized “task horizon” progression as a useful metric (how long/complex tasks an agent can complete autonomously); existing benchmarks (PaperBench, REBench, MLEBench, METR) are relevant but incomplete.
Data & Methods
- Recruitment and sample:
- 182 researchers invited; 25 agreed to in-depth interviews (Aug–Sept 2025).
- Recruitment channels: literature-based (7), conference/workshop (8), network/snowball sampling (10).
- Participant mix: frontier lab researchers (current and former), academics (professors, postdocs, PhD students), industry researchers, startup founder(s), nonprofit/forecasting researchers.
- Interview protocol:
- Semi-structured interviews, 40–60 minutes each; topics: expectations for ASARA, deployment dynamics, risks, governance, red lines.
- Definitions (e.g., Intelligence Explosion, ASARA, Frontier AI Company, Red Lines) provided to participants.
- Analysis:
- Transcription (OpenAI Whisper or Otter.AI), de-identification, inductive coding by the lead author.
- AI-assisted categorical coding (Claude) used for some categorical variables.
- Themes developed from repeated codes across transcripts.
- Limitations to note:
- Small, non-representative sample biased toward those already engaged with the topic; selection and survivorship biases possible.
- Self-report and hypothetical scenario limitations; views reflect perceptions at time of interviews.
- Coding/interpretation subjective; AI-assisted coding may introduce additional artifacts.
Implications for AI Economics
- Market structure and concentration:
- If frontier labs keep ASARA internal, rent capture by a few firms increases: higher returns to compute, proprietary models, and top AI talent. This intensifies winner-take-most dynamics and raises barriers to entry.
- Internalization reduces observable signals about capability progression, hindering market discipline and public valuation accuracy; investors face higher uncertainty and tail risks.
- Returns to factors of production:
- ASARA would shift returns from human AI researchers toward compute and proprietary algorithms. Labor demand for routine research tasks may fall, while demand for compute capacity, specialist safety roles, and institutional oversight roles rises.
- Firms may reallocate compute from external APIs/productization to internal R&D if internal ASARA yields higher marginal returns, affecting API market revenues and consumer‑facing growth.
- R&D productivity and externalities:
- ASARA promises large productivity uplifts in R&D (uplift and task‑horizon extension). Positive private returns may not internalize systemic risks (misuse, rapid capability diffusion, destabilizing economic/safety externalities).
- Market incentives can produce races to internalize and exploit ASARA capabilities, potentially accelerating socially harmful outcomes absent coordination.
- Observability, information asymmetry, and policy:
- Reduced observability (internal-only deployments) creates information asymmetries that complicate regulatory supervision, risk assessment, and market correction.
- Transparency-based policies (model registries, mandatory capability disclosures, independent audits, reporting of task-horizon metrics) are consistent with researcher preferences and could improve market functioning and risk pricing.
- Policy and industrial-policy levers relevant to economists:
- Antitrust and competition policy: consider remedies if concentration creates systemic risks (e.g., forced interoperability, compute-market regulation).
- Compute access and funding: public provision or subsidized access to compute for independent researchers could reduce concentration and improve external monitoring.
- Tax/subsidy design: calibrate to discourage dangerous private races and internal withholding of safety‑relevant information (e.g., conditional R&D subsidies, tax incentives for verified transparency).
- Regulation of disclosure and audits: require independent third‑party evaluation and capability registries to reduce information frictions and enable better private and public decision‑making.
- Investment and risk management:
- Investors should price increased tail risk and uncertainty around timelines; portfolios may need reweighting toward firms with demonstrable governance and transparency practices.
- Startups and non-frontier firms may have niche opportunities (specialized deployments, distillation services), but face competitive pressure from internalized ASARA at large labs.
- Labor-market transition dynamics:
- Short-to-medium term: upskilling toward oversight, safety engineering, and interdisciplinary governance roles; possible displacement of junior research roles.
- Long-term: need to model equilibrium impacts on wages for top AI talent (likely higher) and broader labor market effects as automation of R&D propagates.
- Research priorities for AI economists:
- Quantify the trade-offs firms face when allocating scarce compute between internal ASARA R&D and external products; model equilibrium with differing firm objectives.
- Design market and regulatory mechanisms that align private incentives with social risk mitigation (e.g., mandated transparency with liability rules).
- Model how information asymmetry from internal deployments affects social welfare and investment patterns.
Suggested immediate policy-relevant actions (consistent with participant preferences): - Implement transparency requirements (model registries, capability disclosures, third‑party audits). - Expand public compute and funding for independent research to reduce asymmetries. - Explore conditional subsidies or tax incentives that reward verified safe and shared deployments rather than secret internalization. - Prioritize models of governance that target observability and coordination rather than blunt international bans likely to be hard to enforce.
If you want, I can: (a) extract direct anonymized exemplar quotes for use in policy briefs; (b) draft a short policy memo translating these findings into concrete regulatory proposals; or (c) produce an economic model sketch quantifying internalization vs. public deployment tradeoffs.
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Many leading AI researchers expect AI development to exceed the transformative impact of all previous technological revolutions. Innovation Output | positive | expectation that AI development will be more transformative than prior technological revolutions |
Reading fidelity
high
Study strength
low
|
n=25
|
| This belief (that AI will exceed prior technological revolutions) is based on the idea that AI will be able to automate the process of AI research itself, leading to a positive feedback loop. Research Productivity | positive | attribution of transformative potential to automation of AI research (recursive improvement) |
Reading fidelity
high
Study strength
medium
|
n=25
|
| In August and September of 2025, we interviewed 25 leading researchers from frontier AI labs and academia, including participants from Google DeepMind, OpenAI, Anthropic, Meta, UC Berkeley, Princeton, and Stanford to understand researcher perspectives on these scenarios. Other | null_result | study sample and data-collection procedure (interviews conducted and institutional affiliations represented) |
Reading fidelity
high
Study strength
high
|
n=25
|
| Though AI systems have not yet been able to recursively improve, 20 of the 25 researchers interviewed identified automating AI research as one of the most severe and urgent AI risks. Ai Safety And Ethics | negative | perceived severity/urgency of automating AI research as an AI risk |
Reading fidelity
high
Study strength
medium
|
n=25
20/25 participants
|
| Participants converged on predictions that AI agents will become more capable at coding, math and eventually AI development, gradually transitioning from 'assistants' or 'tools' to 'autonomous AI developers,' after which point, predictions diverge. Research Productivity | positive | predicted capability trajectory of AI agents (coding, math, AI development; transition from assistants to autonomous developers) |
Reading fidelity
high
Study strength
medium
|
n=25
|
| Researchers agreed upon the possibility of recursive improvement, but they disagreed on basic questions of timelines or appropriate governance mechanisms. Governance And Regulation | mixed | agreement on possibility of recursive improvement and disagreement on timelines/governance |
Reading fidelity
high
Study strength
medium
|
n=25
|
| An epistemic divide emerged between frontier lab researchers and academic researchers, the latter of which expressed more skepticism about explosive growth scenarios. Research Productivity | negative | difference in skepticism about explosive growth scenarios between frontier-lab and academic researchers |
Reading fidelity
high
Study strength
medium
|
n=25
|
| 17/25 participants expected AI systems with advanced coding or R&D capabilities to be increasingly reserved for internal use at AI companies or governments, unseen by the public. Adoption Rate | negative | expectation of restricted/internal access to advanced AI systems |
Reading fidelity
high
Study strength
medium
|
n=25
17/25 participants
|
| Participants were split as to whether setting regulatory 'red lines' was a good idea. Governance And Regulation | mixed | support vs. opposition among participants for regulatory 'red lines' |
Reading fidelity
high
Study strength
medium
|
n=25
|
| Almost all participants favored transparency-based mitigations. Governance And Regulation | positive | endorsement of transparency-based mitigation measures |
Reading fidelity
high
Study strength
medium
|
n=25
|