The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A Gemini-based AI substantially speeds complex scientific synthesis: 13 climate scientists produced a 79-paper synthesis in roughly 46 person-hours with most AI-generated content retained; yet expert oversight remained essential, and humans supplied much of the final rigor and content.

AI-Assisted Scientific Assessment: A Case Study on Climate Change
Buck, Christian, Caesar, Levke, Huebscher, Michelle Chen, Ciaramita, Massimiliano, Fischer, Erich M., Hausfather, Zeke, Tokmak, Özge Kart, Knutti, Reto, Leippold, Markus, Ludescher, Joseph, Mach, Katharine J., Corner, Sofia Palazzo, Shahi, Kasra Rafiezadeh, Rockström, Johan, Rogelj, Joeri, Sakschewski, Boris · February 10, 2026 · arXiv (Cornell University)
openalex descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Buck, Christian provider ID
  2. Caesar, Levke provider ID
  3. Huebscher, Michelle Chen provider ID
  4. Ciaramita, Massimiliano provider ID
  5. Fischer, Erich M. provider ID
  6. Hausfather, Zeke provider ID
  7. Tokmak, Özge Kart provider ID
  8. Knutti, Reto provider ID
  9. Leippold, Markus provider ID
  10. Ludescher, Joseph provider ID
  11. Mach, Katharine J. provider ID
  12. Corner, Sofia Palazzo provider ID
  13. Shahi, Kasra Rafiezadeh provider ID
  14. Rockström, Johan provider ID
  15. Rogelj, Joeri provider ID
  16. Sakschewski, Boris provider ID

Semantic Scholar

Latest observation:

  1. Christian Buck provider ID
  2. L. Caesar provider ID
  3. Michelle Chen Huebscher provider ID
  4. Massimiliano Ciaramita provider ID
  5. Erich Fischer provider ID
  6. Z. Hausfather provider ID
  7. Ö. Tokmak provider ID
  8. R. Knutti provider ID
  9. Markus Leippold provider ID
  10. J. Ludescher provider ID
  11. Katharine Mach provider ID
  12. Sofia Palazzo Corner provider ID
  13. Kasra Rafiezadeh Shahi provider ID
  14. Johan Rockström provider ID
  15. J. Rogelj provider ID
  16. B. Sakschewski provider ID
In a 13-person trial using a Gemini-based AI integrated into workflow, AI substantially accelerated a complex literature synthesis—producing a 79-paper report over 104 revisions in ~46 person-hours with most AI text retained—while experts provided critical oversight and authored a sizeable portion of the final report.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The emerging paradigm of AI co-scientists focuses on tasks characterized by repeatable verification, where agents explore search spaces in 'guess and check' loops. This paradigm does not extend to problems where repeated evaluation is impossible and ground truth is established by the consensus synthesis of theory and existing evidence. We evaluate a Gemini-based AI environment designed to support collaborative scientific assessment, integrated into a standard scientific workflow. In collaboration with a diverse group of 13 scientists working in the field of climate science, we tested the system on a complex topic: the stability of the Atlantic Meridional Overturning Circulation (AMOC). Our results show that AI can accelerate the scientific workflow. The group produced a comprehensive synthesis of 79 papers through 104 revision cycles in just over 46 person-hours. AI contribution was significant: most AI-generated content was retained in the report. AI also helped maintain logical consistency and presentation quality. However, expert additions were crucial to ensure its acceptability: less than half of the report was produced by AI. Furthermore, substantial oversight was required to expand and elevate the content to rigorous scientific standards.

Summary

Main Finding

An integrated Gemini-based AI Assistant substantially accelerated a verification-poor scientific assessment (AMOC stability) while leaving human expertise central to rigor. In a 5-week, 13-expert case study the hybrid workflow produced an 8,000-word synthesis of 79 papers through 104 revision cycles in ~46.5 logged person-hours (≈3.5 h/author). Participants estimated the same task typically takes 100–200 person-hours; all reported at least 2–3× speedups (38% estimated ≥7×). Most AI-proposed content was retained, but scientists contributed the majority (58.3%) of final content and were essential to elevate the draft to rigorous assessment. The study introduces simple, sentence-level metrics (Retention and Intervention) and embedding-based convergence measures to quantify human–AI co-authoring dynamics.

Key Points

  • Task and setting

    • Topic: Atlantic Meridional Overturning Circulation (AMOC) stability — deliberately verification-poor and contested.
    • Participants: 13 climate scientists with diverse subfield backgrounds and institutional-report experience.
    • Corpus: curated AMOC collection (~1,660 papers) plus search via OpenAlex and web.
    • Tooling: Gemini LLMs (Gemini 2.5 Pro then Gemini 3 Pro Preview) via Vertex AI embedded into an editor + evidence panel + chatbox workflow.
  • Workflow and outcomes

    • Three phases: outline → section drafting/curation → final collation and sign-off.
    • Output: ~8,000-word report synthesizing 79 papers; 104 version iterations; 46:33 total logged hours.
    • Efficiency: Participants judged substantial speedups (≥2–3× common; some much larger).
  • Quantitative contributions

    • Human vs AI content: scientists authored ~58.3% of final content; AI contributed substantially and much AI text was retained.
    • Revision dynamics: AI-generated revisions were used ~3× more often than manual revisions; AI credited for ~25% of revisions content.
    • Convergence: embedding-based similarity between intermediate versions and final version showed monotonic convergence with identifiable drafting → refinement → finalization stages.
  • Evaluation metrics & methods

    • Sentence-level alignment used to define:
      • Retention (precision-like): fraction of AI-proposed sentences/themes present in final text.
      • Intervention (recall-oriented / 1 − recall): fraction of final content that is new human-authored material.
    • Document convergence measured via embeddings (gemini-embedding-001) and bounded inverse Euclidean distance.
  • Limitations & failure modes

    • LLM issues: hallucinations (e.g., citation numbering), difficulty synthesizing quantitative material, sycophancy/obsequiousness under adversarial prompting, and limited holistic reasoning (identifying caveats).
    • Study limits: single case study, proprietary tool, model switch during experiment (Gemini 2.5 → 3 Pro) — potential confound for pure model-performance attribution.

Data & Methods

  • Data

    • Precompiled AMOC corpus (~1,660 papers) converted to Markdown; assistant had access to this corpus and to external search/OpenAlex metadata.
    • Evidence panel entries included auto-generated summary snippets, bibliographic metadata, and retrieval scores (BM25 + dense retrieval).
    • Final report references: 79 papers synthesized, final reference list consolidated in Phase 3.
  • Models & system

    • LLMs: Gemini 2.5 Pro (Phase 1) and Gemini 3 Pro Preview (Phases 2–3).
    • Embeddings: gemini-embedding-001 for semantic similarity/convergence measures.
    • Retrieval: BM25 and dense retrieval (Karpukhin-style) over document chunks; evidence selection iterated during drafting.
  • Workflow mechanics

    • Editor supported manual and AI edits; AI edits could be triggered at document- or text-span-level.
    • Assistant functions: draft generation, source selection/summarization, argumentative/rhetorical review (Toulmin + discourse frameworks), conflict resolution suggestions, version-history summaries.
    • Human curation: experts edited, expanded, checked numeric and conceptual claims, and performed final sign-off.
  • Evaluation & metrics

    • Time: logged-in elapsed time with a timeout parameter; interpreted as approximate effort.
    • Convergence: sequence of document embeddings → bounded inverse Euclidean distance to final version; step-wise similarity between successive versions.
    • Sentence alignment: many-to-many matching between draft sentences (D) and final sentences (F) used to compute Retention = |D_M|/|D| and Intervention = 1 − |F_M|/|F| (where M is alignment set).

Implications for AI Economics

  • Productivity shock in expert synthesis tasks

    • The case suggests a meaningful productivity gain in verification-poor, high-expertise tasks (reports, assessments, meta-analyses). If generalizable, lower effective costs and faster turnaround for consensus products (e.g., intergovernmental assessments, institutional white papers) are plausible.
    • Effect on labor demand: complementary shift — less time on first-pass drafting, more demand for high-skill curatorial, evaluative, and adjudicative labor (expert oversight, cross-validation, methodological critique). Routine drafting tasks could be partially automated, reducing low-level drafting hours but raising value of verification/credibility roles.
  • Valuation and attribution

    • The paper’s sentence-level Retention and Intervention metrics provide a practical framework to apportion credit/value between human and AI contributions. These can inform payment, authorship conventions, and productivity accounting in institutions and funding agencies.
    • Firm-level advantages: organizations with privileged access to high-performing LLMs and curated corpora may capture disproportionate productivity gains, increasing concentration in research-production capacity.
  • Institutional and market effects

    • Faster, cheaper assessments could increase the supply of syntheses and policy-relevant reports, altering competition among consultancies, NGOs, and academic groups. This may compress timelines for policy cycles and reshape consultancy pricing.
    • Reputation and trust capital become more valuable: because AI can produce fluent but potentially flawed drafts, demand will grow for high-reliability human validators and for institutional certifications that a synthesis was rigorously vetted.
  • Risks, externalities and regulation

    • Quality and signaling: greater throughput risks dilution of scrutiny if oversight is weak; markets may discount outputs unless provenance, traceability, and human-signoff conventions are standardized.
    • Bias and informational cascades: AI-assisted convergence toward consensus could amplify stylistic or evidentiary homogeneity (echo chambers). Economic models of information aggregation should account for reduced diversity of initial drafts when many actors use similar LLMs.
    • Access inequality: proprietary models and curated datasets create barriers; unequal access can generate asymmetric productivity gains across institutions and countries.
  • Metrics for economic assessment and policy

    • Use embedding-based convergence and sentence-level attribution to:
      • Measure productivity gains attributable to AI at project and sector levels.
      • Inform compensation models that allocate value between human expertise and AI tooling.
      • Monitor quality externalities (e.g., rate of hallucination corrections per unit time) as a regulatory metric.
  • Open questions for economic research

    • Generalizability: How do observed gains scale across domains and larger teams? Are speedups similar in other verification-poor fields (macroeconomics, epidemiology, legal synthesis)?
    • Labor reallocation dynamics: What is the long-run impact on demand for mid-skill vs high-skill scientific labor? Will labor markets bifurcate into AI-supervisory roles vs model operators?
    • Market structure: How will concentration of access to high-quality LLMs affect competition among research producers and the pricing of assessment services?

Summary conclusion The study demonstrates a practical hybrid model where LLMs materially increase efficiency in producing high-level, verification-poor scientific assessments, while retaining the necessity of human expertise for rigor. For AI economics, this implies a productivity shock concentrated in drafting-intensive expert goods, leading to complementarities that raise the value of human verification and institutional trust mechanisms, create new metrics for attributing AI value, and produce distributional and regulatory challenges that merit further empirical and theoretical study.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Provides direct, measured outputs (time, revision cycles, retained AI content) from a real-world collaborative trial, showing that AI materially contributed to and accelerated a complex synthesis; however, the study lacks a counterfactual or control group, has a small non-random sample (13 domain experts), and tests a single task/domain, limiting causal claims and external validity. Methods Rigormedium — Systematic measurement of workflow metrics (papers synthesized, revision cycles, person-hours, retention of AI content) and use of a diverse expert group are strengths, but rigor is limited by no randomized or matched comparison, no pre-registered protocol reported here, potential selection and observer biases, and qualitative judgments about 'acceptability' and 'rigor' that required substantial expert intervention. SampleA convenience sample of 13 climate scientists collaborated using a Gemini-based AI environment integrated into a standard scientific workflow to synthesize literature on the stability of the Atlantic Meridional Overturning Circulation (AMOC); the group produced a synthesis covering 79 papers across 104 revision cycles in ~46 person-hours, with metrics on AI-generated content retention and expert edits recorded. Themesproductivity human_ai_collab GeneralizabilitySmall sample of 13 participants limits statistical generalizability, Single substantive domain (climate science, AMOC) — results may not transfer to other scientific fields, Single AI system (Gemini-based environment) — findings may not hold for other models or tool designs, Task-specific: literature synthesis with repeatable verification, not open-ended theoretical discovery or experimental design, Potential selection bias: participants likely self-selected and may be unusually collaborative or AI-savvy, No control/comparison arm to separate AI effects from group process or novelty effects

Claims (9)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The emerging paradigm of AI co-scientists focuses on tasks characterized by repeatable verification, where agents explore search spaces in 'guess and check' loops. Research Productivity positive applicability of AI co-scientist paradigm to repeatable-verification tasks
Reading fidelity high
Study strength speculative
not reported
0.03
This paradigm does not extend to problems where repeated evaluation is impossible and ground truth is established by the consensus synthesis of theory and existing evidence. Research Productivity negative applicability of AI co-scientist paradigm to problems requiring consensus synthesis
Reading fidelity high
Study strength speculative
not reported
0.03
We evaluated a Gemini-based AI environment designed to support collaborative scientific assessment in collaboration with a diverse group of 13 scientists working in the field of climate science. Other null_result study implementation / evaluation (deployment with N=13 participants)
Reading fidelity high
Study strength medium
n=13
0.18
AI can accelerate the scientific workflow. Task Completion Time positive task completion time / acceleration of workflow
Reading fidelity high
Study strength medium
n=13
0.18
The group produced a comprehensive synthesis of 79 papers through 104 revision cycles in just over 46 person-hours. Research Productivity positive research productivity (papers synthesized, revision cycles, person-hours)
Reading fidelity high
Study strength medium
n=13
79 papers through 104 revision cycles in just over 46 person-hours
0.18
AI contribution was significant: most AI-generated content was retained in the report. Output Quality positive output quality / acceptability of AI-generated content (retention in final report)
Reading fidelity medium
Study strength medium
most AI-generated content was retained
0.11
AI also helped maintain logical consistency and presentation quality. Output Quality positive presentation quality and logical consistency of the report
Reading fidelity medium
Study strength low
n=13
0.05
Expert additions were crucial to ensure acceptability: less than half of the report was produced by AI. Task Allocation negative share of report content produced by AI (contribution proportion)
Reading fidelity high
Study strength medium
less than half of the report was produced by AI
0.18
Substantial oversight was required to expand and elevate the content to rigorous scientific standards. Output Quality negative degree of oversight required to achieve scientific rigor / output quality
Reading fidelity high
Study strength low
n=13
0.09

Notes