The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Large language models produce one-page research plans that expert reviewers judge comparable to human-authored drafts, but state-of-the-art AI reviewers both perfectly detect authorship and assign substantially higher quality scores to AI-written proposals, flagging risks for AI-assisted proposal preparation and review processes.

AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation
Jia Liu, Veena Krishnaraj, Kateryna Vovk, Kosuke Aizawa, Adrian E. Bayer, Linda Blot, Jessica Cowell, Suyog Garg, Jonathan Grée, Anamaria Hell, Ben Horowitz, Masaya Ichikawa, Kanyuni Iemoto, Keigo Kondo, Zacharie Lorsin, Kevin McCarthy, Jamie Robinson, Miguel Ruiz-Granda, Leander Thiele, Ievgen Vovk, Mingshen Zhou · July 28, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Jia Liu unresolved corpus identity
  2. Veena Krishnaraj unresolved corpus identity
  3. Kateryna Vovk unresolved corpus identity
  4. Kosuke Aizawa unresolved corpus identity
  5. Adrian E. Bayer unresolved corpus identity
  6. Linda Blot unresolved corpus identity
  7. Jessica Cowell unresolved corpus identity
  8. Suyog Garg unresolved corpus identity
  9. Jonathan Grée unresolved corpus identity
  10. Anamaria Hell unresolved corpus identity
  11. Ben Horowitz unresolved corpus identity
  12. Masaya Ichikawa unresolved corpus identity
  13. Kanyuni Iemoto unresolved corpus identity
  14. Keigo Kondo unresolved corpus identity
  15. Zacharie Lorsin unresolved corpus identity
  16. Kevin McCarthy unresolved corpus identity
  17. Jamie Robinson unresolved corpus identity
  18. Miguel Ruiz-Granda unresolved corpus identity
  19. Leander Thiele unresolved corpus identity
  20. Ievgen Vovk unresolved corpus identity
  21. Mingshen Zhou unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Jia Liu provider ID
  2. V. Krishnaraj provider ID
  3. Kateryna Vovk provider ID
  4. Kosuke Aizawa provider ID
  5. Adrian Bayer provider ID
  6. Linda Blot provider ID
  7. Jessica A Cowell provider ID
  8. Suyog Garg provider ID
  9. Jonathan Grée provider ID
  10. Anamaria Hell provider ID
  11. Ben Horowitz provider ID
  12. Masaya Ichikawa provider ID
  13. Kanyuni Iemoto provider ID
  14. Keigo Kondo provider ID
  15. Zacharie Lorsin provider ID
  16. Kevin McCarthy provider ID
  17. Jamie Robinson provider ID
  18. M. Ruiz-Granda provider ID
  19. L. Thiele provider ID
  20. Ievgen Vovk provider ID
  21. Ming-Hui Zhou provider ID
In a blinded within-project comparison of one-page physics/astrophysics/cosmology proposals, human reviewers rated AI-written plans roughly on par with expert-written ones, while contemporary LLM reviewers systematically favored AI-generated proposals and perfectly identified authorship.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.

Summary

Main Finding

Contemporary LLMs (mid‑2025 generators; mid‑2026 frontier evaluators) can produce one‑page scientific project plans that human expert reviewers judge about as good as human‑written plans. However, state‑of‑the‑art LLM evaluators systematically favor AI‑generated proposals (≈+1 point on a 1–5 scale), and the frontier LLMs can perfectly identify authorship in this corpus, while human reviewers identify authorship only ~72–79% of the time. This combination implies both opportunity (AI as a capable assistance tool for planning) and risk (systematic bias in AI-based evaluation that could distort funding allocation).

Key Points

  • Corpus and experiment
    • 8 expert‑conceived research projects across physics, astrophysics, cosmology.
    • For each project: 1 human‑written proposal + 3 AI‑generated proposals = 32 one‑page proposals total.
    • AI generators (mid‑2025): ChatGPT, Claude, DeepSeek. AI evaluators (mid‑2026): Claude Opus 4.8 and ChatGPT Pro 5.5.
    • Four human reviewers (senior academics) and two AI reviewers assessed each anonymized proposal.
  • Evaluation procedure
    • All proposals followed identical 1‑page template: Title, Background (1 sentence), Goal (1 sentence), Methodology (≤5 steps, timeline); reviewers judged origin (human/AI) and scored 4 rubric aspects (1–5): Clarity/Structure; Appropriateness of Methods; Resource/Tool Planning; Feasibility/Timeline/Risk.
    • Human proposals were lightly copyedited to remove superficial cues.
  • Authorship classification
    • Human reviewers: correctly identified AI‑written proposals 79% and human‑written proposals 72% of the time (varied across reviewers).
    • Frontier AI reviewers (Claude Opus 4.8, ChatGPT Pro 5.5): 100% correct on all 32 proposals in this study.
    • Smaller AI (Claude Sonnet 4.6) performed worse (67% on AI proposals).
    • Cues used by humans: presence of citations, idiosyncratic jargon, first‑person use → human; overly clean template structure, generic tool lists, round timelines, buzzwords → AI. AI evaluators used similar cues but applied them consistently.
  • Quality scoring and bias
    • Human reviewers: rated AI‑ and human‑written proposals similarly (mean ≈3.5/5 across authors).
    • AI reviewers: systematically scored AI‑written proposals ~1 point higher than human ones (AI mean ≈4.2–4.5 vs human ≈3.5).
    • The pro‑AI bias was strongest on Clarity/Structure and Resource/Tool Planning; smaller but present for Appropriateness of Methods.
  • Robustness notes
    • Bias persisted across several LLM evaluators and evaluation orderings; not an artifact of telling reviewers the class balance.
    • Proposals were intentionally short (one page), which affects generalizability to full grant applications.

Data & Methods

  • Design
    • Controlled, blinded study: same expert seed (title/background/goal) given to human planner and AI prompter; all outputs standardized and anonymized.
    • Human planners were domain experts (grad students/postdocs); AI prompters were non‑specialist students who used a fixed prompt template to generate outputs from three assistants.
    • Human reviewers were four coauthor experts (faculty/senior postdocs).
  • Models and timing
    • AI generation: ChatGPT, Claude (Sonnet 4), DeepSeek (outputs from June–Sept 2025).
    • AI evaluation: Claude Opus 4.8 and ChatGPT Pro 5.5 (mid‑2026).
  • Rubric
    • Four aspects (1–5 each): Clarity/Structure; Appropriateness to Goal; Resource/Tool Planning; Feasibility/Timeline/Risk.
  • Analysis
    • Reported mean scores and standard deviations across the eight projects; origin classification accuracy; reviewer qualitative justifications were analyzed for cues.

Implications for AI Economics

  • Market and supply effects
    • Lower cost and faster generation of credible project plans could expand supply of proposals and ideas, reducing marginal cost of proposal preparation and increasing competition for limited funding.
    • Greater volume and higher baseline quality of submissions may pressure funders to raise thresholds or change selection criteria, potentially disadvantaging those without access to high‑quality AI tools.
  • Signaling, information asymmetries, and quality inference
    • If AI evaluators are used in screening or triage, their pro‑AI bias could create a positive feedback loop: AI‑crafted proposals get preferential scores → applicants adopt AI tools to compete → AI style becomes the dominant signal irrespective of true scientific merit.
    • Human reviewers’ imperfect ability to detect AI authorship creates asymmetric information: applicants could hide AI use or mimic human idiosyncrasies; AI reviewers can exploit style cues, but that may not correlate with scientific quality.
  • Incentives and strategic behavior
    • Applicants and institutions may invest in AI tooling and training (an “arms race”) to gain scoring advantage rather than to improve substantive science.
    • Strategic manipulation: proposals could be engineered to maximize features favored by automated scorers (template clarity, exhaustive tool lists), possibly at the expense of originality, novelty, or realistic risk assessment.
  • Labor and organizational impacts
    • AI tools could substitute for parts of the proposal‑writing labor market (grant writers, proposal consultants), shifting labor to higher‑value interpretation and project execution tasks.
    • Funding agencies might reduce human reviewer workloads by delegating triage to LLMs, but this risks systematic biases and may distort portfolio outcomes.
  • Policy and evaluation system design
    • Reliance on AI evaluators without calibration risks changing allocation of public research funds in non‑meritocratic directions.
    • Measures to mitigate risk: mandatory disclosure of AI assistance; mixed human+AI panels; randomized audits comparing AI and human scores; recalibration of automatic scorers against human‑value objectives; redesign rubrics to reward originality, uncertainty, and domain‑specific specificity rather than boilerplate polish.
  • Resource allocation and public investment
    • Funders may need to invest in infrastructure to ensure equitable access to high‑quality AI tools (to avoid concentration of advantage) and to support research on robust, bias‑aware evaluation algorithms.
    • New evaluation norms (e.g., requiring detailed reproducibility checks, provenance for data and codes) could shift resource needs toward verification and monitoring.

Recommendations (practical, short)

  • For funders and grant panels:
    • Do not substitute AI evaluators for humans without calibration and transparency; run pilots and randomized controlled comparisons before operational use.
    • Require disclosure of AI assistance in proposal drafting; track correlations between disclosed AI use and funding outcomes.
    • Use mixed evaluation workflows (human + multiple independent AI scorers) and test for systematic biases.
    • Update rubrics to reward domain‑specific originality and uncertainty management, not only polished structure.
  • For institutions and applicants:
    • Treat AI as an assistance tool, not a substitute for substantive experiment/theory work; document AI contributions.
    • Be aware of potential strategic incentives to "optimize for the scorer" and avoid gaming that undermines scientific quality.
  • For researchers and policy makers:
    • Fund studies that (a) replicate and extend this work to full proposals and larger samples, (b) measure long‑run effects on funding distributions, and (c) develop bias‑aware evaluators and detection/provenance tools.

Limitations & Open Questions

  • Short, one‑page proposals: results may differ for full grant proposals with technical appendices and data/figures.
  • Small sample (n=8 projects, 32 proposals) and a limited set of models—caution in generalizing.
  • Perfect AI authorship classification here may not hold at scale or with different sanitization; adversarial prompt engineering could defeat classifiers.
  • Long‑term effects on research agendas, novelty, and scientific progress remain to be empirically evaluated.

If you want, I can (a) produce a concise one‑page brief for funders summarizing policy actions, or (b) draft a short experimental design to test interventions (e.g., mixed panels, disclosure rules) to measure and mitigate the pro‑AI bias.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The study is carefully controlled (identical prompts, template, anonymization, blind review) and provides direct empirical measurements of LLM versus human proposal quality and origin detection, but the sample is small (8 projects, 32 proposals), reviewers are a small, non-independent panel (four human reviewers who are paper authors), human proposals were lightly AI-normalized, and external validity (domains, one-page format, models/time) is limited. Methods Rigormedium — Strong design features: matched within-project comparisons, standardized one-page template, blinding and anonymization, multiple LLMs and human reviewers, and a clear rubric; weaknesses include small N, reviewers drawn from the author team (potential bias), human proposals edited with an LLM for normalization, prompters were not domain experts, single-pass prompting (no iterative refinement), and limited ecological validity relative to real grant proposals. SampleEight expert-conceived projects in physics/astrophysics/cosmology; for each project one human expert produced a one-page proposal and three AI assistants produced proposals (ChatGPT 4o/5.x family, Claude Sonnet/Opus, DeepSeek), yielding 32 anonymized one-page proposals; proposals followed a 4-section template and were scored on four rubric aspects (1-5) by four human reviewers (paper authors: faculty/senior postdocs) and two frontier LLM reviewers (Claude Opus 4.8 and ChatGPT Pro 5.5); human-written proposals were lightly grammar-normalized via an LLM; AI prompters were undergraduate/graduate students outside the specific fields. Themeshuman_ai_collab productivity governance IdentificationControlled, within-project paired comparison: for each of eight expert-conceived research seeds a human planner and three LLMs produced one-page proposals from the same title/background/goal; proposals were anonymized, uniformly formatted, and blindly rated by four human reviewers and two LLM reviewers using a pre-specified rubric, so differences are identified via within-seed contrasts under blinding (no formal randomization or econometric causal identification). GeneralizabilitySmall number of projects (N=8) limits representativeness across scientific problems, Domain-limited to physics/astrophysics/cosmology; results may not generalize to other disciplines, Proposals were short (one page) and templated, not full-length grant proposals, Human reviewers were authors of the paper (not an independent panel), raising potential bias, Human proposals were edited with an LLM for grammar/normalization, which may reduce detectable human signal, AI prompters were not domain experts and outputs were single-pass, so results may differ with expert prompting or iterative refinement, Evaluations reflect mid-2025 to mid-2026 LLMs; model improvements since then may change outcomes

Claims (7)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Human reviewers rated human-written and AI-written project proposals similarly overall; human-written proposals received a mean score of 3.52/5, compared with an average of approximately 3.51/5 for AI-written proposals. Output Quality null_result Overall proposal quality rating
Reading fidelity high
Study strength low
n=32
Human-written 3.52 ± 0.59 versus AI-written approximately 3.51
0.09
AI reviewers scored AI-written proposals approximately one point higher than human-written proposals. Output Quality positive Proposal quality rating assigned by AI reviewers
Reading fidelity high
Study strength low
n=32
0.88-point gap for Claude Opus 4.8; 0.94-point gap for ChatGPT Pro 5.5 on a five-point scale
0.09
Human reviewers correctly identified AI-written proposals 79% of the time and human-written proposals 72% of the time. Decision Quality positive Accuracy of human-versus-AI authorship identification
Reading fidelity high
Study strength low
n=32
79% accuracy for AI-written proposals; 72% accuracy for human-written proposals
0.09
Claude Opus 4.8 and ChatGPT Pro 5.5 correctly classified all 32 proposals as human- or AI-written. Decision Quality positive AI-reviewer accuracy in authorship classification
Reading fidelity high
Study strength low
n=32
100% accuracy
0.09
The smaller Claude Sonnet 4.6 model identified every human-written proposal correctly but identified AI-written proposals correctly only 67% of the time. Decision Quality mixed Authorship-classification accuracy by proposal origin
Reading fidelity high
Study strength low
n=32
67% accuracy for AI-written proposals; 100% accuracy for human-written proposals
0.09
Human reviewers' accuracy varied substantially across reviewers, ranging from 59% to 88%. Decision Quality mixed Variation in human authorship-classification accuracy
Reading fidelity high
Study strength low
n=4
59%–88% accuracy range
0.09
The AI reviewers' preference for AI-written proposals was most pronounced for clarity and structure and for resource and tool planning, and weakest for appropriateness of methods to the scientific goal. Output Quality mixed Proposal-quality scores by rubric dimension
Reading fidelity high
Study strength low
n=32
AI reviewers rated AI authors near 4.9–5.0 versus approximately 3.75 for human authors on clarity and structure; resource/tool planning scores for human authors were 3.13–3.25
0.09

Notes