0 cumulative citations
View corpus contextLarge language models produce one-page research plans that expert reviewers judge comparable to human-authored drafts, but state-of-the-art AI reviewers both perfectly detect authorship and assign substantially higher quality scores to AI-written proposals, flagging risks for AI-assisted proposal preparation and review processes.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.
Summary
Main Finding
Contemporary LLMs (mid‑2025 generators; mid‑2026 frontier evaluators) can produce one‑page scientific project plans that human expert reviewers judge about as good as human‑written plans. However, state‑of‑the‑art LLM evaluators systematically favor AI‑generated proposals (≈+1 point on a 1–5 scale), and the frontier LLMs can perfectly identify authorship in this corpus, while human reviewers identify authorship only ~72–79% of the time. This combination implies both opportunity (AI as a capable assistance tool for planning) and risk (systematic bias in AI-based evaluation that could distort funding allocation).
Key Points
- Corpus and experiment
- 8 expert‑conceived research projects across physics, astrophysics, cosmology.
- For each project: 1 human‑written proposal + 3 AI‑generated proposals = 32 one‑page proposals total.
- AI generators (mid‑2025): ChatGPT, Claude, DeepSeek. AI evaluators (mid‑2026): Claude Opus 4.8 and ChatGPT Pro 5.5.
- Four human reviewers (senior academics) and two AI reviewers assessed each anonymized proposal.
- Evaluation procedure
- All proposals followed identical 1‑page template: Title, Background (1 sentence), Goal (1 sentence), Methodology (≤5 steps, timeline); reviewers judged origin (human/AI) and scored 4 rubric aspects (1–5): Clarity/Structure; Appropriateness of Methods; Resource/Tool Planning; Feasibility/Timeline/Risk.
- Human proposals were lightly copyedited to remove superficial cues.
- Authorship classification
- Human reviewers: correctly identified AI‑written proposals 79% and human‑written proposals 72% of the time (varied across reviewers).
- Frontier AI reviewers (Claude Opus 4.8, ChatGPT Pro 5.5): 100% correct on all 32 proposals in this study.
- Smaller AI (Claude Sonnet 4.6) performed worse (67% on AI proposals).
- Cues used by humans: presence of citations, idiosyncratic jargon, first‑person use → human; overly clean template structure, generic tool lists, round timelines, buzzwords → AI. AI evaluators used similar cues but applied them consistently.
- Quality scoring and bias
- Human reviewers: rated AI‑ and human‑written proposals similarly (mean ≈3.5/5 across authors).
- AI reviewers: systematically scored AI‑written proposals ~1 point higher than human ones (AI mean ≈4.2–4.5 vs human ≈3.5).
- The pro‑AI bias was strongest on Clarity/Structure and Resource/Tool Planning; smaller but present for Appropriateness of Methods.
- Robustness notes
- Bias persisted across several LLM evaluators and evaluation orderings; not an artifact of telling reviewers the class balance.
- Proposals were intentionally short (one page), which affects generalizability to full grant applications.
Data & Methods
- Design
- Controlled, blinded study: same expert seed (title/background/goal) given to human planner and AI prompter; all outputs standardized and anonymized.
- Human planners were domain experts (grad students/postdocs); AI prompters were non‑specialist students who used a fixed prompt template to generate outputs from three assistants.
- Human reviewers were four coauthor experts (faculty/senior postdocs).
- Models and timing
- AI generation: ChatGPT, Claude (Sonnet 4), DeepSeek (outputs from June–Sept 2025).
- AI evaluation: Claude Opus 4.8 and ChatGPT Pro 5.5 (mid‑2026).
- Rubric
- Four aspects (1–5 each): Clarity/Structure; Appropriateness to Goal; Resource/Tool Planning; Feasibility/Timeline/Risk.
- Analysis
- Reported mean scores and standard deviations across the eight projects; origin classification accuracy; reviewer qualitative justifications were analyzed for cues.
Implications for AI Economics
- Market and supply effects
- Lower cost and faster generation of credible project plans could expand supply of proposals and ideas, reducing marginal cost of proposal preparation and increasing competition for limited funding.
- Greater volume and higher baseline quality of submissions may pressure funders to raise thresholds or change selection criteria, potentially disadvantaging those without access to high‑quality AI tools.
- Signaling, information asymmetries, and quality inference
- If AI evaluators are used in screening or triage, their pro‑AI bias could create a positive feedback loop: AI‑crafted proposals get preferential scores → applicants adopt AI tools to compete → AI style becomes the dominant signal irrespective of true scientific merit.
- Human reviewers’ imperfect ability to detect AI authorship creates asymmetric information: applicants could hide AI use or mimic human idiosyncrasies; AI reviewers can exploit style cues, but that may not correlate with scientific quality.
- Incentives and strategic behavior
- Applicants and institutions may invest in AI tooling and training (an “arms race”) to gain scoring advantage rather than to improve substantive science.
- Strategic manipulation: proposals could be engineered to maximize features favored by automated scorers (template clarity, exhaustive tool lists), possibly at the expense of originality, novelty, or realistic risk assessment.
- Labor and organizational impacts
- AI tools could substitute for parts of the proposal‑writing labor market (grant writers, proposal consultants), shifting labor to higher‑value interpretation and project execution tasks.
- Funding agencies might reduce human reviewer workloads by delegating triage to LLMs, but this risks systematic biases and may distort portfolio outcomes.
- Policy and evaluation system design
- Reliance on AI evaluators without calibration risks changing allocation of public research funds in non‑meritocratic directions.
- Measures to mitigate risk: mandatory disclosure of AI assistance; mixed human+AI panels; randomized audits comparing AI and human scores; recalibration of automatic scorers against human‑value objectives; redesign rubrics to reward originality, uncertainty, and domain‑specific specificity rather than boilerplate polish.
- Resource allocation and public investment
- Funders may need to invest in infrastructure to ensure equitable access to high‑quality AI tools (to avoid concentration of advantage) and to support research on robust, bias‑aware evaluation algorithms.
- New evaluation norms (e.g., requiring detailed reproducibility checks, provenance for data and codes) could shift resource needs toward verification and monitoring.
Recommendations (practical, short)
- For funders and grant panels:
- Do not substitute AI evaluators for humans without calibration and transparency; run pilots and randomized controlled comparisons before operational use.
- Require disclosure of AI assistance in proposal drafting; track correlations between disclosed AI use and funding outcomes.
- Use mixed evaluation workflows (human + multiple independent AI scorers) and test for systematic biases.
- Update rubrics to reward domain‑specific originality and uncertainty management, not only polished structure.
- For institutions and applicants:
- Treat AI as an assistance tool, not a substitute for substantive experiment/theory work; document AI contributions.
- Be aware of potential strategic incentives to "optimize for the scorer" and avoid gaming that undermines scientific quality.
- For researchers and policy makers:
- Fund studies that (a) replicate and extend this work to full proposals and larger samples, (b) measure long‑run effects on funding distributions, and (c) develop bias‑aware evaluators and detection/provenance tools.
Limitations & Open Questions
- Short, one‑page proposals: results may differ for full grant proposals with technical appendices and data/figures.
- Small sample (n=8 projects, 32 proposals) and a limited set of models—caution in generalizing.
- Perfect AI authorship classification here may not hold at scale or with different sanitization; adversarial prompt engineering could defeat classifiers.
- Long‑term effects on research agendas, novelty, and scientific progress remain to be empirically evaluated.
If you want, I can (a) produce a concise one‑page brief for funders summarizing policy actions, or (b) draft a short experimental design to test interventions (e.g., mixed panels, disclosure rules) to measure and mitigate the pro‑AI bias.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Human reviewers rated human-written and AI-written project proposals similarly overall; human-written proposals received a mean score of 3.52/5, compared with an average of approximately 3.51/5 for AI-written proposals. Output Quality | null_result | Overall proposal quality rating |
Reading fidelity
high
Study strength
low
|
n=32
Human-written 3.52 ± 0.59 versus AI-written approximately 3.51
|
| AI reviewers scored AI-written proposals approximately one point higher than human-written proposals. Output Quality | positive | Proposal quality rating assigned by AI reviewers |
Reading fidelity
high
Study strength
low
|
n=32
0.88-point gap for Claude Opus 4.8; 0.94-point gap for ChatGPT Pro 5.5 on a five-point scale
|
| Human reviewers correctly identified AI-written proposals 79% of the time and human-written proposals 72% of the time. Decision Quality | positive | Accuracy of human-versus-AI authorship identification |
Reading fidelity
high
Study strength
low
|
n=32
79% accuracy for AI-written proposals; 72% accuracy for human-written proposals
|
| Claude Opus 4.8 and ChatGPT Pro 5.5 correctly classified all 32 proposals as human- or AI-written. Decision Quality | positive | AI-reviewer accuracy in authorship classification |
Reading fidelity
high
Study strength
low
|
n=32
100% accuracy
|
| The smaller Claude Sonnet 4.6 model identified every human-written proposal correctly but identified AI-written proposals correctly only 67% of the time. Decision Quality | mixed | Authorship-classification accuracy by proposal origin |
Reading fidelity
high
Study strength
low
|
n=32
67% accuracy for AI-written proposals; 100% accuracy for human-written proposals
|
| Human reviewers' accuracy varied substantially across reviewers, ranging from 59% to 88%. Decision Quality | mixed | Variation in human authorship-classification accuracy |
Reading fidelity
high
Study strength
low
|
n=4
59%–88% accuracy range
|
| The AI reviewers' preference for AI-written proposals was most pronounced for clarity and structure and for resource and tool planning, and weakest for appropriateness of methods to the scientific goal. Output Quality | mixed | Proposal-quality scores by rubric dimension |
Reading fidelity
high
Study strength
low
|
n=32
AI reviewers rated AI authors near 4.9–5.0 versus approximately 3.75 for human authors on clarity and structure; resource/tool planning scores for human authors were 3.13–3.25
|