Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review.
How this is built →
2Distinct papers
27Unique collaborators
2/2Semantic Scholar citation coverage
Publication span: 2026. Corpus fetch span: 2026.
Identity provenance
Provider IDs
- Semantic Scholar:
2064013115
ORCID evidence
No valid ORCID is stored.
Observed aliases (1)
- Arvind Narayanan (semantic scholar, provider refresh)
Topics and outcomes in this view
Assessment themes
- Productivity: 2 papers
- Adoption: 1 paper
- Human Ai Collab: 1 paper
Claim outcomes
- Other: 1 paper
- Research Productivity: 1 paper
- Task Completion Time: 1 paper
- Developer Productivity: 1 paper
Papers in the Semantic Scholar view
Latest stored Semantic Scholar author observations only. Citation counts below are from the same provider and are not combined with other services.
Scroll the table horizontally to see every column.
| Paper | Author evidence | Date | Provider citations |
|---|---|---|---|
| Human–AI collaboration roughly halves the time required to reproduce scientific code in a randomized trial, even after benchmark accuracy has saturated; improving the benchmark and adding an OOD suite reveals construct-validity and robustness issues that accuracy alone would miss.arxiv | Arvind Narayanan provider id |
2026-06-23 | 0 |
| Open-world tests find real-world AI reach: in a CRUX pilot, an AI agent built and published a simple iOS app with only one avoidable human intervention, indicating that messy, long-horizon evaluations can reveal capabilities that standard benchmarks understate.arxiv | Arvind Narayanan provider id |
2026-05-19 | 3 |
Citation observation summary
Semantic Scholar supplied counts for 2 of 2 papers in this view; 0 are missing. The observed paper counts sum to 3 cumulative citations. This is a coverage summary, not an author score or h-index.