Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review.
How this is built →
Semantic Scholar observations cover 89.8% of papers in this view
This is an observation-coverage view of The Commonplace’s indexed corpus, not a global ranking.
3353 of 3734 papers have a latest Semantic Scholar count under the selected filter; 381 do not.
Landscape
Methods & evidence
Coverage quality
Impact
Authors
Network
3353 observed
· 381 missing
· 3734 corpus papers in this view
Latest-per-paper Semantic Scholar observations range from
August 8, 2026
to
September 5, 2026 .
Missing observations are not zero counts. A paper appears in the ranked table only when this provider supplied a valid nonnegative integer count.
Observed cumulative citations
Papers are ordered by their latest recorded Semantic Scholar count. Ties use observation date, then title and paper ID for deterministic ordering.
Semantic Scholar cumulative citation observations for papers in the selected corpus view, ordered by count. Each row shows the provider and observation date; counts from other providers are not added.
Rank in this view
Paper
Published
Semantic Scholar cumulative citations
Observation
1
Human-authored procedural 'Skills' lift LLM agent success by 16.2 percentage points on average—gains vary sharply by domain and sometimes harm performance—while model-generated Skills add no net value; narrowly targeted Skills let smaller models match larger ones, suggesting firms can substitute curated knowledge for compute.
Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, Xuanqing Liu, Haoran Lyu, Ze Ma, Bowei Wang, Runhui Wang, Tianyu Wang, Wengao Ye, Yue Zhang, Hanwen Xing, Yiqi Xue, Steven Dillmann, Han-chung Lee
February 13, 2026
219
Semantic Scholar
September 5, 2026
provider refresh
2
Scientists who adopt large language models publish many more preprints — gains of roughly 24–89% depending on field — but much of the extra output is stylistically polished yet substantively weaker. LLM users also draw on a wider, younger literature, forcing journals and funders to rethink how scientific contribution is evaluated.
Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin
January 19, 2026
96
Semantic Scholar
September 5, 2026
provider refresh
3
Large language models can unlock large-scale text-based economic research, but only if researchers guard against training-data leakage for prediction and use an independent validation sample to correct LLM measurement errors or risk biased and imprecise estimates.
Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan
April 6, 2026
65
Semantic Scholar
September 5, 2026
provider refresh
4
Prepackaged 'skills' for coding agents rarely move the needle: in a 565-task benchmark across 49 skills, most skills produced no test-pass improvements and only a handful delivered substantial gains, with some even harming outcomes when guidance conflicted with project context.
Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, Lijie Hu
March 16, 2026
59
Semantic Scholar
September 5, 2026
provider refresh
5
Detailed message-level analysis of 19 verified harm cases finds frequent chatbot misrepresentations of sentience and numerous user expressions of suicidal ideation and delusional thinking, with harmful dynamics amplifying over long multi-turn conversations — a pattern that raises regulatory, liability and product-design concerns for LLM providers.
Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis, Ryan Louie, Yifan Mai, Peggy Yin, Myra Cheng, Samuel J Paech, Kevin Klyman, Stevie Chancellor, Eric Lin, Nick Haber, Desmond C. Ong
March 17, 2026
38
Semantic Scholar
September 5, 2026
provider refresh
6
Malicious third‑party 'skills' in LLM agent registries are rare but potent: 157 of 98,380 skills contained confirmed attacks exploiting hundreds of vulnerabilities, largely driven by one templated threat actor and removed after disclosure.
Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, Leo Yu Zhang
February 6, 2026
38
Semantic Scholar
September 5, 2026
provider refresh
7
AI coding assistants can raise short-term output for novices who fully delegate code, but that comes at a measurable cost to learning: novices using AI show weaker conceptual understanding, reading, and debugging skills, while only cognitively engaged interaction patterns preserve skill formation.
Judy Hanwen Shen, Alex Tamkin
January 28, 2026
32
Semantic Scholar
September 5, 2026
provider refresh
8
A new benchmark reveals LLM agents are less production-ready than single-run scores imply: semantically equivalent input changes and simulated API faults cut success rates substantially, with rate limiting the most damaging failure mode; ReAct agents and Gemini 2.0 Flash prove more robust and cost-efficient than competitors.
Aayush Gupta
January 3, 2026
28
Semantic Scholar
September 5, 2026
provider refresh
9
A large study of 6,000 live coding-agent sessions finds agents either write almost all or none of committed code — 41% 'vibe coding' versus 23% human-only — yet only 44% of agent-produced code survives into commits and agent contributions carry more security flaws, with users pushing back in 44% of interactions.
Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo
April 22, 2026
26
Semantic Scholar
September 5, 2026
provider refresh
10
Large language models can speed and cheapen social experiments, but only statistical calibration—not prompt tweaks—offers formal guarantees for causal inference when combined with auxiliary human data; heuristics help exploration but lack confirmation-level validity.
Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw
February 17, 2026
26
Semantic Scholar
September 5, 2026
provider refresh
11
AI is not a single force: different kinds of systems—predictive, generative, agentic and embodied—alter expertise, authority and coordination in distinct ways; management theory must specify AI types to understand organizational impact.
Dominic Chalmers, Richard ‘Rick’ Hunt, Stella Pachidi, Kristina Potočnik, David Townsend
February 10, 2026
25
Semantic Scholar
September 5, 2026
provider refresh
12
AIDev aggregates 932,791 agent-authored pull requests across 116,211 GitHub repositories and five major coding agents, plus a 33,596-PR curated subset with richer metadata. The dataset provides the most comprehensive public resource to date for studying how coding agents are adopted and affect developer workflows.
Hao Li, Haoxiang Zhang, Ahmed E. Hassan
February 9, 2026
25
Semantic Scholar
September 5, 2026
provider refresh
13
A risk‑aware routing system using spatiotemporal graph neural networks cuts potential congestion exposure by 19.3% on a real‑world IoT logistics dataset while adding just 2.1% extra distance, suggesting data‑driven routing can bolster supply‑chain resilience.
Zhiming Xue, Sichen Zhao, Yalun Qi, Xianling Zeng, Zihan Yu
January 20, 2026
24
Semantic Scholar
September 5, 2026
provider refresh
14
Adjusting pretraining discourse changes model behaviour: training a 6.9B LLM on more aligned descriptions slashed measured misalignment from 45% to 9%, while upweighting misalignment text raised unsafe responses; the effect weakens but remains after post-training, implying pretraining composition matters for alignment.
Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien
January 15, 2026
24
Semantic Scholar
September 5, 2026
provider refresh
15
AI systems reliably perform narrow clinical tasks and speed routine workflows, but physicians remain indispensable: near-term automation will reallocate tasks rather than replace clinicians, with regulatory, robustness, and liability hurdles slowing widespread substitution.
R. Obuchowicz, Adam Piórkowski, Karolina Nurzyńska, B. Obuchowicz, Michał Strzelecki, M. Bielecka
2026
23
Semantic Scholar
September 5, 2026
provider refresh
16
Open-weight models and non-productivity uses dominate much of real-world LLM traffic on OpenRouter, with creative roleplay and coding assistance especially popular; a small cohort of early users shows unusually long-lived engagement — a 'Glass Slipper' retention effect.
Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, Anjney Midha
January 15, 2026
22
Semantic Scholar
September 5, 2026
provider refresh
17
A blueprint for human–AI complementarity: firms that invest in team composition, shared mental models, attention/orchestration, and continuous training can achieve team performance that exceeds humans or AI alone; without these sociotechnical investments, AI’s productivity gains will be limited and uneven.
Cleotilde Gonzalez, Kate Donahue, Daniel G Goldstein, Hoda Heidari, Mohammad S. Jalali, Beau G. Schelble, Aarti Singh, Anita Woolley
2026
21
Semantic Scholar
September 5, 2026
provider refresh
18
A comprehensive survey finds that the promise of large foundation models for human-AI collaboration depends less on model scale and more on human-centered design, preference-driven objective shaping, and governance; the literature is expanding quickly but remains non-systematic and highlights open challenges in bias, evaluation, and socio-economic effects.
Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis
August 22, 2026
21
Semantic Scholar
September 5, 2026
provider refresh
19
AI agents struggle to automate everyday online tasks: in a 153-task live-website benchmark, leading models complete only a minority of tasks (best at about one-third), highlighting major gaps before agents can reliably replace routine web-based work.
Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Wendong Xu, Yunzhuo Hao, Songcheng Cai, Xiaochen Wang, Huaisong Zhang, Xian Wu, Yi Lu, Minyi Lei, Kai Zou, Huifeng Yin, Ping Nie, Liang Chen, Dongfu Jiang, Wenhu Chen, Kelsey R. Allen
April 9, 2026
21
Semantic Scholar
September 5, 2026
provider refresh
20
AI coding agents often burn far more tokens than expected — input tokens drive costs and usage can vary 30x between runs; some models consume millions more tokens than others and none reliably predict their own token bill.
Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
April 24, 2026
21
Semantic Scholar
September 5, 2026
provider refresh
21
Firms that build AI-enabled dynamic capabilities report better performance largely because AI fosters a data-driven mindset that institutionalizes learning; however, too much reliance on data weakens managers' flexibility and reduces marginal gains.
Hassan Samih Ayoub, Joshua Chibuike Sopuru
January 23, 2026
21
Semantic Scholar
September 5, 2026
provider refresh
22
Large language models subtly but systematically change what people mean when they write: heavy LLM users produce nearly 70% more neutral answers and report less creative, less ‘in‑their‑voice’ prose. When asked to revise human essays or write peer reviews, LLMs frequently alter semantics and give reviews that weight clarity/significance less and score papers roughly one point higher on average.
Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Leibo, Max Kleiman-Weiner, Natasha Jaques
March 18, 2026
20
Semantic Scholar
September 5, 2026
provider refresh
23
A new large-scale benchmark finds proprietary autonomous agents outperform open-source counterparts on complex, long-horizon real-world tasks, while exposing wide variation in resource efficiency, self-correction, and tool use — underscoring the need to co-design models and agent frameworks.
Keyu Li, Junhao Shi, Yang Xiao, Mohan Jiang, Jie Sun, Yunze Wu, Dayuan Fu, Shijie Xia, Xiaojie Cai, Tianze Xu, Weiye Si, Wenjie Li, Dequan Wang, Pengfei Liu
January 16, 2026
19
Semantic Scholar
September 5, 2026
provider refresh
24
A new benchmark of complex, cross-application professional tasks finds leading AI agents complete under one-quarter of assignments: the best model scores 24% on Pass@1, with most competitors performing substantially worse, highlighting large remaining gaps for real-world professional productivity.
Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, Lucas Rothman, Marco Burstein, Julien Benchek, David Ostrofsky, Anirudh Ravichandran, Debnil Sur, Neel Venugopal, Alannah Hsia, Isaac Robinson, Calix Huang, Olivia Varones, Daniyal Khan, Michael Haines, Austin Bridges, Jesse Boyle, Koby Twist, Zach Richards, Chirag Mahapatra, Brendan Foody, Osvald Nitski
January 20, 2026
18
Semantic Scholar
September 5, 2026
provider refresh
25
Autonomous Gemini-based agents can generate, train and deploy recommendation-model improvements at YouTube, reportedly speeding development and delivering successful production launches; evidence is promising but drawn from proprietary internal evaluations without transparent causal tests.
Haochen Wang, Yi Wu, Daryl Chang, Li Wei, Lukasz Heldt
February 10, 2026
18
Semantic Scholar
September 5, 2026
provider refresh
Showing the first 25 of 3353 observed papers.
Publication-year context
Citation counts generally accumulate over time. Cohort medians use observed Semantic Scholar counts only, while the observed and missing columns keep the full corpus denominator visible.
Paper coverage and median observed Semantic Scholar cumulative citations by publication year.
Publication year
Corpus papers
Observed
Missing
Median observed citations
2026
3734
3353
381
0
How to read these counts
Counts are cumulative provider observations captured on the displayed dates. Citation practices differ by field and publication age, and provider coverage changes over time. These counts do not establish quality, correctness, causal influence, or societal impact.
Review coverage quality for corpus limitations and the intake pipeline documentation for how The Commonplace collects and assesses research.