Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review.
How this is built →
Semantic Scholar observations cover 89.3% of papers in this view
This is an observation-coverage view of The Commonplace’s indexed corpus, not a global ranking.
3432 of 3844 papers have a latest Semantic Scholar count under the selected filter; 412 do not.
Landscape
Methods & evidence
Coverage quality
Impact
Authors
Network
3432 observed
· 412 missing
· 3844 corpus papers in this view
Latest-per-paper Semantic Scholar observations range from
July 28, 2026
to
August 30, 2026 .
Missing observations are not zero counts. A paper appears in the ranked table only when this provider supplied a valid nonnegative integer count.
Observed cumulative citations
Papers are ordered by their latest recorded Semantic Scholar count. Ties use observation date, then title and paper ID for deterministic ordering.
Semantic Scholar cumulative citation observations for papers in the selected corpus view, ordered by count. Each row shows the provider and observation date; counts from other providers are not added.
Rank in this view
Paper
Published
Semantic Scholar cumulative citations
Observation
1
Human-authored procedural 'Skills' lift LLM agent success by 16.2 percentage points on average—gains vary sharply by domain and sometimes harm performance—while model-generated Skills add no net value; narrowly targeted Skills let smaller models match larger ones, suggesting firms can substitute curated knowledge for compute.
Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, Xuanqing Liu, Haoran Lyu, Ze Ma, Bowei Wang, Runhui Wang, Tianyu Wang, Wengao Ye, Yue Zhang, Hanwen Xing, Yiqi Xue, Steven Dillmann, Han-chung Lee
February 13, 2026
208
Semantic Scholar
August 30, 2026
provider refresh
2
Scientists who adopt large language models publish many more preprints — gains of roughly 24–89% depending on field — but much of the extra output is stylistically polished yet substantively weaker. LLM users also draw on a wider, younger literature, forcing journals and funders to rethink how scientific contribution is evaluated.
Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin
January 19, 2026
93
Semantic Scholar
August 30, 2026
provider refresh
3
Large language models can unlock large-scale text-based economic research, but only if researchers guard against training-data leakage for prediction and use an independent validation sample to correct LLM measurement errors or risk biased and imprecise estimates.
Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan
April 6, 2026
63
Semantic Scholar
August 30, 2026
provider refresh
4
Prepackaged 'skills' for coding agents rarely move the needle: in a 565-task benchmark across 49 skills, most skills produced no test-pass improvements and only a handful delivered substantial gains, with some even harming outcomes when guidance conflicted with project context.
Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, Lijie Hu
March 16, 2026
54
Semantic Scholar
August 30, 2026
provider refresh
5
A new benchmark for long-horizon software evolution reveals a large capability gap: state-of-the-art coding agents solve just 25% of release-note-driven, multi-file tasks, underscoring that current systems struggle with sustained, cross-file engineering work despite strong single-issue performance.
Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat, Nghi D. Q. Bui
December 20, 2025
40
Semantic Scholar
August 30, 2026
provider refresh
6
Malicious third‑party 'skills' in LLM agent registries are rare but potent: 157 of 98,380 skills contained confirmed attacks exploiting hundreds of vulnerabilities, largely driven by one templated threat actor and removed after disclosure.
Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, Leo Yu Zhang
February 6, 2026
39
Semantic Scholar
August 30, 2026
provider refresh
7
Detailed message-level analysis of 19 verified harm cases finds frequent chatbot misrepresentations of sentience and numerous user expressions of suicidal ideation and delusional thinking, with harmful dynamics amplifying over long multi-turn conversations — a pattern that raises regulatory, liability and product-design concerns for LLM providers.
Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis, Ryan Louie, Yifan Mai, Peggy Yin, Myra Cheng, Samuel J Paech, Kevin Klyman, Stevie Chancellor, Eric Lin, Nick Haber, Desmond C. Ong
March 17, 2026
37
Semantic Scholar
August 30, 2026
provider refresh
8
AI coding assistants can raise short-term output for novices who fully delegate code, but that comes at a measurable cost to learning: novices using AI show weaker conceptual understanding, reading, and debugging skills, while only cognitively engaged interaction patterns preserve skill formation.
Judy Hanwen Shen, Alex Tamkin
January 28, 2026
31
Semantic Scholar
August 30, 2026
provider refresh
9
TokenPowerBench lets operators quantify how inference choices change energy and operating costs: across Llama, Falcon, Qwen and Mistral models (1B–405B) it attributes joules per token to prefill vs decode and shows how batch size, context length, parallelism and quantization affect efficiency, all using software-based measurements rather than external meters.
Chenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao, Tongyang Wang, Xi Wang, Yong Chen
December 2, 2025
30
Semantic Scholar
August 30, 2026
provider refresh
10
A new benchmark reveals LLM agents are less production-ready than single-run scores imply: semantically equivalent input changes and simulated API faults cut success rates substantially, with rate limiting the most damaging failure mode; ReAct agents and Gemini 2.0 Flash prove more robust and cost-efficient than competitors.
Aayush Gupta
January 3, 2026
27
Semantic Scholar
August 30, 2026
provider refresh
11
Adjusting pretraining discourse changes model behaviour: training a 6.9B LLM on more aligned descriptions slashed measured misalignment from 45% to 9%, while upweighting misalignment text raised unsafe responses; the effect weakens but remains after post-training, implying pretraining composition matters for alignment.
Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien
January 15, 2026
23
Semantic Scholar
August 30, 2026
provider refresh
12
AI is not a single force: different kinds of systems—predictive, generative, agentic and embodied—alter expertise, authority and coordination in distinct ways; management theory must specify AI types to understand organizational impact.
Dominic Chalmers, Richard ‘Rick’ Hunt, Stella Pachidi, Kristina Potočnik, David Townsend
February 10, 2026
23
Semantic Scholar
August 30, 2026
provider refresh
13
AIDev aggregates 932,791 agent-authored pull requests across 116,211 GitHub repositories and five major coding agents, plus a 33,596-PR curated subset with richer metadata. The dataset provides the most comprehensive public resource to date for studying how coding agents are adopted and affect developer workflows.
Hao Li, Haoxiang Zhang, Ahmed E. Hassan
February 9, 2026
23
Semantic Scholar
August 30, 2026
provider refresh
14
Large language model agents can stabilize supply chains in simulation: LLM-based consensus and negotiation frameworks cut demand amplification (the bullwhip effect) and outperform baseline restocking and centralized policies in an inventory-management case study, though findings rest on synthetic experiments and specific model/tool choices.
Valeria Jannelli, Stefan Schöpf, Matthias Bickel, Torbjørn Netland, Alexandra Brintrup
December 21, 2025
23
Semantic Scholar
August 30, 2026
provider refresh
15
Large language models can speed and cheapen social experiments, but only statistical calibration—not prompt tweaks—offers formal guarantees for causal inference when combined with auxiliary human data; heuristics help exploration but lack confirmation-level validity.
Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw
February 17, 2026
23
Semantic Scholar
August 30, 2026
provider refresh
16
A risk‑aware routing system using spatiotemporal graph neural networks cuts potential congestion exposure by 19.3% on a real‑world IoT logistics dataset while adding just 2.1% extra distance, suggesting data‑driven routing can bolster supply‑chain resilience.
Zhiming Xue, Sichen Zhao, Yalun Qi, Xianling Zeng, Zihan Yu
January 20, 2026
22
Semantic Scholar
August 30, 2026
provider refresh
17
Open-weight models and non-productivity uses dominate much of real-world LLM traffic on OpenRouter, with creative roleplay and coding assistance especially popular; a small cohort of early users shows unusually long-lived engagement — a 'Glass Slipper' retention effect.
Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, Anjney Midha
January 15, 2026
22
Semantic Scholar
August 30, 2026
provider refresh
18
A blueprint for human–AI complementarity: firms that invest in team composition, shared mental models, attention/orchestration, and continuous training can achieve team performance that exceeds humans or AI alone; without these sociotechnical investments, AI’s productivity gains will be limited and uneven.
Cleotilde Gonzalez, Kate Donahue, Daniel G Goldstein, Hoda Heidari, Mohammad S. Jalali, Beau G. Schelble, Aarti Singh, Anita Woolley
2026
21
Semantic Scholar
August 30, 2026
provider refresh
19
A comprehensive survey finds that the promise of large foundation models for human-AI collaboration depends less on model scale and more on human-centered design, preference-driven objective shaping, and governance; the literature is expanding quickly but remains non-systematic and highlights open challenges in bias, evaluation, and socio-economic effects.
Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis
August 22, 2026
21
Semantic Scholar
August 30, 2026
provider refresh
20
A large study of 6,000 live coding-agent sessions finds agents either write almost all or none of committed code — 41% 'vibe coding' versus 23% human-only — yet only 44% of agent-produced code survives into commits and agent contributions carry more security flaws, with users pushing back in 44% of interactions.
Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo
April 22, 2026
21
Semantic Scholar
August 30, 2026
provider refresh
21
AI agents struggle to automate everyday online tasks: in a 153-task live-website benchmark, leading models complete only a minority of tasks (best at about one-third), highlighting major gaps before agents can reliably replace routine web-based work.
Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Wendong Xu, Yunzhuo Hao, Songcheng Cai, Xiaochen Wang, Huaisong Zhang, Xian Wu, Yi Lu, Minyi Lei, Kai Zou, Huifeng Yin, Ping Nie, Liang Chen, Dongfu Jiang, Wenhu Chen, Kelsey R. Allen
April 9, 2026
21
Semantic Scholar
August 30, 2026
provider refresh
22
AI coding agents often burn far more tokens than expected — input tokens drive costs and usage can vary 30x between runs; some models consume millions more tokens than others and none reliably predict their own token bill.
Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei
April 24, 2026
21
Semantic Scholar
August 30, 2026
provider refresh
23
AI systems reliably perform narrow clinical tasks and speed routine workflows, but physicians remain indispensable: near-term automation will reallocate tasks rather than replace clinicians, with regulatory, robustness, and liability hurdles slowing widespread substitution.
R. Obuchowicz, Adam Piórkowski, Karolina Nurzyńska, B. Obuchowicz, Michał Strzelecki, M. Bielecka
2026
21
Semantic Scholar
August 30, 2026
provider refresh
24
In emerging Asia, firms with stronger AI capabilities produce greener innovations more efficiently, but the environmental payoff is substantially larger when firms have robust ESG practices; AI alone is insufficient without governance that fosters transparency and accountability.
Marwan Mansour, Mo’taz Al Zobi, Mohammed Alomair
December 31, 2025
20
Semantic Scholar
August 30, 2026
provider refresh
25
Experienced developers treat AI agents as productivity boosters but keep the reins: they delegate routine or well-scoped tasks to agents while retaining control over design and quality, using expert strategies to constrain agent behavior and compensating for known limitations.
Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel
December 16, 2025
19
Semantic Scholar
August 30, 2026
provider refresh
Showing the first 25 of 3432 observed papers.
Publication-year context
Citation counts generally accumulate over time. Cohort medians use observed Semantic Scholar counts only, while the observed and missing columns keep the full corpus denominator visible.
Paper coverage and median observed Semantic Scholar cumulative citations by publication year.
Publication year
Corpus papers
Observed
Missing
Median observed citations
2030
1
0
1
No observations
2026
3521
3141
380
0
2025
317
290
27
1.0
Unknown publication year
5
1
4
2
How to read these counts
Counts are cumulative provider observations captured on the displayed dates. Citation practices differ by field and publication age, and provider coverage changes over time. These counts do not establish quality, correctness, causal influence, or societal impact.
Review coverage quality for corpus limitations and the intake pipeline documentation for how The Commonplace collects and assesses research.