The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Semantic Scholar observations cover 89.3% of papers in this view

This is an observation-coverage view of The Commonplace’s indexed corpus, not a global ranking. 3432 of 3844 papers have a latest Semantic Scholar count under the selected filter; 412 do not.

Clear

Observed cumulative citations

Papers are ordered by their latest recorded Semantic Scholar count. Ties use observation date, then title and paper ID for deterministic ordering.

Semantic Scholar cumulative citation observations for papers in the selected corpus view, ordered by count. Each row shows the provider and observation date; counts from other providers are not added.
Rank in this view Paper Published Semantic Scholar cumulative citations Observation
1 Human-authored procedural 'Skills' lift LLM agent success by 16.2 percentage points on average—gains vary sharply by domain and sometimes harm performance—while model-generated Skills add no net value; narrowly targeted Skills let smaller models match larger ones, suggesting firms can substitute curated knowledge for compute. Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, Xuanqing Liu, Haoran Lyu, Ze Ma, Bowei Wang, Runhui Wang, Tianyu Wang, Wengao Ye, Yue Zhang, Hanwen Xing, Yiqi Xue, Steven Dillmann, Han-chung Lee 208 Semantic Scholar

provider refresh
2 Scientists who adopt large language models publish many more preprints — gains of roughly 24–89% depending on field — but much of the extra output is stylistically polished yet substantively weaker. LLM users also draw on a wider, younger literature, forcing journals and funders to rethink how scientific contribution is evaluated. Keigo Kusumegi, Xinyu Yang, Paul Ginsparg, Mathijs de Vaan, Toby Stuart, Yian Yin 93 Semantic Scholar

provider refresh
3 Large language models can unlock large-scale text-based economic research, but only if researchers guard against training-data leakage for prediction and use an independent validation sample to correct LLM measurement errors or risk biased and imprecise estimates. Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan 63 Semantic Scholar

provider refresh
4 Prepackaged 'skills' for coding agents rarely move the needle: in a 565-task benchmark across 49 skills, most skills produced no test-pass improvements and only a handful delivered substantial gains, with some even harming outcomes when guidance conflicted with project context. Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, Lijie Hu 54 Semantic Scholar

provider refresh
5 A new benchmark for long-horizon software evolution reveals a large capability gap: state-of-the-art coding agents solve just 25% of release-note-driven, multi-file tasks, underscoring that current systems struggle with sustained, cross-file engineering work despite strong single-issue performance. Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat, Nghi D. Q. Bui 40 Semantic Scholar

provider refresh
6 Malicious third‑party 'skills' in LLM agent registries are rare but potent: 157 of 98,380 skills contained confirmed attacks exploiting hundreds of vulnerabilities, largely driven by one templated threat actor and removed after disclosure. Yi Liu, Zhihao Chen, Yanjun Zhang, Gelei Deng, Yuekang Li, Jianting Ning, Leo Yu Zhang 39 Semantic Scholar

provider refresh
7 Detailed message-level analysis of 19 verified harm cases finds frequent chatbot misrepresentations of sentience and numerous user expressions of suicidal ideation and delusional thinking, with harmful dynamics amplifying over long multi-turn conversations — a pattern that raises regulatory, liability and product-design concerns for LLM providers. Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis, Ryan Louie, Yifan Mai, Peggy Yin, Myra Cheng, Samuel J Paech, Kevin Klyman, Stevie Chancellor, Eric Lin, Nick Haber, Desmond C. Ong 37 Semantic Scholar

provider refresh
8 AI coding assistants can raise short-term output for novices who fully delegate code, but that comes at a measurable cost to learning: novices using AI show weaker conceptual understanding, reading, and debugging skills, while only cognitively engaged interaction patterns preserve skill formation. Judy Hanwen Shen, Alex Tamkin 31 Semantic Scholar

provider refresh
9 TokenPowerBench lets operators quantify how inference choices change energy and operating costs: across Llama, Falcon, Qwen and Mistral models (1B–405B) it attributes joules per token to prefill vs decode and shows how batch size, context length, parallelism and quantization affect efficiency, all using software-based measurements rather than external meters. Chenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao, Tongyang Wang, Xi Wang, Yong Chen 30 Semantic Scholar

provider refresh
10 A new benchmark reveals LLM agents are less production-ready than single-run scores imply: semantically equivalent input changes and simulated API faults cut success rates substantially, with rate limiting the most damaging failure mode; ReAct agents and Gemini 2.0 Flash prove more robust and cost-efficient than competitors. Aayush Gupta 27 Semantic Scholar

provider refresh
11 Adjusting pretraining discourse changes model behaviour: training a 6.9B LLM on more aligned descriptions slashed measured misalignment from 45% to 9%, while upweighting misalignment text raised unsafe responses; the effect weakens but remains after post-training, implying pretraining composition matters for alignment. Cameron Tice, Puria Radmard, Samuel Ratnam, Andy Kim, David Africa, Kyle O'Brien 23 Semantic Scholar

provider refresh
12 AI is not a single force: different kinds of systems—predictive, generative, agentic and embodied—alter expertise, authority and coordination in distinct ways; management theory must specify AI types to understand organizational impact. Dominic Chalmers, Richard ‘Rick’ Hunt, Stella Pachidi, Kristina Potočnik, David Townsend 23 Semantic Scholar

provider refresh
13 AIDev aggregates 932,791 agent-authored pull requests across 116,211 GitHub repositories and five major coding agents, plus a 33,596-PR curated subset with richer metadata. The dataset provides the most comprehensive public resource to date for studying how coding agents are adopted and affect developer workflows. Hao Li, Haoxiang Zhang, Ahmed E. Hassan 23 Semantic Scholar

provider refresh
14 Large language model agents can stabilize supply chains in simulation: LLM-based consensus and negotiation frameworks cut demand amplification (the bullwhip effect) and outperform baseline restocking and centralized policies in an inventory-management case study, though findings rest on synthetic experiments and specific model/tool choices. Valeria Jannelli, Stefan Schöpf, Matthias Bickel, Torbjørn Netland, Alexandra Brintrup 23 Semantic Scholar

provider refresh
15 Large language models can speed and cheapen social experiments, but only statistical calibration—not prompt tweaks—offers formal guarantees for causal inference when combined with auxiliary human data; heuristics help exploration but lack confirmation-level validity. Jessica Hullman, David Broska, Huaman Sun, Aaron Shaw 23 Semantic Scholar

provider refresh
16 A risk‑aware routing system using spatiotemporal graph neural networks cuts potential congestion exposure by 19.3% on a real‑world IoT logistics dataset while adding just 2.1% extra distance, suggesting data‑driven routing can bolster supply‑chain resilience. Zhiming Xue, Sichen Zhao, Yalun Qi, Xianling Zeng, Zihan Yu 22 Semantic Scholar

provider refresh
17 Open-weight models and non-productivity uses dominate much of real-world LLM traffic on OpenRouter, with creative roleplay and coding assistance especially popular; a small cohort of early users shows unusually long-lived engagement — a 'Glass Slipper' retention effect. Malika Aubakirova, Alex Atallah, Chris Clark, Justin Summerville, Anjney Midha 22 Semantic Scholar

provider refresh
18 A blueprint for human–AI complementarity: firms that invest in team composition, shared mental models, attention/orchestration, and continuous training can achieve team performance that exceeds humans or AI alone; without these sociotechnical investments, AI’s productivity gains will be limited and uneven. Cleotilde Gonzalez, Kate Donahue, Daniel G Goldstein, Hoda Heidari, Mohammad S. Jalali, Beau G. Schelble, Aarti Singh, Anita Woolley 2026 21 Semantic Scholar

provider refresh
19 A comprehensive survey finds that the promise of large foundation models for human-AI collaboration depends less on model scale and more on human-centered design, preference-driven objective shaping, and governance; the literature is expanding quickly but remains non-systematic and highlights open challenges in bias, evaluation, and socio-economic effects. Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis 21 Semantic Scholar

provider refresh
20 A large study of 6,000 live coding-agent sessions finds agents either write almost all or none of committed code — 41% 'vibe coding' versus 23% human-only — yet only 44% of agent-produced code survives into commits and agent contributions carry more security flaws, with users pushing back in 44% of interactions. Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo 21 Semantic Scholar

provider refresh
21 AI agents struggle to automate everyday online tasks: in a 153-task live-website benchmark, leading models complete only a minority of tasks (best at about one-third), highlighting major gaps before agents can reliably replace routine web-based work. Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Wendong Xu, Yunzhuo Hao, Songcheng Cai, Xiaochen Wang, Huaisong Zhang, Xian Wu, Yi Lu, Minyi Lei, Kai Zou, Huifeng Yin, Ping Nie, Liang Chen, Dongfu Jiang, Wenhu Chen, Kelsey R. Allen 21 Semantic Scholar

provider refresh
22 AI coding agents often burn far more tokens than expected — input tokens drive costs and usage can vary 30x between runs; some models consume millions more tokens than others and none reliably predict their own token bill. Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei 21 Semantic Scholar

provider refresh
23 AI systems reliably perform narrow clinical tasks and speed routine workflows, but physicians remain indispensable: near-term automation will reallocate tasks rather than replace clinicians, with regulatory, robustness, and liability hurdles slowing widespread substitution. R. Obuchowicz, Adam Piórkowski, Karolina Nurzyńska, B. Obuchowicz, Michał Strzelecki, M. Bielecka 2026 21 Semantic Scholar

provider refresh
24 In emerging Asia, firms with stronger AI capabilities produce greener innovations more efficiently, but the environmental payoff is substantially larger when firms have robust ESG practices; AI alone is insufficient without governance that fosters transparency and accountability. Marwan Mansour, Mo’taz Al Zobi, Mohammed Alomair 20 Semantic Scholar

provider refresh
25 Experienced developers treat AI agents as productivity boosters but keep the reins: they delegate routine or well-scoped tasks to agents while retaining control over design and quality, using expert strategies to constrain agent behavior and compensating for known limitations. Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel 19 Semantic Scholar

provider refresh

Showing the first 25 of 3432 observed papers.

Publication-year context

Citation counts generally accumulate over time. Cohort medians use observed Semantic Scholar counts only, while the observed and missing columns keep the full corpus denominator visible.

Paper coverage and median observed Semantic Scholar cumulative citations by publication year.
Publication year Corpus papers Observed Missing Median observed citations
2030 1 0 1 No observations
2026 3521 3141 380 0
2025 317 290 27 1.0
Unknown publication year 5 1 4 2

How to read these counts

Counts are cumulative provider observations captured on the displayed dates. Citation practices differ by field and publication age, and provider coverage changes over time. These counts do not establish quality, correctness, causal influence, or societal impact.

Review coverage quality for corpus limitations and the intake pipeline documentation for how The Commonplace collects and assesses research.