The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Semantic Scholar observations cover 91.0% of papers in this view

This is an observation-coverage view of The Commonplace’s indexed corpus, not a global ranking. 1585 of 1741 papers have a latest Semantic Scholar count under the selected filter; 156 do not.

Clear

Observed cumulative citations

Papers are ordered by their latest recorded Semantic Scholar count. Ties use observation date, then title and paper ID for deterministic ordering.

Semantic Scholar cumulative citation observations for papers in the selected corpus view, ordered by count. Each row shows the provider and observation date; counts from other providers are not added.
Rank in this view Paper Published Semantic Scholar cumulative citations Observation
1 Human-authored procedural 'Skills' lift LLM agent success by 16.2 percentage points on average—gains vary sharply by domain and sometimes harm performance—while model-generated Skills add no net value; narrowly targeted Skills let smaller models match larger ones, suggesting firms can substitute curated knowledge for compute. Xiangyi Li, Wenbo Chen, Yimin Liu, Shenghan Zheng, Xiaokun Chen, Yifeng He, Yubo Li, Bingran You, Haotian Shen, Jiankai Sun, Shuyi Wang, Binxu Li, Qunhong Zeng, Di Wang, Xuandong Zhao, Yuanli Wang, Roey Ben Chaim, Zonglin Di, Yipeng Gao, Junwei He, Yizhuo He, Liqiang Jing, Luyang Kong, Xin Lan, Jiachen Li, Songlin Li, Yijiang Li, Yueqian Lin, Xinyi Liu, Xuanqing Liu, Haoran Lyu, Ze Ma, Bowei Wang, Runhui Wang, Tianyu Wang, Wengao Ye, Yue Zhang, Hanwen Xing, Yiqi Xue, Steven Dillmann, Han-chung Lee 150 Semantic Scholar

provider refresh
2 Large language models can unlock large-scale text-based economic research, but only if researchers guard against training-data leakage for prediction and use an independent validation sample to correct LLM measurement errors or risk biased and imprecise estimates. Jens Ludwig, Sendhil Mullainathan, Ashesh Rambachan 54 Semantic Scholar

provider refresh
3 Prepackaged 'skills' for coding agents rarely move the needle: in a 565-task benchmark across 49 skills, most skills produced no test-pass improvements and only a handful delivered substantial gains, with some even harming outcomes when guidance conflicted with project context. Tingxu Han, Yi Zhang, Wei Song, Chunrong Fang, Zhenyu Chen, Youcheng Sun, Lijie Hu 36 Semantic Scholar

provider refresh
4 Detailed message-level analysis of 19 verified harm cases finds frequent chatbot misrepresentations of sentience and numerous user expressions of suicidal ideation and delusional thinking, with harmful dynamics amplifying over long multi-turn conversations — a pattern that raises regulatory, liability and product-design concerns for LLM providers. Jared Moore, Ashish Mehta, William Agnew, Jacy Reese Anthis, Ryan Louie, Yifan Mai, Peggy Yin, Myra Cheng, Samuel J Paech, Kevin Klyman, Stevie Chancellor, Eric Lin, Nick Haber, Desmond C. Ong 24 Semantic Scholar

provider refresh
5 AI agents struggle to automate everyday online tasks: in a 153-task live-website benchmark, leading models complete only a minority of tasks (best at about one-third), highlighting major gaps before agents can reliably replace routine web-based work. Yuxuan Zhang, Yubo Wang, Yipeng Zhu, Penghui Du, Junwen Miao, Xuan Lu, Wendong Xu, Yunzhuo Hao, Songcheng Cai, Xiaochen Wang, Huaisong Zhang, Xian Wu, Yi Lu, Minyi Lei, Kai Zou, Huifeng Yin, Ping Nie, Liang Chen, Dongfu Jiang, Wenhu Chen, Kelsey R. Allen 16 Semantic Scholar

provider refresh
6 Large language models subtly but systematically change what people mean when they write: heavy LLM users produce nearly 70% more neutral answers and report less creative, less ‘in‑their‑voice’ prose. When asked to revise human essays or write peer reviews, LLMs frequently alter semantics and give reviews that weight clarity/significance less and score papers roughly one point higher on average. Marwa Abdulhai, Isadora White, Yanming Wan, Ibrahim Qureshi, Joel Leibo, Max Kleiman-Weiner, Natasha Jaques 16 Semantic Scholar

provider refresh
7 A large study of 6,000 live coding-agent sessions finds agents either write almost all or none of committed code — 41% 'vibe coding' versus 23% human-only — yet only 44% of agent-produced code survives into commits and agent contributions carry more security flaws, with users pushing back in 44% of interactions. Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo 15 Semantic Scholar

provider refresh
8 A blueprint for human–AI complementarity: firms that invest in team composition, shared mental models, attention/orchestration, and continuous training can achieve team performance that exceeds humans or AI alone; without these sociotechnical investments, AI’s productivity gains will be limited and uneven. Cleotilde Gonzalez, Kate Donahue, Daniel G Goldstein, Hoda Heidari, Mohammad S. Jalali, Beau G. Schelble, Aarti Singh, Anita Woolley 2026 14 Semantic Scholar

provider refresh
9 AI systems reliably perform narrow clinical tasks and speed routine workflows, but physicians remain indispensable: near-term automation will reallocate tasks rather than replace clinicians, with regulatory, robustness, and liability hurdles slowing widespread substitution. R. Obuchowicz, Adam Piórkowski, Karolina Nurzyńska, B. Obuchowicz, Michał Strzelecki, M. Bielecka 2026 14 Semantic Scholar

provider refresh
10 A survey of 627 studies identifies four distinct AI–human decision-making paradigms driven by AI–human dynamics and decision typologies; the framework helps organizations choose between intuitive, algorithmic, analytical and hybrid decision modes as they integrate AI into operations. Han Li, Feng Tian 12 Semantic Scholar

provider refresh
11 Co-designed quantum–classical supercomputers are the missing link to scale useful hybrid algorithms: integrating QPUs with GPUs/CPUs and middleware can slash manual orchestration and speed materials and drug discovery, but the capital- and skill-intensity will favor well-resourced firms and cloud or national testbeds. Seetharami Seelam, Jerry M. Chow, Antonio Córcoles, Sarah Sheldon, Tushar Mittal, Abhinav Kandala, Sean Dague, Ian Hincks, Hiroshi Horii, Blake Johnson, Michael Le, Hani Jamjoom, Jay M. Gambetta 12 Semantic Scholar

provider refresh
12 A new pipeline turns real software into 10,000+ long-horizon agent environments tied to U.S. occupations, creating CUA-World for evaluating computer-use assistants; distilled 2B vision-language models and reviewer-audit loops yield measurable but modest performance gains on these economically relevant tasks. Pranjal Aggarwal, Graham Neubig, Sean Welleck 11 Semantic Scholar

provider refresh
13 AI agents’ risks hinge on their execution histories, so static prompts and access controls cannot reliably enforce path-dependent rules; firms must evaluate actions at runtime, a change that raises latency, engineering and compliance costs and reshapes markets for governance, insurance and enterprise adoption. Maurits Kaptein, Vassilis-Javed Khan, Andriy Podstavnychy 11 Semantic Scholar

provider refresh
14 AI coding agents often burn far more tokens than expected — input tokens drive costs and usage can vary 30x between runs; some models consume millions more tokens than others and none reliably predict their own token bill. Longju Bai, Zhemin Huang, Xingyao Wang, Jiao Sun, Rada Mihalcea, Erik Brynjolfsson, Alex Pentland, Jiaxin Pei 11 Semantic Scholar

provider refresh
15 Contracts and neutral mediators restore cooperation between powerful LLM agents, while repeated play and reputations often do not; cooperation that appears under repeated interaction collapses when partners change, but mechanisms that enforce contingent payments or delegate decisions sustain cooperative equilibria. Emanuel Tewolde, Xiao Zhang, David Guzman Piedrahita, Vincent Conitzer, Zhijing Jin 11 Semantic Scholar

provider refresh
16 A new Pokemon-based benchmark exposes large capability gaps on multi-agent, partial-observability and long-horizon planning tasks: specialist RL systems and human experts outperform generalist LLMs, and a 20M+ trajectory dataset plus a NeurIPS competition confirm strong community interest and reproducible evaluation. Seth Karten, Jake Grigsby, Tersoo Upaa, Junik Bae, Seonghun Hong, Hyunyoung Jeong, Jaeyoon Jung, Kun Kerdthaisong, Gyungbo Kim, Hyeokgi Kim, Yujin Kim, Eunju Kwon, Dongyu Liu, Patrick Mariglia, Sangyeon Park, Benedikt Schink, Xianwei Shi, Anthony Sistilli, Joseph Twin, Arian Urdu, Matin Urdu, Qiao Wang, Ling Wu, Wenli Zhang, Kunsheng Zhou, Stephanie Milani, Kiran Vodrahalli, Amy Zhang, Fei Fang, Yuke Zhu, Chi Jin 10 Semantic Scholar

provider refresh
17 AI and other new digital skills are appearing in roughly one in ten vacancies in advanced economies and pay a clear wage premium. Yet their spread is linked to sharper labor-market polarization—helping high-skilled workers while hollowing out middle-skilled roles and reducing employment in AI‑exposed occupations with low worker complementarity, with young workers hit especially hard. Florence Jaumotte, Jaden Kim, David Koll, Elmer Li, Longji Li, Giovanni Melina, Alina Song, Marina Mendes Tavares 2026 10 Semantic Scholar

provider refresh
18 An AI pipeline produced a machine-verified mathematical formalization in 10 days under a single supervisor, automating coding and lemma-proving but still requiring expert oversight; the case shows high-skill scientific tasks can be substantially automated in practice, though verification and definition-alignment remain critical bottlenecks. Vasily Ilin 10 Semantic Scholar

provider refresh
19 A predictive prompt-selection method halves (or more) the costly rollouts needed for RL finetuning and improves reasoning accuracy across math, planning and visual-geometry benchmarks; by making iterative finetuning cheaper and faster, the technique could materially lower the marginal compute cost of model improvement for practitioners. Yixiu Mao, Yun Qu, Qi Wang, Heming Zou, Xiangyang Ji 9 Semantic Scholar

provider refresh
20 A simple evaluation-driven scaling strategy (SimpleTES) lets relatively modest LLMs outperform larger baselines across 21 scientific tasks — doubling LASSO speed, cutting quantum gate overhead by 24.5%, and finding new combinatorial constructions — and trajectory-based retraining improves generalization to new problems. Haotian Ye, Haowei Lin, Jingyi Tang, Yizhen Luo, Caiyin Yang, Chang Su, Rahul Thapa, Rui Yang, Ruihua Liu, Zeyu Li, Chong Gao, Dachao Ding, Guangrong He, Miaolei Zhang, Lina Sun, Wenyang Wang, Yuchen Zhong, Zhuohao Shen, Di He, Jianzhu Ma, Stefano Ermon, Tongyang Li, Xiaowen Chu, James Zou, Yuzhi Xu 9 Semantic Scholar

provider refresh
21 Formalizing developer intent into checkable specifications is the linchpin for reliable AI-generated code; without it, abundant code risks being incorrect. The paper sets out a research agenda—specification validation, interactive test-driven workflows, and verified synthesis—to make AI coding tools dependable in practice. Shuvendu K. Lahiri 9 Semantic Scholar

provider refresh
22 Current agent 'memory' is mostly lookup, not learning—this matters: without slow weight-based consolidation, agents cannot develop expertise, face a provable ceiling on generalizing to novel compositions, and remain structurally vulnerable to persistent poisoning; pairing fast exemplar storage with slower consolidation is required. Binyan Xu, Xilin Dai, Kehuan Zhang 8 Semantic Scholar

provider refresh
23 Current AI architectures lack the mechanisms for sustained, autonomous learning; a three-part design that integrates observation, active experimentation and an internal meta-controller could unlock more adaptable, sample-efficient agents and accelerate automation in embodied and social tasks. Emmanuel Dupoux, Yann LeCun, Jitendra Malik 8 Semantic Scholar

provider refresh
24 LLM coding assistants speed up developers and cut routine work, but their effect on code quality and teamwork remains unresolved; most studies are short-term and exploratory, leaving long-run and team-level impacts unclear. Amr Mohamed, Maram Assi, Mariam Guizani 8 Semantic Scholar

provider refresh
25 A live benchmark shows state-of-the-art LLM agents complete at most two-thirds of realistic workflow tasks: the top model passes 66.7% of 105 controlled tasks. Failures cluster in HR, management and multi-system business workflows, indicating end-to-end workflow automation remains far from solved. Chenxin Li, Zhengyang Tang, Huangxin Lin, Yunlong Lin, Shijue Huang, Shengyuan Liu, Bowen Ye, Rang Li, Lei Li, Benyou Wang, Yixuan Yuan 7 Semantic Scholar

provider refresh

Showing the first 25 of 1585 observed papers.

Publication-year context

Citation counts generally accumulate over time. Cohort medians use observed Semantic Scholar counts only, while the observed and missing columns keep the full corpus denominator visible.

Paper coverage and median observed Semantic Scholar cumulative citations by publication year.
Publication year Corpus papers Observed Missing Median observed citations
2026 1741 1585 156 0

How to read these counts

Counts are cumulative provider observations captured on the displayed dates. Citation practices differ by field and publication age, and provider coverage changes over time. These counts do not establish quality, correctness, causal influence, or societal impact.

Review coverage quality for corpus limitations and About & Methodology for how The Commonplace collects and assesses research.