The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests About 🎲 Workforce Futures
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Semantic Scholar observations cover 91.5% of papers in this view

This is an observation-coverage view of The Commonplace’s indexed corpus, not a global ranking. 290 of 317 papers have a latest Semantic Scholar count under the selected filter; 27 do not.

Clear

Observed cumulative citations

Papers are ordered by their latest recorded Semantic Scholar count. Ties use observation date, then title and paper ID for deterministic ordering.

Semantic Scholar cumulative citation observations for papers in the selected corpus view, ordered by count. Each row shows the provider and observation date; counts from other providers are not added.
Rank in this view Paper Published Semantic Scholar cumulative citations Observation
1 A new benchmark for long-horizon software evolution reveals a large capability gap: state-of-the-art coding agents solve just 25% of release-note-driven, multi-file tasks, underscoring that current systems struggle with sustained, cross-file engineering work despite strong single-issue performance. Tue Le, Minh V. T. Thai, Dung Nguyen Manh, Huy Phan Nhat, Nghi D. Q. Bui 24 Semantic Scholar

provider refresh
2 TokenPowerBench lets operators quantify how inference choices change energy and operating costs: across Llama, Falcon, Qwen and Mistral models (1B–405B) it attributes joules per token to prefill vs decode and shows how batch size, context length, parallelism and quantization affect efficiency, all using software-based measurements rather than external meters. Chenxu Niu, Wei Zhang, Jie Li, Yongjian Zhao, Tongyang Wang, Xi Wang, Yong Chen 22 Semantic Scholar

provider refresh
3 Large language model agents can stabilize supply chains in simulation: LLM-based consensus and negotiation frameworks cut demand amplification (the bullwhip effect) and outperform baseline restocking and centralized policies in an inventory-management case study, though findings rest on synthetic experiments and specific model/tool choices. Valeria Jannelli, Stefan Schöpf, Matthias Bickel, Torbjørn Netland, Alexandra Brintrup 19 Semantic Scholar

provider refresh
4 Turning old landfills into modular waste-to-energy assets could cut methane emissions by up to 70% and deliver ~20 MW of islandable power per site, offering crucial microgrid resilience for AI data centres; in the U.S., fragmented governance and accounting rules make such de‑landfilling a politically and institutionally distinct challenge, so remediation should be valued as a resilience buffer rather than a substitute for bulk renewables. Qi He, Chunyu Qu 18 Semantic Scholar

provider refresh
5 Experienced developers treat AI agents as productivity boosters but keep the reins: they delegate routine or well-scoped tasks to agents while retaining control over design and quality, using expert strategies to constrain agent behavior and compensating for known limitations. Ruanqianqian Huang, Avery Reyna, Sorin Lerner, Haijun Xia, Brian Hempel 16 Semantic Scholar

provider refresh
6 In emerging Asia, firms with stronger AI capabilities produce greener innovations more efficiently, but the environmental payoff is substantially larger when firms have robust ESG practices; AI alone is insufficient without governance that fosters transparency and accountability. Marwan Mansour, Mo’taz Al Zobi, Mohammed Alomair 16 Semantic Scholar

provider refresh
7 Large language models match human persuaders on average, but results swing widely by context; a meta-analysis of seven studies finds no average advantage for humans or LLMs, while model choice, message design and domain jointly explain most of the variation. Lukas Hölbling, Sebastian Maier, Stefan Feuerriegel 15 Semantic Scholar

provider refresh
8 Naive deployment-weighted 'world prices' for data-center SKUs can invert cost rankings and mislead procurement; two cost-preserving operators—a normalized two-way fixed-effects decomposition and a convex common-weight benchmark—cut ranking violations markedly while preserving total costs. Qi He 14 Semantic Scholar

provider refresh
9 An AI agent nearly rivalled human pentesters in a live campus network: ARTEMIS discovered nine valid vulnerabilities with an 82% validation rate and outperformed nine of ten professionals, with some variants operating at about $18/hour versus $60/hour for human testers; existing AI scaffolds lagged, and agents still suffer higher false positives and struggle with GUI-driven tasks. Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, Anna Wu, Arnold Tianyi Yang, Neil Perry, Andy Zou, Matt Fredrikson, J. Zico Kolter, Percy Liang, Dan Boneh, Daniel E. Ho 13 Semantic Scholar

provider refresh
10 AI promises to make HR leaner and more personalized—automating hiring, standardizing reviews and tailoring development—but the evidence base lacks robust field evaluations and flags serious privacy and ethical risks, underscoring the need for regulation and more empirical research. Muhammad Asif, Asif Ali, Fayyaz Ahmad Shaheen 11 Semantic Scholar

provider refresh
11 AI browser agents are primarily used by wealthier, better-educated and knowledge-sector users and mostly for productivity and learning tasks. Personal queries make up over half of interactions, but usage becomes more cognitively oriented over time, with the top 10 tasks accounting for more than half of activity. Jeremy Yang, Noah Yonack, Kate Zyskowski, Denis Yarats, Johnny Ho, Jerry Ma 10 Semantic Scholar

provider refresh
12 Aligning an AI to its operator’s wishes is not enough: without institutions that embody shared values, even 'perfectly aligned' systems can produce harmful societal outcomes. The authors propose 'full‑stack alignment'—thick, structured models of value embedded in institutions and systems—to enable normatively competent agents, stewardship, win‑win negotiation, meaning‑preserving economic mechanisms, and democratic regulation. Edelman, Joe, Zhi-Xuan, Tan, Lowe, Ryan, Klingefjord, Oliver, Wang-Mascianica, Vincent, Franklin, Matija, Kearns, Ryan Othniel, Hain, Ellie, Sarkar, Atrisha, Bakker, Michiel, Barez, Fazl, Duvenaud, David, Foerster, Jakob, Gabriel, Iason, Gubbels, Joseph, Goodman, Bryce, Haupt, Andreas, Heitzig, Jobst, Jara-Ettinger, Julian, Kasirzadeh, Atoosa, Kirkpatrick, James Ravi, Koh, Andrew, Knox, W. Bradley, Koralus, Philipp, Lehman, Joel, Levine, Sydney, Marro, Samuele, Revel, Manon, Shorin, Toby, Sutherland, Morgan, Tessler, Michael Henry, Vendrov, Ivan, Wilken-Smith, James 10 Semantic Scholar

provider refresh
13 Generative AI adoption is linked to more opportunistic ESG behavior: firms using generative models show stronger environmental scores but weaker social and governance performance, a pattern amplified across supply chains and by strict regulation and green investor pressure but mitigated by analyst scrutiny and better disclosures. Zhe Sun, Lei Liu, Liang Zhao, Hind Alofaysan, Bhumika Gupta 10 Semantic Scholar

provider refresh
14 Machine learning can cut enterprise forecasting errors by roughly 15–40% compared with traditional rule-based systems, but the payoff depends on firm data, talent and governance. Emerging solutions—explainable AI, AutoML and federated learning—promise to reduce implementation and regulatory frictions that currently limit broad adoption. Zifan Chen, Jingyi Liu, Jiaying Chen 10 Semantic Scholar

provider refresh
15 Saudi banks that deploy AI-powered FinTech tools show stronger returns, higher valuations and greater stability, alongside improved ESG metrics; results hold across panel and dynamic models but stem from a small, single-country sample. Amina Hamdouni 10 Semantic Scholar

provider refresh
16 AI 'co-pilots' can lift market-disruption prediction accuracy by roughly a third to a half and speed strategic responses, but firms capture those gains only when they pair algorithms with accountability, calibrated trust, and preserved human oversight. Simon Suwanzy Dzreke 9 Semantic Scholar

provider refresh
17 Tailored explainable-AI boosts gig workers' acceptance, but too much explanation backfires: local or counterfactual reasons raise trust and improve manager-worker relations, while combining both overwhelms workers and reduces acceptance. Miles M. Yang, Ying Lu, Fang Lee Cooke 9 Semantic Scholar

provider refresh
18 Large language models are already being deployed inside romance-baiting crime rings and can outperform humans at eliciting trust and compliance — in a week-long blinded study an LLM secured 46% compliance versus 18% for humans (p=0.007). Commercial safety filters tested detected none of the romance-baiting dialogues, suggesting current defenses may not prevent automated expansion. Gilad Gressel, Rahul Pankajakshan, Shir Rozenfeld, Ling Li, Ivan Franceschini, Krishnashree Achuthan, Yisroel Mirsky 8 Semantic Scholar

provider refresh
19 A new 142,808-conversation corpus shows that commercial LLM platforms produce systematically different user interactions and outputs when their native affordances (citations, thinking traces, code) are preserved; these differences affect intent satisfaction, citation strategies, and latency dynamics and are invisible to homogenized, single-platform benchmarks. Yueru Yan, Tuc Nguyen, Bo Su, Melissa Lieffers, Thai Le 7 Semantic Scholar

provider refresh
20 Algorithmic hiring tools risk hiding and amplifying workplace inequalities: while promised as efficient and objective, recruitment AIs can reproduce bias through opaque models and institutional legitimation, requiring interdisciplinary scrutiny and shared regulatory responsibility. Karen D Hughes, Alla Konnikov, Nicole Denier, Yang Hu 7 Semantic Scholar

provider refresh
21 AGI might arise as a patchwork of cooperating specialized agents rather than a single superintelligence; policymakers and developers should build regulated virtual 'agent economies' with market rules, audit trails and reputation systems to detect and constrain dangerous collective behaviors. Nenad Tomašev, Matija Franklin, Julian Jacobs, Sébastien Krier, Simon Osindero 6 Semantic Scholar

provider refresh
22 A bandit-driven method for composing specialized LLM sub-agents (BOAD) materially improves performance on challenging software-engineering tasks; a tuned 36B BOAD system placed second on the SWE-bench-Live leaderboard, outperforming larger models including GPT-4 and Claude. Iris Xu, Guangtao Zeng, Zexue He, Charles Jin, Aldo Pareja, Dan Gutfreund, Chuang Gan, Zhang-Wei Hong 5 Semantic Scholar

provider refresh
23 A new benchmark of real-world finance workflows finds top AI agents complete fewer than 40% of tasks: GPT-5.1 spends nearly 17 minutes per workflow yet passes only 38.4% of cases, exposing persistent failure modes on messy, multimodal enterprise work. Dong, Haoyu, Zhang, Pengkun, Gao, Yan, Dong, Xuanyu, Cheng, Yilin, Lu, Mingzhe, Zhu, Zikun, Yakefu, Adina, Zheng, Shuxin 5 Semantic Scholar

provider refresh
24 Embedding ad allocation into language model outputs can improve auction efficiency: a proposed 'LLM-Auction' trains models to balance advertiser value and user experience and, in simulation, outperforms prior methods while retaining favorable incentive properties under a first-price rule. Chujie Zhao, Qun Hu, Shiping Song, Dagui Chen, Han Zhu, Jian Xu, Bo Zheng 5 Semantic Scholar

provider refresh
25 Insurer agents could make autonomous AI agents economically accountable by staking collateral and auditing behavior; competitive underwriting and TEE-mediated audits decentralize verification and create incentive-compatible dispute resolution without sole reliance on brittle reputations. Botao 'Amber' Hu, Bangdao Chen 5 Semantic Scholar

provider refresh

Showing the first 25 of 290 observed papers.

Publication-year context

Citation counts generally accumulate over time. Cohort medians use observed Semantic Scholar counts only, while the observed and missing columns keep the full corpus denominator visible.

Paper coverage and median observed Semantic Scholar cumulative citations by publication year.
Publication year Corpus papers Observed Missing Median observed citations
2025 317 290 27 0.0

How to read these counts

Counts are cumulative provider observations captured on the displayed dates. Citation practices differ by field and publication age, and provider coverage changes over time. These counts do not establish quality, correctness, causal influence, or societal impact.

Review coverage quality for corpus limitations and About & Methodology for how The Commonplace collects and assesses research.