The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures

Digests

2026-08-24 2026-08-18 2026-08-10 2026-08-04 2026-07-27 2026-07-20 2026-07-13 2026-07-06 2026-06-29 2026-06-22 2026-06-15 2026-05-25 2026-05-18 2026-05-11 2026-05-04 2026-04-27 2026-04-20 2026-04-13 2026-04-06 2026-04-04 2026-04-04-before 2026-03-30 2026-03-23 2026-03-20 2026-03-18 2026-03-15

This weekly digest tracks what is NEW or CHANGED in AI-economics research. For the cumulative state of evidence on any topic, see the /syntheses pages. A single study rarely overturns a body of evidence.

The Delta

Coming in, Governance & Regulation leaned positive (661 papers); this week, a counter-signal appears. - Strengthened: two large field experiments find generative AI (GenAI) assistants and automated interviewers raise throughput, ratings, offers, and starts, with the biggest gains for lower performers and no short‑run productivity penalty among hires. - Better measured: behavioral effects of AI advice sharpened; sycophancy can still depolarize choices on average, while perceived political bias reduces large language model (LLM) persuasiveness. - Challenged: single‑agent price audits’ ability to catch collusion, with formal results indicating that marginals‑preserving conspiracies are invisible to such tests.

What Moved & What Held

Coming in, the standing view was: GenAI reliably boosts speed and standardizes work in structured, repeatable tasks, disproportionately helping lower‑performing workers; decision quality and persuasion effects are mixed and design‑sensitive; audit and governance tools often miss joint behavior; and infra advances keep shifting deployment costs.

This week adds two strong field experiments that move the productivity story toward established in the studied service and hiring funnels, plus cleaner estimates that AI tone and credibility shape downstream choices in ways not captured by capability benchmarks. On governance, formal results clarify why single‑agent price‑level audits miss communication‑free collusion that preserves marginals. Still holds this week: benefits concentrate in routine workflows, quality effects vary by task and user, learning can erode without engagement, and audit/provenance design remains a first‑order risk.

Top Papers

  • Confirms · established Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations (Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, Congyi Zhou; Alibaba collaboration; randomized field experiment) - In Alibaba’s China-based customer support operation, random access to a GenAI assistant speeds chats and raises customer ratings, with the largest gains among low‑baseline agents; objective resolution quality does not fall. This randomized controlled trial (RCT) in live production aligns with prior lab-in-the-field findings that GenAI reduces variance and lifts the lower tail. - So what: If this holds, the risk you own is overindexing on averages and missing that returns come from variance reduction, not across‑the‑board uplift. - Full numbers

  • Confirms · established Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews (Brian Jabarian, Luca Henkel; natural field experiment with randomized interviewer assignment) - Among 70,884 applications, assignment to automated voice interviews increases offers by about 12% and starts or early retention by about 18%, with no detectable productivity drop among hires who came through AI‑led screens; human reviewers still made final decisions. This strengthens the case that structured, consistent elicitation improves selection without short‑run output penalties in the studied funnels. - So what: If this holds, the risk you own is misattributing better funnel yield to looser standards when the gain is coming from process consistency. - Full numbers

  • Extends · established AI Sycophancy and Decisions (John Conlon, Peter Schwardmann; preregistered randomized experiment) - In an RCT with 1,500 participants across 30 incentivized decision tasks, a sycophantic LLM still depolarizes choices on average (about 0.22 standard deviations toward center) relative to no‑chat controls, with higher sycophancy attenuating that depolarization. This nuances the standing view that agreement‑seeking AI primarily amplifies priors. - So what: If this generalizes, the risk you own is assuming any one alignment tweak (less sycophancy) monotonically improves decision quality when effects depend on task and baseline tilt. - Full numbers

Also Notable

What Moved

  • Field productivity and hiring: Relative to the baseline of lab and pilot studies suggesting GenAI lifts throughput and low performers, the Alibaba RCT and the 70k‑applicant hiring experiment expand external validity to at‑scale operations in the studied contexts and find no short‑run productivity penalty among hires. Together they increase confidence for variance‑reduction gains in routine service and structured interviews.

  • Decision influence and credibility: The sycophancy RCT finds that even agreement‑seeking assistants can depolarize choices on average, while a preregistered US survey experiment quantifies how bias warnings cut persuasion by roughly a quarter. This refines the standing view from “AI advice can sway users” to “effects hinge on perceived neutrality and baseline tilt,” tightening the design space for credible decision support.

  • Audit design and collusion risk: Formal results argue single‑agent price‑level audits are blind by construction to profitable coordination that preserves marginals, sharpening prior concerns about audit blind spots. This raises the bar for enforcement evidence, from marginal checks to joint‑dependence testing.

Contested & Watch

  • Short‑run productivity gains vs long‑run skill formation - Finding: RCTs show throughput gains and higher ratings or offers in customer service and hiring at scale, with larger lifts for low performers (N≈70k applicants; large‑scale service RCT). - Standing evidence: Multiple RCTs and lab‑in‑the‑field studies indicate AI assistance can erode novices’ conceptual learning and debugging skill; strength: several high‑evidence studies leaning negative on learning. - Watch: Longitudinal field data on post‑adoption skill trajectories and error recovery, with task‑level engagement instrumentation.

  • Do explanations help teams or just add confidence? - Finding: Controlled studies report explanations raise confidence and reliance without improving accuracy and can hinder error recovery (multiple samples). - Standing evidence: A smaller set of education‑context RCTs show AI‑mediated suggestions can improve revisions when humans filter or adopt them; strength: established but context‑specific. - Watch: Cross‑domain trials that randomize explanation style versus selective automation with common outcome metrics and calibrated uncertainty displays.

  • Sycophancy: polarizer or equalizer? - Finding: In 1,500‑person RCTs across 30 tasks, a sycophantic model still depolarizes average choices, with more sycophancy dampening the effect. - Standing evidence: Prior work documents agreement bias and persuasion risks; strength: several lab studies, mixed on direction across domains. - Watch: Field experiments in political and organizational decisions that vary baseline polarization and model tone, measuring downstream actions as opposed to stated choices.

  • Can single‑agent audits deter algorithmic collusion? - Finding: Theory plus demonstrations show price‑level audits focused on marginals cannot detect profitable, communication‑free coordination that preserves those marginals. - Standing evidence: Enforcement and audit literatures flag detection gaps but rely on marginal and unilateral tests; strength: descriptive and theoretical, limited joint‑test adoption. - Watch: Regulator‑run pilots of pairwise or network dependence tests and time‑series experiments on coordinated shocks.

  • Are synthetic users viable for evaluation and targeting? - Finding: Benchmarks show instruction‑tuned models collapse sampling and simulated respondents underperform demographic baselines, overstating between‑group differences (N across multiple domains). - Standing evidence: Mixed descriptive papers claim face‑validity in narrow tasks; strength: small‑N, protocol‑sensitive. - Watch: Protocols that enforce proper sampling, plus head‑to‑head forecasting of human choices with preregistered out‑of‑sample tests.

Methods Spotlight

  • Fault‑tolerant hybrid shared data parallelism (FT‑HSDP): Training LLMs with Fault Tolerant HSDP on 100,000 GPUs reports roughly double effective utilization at extreme scale in testing, with no measured accuracy loss, shifting the economics of large‑model training.

  • Condensate manifold projection for attention: The Condensate Theorem proposes a provable approach, with demonstrations suggesting bit‑exact long‑context speedups by projecting attention onto a learned topology.

  • Closed‑loop LP repair with solver‑verified rewards: OptiRepair couples infeasibility diagnosis with targeted training of an 8B model, materially outperforming generic APIs on repairing supply‑chain optimization models in evaluated cases.