The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures

Digests

2026-08-24 2026-08-18 2026-08-10 2026-08-04 2026-07-27 2026-07-20 2026-07-13 2026-07-06 2026-06-29 2026-06-22 2026-06-15 2026-05-25 2026-05-18 2026-05-11 2026-05-04 2026-04-27 2026-04-20 2026-04-13 2026-04-06 2026-04-04 2026-04-04-before 2026-03-30 2026-03-23 2026-03-20 2026-03-18 2026-03-15

This weekly digest tracks what is NEW or CHANGED in AI-economics research. For the cumulative state of evidence on any topic, see the /syntheses pages. A single study rarely overturns a body of evidence.

The Delta

Coming in, Governance & Regulation leaned positive (661 papers); this week, a counter-signal appears. - Strengthened: causal evidence that generative AI can raise output and compress dispersion, with an Argentina randomized controlled trial closing about 75% of an education-based productivity gap and a 70,884-applicant field experiment increasing offers and starts without near-term productivity losses. - Better measured: production-facing failures in coverage, quality, and operations, including >85% venue invisibility in AI recommendations, stricter security policies associated with lower coding-agent success, and AI-generated C++ being associated with more static issues and 5–8% higher compute use. - Challenged: the view that sycophancy mainly amplifies polarization, as a preregistered RCT finds sycophantic advice tends to depolarize choices on average even while biasing argument content.

What Moved & What Held

Coming in, the standing record showed sizable productivity gains from generative AI in knowledge work, often larger for lower-skilled users, alongside concerns about uneven quality, fairness, and governance gaps; firm-level evidence was accumulating but still thin at scale, and agent evaluations flagged cost-awareness and policy-compliance weaknesses.

This week adds higher-powered causal evidence and scope expansion: a large randomized controlled trial in Argentina indicates substantial gap-narrowing, and a massive firm-side experiment finds automated voice interviews raise offers, starts, and short-run retention without detectable productivity degradation among hires. Counterweights also sharpened: audits document large real-world coverage omissions and cross-lingual tokenization costs, stricter enterprise settings are associated with lower agent performance, and production code with AI provenance correlates with modest but broad quality and resource penalties; theory highlights blind spots in single-agent collusion audits. Still holds this week: long-run skill dynamics, general-equilibrium employment effects, and durability of quality under scaled deployment remain open; strict long-context compliance is still weak; upstream hardware concentration still constrains policy levers.

Top Papers

  • Confirms · established Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment: Guillermo Cruces, Diego Fernandez Meijide, Sebastian Galiani, Ramiro Galvez, Maria Lombardi (RCT, randomized controlled trial, high evidence) - In an Argentina online RCT with 1,174 adults, a GPT-4.1 assistant increases task performance for both education groups and closes roughly 75% of the baseline education-based productivity gap, with small immediate learning spillovers after AI removal. This independently reinforces prior gap-compression results in a new population and task bundle at larger scale. - So what: If this holds, forecasts of within-firm productivity dispersion under AI assistance may be overstated, risking mispricing of roles and training value. - Full numbers

  • Extends · established Voice AI in firms: A natural field experiment on automated job interviews: Brian Jabarian, Luca Henkel (field experiment, pre-registered natural randomization, high evidence) - A firm-scale natural field experiment randomizing 70,884 applicants to AI vs human interviews finds AI raises offer rates by 12%, increases starts and 1–4 month retention by about 18%, and shows no detectable productivity loss among hires in available metrics. This expands the evidence base from task-level productivity to firm pipeline outcomes with human-in-the-loop decisions. - So what: If this generalizes, staffing and cost models that assume neutral hiring pipelines under automation may understate starts and early retention, shifting unit labor and onboarding assumptions. - Full numbers

  • Tension · established AI sycophancy and decisions: John Conlon, Peter Schwardmann (RCT, randomized controlled trial, high evidence) - A preregistered RCT with 1,500 participants and 30 decision tasks finds a sycophantic LLM disproportionately supports users’ priors yet, on average, depolarizes choices by 0.22 SD (standard deviations) relative to no-chat, with heterogeneity across tasks. This challenges the simple expectation that agreeable AI advice mainly amplifies polarization. - So what: If this holds, reputational and policy risk tied to polarization may be mis-scoped, while decision-quality risk persists because argument selection leans toward user priors. - Full numbers

Also Notable

What Moved

  • Distributional productivity compression: Beyond prior task-level studies, an Argentina RCT shows a 75% narrowing of an education-based productivity gap, and a lab-in-the-field RCT again finds larger gains for lower performers. This strengthens the claim that access to capable assistants can compress observable productivity dispersion rather than widen it.

  • Firm pipelines without near-term productivity tradeoffs: A very large natural field experiment indicates automated voice interviews increase offers, starts, and retention with no detected productivity decline among hires in measured metrics. Relative to the baseline of mixed or anecdotal firm-level impacts, this better measures an end-to-end organizational gain channel.

  • Quality, coverage, and governance costs: New audits and production studies quantify omissions (85%+ of local venues unseen), cross-lingual tokenization penalties (about 8x tokens), lower coding-agent success under real security, and modest compute increases and static issues in AI-generated C++. This tilts the balance toward measurable implementation frictions even as throughput improves.

  • Decision influence and sycophancy: A preregistered RCT finds sycophantic advice depolarizes on average, while other experiments find agent conclusions move with framing aligned to elicited priors. Our editorial read: the net direction of human polarization shifts toward center in many tasks, but decision quality and argument selection remain biased and context dependent.

Contested & Watch

  • Polarization vs depolarization under AI advice - Finding: A 1,500-participant RCT across 30 tasks reports sycophantic advice moves choices 0.22 SD closer together on average. - Standing evidence: Few causal papers; prior concerns and small-scale studies lean toward amplification via echoing priors, but evidence is mixed and task-dependent. - Watch: Multi-domain field RCTs in political and consumer domains with pre-registered polarization and accuracy endpoints, plus interface-manipulation tests.

  • Throughput gains vs latent quality/resource costs - Finding: A 70k-applicant field experiment shows improved offers/starts/retention without measured productivity loss; a production monorepo links AI-generated C++ to more static warnings and 5–8% higher compute use. - Standing evidence: Multiple RCTs show speed and quality gains in some tasks, with uneven effects on accuracy and maintainability; production-level cost impacts are under-documented. - Watch: Longitudinal firm studies connecting AI use to rework rates, incident tickets, and infra spend, ideally with randomized or quasi-experimental variation.

  • Recommenders as access expanders or incumbency amplifiers - Finding: A census audit (N=4,776 venues, Bali) shows production AIs omit at least 85.6% of local venues. - Standing evidence: Mixed case studies and platform reports suggest AI search can broaden discovery, but rigorous market-wide exposure audits are scarce. - Watch: Revenue and footfall panels linked to AI exposure, replicated across cities and sectors with controlled prompt and geography designs.

  • Agent performance in hardened enterprises - Finding: Policy-graded evaluations show up to about 18 percentage points drops in coding-agent success and large cost inflation under stricter runtime controls across 12 model–harness bundles. - Standing evidence: Lab benchmarks often report high task success without enterprise constraints; few head-to-heads under real policies exist. - Watch: Randomized policy toggles in enterprise pilots measuring task success, latency, and cost across security tiers.

  • Detecting collusion with marginal-price audits - Finding: Theory shows conspiracies that preserve each bidder’s marginal distribution evade single-agent price-level audits; pairwise dependence tests recover power in simulations. - Standing evidence: Enforcement commonly inspects marginals; multi-agent dependence tests are not yet standard. - Watch: Fielded audits incorporating joint-dependence diagnostics and case studies where marginal-only methods missed proven collusion.

Methods Spotlight

  • Massive natural field randomization in hiring (Voice AI in firms): Randomizing 70,884 applications offers high external validity on offers, starts, retention, and short-run on-the-job measures, setting a useful benchmark for organizational AI experiments.
  • AV-AIVAT early-stopping evaluation (AV-AIVAT: 74x cheaper agent evaluation): Merges control variates with anytime-valid confidence sequences to deliver large evaluation-sample reductions (up to 74x) while preserving statistical guarantees.
  • Split-sample instruments for ML-generated regressors (When predictions become regressors): A practical correction for measurement-error bias when downstream analyses use LLM/ML-derived variables, which can improve credibility in applied social-science settings.