Digests
This weekly digest tracks what is NEW or CHANGED in AI-economics research. For the cumulative state of evidence on any topic, see the /syntheses pages. A single study rarely overturns a body of evidence.
The Delta
Coming in, Governance & Regulation leaned positive (661 papers); this week, a counter-signal appears. - Better measured: An incentivized pension randomized controlled trial (RCT) estimates about one-third pass-through (share of recommendation incorporated) from AI advice into allocations, with no gain in risk-adjusted performance (return per unit of risk). - Newly observed: A pre-registered multi-model audit finds LLM (large language model) physician recommenders respond to randomized ordering and name cues with fee-equivalent magnitudes. - Strengthened: Production A/B tests (randomized experiments) at a large platform find uplift-targeted models (optimizing for incremental impact) and agentic workflows (multi-step AI agents that take actions) raise firm metrics, while a field quasi-experiment estimates time costs from guardrails (policy and safety filters).
What Moved & What Held
Coming in, the standing view was that LLMs and agents change choices and workflows, with partial behavioral pass-through and measurable firm-side gains in field A/Bs; governance can reduce visible harms but adds friction; and one-shot audits and thin disclosures often misstate safety, fairness, or procurement risk.
This week adds causal magnitudes in two high-stakes settings and sharpens trade-offs: a lab-in-the-field pension experiment estimates pass-through at 37% without Sharpe improvements (return per unit of risk), and an LLM doctor-recommender audit estimates dollar-equivalent sizes on ordering and demographic effects; on the production side, a causal targeting system increases incremental value while a guardrail rollout reports large reductions in hallucinations at the cost of slower tasks. Still holds this week: AI steers decisions and can lift firm outcomes, but governance that relies on single snapshots or unspecified metrics can be unreliable.
Top Papers
-
New · established Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice - Syeda Anshrah Gillani, Mirza Samad Ahmed Baig (pre-registered randomized conjoint, a choice experiment with randomized attributes) - Across seven commercial LLMs and 3,024 randomized choice sets (40,068 model responses), higher ratings and lower fees increase recommendation probability, while randomized ordering and name cues induce smaller but consistent shifts worth roughly an $11 fee change. This identifies causal attribute effects on LLM-mediated recommendations in a synthetic US-style physician-card setting, filling in magnitudes the baseline lacked. - So what: If this generalizes, LLM-mediated referrals may amplify reputational inequality and embed subtle ordering and demographic skews that change who gets work. - Full numbers
-
Extends · established Do People Follow AI Advice? Evidence from a Pension Portfolio Choice Experiment - Hongseok Choi, Jeongbin Kim, Matthew Kovach, Kyu-Min Lee, Euncheol Shin, Hector Tzavellas (incentivized RCT) - In an RCT with N=400 pension participants, randomized aggressive vs conservative AI recommendations shift about 37% of the allocation gap into final portfolios, increasing expected returns and risk but leaving Sharpe ratios unchanged and reducing diversification. This extends prior pass-through evidence to retirement finance with causal estimates on risk exposure. - So what: If this holds, AI advice can move household risk loads without improving risk-adjusted performance, exposing providers and plans to distributional risk they may be undercounting. - Full numbers
-
Extends · established From prediction to incrementality: Causal optimization for large-scale targeting and recommendation - Changshuai Wei, John Bencina, Phuc Nguyen, Andre Assuncao Silva T Ribeiro, Benjamin Zelditch (LinkedIn, production A/B test) - A production system at US-based LinkedIn that combines causal lift estimation, exploration, and constrained allocation increases incremental marketing long-term value by 7.20% in an A/B test, outperforming prediction-only baselines. This adds to firm-side evidence that optimizing for incremental impact can deliver gains beyond engagement proxies. - So what: If this holds, firms relying on predictive scores risk wasting spend on users who would convert anyway and understate true return on model-driven outreach. - Full numbers
Also Notable
- New · suggestive Governing generative AI in organizations: a design theory and quasi-experimental field study of sociotechnical guardrails - Maikel Leon - A stepped-wedge rollout (phased rollout across units over time) at a Fortune 500 firm associates layered guardrails with large drops in hallucinations and stronger audit trails, alongside a measurable time penalty and some circumvention.
- New · established Self-evolving agentic customer support system at LinkedIn - Chih Hui Wang, Mengdie Tu, Qianyun Zhang, Wei Wu, Lili Zhou, Mingqi Shen, Changshuai Wei (LinkedIn) - A two-week production A/B finds modular agentic workflows increase self-serve and measurably improve routing accuracy in this platform's sample.
- New · descriptive No task fails every time: Why one-shot audits are structurally blind to agent damage - Shiven Khurdi - Repeated state-diff audits (comparing system or environment state before and after runs) find stochastic, irreversible damage events across models, implying single-run checks will often miss dangerous behaviors.
- New · established Toward meaningful transparency for AI chatbots: Disclosing persuasive intent reduces persuasion - Adrian Rauchfleisch, Andreas Jungherr - A 1,500-person RCT finds intent disclosure halves a chatbot's persuasive effect, while generic AI identity labels do not.
- New · descriptive Same system, opposite verdicts: Metric discretion in AI ethics audits and the limits of disclosure - Shay Tsaban - A multiverse audit (analyzing many plausible metric and sample choices) documents how plausible metric and population choices flip fairness verdicts and are rarely disclosed.
- New · descriptive VAKRA: Evaluating multi-hop reasoning across APIs and retrieval under tool-use policies - Ankita Rajaram Naik, Anupama Murthi, Benjamin Elder, Siyu Huo, Raavi Gupta, Abhinav Jain, Praveen Venkateswaran, Abdulhamid Adebayo, Danish Contractor - Accuracy drops with compositional depth and policy constraints, flagging limits for enterprise agents.
- New · framework Deployment decision reliability: A generalizability-theory framework for sizing long-horizon agent evaluations - Vasundra Srinivasan - Variance decompositions suggest agent-by-task interactions may dominate, so leaderboards may reflect specialization rather than general capability in this framework's samples.
- New · descriptive Frontier AI forecasting has a measurement problem: An audit of progress evidence - Fabricio F Costa - Public records on compute and capabilities are sparse, concentrated, and rarely linked, which weakens naive extrapolation.
- New · descriptive Pricing the risk of runtime compression: Anytime-valid admission and a served-output law for compressed serving state - Fanzhe Wei, Li Liu - In this deployment, an admission ledger makes compression risk auditable and fallbacks run at about half at matched loss on a live mixture-of-experts stack (a model with multiple specialized subnetworks).
- New · descriptive Human versus computer vision - Elena Sirotkina - On news photos, a simple center bias beats trained saliency models, and residuals vary with viewer demographics.
What Moved
-
Pass-through and portfolio risk: Relative to a general claim that AI nudges shift choices, the pension RCT estimates the magnitude at 37% in this population and shows no Sharpe improvement, so the baseline view now carries a quantified risk-exposure shift without performance gain. This narrows the plausible range for financial-advice pass-through in policy and supervisory models.
-
LLM-mediated market allocation: The physician recommender audit newly documents ordering and demographic name effects with dollar-equivalent sizes, extending prior anecdotal concerns into causal magnitudes across seven LLMs. That sharpens the baseline from "possible bias" to "priced bias" in referral-style use.
-
Production impact and governance trade-offs: A causal targeting win at LinkedIn strengthens the case that uplift-centric systems can deliver firm outcomes, while a field guardrail rollout reports large error reductions with slower work. The net productivity-safety trade-off is better parameterized than last week, though still sample-local.
Contested & Watch
-
Net effect of guardrails on productivity - Finding: A stepped-wedge field study at a Fortune 500 firm reports halved hallucinations and higher audit completeness alongside a measurable task-time penalty (firm-wide rollout, N not stated). - Standing evidence: Several A/Bs and quasi-experiments show productivity gains from copilots and agents, but few quantify governance-induced slowdowns (handful of studies, medium-to-high strength, leaning positive on productivity). - Watch: Multi-site RCTs that jointly estimate error rates and time costs over months, with variance by role and task criticality.
-
Do AI advice systems improve risk-adjusted performance? - Finding: In an RCT with 400 pension participants, aggressive AI recommendations raise expected returns and risk but leave Sharpe ratios unchanged and reduce diversification. - Standing evidence: Sparse causal evidence on realized portfolio performance from AI advice (few studies, mostly lab or short-horizon, mixed direction). - Watch: Long-horizon field trials linking AI advice to realized returns, drawdowns, and concentration at the account level.
-
Are one-shot audits sufficient for action-taking agents? - Finding: Repeated state-diff audits show stochastic, irreversible damage events that single-run checks usually miss (multiple model families, many runs). - Standing evidence: Pre-deployment checks are common but often single-pass or task-limited (several descriptive studies, low-to-medium strength on sufficiency). - Watch: Standardized multi-run audit protocols tied to deployment gates, with cross-lab replication.
-
Can systems route collaboration protocols cost-effectively? - Finding: Models can predict failure risk yet struggle to choose which collaboration protocol pays off across reasoning tasks in this benchmark (7 agent-model configurations, 7,308 runs). - Standing evidence: Case studies show multi-agent gains but selection and routing rules are ad hoc (several descriptive papers, suggestive, leaning optimistic). - Watch: Live A/Bs of cost-aware protocol routing with explicit spend and accuracy targets.
-
Do LLM-mediated recommenders shift real traffic via ordering and demographic cues? - Finding: Randomized conjoint estimates imply a first-position advantage worth about an $11 fee change and small name-based effects across seven models (40,068 responses). - Standing evidence: Platform recommender biases are documented, but LLM-mediated referral bias is under-measured (few studies, suggestive). - Watch: Field tests on real platforms linking LLM referral exposure to click-throughs and bookings, by position and demographic signal.
Methods Spotlight
- Pre-registered randomized conjoint audit across multiple LLMs: Gillani and Baig. Causally identifies how ratings, fees, names, and ordering drive LLM recommendations at scale, yielding transparent marginal effects with clustered uncertainty.
- End-to-end causal uplift optimization with exploration and constrained allocation: Wei et al. (LinkedIn). Puts modern causal inference into production, showing incremental value gains over prediction-only systems in a live A/B.
- Repeated ground-truth state-diff auditing for agents: Khurdi. Captures stochastic, irreversible harms that one-shot audits miss, a practical template for safety-critical evaluations.