The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures

Completed three-study editorial pilot

Three studies on AI and work in 2026

By Alex Farach ·

The editorial pilot is complete. A broader 2026 review has no announced publication date.

This article discusses three papers posted in 2026 on AI-assisted task performance, code comprehension and hiring. The larger paper counts below describe the source collection, not studies individually discussed here.

Source window and review details

This dated editorial pilot uses the database published on September 2, 2026 and its subset of papers dated January 1 through September 2. Identities, abstracts and relevant stored full-text passages for the examples were inspected. No new pipeline model calls or historical backfill was run, and collection eligibility was not changed.

The four cited arXiv records were checked on September 5, 2026. Each listed version 1, with no withdrawal notice observed on those records. Journal publication and peer-review status were not independently verified.

Where are the other papers?

Browse or search the live collection to find papers beyond these examples. It contains the current public paper records and may include additions after the September 2 snapshot. Its counts can therefore differ from the dated counts in this article.

Individual records contain automated assessments and extracted claims; inclusion does not mean a person has reviewed each paper. The annual evidence map and bulk downloads remain closed while publication rights and provenance are reviewed. Ordinary paper browsing remains available.

When will a broader 2026 review be available?

No publication date has been announced. The paper-intake pipeline is scheduled daily, and the weekly digest has a separate review process. Neither automatically produces a full-year synthesis or updates this article.

A broader review would first need defined questions and study-selection rules. The selected papers would then need source checking and consistent outcome coding, with an account of collection gaps and missing dates. That work would have to cover the full comparison period without changing the original 2026 collection rules. No publication timetable has been set for it.

What the September 2 collection contains

The source collection's January 1 through September 2 subset contains 3,234 papers and 29,466 extracted claims. These are the records counted in the table, not the number of studies discussed in this pilot. Broad organizational and governance tags occur more often than direct employment, wage and labor-share tags. This describes the collection and its automated coding, not the allocation of attention in the entire research field.

Selected automated outcome tags in the dated 3,234-paper subset. Counts describe the collection, not the three study discussions.
Outcome tagDistinct papers
Organizational efficiency1,376
Governance and regulation1,268
Firm productivity499
Employment level194
Wages and compensation118
Labor share of income45

A paper can appear in several categories, and some categories cover more subjects than others. The counts cannot be added together or used to rank research importance. In this cohort, 1,509 papers have at least one claim tagged “Other,” which is particularly hard to interpret.

Assisted performance and learning are different outcomes

Cruces and colleagues report an online randomized study of adults completing a workplace-style task. AI assistance improved performance, with larger gains among less-educated participants. Lower-education participants retained some improvement in the unassisted follow-up. The experiment covered one task and a short follow-up; it did not measure longer-term learning at work.

Balepur and colleagues studied students building a website. Both groups used AI: one used a coding agent and the other a chatbot. Agent users did better on initial completion but understood their code less well.

Cruces and colleagues compared AI assistance with no assistance; Balepur and colleagues compared two AI workflows. Their tasks, samples and outcomes also differ, so the results should not be pooled or treated as contradictory estimates. Future learning studies should measure unaided understanding and performance on new tasks alongside assisted output.

Human recruiters kept hiring authority

Jabarian and Henkel report a natural field experiment assigning applicants to human or voice-AI interviews. The paper reports improved offer, start and retention outcomes in the AI-interview condition. Human recruiters still evaluated the interviews and made hiring decisions.

The experiment tested AI for gathering information before a human decision. Its findings do not establish what happens with autonomous hiring, in other occupations or across the economy.

Four questions to investigate going into 2027

Each question needs its own literature review, including earlier research. Some may already have answers in work outside this collection.

1. Which workflows preserve learning after assistance is removed?

Test unaided performance and transfer to new tasks after AI access ends. Report professional, student and online samples separately, and record the length of follow-up before drawing conclusions about durable learning.

2. Which gains travel across employers, tasks and decision arrangements?

Record what AI does in each study and who makes the final decision. Evidence about collecting interview information cannot stand in for a test of automated hiring. Compare those roles and the employers' settings before applying a finding elsewhere.

3. When do productivity gains become better pay or careers?

A smaller performance gap in an experiment cannot show how employers will set pay. Studies of distributional effects need to follow wages, hours and career changes after AI adoption, including losses from displacement. Before interpreting the small wage and labor-share categories as a research gap, check how much relevant work the collection missed.

4. How do organizational gains aggregate into employment and worker transitions?

The interview trial's increase in job starts does not measure net job creation. A policy review needs to distinguish firm-level outcomes from worker reallocation and economy-wide effects. Identify data that follow displaced workers into subsequent jobs.

Human training and model training need separate labels

A claim from ASIL about supervised fine-tuning of Qwen models is tagged “Training effectiveness.” It concerns model training and cannot be used as evidence about training workers. Check what a study actually examines before using its category to answer a policy question.

Use a separate, versioned set of analytical annotations to distinguish humans from computer agents and observations from simulations. Record whether evidence addresses the question directly. Apply the same rules across the entire 2026 window, keeping the original records and labels. Intake criteria should remain unchanged.

For 8,943 of the 29,466 included claims, the database records both a quote and a locator. That count measures the availability of locators, not the accuracy of the claims. A missing locator may reflect incomplete extraction even when the paper contains relevant evidence.

Publication and collection dates differ

All 442 January papers and all 333 February papers in this cohort were collected in a later month. In July, 165 of 337 were collected later. These dates show when papers entered Commonplace. They do not tell us whether a delayed entry came from a backfill or ordinary indexing lag. Analyses of publication trends should use publication dates.

The September 2 source contained 3,887 papers. The publication-date subset excludes 317 published before 2026, 332 without a usable publication date, and four dated after the cutoff. Those records remain in the database. Report these exclusions and the source/model composition with any year-to-year comparison.