The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Task-level analysis of 193,497 UK Civil Service vacancies finds wide variation in AI exposure even within identical job titles; LLM-driven redesigns point to augmentation and productivity gains—strategic leadership, complex problem-solving and stakeholder management remain human strengths rather than roles being broadly automated.

Beyond Automation: Redesigning Jobs with LLMs to Enhance Productivity
Andrew Ledingham, Michael Hollins, Matthew Lyon, David Gillespie, Umar Yunis-Guerra, Jamie Siviter, David Duncan, Oliver P. Hauser · December 05, 2025
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Andrew Ledingham unresolved corpus identity
  2. Michael Hollins unresolved corpus identity
  3. Matthew Lyon unresolved corpus identity
  4. David Gillespie unresolved corpus identity
  5. Umar Yunis-Guerra unresolved corpus identity
  6. Jamie Siviter unresolved corpus identity
  7. David Duncan unresolved corpus identity
  8. Oliver P. Hauser unresolved corpus identity

Semantic Scholar

Latest observation:

  1. A. Ledingham provider ID
  2. Michael Hollins provider ID
  3. M. Lyon provider ID
  4. D. Gillespie provider ID
  5. Umar Yunis-Guerra provider ID
  6. Jamie Siviter provider ID
  7. David Duncan provider ID
  8. Oliver P. Hauser provider ID
Using LLM-assigned AI-exposure scores on 1.54M tasks from 193,497 UK Civil Service vacancies, the paper finds substantial within-role heterogeneity in AI exposure and LLM-driven job redesigns that emphasize augmentation and productivity gains—retaining human comparative advantages in leadership, complex problem-solving and stakeholder management rather than widespread displacement.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

The adoption of generative artificial intelligence (AI) is predicted to lead to fundamental shifts in the labour market, resulting in displacement or augmentation of AI-exposed roles. To investigate the impact of AI across a large organisation, we assessed AI exposure at the task level within roles at the UK Civil Service (UKCS). Using a novel dataset of UKCS job adverts, covering 193,497 vacancies over 6 years, our large language model (LLM)-driven analysis estimated AI exposure scores of 1,542,411 tasks. By aggregating AI exposure scores for tasks within each role, we calculated the mean and variance of job-level exposure to AI, highlighting the heterogeneous impacts of AI, even for seemingly identical jobs. We then use an LLM to redesign jobs, focusing on task automation, task optimisation, and task reallocation. We find that the redesign process leads to tasks where humans have comparative advantage over AI, including strategic leadership, complex problem resolution, and stakeholder management. Overall, automation and augmentation are expected to have nuanced effects across all levels of the organisational hierarchy. Most economic value of AI is expected to arise from productivity gains rather than role displacement. We contribute to the automation, augmentation and productivity debates as well as advance our understanding of job redesign in the age of AI.

Summary

Main Finding

Using a large LLM-driven task-level analysis of 193,497 UK Civil Service job adverts (1,542,411 extracted tasks), the authors show that AI’s labour-market impact is heterogeneous: most jobs are only partially exposed to generative AI, and the largest economic value will likely come from productivity gains via job redesign (reallocating freed-up time to higher‑value human tasks) rather than mass role displacement.

Key Points

  • Dataset and scope
    • Nearly all UK Civil Service (UKCS) job adverts from 16 Jan 2019 to 3 Dec 2024: 193,497 vacancies, 1,542,411 tasks.
    • Covers 37 departments, 28 professions, and all grades; sample over‑represents mid/high grades (HEO, SEO, G6/G7).
  • Task-level AI exposure
    • Tasks were extracted from job adverts and each scored for AI exposure with an LLM.
    • Aggregated to job-level by mean and variance of task exposures, exposing substantial heterogeneity even among similar job titles.
    • Distribution of job exposure: >60% medium exposure, 20% low exposure (concentrated among senior roles — ~47% of SCS), ≈18% high exposure (could be fully automated).
  • Job redesign via LLM
    • For jobs not declared fully automatable, an LLM removed automatable tasks and proposed redesigned task mixes that reinvest saved time into activities where humans have comparative advantage.
    • New task composition after redesign: 26% strategic leadership, 18% complex problem resolution, 17% stakeholder management/communication; declines in administrative support, records management, and some data‑processing tasks.
    • Higher-grade roles skew toward leadership/complex problem solving; lower grades see more stakeholder/communication tasks.
  • Thresholding and economic magnitudes
    • Introducing a conservative automation threshold θ ≥ 80: ~145,864 jobs (≈75% of dataset) qualify for some degree of redesign rather than full automation.
    • Estimated potential benefits for UKCS (under their model and assumptions): ~£5.2bn in productivity gains from redesigned roles and ~£1.1bn in cost reductions from roles dominated by highly automatable tasks.
  • Robustness and context
    • Authors vary the redesign process and thresholds; results are robust in that new tasks generally emphasize human‑comparative strengths.
    • Results are context‑specific (public sector, advertised tasks) and depend on LLM judgments and assumptions about reinvestment of time.

Data & Methods

  • Data sources
    • GRID (Government Recruitment Information Database): job adverts (193,497 vacancies).
    • UK Civil Service Statistics (UKCSS): departmental and workforce context; used to map grades, estimate median salaries by grade/department.
  • Pipeline
  • Task segmentation: extract discrete tasks from each job advert (LLM-assisted).
  • AI exposure scoring: use an LLM to assign an exposure score to each task (1,542,411 tasks scored).
  • Job aggregation: compute job-level mean and variance of task exposure; identify tasks above exposure thresholds.
  • Thresholding: define a critical threshold θ above which tasks/jds are considered automatable; jobs with sufficient mass above θ are treated as fully automatable, otherwise eligible for redesign.
  • Job redesign: apply an LLM to remove automatable tasks and propose replacements or reallocated tasks focused on human comparative advantage (leadership, complex problem solving, stakeholder engagement, risk/quality management).
  • Economic simulation: combine redesigned task time reallocations with salary bands and staffing counts to estimate productivity gains and cost reductions.
  • Adjustments and checks
    • Iterative proportional fitting to adjust sample representation to the overall UKCS distribution by department/grade/profession.
    • Robustness: alternative redesign constraints and variations in θ to test sensitivity.
  • Limitations (methodological caveats noted by authors)
    • Reliance on vacancy text (may differ from on‑the‑job tasks).
    • Exposure and redesign depend on LLM judgments and prompt design.
    • Salary/cost estimates approximate (using grade/department median salaries).
    • Focused on one large public-sector employer; generalisability requires applying the pipeline to other datasets.

Implications for AI Economics

  • Granularity matters: task-level measurement reveals large within-occupation heterogeneity. Aggregate/occupation-level exposure indices risk overstating or misstating displacement risk.
  • Productivity vs displacement trade-off:
    • Many exposed jobs are only partially automatable; firms and public organisations may capture most value by redesigning jobs and reallocating human effort toward complementary tasks, producing productivity gains rather than mass layoffs.
    • Policy and firm strategies should prioritise redesign and reallocation mechanisms (training, workflow redesign, incentives) to capture gains.
  • Labour reallocation and skill demand:
    • Demand will shift toward human‑centric activities (leadership, complex problem solving, stakeholder management, risk/quality oversight). Upskilling and career‑path design should anticipate these shifts.
    • Senior/expert roles appear less automatable on average; complementarities may amplify returns to expertise, affecting wage and inequality dynamics.
  • Measurement and policy evaluation:
    • LLMs can be used both as measurement tools and prescriptive designers for job redesign at scale, but their outputs need validation (field trials, task-time studies) before large policy or staffing decisions.
    • Threshold choice (θ) for deciding between automation versus redesign is economically consequential; organisations should consider conservative thresholds and pilot trials to assess net welfare effects.
  • Research agenda:
    • Extend within‑firm, task-level analyses across sectors and countries to map heterogeneity in exposure and redesign potential.
    • Empirically evaluate realized productivity gains from redesign (randomised or quasi-experimental deployments) to validate monetary estimates and measure spillovers (task creation, new occupations).

Assessment

Paper Typedescriptive Evidence Strengthmedium — Large, novel dataset (193,497 job adverts, 1,542,411 tasks) gives comprehensive coverage of UK Civil Service vacancies and enables fine-grained, task-level analysis; however, all AI-exposure measures and job-redesign outputs are generated by LLM inference rather than observed outcomes, with limited or no validation against real-world productivity, displacement, or employer-verified task lists, reducing causal and empirical strength. Methods Rigormedium — Methodologically innovative: systematic extraction of tasks from adverts and use of an LLM to score tasks and propose redesigns; strengths include scale and task-level granularity, but rigor is limited by reliance on a single modelling approach (LLM-based labels), potential prompt/labeler sensitivity, possible measurement error in mapping adverts to tasks, limited validation or inter-rater checks reported, and aggregation choices (mean/variance at job level) that may mask important heterogeneity. SampleNovel dataset of UK Civil Service job adverts (193,497 vacancies) spanning six years, from which 1,542,411 tasks were extracted and scored for AI exposure using a large language model; analysis restricted to advertised roles in the UK public sector rather than private firms or other countries. Themeshuman_ai_collab productivity org_design GeneralizabilitySingle organisation and sector: UK Civil Service (public sector) — may not generalise to private sector firms or other institutional contexts, Country-specific: UK labour market, regulatory and job-structure features may limit applicability to other countries, Job adverts vs actual task content: advertised duties may be incomplete, stylised, or outdated compared with daily work, LLM-dependent measures: exposure scores and redesigns depend on model, prompts, and training data, which may embed biases or misestimate AI capabilities, Time-bound: six-year window and rapid AI progress mean findings could age quickly, No observed outcomes: lack of measured productivity, employment, wages, or displacement limits external validity for economic impact

Claims (8)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We assessed AI exposure at the task level within roles at the UK Civil Service (UKCS). Automation Exposure null_result AI exposure scores at the task level
Reading fidelity high
Study strength medium
not reported
0.18
We used a novel dataset of UKCS job adverts covering 193,497 vacancies over 6 years. Other null_result number of job adverts / dataset coverage
Reading fidelity high
Study strength high
n=193497
0.3
Our LLM-driven analysis estimated AI exposure scores for 1,542,411 tasks. Automation Exposure null_result count of tasks with estimated AI exposure scores
Reading fidelity high
Study strength medium
n=1542411
0.18
Aggregating task-level AI exposure scores to the job-level reveals heterogeneity in AI exposure (mean and variance) even for seemingly identical jobs. Automation Exposure mixed mean and variance of job-level AI exposure
Reading fidelity high
Study strength medium
not reported
0.18
We used an LLM to redesign jobs, focusing on task automation, task optimisation, and task reallocation. Task Allocation null_result job redesign outputs (tasks reclassified into automation, optimisation, reallocation)
Reading fidelity high
Study strength medium
not reported
0.18
The redesign process produces tasks where humans have comparative advantage over AI, including strategic leadership, complex problem resolution, and stakeholder management. Task Allocation positive types/categories of tasks where humans retain comparative advantage
Reading fidelity medium
Study strength medium
not reported
0.11
Automation and augmentation are expected to have nuanced effects across all levels of the organisational hierarchy. Organizational Efficiency mixed heterogeneity of AI impacts across organisational levels
Reading fidelity high
Study strength speculative
not reported
0.03
Most economic value of AI is expected to arise from productivity gains rather than role displacement. Firm Productivity positive source of economic value from AI (productivity gains versus role displacement)
Reading fidelity medium
Study strength speculative
not reported
0.02

Notes