0 cumulative citations
View corpus contextAI coding assistants are reshaping software engineering rather than replacing it: firms are substituting automation for routine junior work—cutting entry-level hiring—while value and risk move toward architecture, verification and governance; unreviewed AI-generated code raises measurable security and technical-debt concerns.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Artificial intelligence has moved rapidly from an experimental add-on to a routine part of software development, with AI coding assistants now writing functions, generating tests, and, in agentic form, completing multi-step engineering tasks with limited human oversight. This raises a familiar question: does automating part of a skilled trade erode the expertise around it, and who bears the cost? Drawing on randomized trials, labour-market studies, industry surveys, and code-security audits, this report examines how AI adoption is reshaping human expertise, employment, and code quality in software engineering. Entry-level employment is contracting even as overall technical demand rises, and AI-written code merged without scrutiny carries more security flaws and technical debt. Software engineering is being reorganized rather than elimin0ated, with value shifting toward architecture, verification, and accountability. The report closes with recommendations for employers, institutions, and regulators.
Summary
Main Finding
AI coding tools are reshaping software engineering rather than eliminating it: routine implementation is increasingly automated, value shifts toward architecture/verification/accountability, entry-level hiring is contracting, and unreviewed AI-generated code raises measurable security and technical-debt risks. The net effect is a reallocation of labor and skills within the software sector with short-term contraction for junior roles and medium-term growth for judgment- and AI-oversight–intensive roles.
Key Points
- Productivity effects are heterogeneous:
- RCTs and trials report results from a ~19% slowdown (early-2025 tools) to an ~18–56% speed-up (newer models or bounded tasks), depending on model generation, task scope, and developer experience.
- AI helps most on routine tasks (reported ~30–60% time savings in some surveys) and least on complex, unfamiliar systems.
- Perceived speed-ups often exceed measured gains.
- Employment and skill formation:
- Payroll data (2021–2025) show ~20% decline in employment for developers aged 22–25 (late-2022 to mid-2025), while experienced developer employment rose (~6–12%).
- Job postings remain below pre-pandemic levels, with sharper falls for junior roles.
- New roles (AI-governance, AI-validation, hybrid oversight) command wage premiums; AI fluency pays more in entry-level roles that require evaluation/directing of AI.
- Risk of skill atrophy: automating code generation can weaken apprenticeship pathways because evaluating AI output often requires hands-on generative experience.
- Code quality, security, and technical debt:
- Large-sample security audits find only ~55% of AI-generated samples free of known security issues; earlier Copilot evaluations found ~40% of generated code contained vulnerabilities.
- Merged AI contributions without adequate review are associated with higher code churn, mismatched dependencies, architectural debt, and more maintenance costs later.
- AI shifts effort from writing to evaluating code, and evaluation is often more error-prone and context-dependent.
- Opportunities:
- Lowers barriers for domain experts to prototype (“vibe coding”).
- Frees experienced engineers to focus on architecture and product judgment if organizations deliberately reallocate freed time.
- Creates premium roles around AI oversight, governance, and review pipelines.
- Policy and organizational recommendations (summarized):
- Mandate human review of AI-generated code for critical systems; redesign junior roles to preserve skill formation; keep foundational CS education AI-independent while adding AI-collaboration training; support transparency, reskilling, and independent research.
Data & Methods
- Approach: structured narrative synthesis combining randomized trials, payroll-based labor analyses, industry surveys, and code-security audits (preprints, peer-reviewed studies, and institutional reports).
- Key empirical sources and findings referenced in the paper:
- METR RCTs with experienced developers: early-2025 model slowed performance (~19%); follow-up with newer models showed ~18% speed-up.
- Peng et al. RCT on GitHub Copilot on bounded tasks: ~56% faster completion for assisted group.
- Industry surveys (McKinsey, Google DORA) reporting sizeable routine-task savings but mixed effects on delivery stability.
- Payroll/administrative data showing ~20% fall in employment among 22–25-year-old developers (late-2022 to mid-2025) and growth for experienced engineers.
- Security audits across >100 LLMs: only ~55% of samples free of known security issues; earlier Copilot studies found ~40% of generated code with vulnerabilities.
- Production-repository analyses linking AI-assisted merges to higher churn and architectural debt.
- Limitations acknowledged:
- Evidence base is young and rapidly changing; many labour-market findings are correlational (not strictly causal).
- Heterogeneity by model generation, task type, developer experience, and organizational practices reduces external comparability.
- Formal meta-analysis was not feasible due to differing designs; recommendations are based on synthesis rather than pooled causal estimates.
Implications for AI Economics
- Labor reallocation, not net elimination (short-to-medium run):
- Demand shifts from routine implementation to higher-skilled roles (architecture, verification, AI oversight). This produces heterogeneous wage and employment effects: contraction for junior, growth and wage premiums for AI-adjacent senior roles.
- Human capital formation risk:
- If junior hiring and hands-on coding opportunities shrink, long-term supply of senior engineers with deep evaluative judgment may fall, raising future labor scarcity and wage pressures for senior talent.
- Hidden/shifted costs and productivity accounting:
- Measured time savings on coding can mask increased downstream maintenance, review costs, and delivery instability. Economic evaluation of AI should pair productivity metrics with quality, churn, and defect costs.
- Market structure and task bundling:
- Firms may substitute AI for junior hires where budget-constrained, concentrating judgment tasks among fewer, more expensive senior staff or new specialist roles—altering team composition and career ladders.
- Policy and redistribution considerations:
- Public investment in reskilling, apprenticeship replacements, and independent security research can mitigate short-term dislocation and longer-term skill-supply risks.
- Regulation or standards that require transparency and mandatory review for AI-generated code in safety-critical domains can internalize risks and change firm incentives.
- Research priorities for economists:
- Causal identification of AI’s effect on employment and wages, longitudinal studies of skill development under AI use, and firm-level evaluations of organizational practices that best capture AI benefits while containing risks.
- Strategic implication for firms and policymakers:
- To realize productivity gains without increasing systemic risk, organizations must invest in verification pipelines, training, and deliberate hiring practices; policymakers should support metrics, transparency, and programs that preserve human capital formation.
If you want, I can (a) extract a one-page bulletable policy brief for employers or (b) produce a short table mapping research gaps to specific empirical designs economists could use. Which would you prefer?
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| In a randomized trial of experienced developers working in large, unfamiliar codebases, early-2025 AI tools reduced measured developer speed by approximately 19%, despite developers predicting a 24% speed-up. Developer Productivity | negative | Developer productivity measured by task completion speed |
Reading fidelity
high
Study strength
high
|
n=16
about 19% slowdown
|
| A follow-up using newer AI model generations reversed the earlier result, producing roughly an 18% speed-up for experienced developers. Developer Productivity | positive | Developer productivity measured by task completion speed |
Reading fidelity
high
Study strength
medium
|
roughly an 18% speed-up
|
| Developers using GitHub Copilot on a tightly bounded coding task completed the task approximately 56% faster than control-group developers. Task Completion Time | positive | Task completion time |
Reading fidelity
high
Study strength
medium
|
about 56% faster
|
| An industry survey reported approximately 46% time savings from AI on routine software-development tasks, but less than 10% savings on complex work. Task Completion Time | mixed | Time savings from AI-assisted development by task complexity |
Reading fidelity
high
Study strength
medium
|
n=4500
about 46% time savings on routine tasks; under 10% on complex work
|
| Google's DORA survey found that individual developer effectiveness increased by 17%, while delivery stability declined by nearly 10%. Developer Productivity | mixed | Individual developer effectiveness and software delivery stability |
Reading fidelity
high
Study strength
medium
|
individual effectiveness +17%; delivery stability -10%
|
| Early evaluations found that approximately 40% of GitHub Copilot-generated code contained security vulnerabilities, and developers frequently judged their own insecure AI-assisted code to be safe. Error Rate | negative | Security vulnerability rate in AI-generated code and developers' security assessment accuracy |
Reading fidelity
high
Study strength
medium
|
about 40% contained vulnerabilities
|
| Across testing of more than 100 large language models, only approximately 55% of AI-generated code samples were free of known security issues. Error Rate | negative | Share of AI-generated code samples free of known security flaws |
Reading fidelity
high
Study strength
medium
|
about 55% free of known security issues
|
| Production-repository studies associate AI-assisted development with higher code churn and increased architectural or technical debt when AI-generated contributions are merged without adequate review. Output Quality | negative | Code churn and technical or architectural debt |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Employment among developers aged 22–25 fell nearly 20% from its late-2022 peak between 2021 and mid-2025, while employment among more experienced developers grew by approximately 6–12%. Employment | mixed | Employment by developer career stage |
Reading fidelity
high
Study strength
medium
|
nearly 20% decline among developers aged 22–25; 6–12% growth among experienced developers
|
| Software-development job postings remained approximately one-third below pre-pandemic or 2020 levels through late 2025, with junior postings declining more sharply than senior postings. Hiring | negative | Availability of software-development job postings, especially junior roles |
Reading fidelity
high
Study strength
medium
|
about one-third below 2020 levels
|
| The World Economic Forum projects a net global gain of 78 million jobs by 2030, alongside 92 million job losses in more automatable categories, with software- and AI-related roles among the fastest-growing. Employment | mixed | Projected employment gains and losses through 2030 |
Reading fidelity
high
Study strength
low
|
78 million net gain; 92 million losses
|
| AI-assisted development can create skill-formation risks: using AI beyond one's current understanding may crowd out learning, while relying heavily on AI for already-mastered tasks may erode hands-on skills. Skill Obsolescence | negative | Development and retention of hands-on software-engineering skills |
Reading fidelity
high
Study strength
low
|
not reported
|