The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

There are roughly a thousand technically-competent ML researchers working in ML consulting firms worldwide — about twice the number of alumni from the MATS program — but only a small fraction of consultancies clear a hands-on research-and-engineering work trial, and no AI model passed that trial by late 2025.

How much technical talent is there? A systematic estimate of the ML research pool among 3 million consultants
Maximilian Schons, Red Bermejo, Florian Aldehoff-Zeidler, Niccolò Zanichelli, Oliver Evans, Gavin Leech, Samuel Härgestam · February 10, 2026
arxiv descriptive medium evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Maximilian Schons unresolved corpus identity
  2. Red Bermejo unresolved corpus identity
  3. Florian Aldehoff-Zeidler unresolved corpus identity
  4. Niccolò Zanichelli unresolved corpus identity
  5. Oliver Evans unresolved corpus identity
  6. Gavin Leech unresolved corpus identity
  7. Samuel Härgestam unresolved corpus identity

Semantic Scholar

Latest observation:

  1. M. Schons provider ID
  2. Red Bermejo provider ID
  3. Florian Aldehoff-Zeidler provider ID
  4. Niccoló Zanichelli provider ID
  5. O. Evans provider ID
  6. Gavin Leech provider ID
  7. Samuel Härgestam provider ID
A systematic search of ML consulting firms estimates roughly 1,121 'highly technical' ML research staff across 403 firms (80% CI 252–3,165), and a small hands-on trial found a handful of firms suitable for advanced AI-safety work while no AI model passed the trial as of late 2025.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

We identify a substantial pool of technically competent ML research talent (in the low thousands) in companies which offer consulting in machine learning. We systematically searched the internet, global business databases, and conference/paper affiliations for ML consulting firms. Employee LinkedIn resumes were then scored by keyword filters and large-language-model (LLM) classifiers; these signals were combined in a bootstrap probit model to estimate technical ML research talent per firm. A subset of companies also completed a 3-day research and engineering work trial. We screened 2121 organizations and found 403 offering broad ML consulting. Our 50th percentile aggregate estimate of 'highly technical' ML research talent across these organizations was 1121 (80% CI: 252-3165) -- i.e. twice as many as all alumni of the MATS training program. For our work trial 97 companies were approached, 20 applied, 8 were invited to participate, and 5 of 8 received at least a conditional recommendation for technical AI safety work. As of late 2025, no AI model was able to pass the work trial.

Summary

Main Finding

A systematic global sweep of IT consultancies finds a non-trivial, though modest, pool of technically capable ML research talent concentrated in a subset of firms. Across 403 firms (3.269 million associated employees) the median aggregate estimate of high-technical ML Research Talent is ~1,121 people (80% CI: 252–3,165). A smaller set of 81 “probable” consultancies account for most of that capacity (~890 median). Targeted work trials show some consultancies can execute demanding ML research tasks at high quality, while contemporary LLM agents could not.

Key Points

  • Sample and scope

    • 2,121 organizations screened → 403 identified as offering ML consultancy services.
    • LinkedIn-derived headcount associated with the sample: 3,269,000 employees.
    • Company-size composition: 284 small (<100), 76 medium (100–999), 23 large (1,000–9,999), 20 giant (≥10,000).
    • Geographic concentration: skewed to North America and Western Europe (130 and 52 firms respectively), but global coverage across regions.
  • Talent estimates

    • Median aggregate technical ML Research Talent: 1,121 (80% CI 252–3,165).
    • 81 firms (20% of the sample) have 80% CIs excluding zero and account for ~79% of estimated talent (~890 median).
    • Overall ML share of total employees is low (~0.01% aggregate), but individual firms vary widely (up to ~20% in some units).
  • Validation & estimator performance

    • Resume evaluation pipeline validated on 585 manually labeled CVs.
    • Final estimator: sensitivity 0.79, specificity 0.926, accuracy 0.89, LR+ = 10.67, LR- = 0.23.
    • Noted systematic underestimates for frontier labs where LLM-based CV assessment was unavailable and synthetic imputation was used (e.g., implausibly low estimates for Anthropic, Amazon, Meta, Microsoft, NVIDIA).
  • Work trials

    • Outreach: 97 companies contacted; 20 applied; 8 invited; 5/8 received at least a conditional recommendation for technical AI safety work (3 recommended, 2 conditional, 3 no recommendation).
    • Trial task: implement and integrate a Sequential Unlearning method over 3-day work trial (with a longer engagement).
    • Price range observed: $45–$350 per hour (two firms volunteered trial work free).
    • LLM agents (GPT-5, Claude Opus/Gemini variants) performed worse than consultancies on the task (agents scored 30–40% → no recommendations).
  • Limitations highlighted by authors

    • English-language, LinkedIn-anchored search → probable underrepresentation of some regions/firms.
    • Resumes are noisy proxies for competence; definitions excluded adjacent-but-relevant roles (e.g., ML Ops).
    • Data gaps for very large companies and instances where LLM-based CV scoring was unavailable forced synthetic imputation that can misrepresent talent density.
    • Small-scale work-trial sample (8 companies) and a single task limit generalizability.

Data & Methods

  • Company identification (June–Aug 2025)

    • Sources: personal networks, web search assisted by LLMs, arXiv affiliation metadata, ICLR/ICML/NeurIPS affiliations (2019+), Crunchbase keyword searches.
    • Exclusions: orgs with <10 staff on Crunchbase/LinkedIn or younger than 2 years.
  • Resume access & processing

    • Data sources: LinkedIn Sales Navigator and Brightdata LinkedIn exports.
    • Two parallel evaluation pipelines:
      • Keyword-based filters (broad_yes, strict_no, broad_yes_strict_no) derived from analysis of 421 labeled CVs and a top-100 keyword selection process.
      • LLM-based classification using prompts validated across models (selected models: Google Gemini-2.5-flash, OpenAI gpt-5-mini-thinking, Anthropic sonnet-4).
    • Validation dataset: 585 CVs manually labeled by two human reviewers using a defined technical ML Research Talent criteria (ability to implement/train architectures, work from specs to code, debug model behavior, engage with research, communicate, public artifacts).
  • Aggregation & estimation

    • Combined keyword and LLM signals.
    • Missing LLM assessments handled with synthetic imputation in some cases.
    • Final counts estimated via a bootstrap probit model producing per-organization distributions (q10, q50, q90), aggregated to produce the global median and confidence intervals.
    • Additional validation: applied estimator to known high-talent AI orgs and known non-AI firms.
  • Work trials

    • Recruitment and selection pipeline (97 approached → 8 trialed).
    • 3-day technical task: implement Sequential Unlearning wrapper + integration, daily updates, final code + write-up.
    • Trials evaluated against explicit technical criteria; same criteria applied to LLM agent runs for comparison.
  • Dates and provenance

    • Crunchbase extraction: 8 August 2025.
    • Work trials: July–August 2025.
    • Study funded by Coefficient Giving; authors report no conflicts.

Implications for AI Economics

  • Latent labor capacity for AI assurance exists outside frontier labs

    • IT consultancies contain a scalable, deployable pool (low thousands by this estimate) of capable ML personnel that could be redirected toward AI safety, evaluation, and alignment tasks more rapidly than expanding academic programs alone.
    • Funders and policymakers seeking to expand applied safety capacity should consider consultancies as a complementary labor channel.
  • Heterogeneity matters for targeting and pricing

    • Talent is highly unevenly distributed across consultancies and within large conglomerates; program designers must target sub-units with high technical density rather than treating “consultancy” as homogeneous.
    • Observed hourly rates ($45–$350) indicate market pricing diversity; budget planning for scaling assurance work should account for this spread.
  • Contracting, coordination, and governance considerations

    • Consultancies often prefer clear scope, an external “vision holder,” and faster/smaller contracting paths for effective engagement—implications for procurement design and program management.
    • Incentive alignment and quality assurance (e.g., work trials, artifact checks, integration of public outputs) will be critical to reliably convert latent capacity into trustworthy safety work.
  • Labor market effects and comparative advantage

    • Engaging consultancies could relieve bottlenecks at frontier labs for certain applied evaluation and engineering tasks, freeing frontier researchers to focus on frontier research.
    • However, consultancies may lack deep research-only capacity for theoretical or frontier alignment problems; matching task types to firm strengths is important.
  • Risks, verification, and scale limits

    • Resume-based identification has limits—verification via public artifacts, code reviews, publication checks, and practical trials will be necessary when awarding important safety work.
    • The measured pool is small relative to total AI labor demand; scaling beyond the low-thousands requires training pipelines, incentives, or new hiring.
  • Recommendations for funders and researchers

    • Use focused outreach to the subset of consultancies with demonstrated technical capability (work-trial validated).
    • Fund targeted short-term contracts with embedded verification (work trials, code artefacts), and fund capacity-building programs (secondments, sponsored projects) to increase durable supply.
    • Invest in improved measurement: broader LinkedIn sweeps, inclusion of non-English sources, linking CVs to GitHub/publications, and domain-specific priors to reduce imputation errors.

Limitations to keep in mind - The central estimate is sensitive to the authors’ operational definition of “technical ML Research Talent,” the LinkedIn-centric sampling frame, and situations where LLM-based CV evaluation was unavailable and imputed. The true global supply—especially in underrepresented geographies and in adjacent skill areas (ML Ops, data engineering)—remains uncertain and likely larger when including adjacent competencies.

If helpful, I can extract the main numeric tables (company-size breakdown, estimator performance, work-trial outcomes) into a concise CSV-style summary or produce a short checklist for funders on how to operationalize outreach to consultancies based on the paper’s findings.

Assessment

Paper Typedescriptive Evidence Strengthmedium — The paper uses a systematic search and quantitative estimation (keyword filters, LLM classifiers, and a bootstrap probit model) to produce a numeric estimate with uncertainty intervals, which provides useful empirical evidence; however, the measurement relies on noisy signals (resumes/LinkedIn and classifiers), the CI is wide, and the small, selective work-trial sample limits confidence in generalizing the trial findings. Methods Rigormedium — The authors applied a reproducible search frame across multiple databases, combined multiple automated signals, and used a bootstrap probit to quantify uncertainty, which are sound methods for a descriptive count; but the approach lacks substantial external validation, depends on potentially biased resume data and classifier decisions, and the hands-on work trial has a small, self-selected sample and limited reporting on trial tasks and scoring. SampleSystematic screening of 2,121 organizations using internet searches, global business databases, and conference/paper affiliations identified 403 firms offering broad ML consulting; employee LinkedIn resumes were scored via keyword filters and LLM classifiers and combined in a bootstrap probit model to estimate counts of 'highly technical' ML research talent (50th percentile = 1,121; 80% CI: 252–3,165); a follow-up work trial approached 97 companies, 20 applied, 8 were invited, and 5 of 8 received at least a conditional recommendation for technical AI safety work; as of late 2025 no AI model passed the trial. Themeslabor_markets skills_training GeneralizabilityLimited to organizations offering ML consulting (excludes in-house R&D at product firms, academia, and government labs), Dependent on online presence and LinkedIn use — geographic and language biases likely (e.g., undercounts in regions with less LinkedIn penetration), Definition of 'highly technical' depends on keywords/classifier thresholds and may not map perfectly to actual research ability, Possible classifier and resume-reporting errors (overstatement or understatement of skills), Work-trial findings are based on a small, self-selected subset and may not generalize to the broader consulting talent pool, Time-specific snapshot (late 2025); talent levels and model capabilities may change rapidly

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
We screened 2121 organizations. Adoption Rate null_result number of organizations screened
Reading fidelity high
Study strength high
n=2121
2121 organizations screened
0.3
We found 403 organizations offering broad ML consulting. Adoption Rate positive number of firms offering broad ML consulting
Reading fidelity high
Study strength high
n=403
403 organizations offering broad ML consulting
0.3
We identify a substantial pool of technically competent ML research talent (in the low thousands) in companies which offer consulting in machine learning. Employment positive count of technically competent ML research personnel across ML consulting firms
Reading fidelity high
Study strength medium
n=403
in the low thousands
0.18
Our 50th percentile aggregate estimate of 'highly technical' ML research talent across these organizations was 1121 (80% CI: 252-3165). Employment positive median estimate of 'highly technical' ML research talent
Reading fidelity high
Study strength medium
n=403
1121 (80% CI: 252-3165)
0.18
The 50th percentile estimate (1121) is twice as many as all alumni of the MATS training program. Employment positive relative size of ML research talent pool versus MATS alumni
Reading fidelity medium
Study strength medium
twice as many
0.11
Employee LinkedIn resumes were scored by keyword filters and large-language-model (LLM) classifiers; these signals were combined in a bootstrap probit model to estimate technical ML research talent per firm. Other null_result method used to estimate per-firm technical ML talent
Reading fidelity high
Study strength high
not reported
0.3
A subset of companies completed a 3-day research and engineering work trial. Training Effectiveness null_result participation in a 3-day work trial
Reading fidelity high
Study strength high
not reported
0.3
For the work trial, 97 companies were approached. Hiring null_result number of firms approached for the work trial
Reading fidelity high
Study strength high
n=97
97 companies approached
0.3
20 companies applied to participate in the work trial. Hiring positive number of firms that applied to the work trial
Reading fidelity high
Study strength high
n=20
20 companies applied
0.3
8 companies were invited to participate in the work trial. Hiring positive number of firms invited to the trial
Reading fidelity high
Study strength high
n=8
8 companies invited
0.3
5 of 8 participating companies received at least a conditional recommendation for technical AI safety work. Hiring positive recommendation outcome for firms following the work trial
Reading fidelity high
Study strength medium
n=8
5 of 8
0.18
As of late 2025, no AI model was able to pass the work trial. Output Quality negative whether AI models can pass the 3-day research & engineering work trial
Reading fidelity high
Study strength medium
not reported
0.18

Notes