0 cumulative citations
View corpus contextThere are roughly a thousand technically-competent ML researchers working in ML consulting firms worldwide — about twice the number of alumni from the MATS program — but only a small fraction of consultancies clear a hands-on research-and-engineering work trial, and no AI model passed that trial by late 2025.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
We identify a substantial pool of technically competent ML research talent (in the low thousands) in companies which offer consulting in machine learning. We systematically searched the internet, global business databases, and conference/paper affiliations for ML consulting firms. Employee LinkedIn resumes were then scored by keyword filters and large-language-model (LLM) classifiers; these signals were combined in a bootstrap probit model to estimate technical ML research talent per firm. A subset of companies also completed a 3-day research and engineering work trial. We screened 2121 organizations and found 403 offering broad ML consulting. Our 50th percentile aggregate estimate of 'highly technical' ML research talent across these organizations was 1121 (80% CI: 252-3165) -- i.e. twice as many as all alumni of the MATS training program. For our work trial 97 companies were approached, 20 applied, 8 were invited to participate, and 5 of 8 received at least a conditional recommendation for technical AI safety work. As of late 2025, no AI model was able to pass the work trial.
Summary
Main Finding
A systematic global sweep of IT consultancies finds a non-trivial, though modest, pool of technically capable ML research talent concentrated in a subset of firms. Across 403 firms (3.269 million associated employees) the median aggregate estimate of high-technical ML Research Talent is ~1,121 people (80% CI: 252–3,165). A smaller set of 81 “probable” consultancies account for most of that capacity (~890 median). Targeted work trials show some consultancies can execute demanding ML research tasks at high quality, while contemporary LLM agents could not.
Key Points
-
Sample and scope
- 2,121 organizations screened → 403 identified as offering ML consultancy services.
- LinkedIn-derived headcount associated with the sample: 3,269,000 employees.
- Company-size composition: 284 small (<100), 76 medium (100–999), 23 large (1,000–9,999), 20 giant (≥10,000).
- Geographic concentration: skewed to North America and Western Europe (130 and 52 firms respectively), but global coverage across regions.
-
Talent estimates
- Median aggregate technical ML Research Talent: 1,121 (80% CI 252–3,165).
- 81 firms (20% of the sample) have 80% CIs excluding zero and account for ~79% of estimated talent (~890 median).
- Overall ML share of total employees is low (~0.01% aggregate), but individual firms vary widely (up to ~20% in some units).
-
Validation & estimator performance
- Resume evaluation pipeline validated on 585 manually labeled CVs.
- Final estimator: sensitivity 0.79, specificity 0.926, accuracy 0.89, LR+ = 10.67, LR- = 0.23.
- Noted systematic underestimates for frontier labs where LLM-based CV assessment was unavailable and synthetic imputation was used (e.g., implausibly low estimates for Anthropic, Amazon, Meta, Microsoft, NVIDIA).
-
Work trials
- Outreach: 97 companies contacted; 20 applied; 8 invited; 5/8 received at least a conditional recommendation for technical AI safety work (3 recommended, 2 conditional, 3 no recommendation).
- Trial task: implement and integrate a Sequential Unlearning method over 3-day work trial (with a longer engagement).
- Price range observed: $45–$350 per hour (two firms volunteered trial work free).
- LLM agents (GPT-5, Claude Opus/Gemini variants) performed worse than consultancies on the task (agents scored 30–40% → no recommendations).
-
Limitations highlighted by authors
- English-language, LinkedIn-anchored search → probable underrepresentation of some regions/firms.
- Resumes are noisy proxies for competence; definitions excluded adjacent-but-relevant roles (e.g., ML Ops).
- Data gaps for very large companies and instances where LLM-based CV scoring was unavailable forced synthetic imputation that can misrepresent talent density.
- Small-scale work-trial sample (8 companies) and a single task limit generalizability.
Data & Methods
-
Company identification (June–Aug 2025)
- Sources: personal networks, web search assisted by LLMs, arXiv affiliation metadata, ICLR/ICML/NeurIPS affiliations (2019+), Crunchbase keyword searches.
- Exclusions: orgs with <10 staff on Crunchbase/LinkedIn or younger than 2 years.
-
Resume access & processing
- Data sources: LinkedIn Sales Navigator and Brightdata LinkedIn exports.
- Two parallel evaluation pipelines:
- Keyword-based filters (broad_yes, strict_no, broad_yes_strict_no) derived from analysis of 421 labeled CVs and a top-100 keyword selection process.
- LLM-based classification using prompts validated across models (selected models: Google Gemini-2.5-flash, OpenAI gpt-5-mini-thinking, Anthropic sonnet-4).
- Validation dataset: 585 CVs manually labeled by two human reviewers using a defined technical ML Research Talent criteria (ability to implement/train architectures, work from specs to code, debug model behavior, engage with research, communicate, public artifacts).
-
Aggregation & estimation
- Combined keyword and LLM signals.
- Missing LLM assessments handled with synthetic imputation in some cases.
- Final counts estimated via a bootstrap probit model producing per-organization distributions (q10, q50, q90), aggregated to produce the global median and confidence intervals.
- Additional validation: applied estimator to known high-talent AI orgs and known non-AI firms.
-
Work trials
- Recruitment and selection pipeline (97 approached → 8 trialed).
- 3-day technical task: implement Sequential Unlearning wrapper + integration, daily updates, final code + write-up.
- Trials evaluated against explicit technical criteria; same criteria applied to LLM agent runs for comparison.
-
Dates and provenance
- Crunchbase extraction: 8 August 2025.
- Work trials: July–August 2025.
- Study funded by Coefficient Giving; authors report no conflicts.
Implications for AI Economics
-
Latent labor capacity for AI assurance exists outside frontier labs
- IT consultancies contain a scalable, deployable pool (low thousands by this estimate) of capable ML personnel that could be redirected toward AI safety, evaluation, and alignment tasks more rapidly than expanding academic programs alone.
- Funders and policymakers seeking to expand applied safety capacity should consider consultancies as a complementary labor channel.
-
Heterogeneity matters for targeting and pricing
- Talent is highly unevenly distributed across consultancies and within large conglomerates; program designers must target sub-units with high technical density rather than treating “consultancy” as homogeneous.
- Observed hourly rates ($45–$350) indicate market pricing diversity; budget planning for scaling assurance work should account for this spread.
-
Contracting, coordination, and governance considerations
- Consultancies often prefer clear scope, an external “vision holder,” and faster/smaller contracting paths for effective engagement—implications for procurement design and program management.
- Incentive alignment and quality assurance (e.g., work trials, artifact checks, integration of public outputs) will be critical to reliably convert latent capacity into trustworthy safety work.
-
Labor market effects and comparative advantage
- Engaging consultancies could relieve bottlenecks at frontier labs for certain applied evaluation and engineering tasks, freeing frontier researchers to focus on frontier research.
- However, consultancies may lack deep research-only capacity for theoretical or frontier alignment problems; matching task types to firm strengths is important.
-
Risks, verification, and scale limits
- Resume-based identification has limits—verification via public artifacts, code reviews, publication checks, and practical trials will be necessary when awarding important safety work.
- The measured pool is small relative to total AI labor demand; scaling beyond the low-thousands requires training pipelines, incentives, or new hiring.
-
Recommendations for funders and researchers
- Use focused outreach to the subset of consultancies with demonstrated technical capability (work-trial validated).
- Fund targeted short-term contracts with embedded verification (work trials, code artefacts), and fund capacity-building programs (secondments, sponsored projects) to increase durable supply.
- Invest in improved measurement: broader LinkedIn sweeps, inclusion of non-English sources, linking CVs to GitHub/publications, and domain-specific priors to reduce imputation errors.
Limitations to keep in mind - The central estimate is sensitive to the authors’ operational definition of “technical ML Research Talent,” the LinkedIn-centric sampling frame, and situations where LLM-based CV evaluation was unavailable and imputed. The true global supply—especially in underrepresented geographies and in adjacent skill areas (ML Ops, data engineering)—remains uncertain and likely larger when including adjacent competencies.
If helpful, I can extract the main numeric tables (company-size breakdown, estimator performance, work-trial outcomes) into a concise CSV-style summary or produce a short checklist for funders on how to operationalize outreach to consultancies based on the paper’s findings.
Assessment
Claims (12)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| We screened 2121 organizations. Adoption Rate | null_result | number of organizations screened |
Reading fidelity
high
Study strength
high
|
n=2121
2121 organizations screened
|
| We found 403 organizations offering broad ML consulting. Adoption Rate | positive | number of firms offering broad ML consulting |
Reading fidelity
high
Study strength
high
|
n=403
403 organizations offering broad ML consulting
|
| We identify a substantial pool of technically competent ML research talent (in the low thousands) in companies which offer consulting in machine learning. Employment | positive | count of technically competent ML research personnel across ML consulting firms |
Reading fidelity
high
Study strength
medium
|
n=403
in the low thousands
|
| Our 50th percentile aggregate estimate of 'highly technical' ML research talent across these organizations was 1121 (80% CI: 252-3165). Employment | positive | median estimate of 'highly technical' ML research talent |
Reading fidelity
high
Study strength
medium
|
n=403
1121 (80% CI: 252-3165)
|
| The 50th percentile estimate (1121) is twice as many as all alumni of the MATS training program. Employment | positive | relative size of ML research talent pool versus MATS alumni |
Reading fidelity
medium
Study strength
medium
|
twice as many
|
| Employee LinkedIn resumes were scored by keyword filters and large-language-model (LLM) classifiers; these signals were combined in a bootstrap probit model to estimate technical ML research talent per firm. Other | null_result | method used to estimate per-firm technical ML talent |
Reading fidelity
high
Study strength
high
|
not reported
|
| A subset of companies completed a 3-day research and engineering work trial. Training Effectiveness | null_result | participation in a 3-day work trial |
Reading fidelity
high
Study strength
high
|
not reported
|
| For the work trial, 97 companies were approached. Hiring | null_result | number of firms approached for the work trial |
Reading fidelity
high
Study strength
high
|
n=97
97 companies approached
|
| 20 companies applied to participate in the work trial. Hiring | positive | number of firms that applied to the work trial |
Reading fidelity
high
Study strength
high
|
n=20
20 companies applied
|
| 8 companies were invited to participate in the work trial. Hiring | positive | number of firms invited to the trial |
Reading fidelity
high
Study strength
high
|
n=8
8 companies invited
|
| 5 of 8 participating companies received at least a conditional recommendation for technical AI safety work. Hiring | positive | recommendation outcome for firms following the work trial |
Reading fidelity
high
Study strength
medium
|
n=8
5 of 8
|
| As of late 2025, no AI model was able to pass the work trial. Output Quality | negative | whether AI models can pass the 3-day research & engineering work trial |
Reading fidelity
high
Study strength
medium
|
not reported
|