The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A 'dynamic O*NET' from 752 million Chinese job ads shows posted job requirements and tasks can be read monthly; since late 2023 AI-exposed occupations have lost posting share mainly because those occupations are posted less often, not because surviving occupations systematically remove AI-exposed tasks.

The Pulse Beneath the Job Title: Monthly Readings of Requirements and Tasks from 750 Million Chinese Job Ads
Qin Chen, Ying Fang, Xiangyu Wang, Leo Yang Yang · August 27, 2026
arxiv descriptive medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Qin Chen unresolved corpus identity
  2. Ying Fang unresolved corpus identity
  3. Xiangyu Wang unresolved corpus identity
  4. Leo Yang Yang unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Qin Chen unresolved corpus identity
  2. Ying Fang provider ID
  3. Xiang-Yu Wang provider ID
  4. Leo Yang Yang unresolved corpus identity
The paper builds a monthly 'dynamic O*NET' for China from 752.6 million job ads, producing standardized catalogs of 20,721 requirements and 44,479 tasks (with task-level LLM-exposure scores) and documents that declines in AI-exposed content in posted demand have occurred mainly because exposed occupations shrink rather than because surviving occupations systematically shed exposed tasks.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

How do we define an occupation? By its job title? An accountant at a small trading company keeps the books; at a listed firm the same title demands a certified-accountant licence, and the week goes to the reports that regulators and the board read. Same title, different bar, different work. What defines an occupation is who it lets in and what it asks them to do. In a rapidly changing labor market, tracking those requirements and tasks is how to take the market's pulse. Yet no instrument reads both at the speed they change. Official occupational directories like O*NET report one national average per occupation, updated every few years. Job postings are timely but unstructured. Research built on them works from job titles plus proprietary skill keywords, which blur what is asked of a candidate into what a candidate is asked to do. The blur matters, because rising requirements and changing tasks are different events with different causes. We separate them. From 752.6 million job ads posted on China's five leading recruitment platforms between 2022 and 2026, we extract the phrases employers write, unify those that name the same thing, and validate the mapping from text back to entry. By doing so we construct two catalogs, 20,721 requirements a candidate must meet and 44,479 tasks the hire will do. With the entries standardized, we annotate them further. Each task, for example, carries a score for how far a language model could absorb it. Matched back onto every ad, the catalogs read the market month by month. Two examples show what the layer beneath the job title buys. First, the occupational registry records one accountant where the ads record a staircase, the junior certificate at the bottom of the wage range and the intermediate one at the top. Second, counting occupations says the work most exposed to language models is disappearing, and counting tasks says far less of it is.

Summary

Main Finding

The authors build a “dynamic O*NET” for China by turning 752.6 million Chinese job ads (Jan 2022–Jun 2026) into a standardized, monthly-updatable vocabulary of what employers require and what hires will do. They construct 20,721 requirement entries and 44,479 task entries, score each entry on multiple dimensions (including exposure to large language models), and map the vocabulary back to every ad (116 million unique descriptions, 12.1 billion item–posting mentions). Key empirical findings: (1) occupational averages hide large within-title heterogeneity (e.g., accountants show a clear task/wage staircase and credential segmentation); (2) observed declines in demand for LLM‑exposed work operate mostly through which occupations are posted (occupational posting-share declines), not by systematic stripping of exposed tasks inside surviving occupations — i.e., tasks tend to outlive the jobs that carried them. Since late 2023, posted task composition has drifted away from LM-absorbable work toward more hands-on, on-site, higher-stakes tasks.

Key Points

  • Scale and products:
    • Data: 752.6 million postings across five leading Chinese platforms (Jan 2022–Jun 2026); 116 million unique description texts after de-duplication.
    • Vocabulary: 20,721 requirements (credentials, identity/physical screens, etc.) and 44,479 tasks (what hires will do).
    • Matches: 12.1 billion item–posting mentions; monthly, sliceable by city, wage, education, employer type.
  • Construction pipeline (cost‑efficient, reusable):
    • LLM-assisted atomic-phrase extraction on a 2.31M-description stratified sample (one action/ability/credential/condition per phrase).
    • Embedding-based clustering; model adjudication of ~520k boundary pairs; connected-component merging to unify surface variants (e.g., “maintain client relationships” absorbed 797 variants).
    • Validation: generate multiple candidate regex matchers per entry, score against benchmark, keep best; drop failures; freeze vocabulary for fast re-scan.
    • Human/model adjudication cost for vocabulary boundaries ~¥3,300 (far cheaper than naive per-title classification, which would cost ~¥9.7M per 100M titles).
  • Measurement features:
    • The system keeps requirements (who is allowed in) distinct from tasks (what is done) — a separation often conflated in commercial taxonomies.
    • Each entry is annotated on many dimensions (eleven of thirteen for tasks; nine for requirements), including a numeric LLM-exposure score.
    • LLM-exposure scores correlate r = 0.73–0.79 with leading published exposure measures, supporting construct validity.
  • Representative evidence and patterns:
    • Accountants: tasks shift across wage bands (core bookkeeping dominates low wages; compliance/data analysis rise at high wages); certification requirements and intermediate-title shares increase sharply with wage.
    • Production workers: average LLM exposure rises from ~0.17 (low wage band) to ~0.48 (high wage band) because higher-paying postings emphasize documentation and coordination rather than assembly.
    • Java developers: task and requirement bundles are stable across experience levels (low certificate incidence 3–5%).
    • Exposure dynamics: occupations composed largely of exposed tasks are losing posting share faster (slope ≈ −1.33) than the broad task categories themselves (slope ≈ −0.83); shift-share decompositions indicate posting-share channel (which occupations are posted) accounts for most of the task-level decline.
  • Data coverage & caveats:
    • Structured fields coverage: posted wages present on 97.4% of ads; education requirement on 65.8%; experience on 50.8%.
    • Platform/selection: five platforms anonymized (A–E); growth in posting counts partly reflects platform coverage expansion — time-series analyses use within-platform shares or constant-platform subsets.
    • Representativeness: online postings overrepresent urban, formal, white‑collar, and service hiring (e.g., agriculture nearly absent). The instrument measures advertised demand, not realized employment or wages.

Data & Methods

  • Data source and scope:
    • Five leading Chinese recruitment platforms (anonymized), Jan 2022–Jun 2026, ~1.8 TB text; 752.6M postings total; 116M unique texts after de-duplication.
    • Each ad: free-text title, description (tasks & requirements), and structured fields (city, wage range, education, experience, employer size, industry tags, major requirements).
  • Vocabulary construction (three-stage, cost‑bounded):
  • Extraction: LLM extracts atomic phrases from a stratified 2.31M-description sample (single atomic unit per phrase).
  • Unification: embedding clustering groups similar phrasings; an adjudication model answers ~520k pairwise boundary questions (are these the same entry?); connected components merge variants.
  • Validation & matching rules: for each entry generate up to five regex matchers, score them on the benchmark sample, keep best-performing pattern; drop entries with failing matches; freeze vocabulary.
  • Matching and annotation:
    • Frozen vocabulary applied to every unique description; matches reattached to postings (posting-weighting reflects advertised demand in city-month-wage slices).
    • Each task and requirement scored on multiple dimensions (machine/people/place/institution relations, LLM exposure, etc.). Exposure measure validated via correlation with published measures.
  • Cost and efficiency:
    • Adjudicating vocabulary boundaries: ~¥3,300.
    • Once vocabulary is frozen, re-scanning the entire historical corpus runs in hours — enabling monthly updates and re-scoring.
  • Validation and robustness:
    • Multiple robustness checks: within-platform time series, unique-description weighting, alternative weighting conventions; reported sensitivity (e.g., occupation gradient steepens under unique-description weighting).
    • Correlations with external measures (r = 0.73–0.79) and illustrative case studies (accountant, production worker, Java developer).

Implications for AI Economics

  • Measurement innovation for AI impact studies:
    • Task-level, monthly-updated exposure scores allow more precise, time‑sensitive measurement of which parts of labor demand are LM‑absorbable versus resilient.
    • Separating requirements from tasks prevents conflation of credential/entry barriers (which can rise independently) with task content; this matters for interpreting posting-based “AI hiring” signals.
  • Rethinking displacement and adjustment pathways:
    • Empirical result: declines in demand for exposed work primarily reflect which occupations employers choose to post (re-weighting across occupations), not systematic within-occupation replacement of exposed tasks. Thus:
      • Forecasts that treat occupations as indivisible atoms may overstate actual task-level disappearance.
      • Policy responses (retraining, mobility assistance) should focus on task portability and occupation reallocation channels, not only on task elimination.
  • Labor-market frictions and credentialing:
    • Rising posted requirements (certificates, credentials) are a distinct margin of adjustment; credential inflation can gate access to higher‑paying, less‑exposed tasks. Monitoring requirements monthly helps detect credential-based labor‑market segmentation driven by technology or supply shifts.
  • Targeting interventions:
    • The dynamic task vocabulary can identify resilient task bundles (hands-on, on-site, higher-stakes) and exposed bundles (routine language-amenable tasks), improving targeting of reskilling and regional policies.
    • Because tasks persist across occupations, interventions that train workers in portable, complementary tasks (coordination, supervision, complex hands-on skills) may be more effective than narrow task retraining.
  • Research & policy tools:
    • The open, documented Chinese task/requirement vocabulary fills a major data gap and enables: (i) monthly monitoring of AI exposure trends in the world’s largest labor market; (ii) more nuanced shift-share and counterfactual decompositions; (iii) linkage opportunities with payroll and outcome data to study realization of posting‑level risk.
    • The approach is computationally and cost scalable (low per-ad cost once vocabulary is frozen), so it can be adapted for ongoing surveillance and cross-country comparisons.
  • Limitations for causal inference:
    • Posting-level signals reflect advertised demand, not realized employment or layoffs; declines in posting share need not equal employment losses absent corroborating payroll data.
    • Platform-selection and urban/sector bias mean that generalization to entire national employment should be cautious; policy inferences should ideally be triangulated with other sources.

Short summary takeaway: a low-cost, LLM-assisted pipeline produces an open, monthly-updatable task and requirement vocabulary for Chinese job postings. This reveals that posted demand is changing mainly by re-weighting which occupations are advertised (not by removing exposed tasks within surviving occupations), and that credential inflation accompanies shifts toward higher‑paid, less‑LM‑absorbable work — findings that materially change how economists should measure and forecast AI’s labor-market effects.

Assessment

Paper Typedescriptive Evidence Strengthmedium — Very large, novel, and well-validated descriptive evidence: 752.6 million postings (116M unique descriptions) with extraction, clustering, human adjudication, and multiple validation checks (including correlation with existing exposure measures). However, the paper is not estimating causal effects and is subject to selection and measurement limitations (online platforms, urban/formal tilt, reliance on LLM extraction and regex matching), so strength is not 'high'. Methods Rigorhigh — The construction pipeline is carefully designed: stratified extraction sample, LLM-based phrase extraction, embedding clustering, human adjudication of boundary pairs, connected-component merging, per-entry regex validation and freezing of the vocabulary, and multiple robustness checks (within-platform shares, unique-description robustness, validation correlations). Remaining risks include selection on platform coverage over time, potential LLM biases in extraction, and the limited scope of human adjudication relative to the full vocabulary. Sample752.6 million job ads from the five leading Chinese recruitment platforms observed January 2022–June 2026 (30M in 2022 up to 287M in 2025), de-duplicated to 116 million unique descriptions; structured metadata include city, wage band, education, experience, employer size and industry tags; 2.31M descriptions used in an extraction sample; final standardized catalogs: 20,721 requirements and 44,479 tasks; 12.1 billion item–posting mentions across the full record. Analyses typically use posting-weighting and within-platform shares to account for changing platform coverage. Themeslabor_markets human_ai_collab adoption GeneralizabilityOver-represents urban, formal, white-collar and service-sector hiring relative to national employment (agriculture and informal work largely absent)., Limited to online recruitment platforms in China; findings may not generalize to offline hiring channels or other countries/languages., Platform coverage expanded over time; although within-platform shares help, selection within platforms may change and affect time trends., Measures depend on LLM-based extraction, embedding clustering, and regex matching which could misclassify or miss novel phrasing; human adjudication was bounded to a sample and reusable artifacts., LLM-exposure scores depend on the scoring rubric and the models used for calibration and may be model-dependent.

Claims (11)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study analyzes 752.6 million job advertisements posted on five leading Chinese recruitment platforms between January 2022 and June 2026. Other other Size and coverage of the job-advertisement dataset
Reading fidelity high
Study strength high
n=752600000
752.6 million job ads
0.3
The paper constructs a standardized vocabulary containing 20,721 job requirements and 44,479 job tasks. Other other Number of standardized requirements and task entries
Reading fidelity high
Study strength medium
n=2310000
20,721 requirements and 44,479 tasks
0.18
Within the accountant occupation, core bookkeeping accounts for 68% of task mentions in the lowest wage band but only 49% in the highest wage band, where compliance and data-analysis tasks become more prominent. Task Allocation negative Share of accountant task mentions devoted to core bookkeeping
Reading fidelity high
Study strength medium
n=752600000
68% in the lowest wage band versus 49% in the highest
0.18
The share of accountant job postings requiring some certificate nearly doubles across the lower wage bands and then saturates. Skill Obsolescence positive Share of accountant postings requiring at least one certificate
Reading fidelity high
Study strength medium
n=752600000
nearly doubles across the lower wage bands
0.18
Within the accountant occupational class, postings requiring an intermediate certificate rise from 1.9% to a peak of 14.1%. Skill Acquisition positive Share of accountant postings requiring an intermediate certificate
Reading fidelity high
Study strength medium
n=752600000
1.9% to 14.1%
0.18
For production-worker postings, estimated language-model exposure is approximately 0.17 in the lowest wage band and approximately 0.48 in the highest wage band. Automation Exposure positive Average task-based LLM exposure of production-worker job postings
Reading fidelity high
Study strength medium
n=752600000
about 0.17 versus about 0.48
0.18
Java developer postings carry nearly the same task bundle across experience levels, while only 3–5% of Java postings request a certificate. Task Allocation null_result Variation in Java-developer task composition and certificate requirements by experience level
Reading fidelity high
Study strength medium
n=752600000
3–5% of Java postings ask for a certificate
0.18
Since late 2023, the task composition of posted work has shifted away from tasks absorbed by language models and toward hands-on, on-site, higher-stakes work. Automation Exposure negative Aggregate composition of posted tasks by LLM exposure
Reading fidelity high
Study strength medium
n=752600000
aggregate exposure drift of −0.012 under one-posting-one-vote
0.18
Occupations composed mostly of language-model-exposed tasks lose posting share more sharply than the task categories those occupations contain. Job Displacement negative Posting-share gradient by occupational exposure and task-category mention-share gradient
Reading fidelity high
Study strength medium
n=752600000
occupation slope −1.33; task-category slope −0.83
0.18
The shift-share decomposition indicates that exposed tasks lose ground mainly through changes in which occupations are posted, while the task mix within surviving occupations remains broadly flat. Task Allocation mixed Between-occupation posting-share changes versus within-occupation task-mix changes
Reading fidelity high
Study strength medium
n=752600000
runs almost entirely through the posting-share channel
0.18
The task-based exposure scores produced by the system correlate with leading published exposure measures at r = 0.73–0.79. Automation Exposure positive Correlation between the paper’s exposure measure and published exposure measures
Reading fidelity high
Study strength medium
r = 0.73–0.79
0.18

Notes