The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

Kaggle medals remain informative but perishable: almost all predictive power sits in the first year, and the apparent collapse of upload-based credentials after the platform phased out upload evaluations is mainly a consequence of a frozen medal stock aging, not clear evidence that individuals became less skilled under AI.

Stranded credentials: how a skill-signaling market absorbed generative AI
Song Yao · August 17, 2026
arxiv quasi_experimental medium evidence 8/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Song Yao unresolved corpus identity

Semantic Scholar

Latest observation:

  1. Song Yao provider ID
Using 2010–2026 Kaggle competition data, the paper shows that medals predict hidden-test performance almost entirely within their first year, that the collapse in informativeness of upload-earned medals after the platform phased out upload competitions is largely explained by the medal stock aging (institutional stranding) rather than measured declines in holders’ ability, and that credential reading (isolated badge vs full-profile) materially alters inferred changes.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

Generative AI can now perform many tasks that credentialing institutions count on to assess skill. During the AI era, do credentials retain their signaling value for subsequent performance? Mostly, yes. We audit the 2010-2026 archive of Kaggle, the largest data science competition platform, which ran two evaluation formats concurrently: upload-competitions, which directly score entrants' predictions computed on published data, and code-competitions, which score predictions by executing entrants' code on hidden data. Across 444,698 participations, competition medals predict subsequent leaderboard performance almost entirely in the first year after being earned, in both formats. Fresh medals retained most of their signaling value through the AI transition; credential stocks are only as informative as their replenishment. Although upload-competition medal stocks lost 82% of their informativeness, institutional stranding explains half to three quarters of the loss: upload-competitions had exited for reasons predating AI, and their frozen medal stock aged out under the pre-existing decay pattern. Old upload-competition medals look more valuable only in isolation, by proxying for the rest of the holder's record (e.g., experience). The measured changes are institutional rather than personal: an AI-like working style predicts performance similarly in both formats. The platform's official credential tiers, based on lifetime medal counts, discard 13-16% of the medals' information; an index weighting recent medals more heavily, built on pre-AI-era data alone, outperforms the official tiers in predicting AI-era performance. In conclusion, credentials are informative, perishable, institution-bound, and interdependent; sustaining their value under AI is a high-stakes, socio-economic problem of institutional design.

Summary

Main Finding

Generative-AI did not make credentials on Kaggle worthless. Medals remain informative about subsequent (hidden-test) performance, but their informativeness is highly perishable (mostly concentrated in the first year), institution-bound (depends on the evaluation format that awarded them), and interdependent with the rest of a participant’s record. The large collapse in the value of upload-earned medals after the platform shifted to code-execution evaluation is driven mostly by institutional “stranding” (the upload format being phased out and the medal stock aging), not by a uniform collapse of personal skill signals. Simple changes in how credentials are read (isolated badge-count vs full-profile) materially change apparent revaluations. Platform lifetime tiers also throw away nontrivial signal (13–16%).

Key Points

  • Data scope and outcome

    • Audit of Kaggle’s full archive (2010–2026) focusing on 444,698 participations; two concurrent evaluation formats: upload-competitions (submit predictions on published data) and code-competitions (platform runs code on hidden data).
    • Outcome measured: one minus final hidden-test percentile (higher = better).
  • Freshness and decay

    • Nearly all predictive power of a medal is in its first year. Pre-AI fresh medals (<1 year) have slopes ≈ 0.15–0.18 per log medal (unconditional) and under full-profile conditioning deliver the vast majority of explained within-competition variance (two under-one-year bands alone explain ~99% pre-AI, ~95% AI-era).
    • Medals older than 1 year carry very little incremental informativeness.
  • Institutional stranding explains most of the upload-medal collapse

    • Before AI: upload- and code-earned medal stocks were similarly informative (stock slopes ≈ 0.067 and 0.065 per log medal, conditional on one another and prior participation).
    • After the AI transition (competitions with deadlines 2023+; 2022Q4 treated as transition and excluded), upload-stock slope fell by 0.055 (SE 0.007, P < 0.001) from 0.067 to 0.012 — an 82% loss. Code-stock rose by 0.022 (to 0.087, P = 0.04).
    • Decomposition of the upload decline (total −0.055): composition (age-mix shift) −0.028; curve shift (band slopes change) −0.015; interaction +0.007; aggregation error −0.019. Aging (composition + plausible shares of interaction/aggregation) explains roughly half to three quarters of the decline (competition-cluster bootstrap: 43–62% under conservative attribution, 64–87% under generous attribution).
    • Interpretation: many upload medals were stranded because the platform largely stopped issuing new upload medals before or during the AI transition; the existing stock aged per the pre-existing decay pattern.
  • Reading lens and proxying effects

    • Two reading lenses: badge-count read (credential in isolation) vs full-profile read (what a medal adds conditional on the rest of the record).
    • Old upload-earned medals appear more informative under a badge-count read in the AI-era, but this inflation disappears under a full-profile read — stale medals proxy for other parts of the holder’s profile (e.g., recent code medals, experience).
    • Old code-earned medals show the opposite sign-change between lenses: full-profile weights for older code medals rise in the AI era because other, previously stronger overlapping signals (upload medals) lost predictive power; this is re-weighting, not intrinsic durability.
  • Person-level patterns

    • Predicted person-level shifts toward an “AI-like working style” were not found. A behavioral index (submissions per competition; share same-day resubmissions; strength of first scored submission) meant to capture AI-like fewer-iteration high-quality first shots predicts performance similarly in code- and upload-competitions.
    • Under person and competition fixed effects, the index’s difference in returns between code and upload is +0.006 per index unit (SE 0.005, P = 0.28) — statistically indistinguishable from zero.
  • Platform tiers and alternative aggregation

    • Kaggle’s official lifetime tiers (Expert, Master, Grandmaster) discard ~13–16% of the information in medal histories.
    • An index that weights recent medals more heavily, estimated on pre-AI data only, outperforms the official tiers at predicting AI-era performance.
  • Robustness and auxiliary findings

    • Community competitions (which award no tier medals) show the same upload fall and code rise (upload −0.046, code +0.074; n = 344,181), supporting that the stock divergence is not an artifact of medal competitions alone.
    • Analyses are associational: results describe informativeness of records under the measurement system, not causal claims about individual skill changes or the amount of AI use.

Data & Methods

  • Data

    • Kaggle competition archive 2010–2026; primary estimation samples split into pre-AI-era (competition deadlines through 2022Q3) and AI-era (deadlines 2023 onward); transition quarter 2022Q4 excluded.
    • 444,698 participant–competition observations in main analyses (teams credited individually when competing jointly).
    • Medal counts by format, medal-age bands (0–1y, 1–2y, 2–3y, 3–5y, 5y+), prior participation counts, submission telemetry.
  • Outcome and predictors

    • Outcome: one minus final percentile on hidden-test leaderboard (task-specific, objective).
    • Predictors: log medal counts by format and age band; total medal stocks; behavioral index (submissions per competition, share same-day resubmits, strength of first scored submission).
  • Estimation

    • Participation-level regressions with competition fixed effects (comparison among entrants in the same competition).
    • Standard errors two-way clustered by user and competition.
    • Decomposition of stock slope changes into composition (age-mix), curve shift (band slopes), interaction, and aggregation error; competition-cluster bootstrap for confidence intervals on decomposition shares.
    • Reading-lens comparisons: unconditional (badge-count) vs full-profile conditioning (control for other-format medals and prior participation); balanced panel checks (users present in both eras) used to rule out survivorship as primary driver of some patterns.
    • Person-level tests use person and competition fixed effects to isolate within-person changes.
  • Identification and limits

    • The study measures informativeness (predictive power for hidden-test outcomes) — a measurement-system property. It does not observe or identify AI usage by individuals, nor does it make causal claims about changes in underlying skills.

Implications for AI Economics

  • Credentials remain useful but are perishable and institution-dependent

    • Credential value depends heavily on recent, validated outputs. Markets that rely on lifetime, aggregated, or frozen credentials risk mispricing candidate ability after institutional evaluation changes or technological shifts.
    • Policy and organizational screening should prioritize recent validated evidence (e.g., executed code, proctored assessments, or time-bound demonstrations) over distant historical badges.
  • Institutional design matters

    • Much of the apparent devaluation after an AI-era transition can stem from institutional choices (which evaluation formats are used, whether credentials are replenished). Design features that preserve ongoing validation (e.g., requiring execution on hidden data, maintaining active credentialing streams) sustain credential informativeness.
    • Platforms and credentialing bodies should consider mechanisms that allow credentials to be refreshed, decay-aware tiering, or explicit dating/expiration of signals.
  • Reading rules and hiring practice

    • Employers and platforms should read credentials in profile context, not in isolation. Simple badge-count heuristics can mislead because stale badges can proxy for other attributes and over- or under-state helpful information.
    • Aggregation rules (like lifetime tiers) should be revisited: weighting recent performance more heavily recovers substantial predictive power.
  • Labor-market dynamics and inequality

    • Short-lived signals favor agents who can continually produce validated evidence (ongoing contest participation, up-to-date portfolios). This could advantage incumbents with resources to continually replenish credentials and disadvantage those whose prior achievements are frozen.
    • Credential stranding is a socio-economic problem: as evaluation standards change (for reasons technological or institutional), holders of previously useful credentials may be asset-stranded unless institutions provide transition paths.
  • Generalizability, measurement, and future research

    • The audit is an end-to-end measurement in one high-quality, objective setting (Kaggle), but institutional structures there are shared broadly (objective scoring, public credentials, screening). Results suggest broader lessons but should be tested across occupations and credential systems (degrees, certifications, platform reputation).
    • Future work should examine policy designs for credential replenishment, expiry, or revalidation; the labor supply responses to short-lived signals; and causal links between AI usage, skill acquisition, and credential informativeness.

Limitations to keep in mind: the analysis is associational (no direct observation of AI assistance), restricted to task performance on Kaggle-style benchmarks, and focuses on measurement/informativeness rather than causal effects on human skill or welfare.

Assessment

Paper Typequasi_experimental Evidence Strengthmedium — Very large, rich administrative dataset and careful, transparent analyses (within-competition FE, balanced panels, pre-commitment of some hypotheses, bootstrap inference) give credible descriptive evidence that credential informativeness shifted in particular ways; but the study is observational, cannot observe individual AI use, and relies on era comparisons that leave open alternative time-varying confounders and selection mechanisms, so causal claims about AI changing individual skill or behavior are limited. Methods Rigorhigh — The paper applies appropriate econometric tools for archival data: competition fixed effects, two-way clustering, balanced-panel checks, pre-registered hypotheses for many tests, decomposition of mechanisms, and robustness to different control sets and lenses; it transparently distinguishes associational from causal statements. Remaining weaknesses are inherent to non-experimental data (unobserved AI use, potential contemporaneous shocks, and decomposition attribution/aggregation error). SampleAdministrative archive of Kaggle competitions (2010–2026), analyzed at the participation level (N reported: 444,698 participations used in many analyses; platform-wide archive contains ~18.7 million entries). Outcomes are hidden-test leaderboard percentiles (one minus percentile used so higher=better). Key covariates: per-user medal counts by format (upload vs code) and age bands, prior participation counts, submission telemetry (used to construct an AI-like working-style index), team attribution (73% solo). Pre-AI era defined as competition deadlines through 2022Q3; AI era from 2023 onward; transition quarter excluded. Various subsamples include balanced panels of users present in both eras and community-competition subsets. Themeshuman_ai_collab labor_markets IdentificationComparative archival audit exploiting concurrent evaluation formats on Kaggle and a platform-level institutional change (phase-out of upload-format competitions) to compare medal informativeness before versus after the AI transition; uses within-competition fixed effects to compare entrants facing the same task, competition-clustered standard errors/bootstraps, balanced-panel checks (users who competed in both eras), band-level decomposition of stock slope changes into composition (aging) vs curve-shift components, and alternative ‘‘reading lenses’’ (badge-count vs full-profile) to probe proxied signals. No randomized assignment; the institutional flip and time (pre/post-ChatGPT) serve as the quasi-exogenous contrast, while person and competition FE and robustness checks address some confounding. GeneralizabilitySingle-platform (Kaggle) setting: findings may not generalize to formal academic degrees, proctored certification, or many workplace hiring screens., Task-specific (data-science competitions with objective hidden-test scoring); other occupations/tasks may behave differently., Public leaderboard and community norms on Kaggle shape behavior; private hiring and credentialing institutions differ., Cannot observe or instrument individual AI usage, so results do not identify within-person causal effects of using generative AI., Results depend on platform-specific institutional design (concurrent formats, how medals/tier systems are displayed); other institutions may respond differently., Sample is self-selected (participants who choose to compete) and international composition may vary over time, limiting representativeness for broader labor markets.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Competition medals predict subsequent leaderboard performance, with nearly all of a medal's predictive power concentrated in the first year after it is earned. Output Quality positive Subsequent hidden-test leaderboard performance, measured as one minus the team's final leaderboard percentile
Reading fidelity high
Study strength medium
99% of within-competition variance explained by the full set of medal-age bands
0.48
Fresh medals are substantially more informative than older medals: a medal less than one year old has a slope of 0.15–0.18 per log medal, whereas medals older than one year have slopes of only 0.01–0.03. Output Quality positive Subsequent hidden-test leaderboard performance
Reading fidelity high
Study strength medium
0.15–0.18 per log medal for medals less than one year old; 0.01–0.03 for medals older than one year
0.48
The informativeness of the upload-earned medal stock fell by 82% in the AI era, from a slope of 0.067 before AI to 0.012 after ChatGPT's release. Output Quality negative Predictive informativeness of the upload-earned credential stock for subsequent leaderboard performance
Reading fidelity high
Study strength medium
82% loss; slope fell by 0.055 from 0.067 to 0.012
0.48
Institutional stranding and aging explain between roughly one half and three quarters of the decline in upload-earned medal-stock informativeness. Output Quality negative Decline in the predictive informativeness of the upload-earned credential stock
Reading fidelity high
Study strength medium
Aging explains 43–62% under the least favorable attribution and 64–87% under the most favorable attribution
0.48
Fresh upload medals lost about one quarter of their predictive slope in the AI era under full-profile conditioning, while fresh code medals did not show a statistically significant decline. Output Quality mixed Predictive informativeness of fresh upload- and code-earned medals for subsequent performance
Reading fidelity high
Study strength medium
Fresh upload slope declined by 0.033, approximately one quarter of 0.122; fresh code slope increased by 0.015
0.48
Old upload-earned medals appear more informative in the AI era only when read in isolation; their apparent gains disappear after conditioning on the holder's full credential profile. Output Quality mixed Predictive informativeness of older upload-earned medals
Reading fidelity high
Study strength medium
+0.062 and +0.036 slope changes under the badge-count read; gains vanish under full-profile conditioning
0.48
Older code-earned medals received higher full-profile weights in the AI era because their pre-AI discount disappeared, rather than because they became strong positive signals. Output Quality positive Conditional predictive weight of older code-earned medals for subsequent performance
Reading fidelity high
Study strength medium
Weights increased by +0.051 and +0.029, from -0.034 and -0.032 to +0.016 and -0.003
0.48
An AI-like working style did not predict performance differently in code competitions relative to upload competitions. Output Quality null_result Difference in performance between code- and upload-competition formats associated with AI-like working style
Reading fidelity high
Study strength medium
+0.006 per index unit; SE 0.005; P = 0.28
0.48
The platform's official lifetime credential tiers discard 13–16% of the information contained in medals, while an index that weights recent medals more heavily and was built using pre-AI data predicts AI-era performance better than the official tiers. Output Quality positive Prediction of AI-era leaderboard performance from credential summaries
Reading fidelity high
Study strength low
Official tiers discard 13–16% of medal information
0.24
The study finds no evidence that AI eroded participants' underlying human skill, because entrants were never observed working without AI access. Skill Obsolescence null_result Human skill change independent of AI assistance
Reading fidelity high
Study strength high
not reported
0.8

Notes