0 cumulative citations
View corpus contextA simple dynamical model finds that unmoderated LLM contributions can invert knowledge flows—LLMs can crowd out human content and weaken human learning—while stricter admission gates and training choices can restore healthy growth; calibration to Wikipedia shows post-ChatGPT increases in LLM additions concurrent with falling human inflow, consistent with a risky regime.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Such human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and systemic risks (e.g., quality dilution, skill reduction, model collapse). To understand such phenomena, we propose a minimal, interpretable dynamical model of the co-evolution of archive size, archive quality, model (LLM) skill, aggregate human skill, and query volume. The model captures two content inflows (human, LLM) controlled by a gate on LLM-content admissions, two learning pathways for humans (archive study vs. LLM assistance), and two LLM-training modalities (corpus-driven scaling vs. learning from human feedback). Through numerical experiments, we identify different growth regimes (e.g., healthy growth, inverted flow, inverted learning, oscillations), and show how platform and policy levers (gate strictness, LLM training, human learning pathways) shift the system across regime boundaries. Two domain configurations (PubMed, GitHub and Copilot) illustrate contrasting steady states under different growth rates and moderation norms. We also fit the model to Wikipedia's knowledge flow during pre-ChatGPT and post-ChatGPT eras separately. We find a rise in LLM additions with a concurrent decline in human inflow, consistent with a regime identified by the model. Our model and analysis yield actionable insights for sustainable growth of human-AI collective knowledge on the Web.
Summary
Main Finding
The paper introduces a compact, interpretable five-variable dynamical model of human–AI collective knowledge (archive size K, archive quality q, LLM capacity θ, aggregate human skill H, and query volume Q). Using simulations, case studies (PubMed; GitHub & Copilot) and calibration to Wikipedia before/after ChatGPT, the authors show that interacting feedbacks among humans, LLMs, archives and platform moderation produce distinct growth regimes (healthy growth, stagnation, oscillation, quality dilution, model collapse, human competence inversion). Platform and policy levers—admission-gate strictness, relative content flows, human learning pathways, and LLM training modality (corpus vs. RLHF)—move systems across these regimes and can be tuned to promote sustainable growth or to mitigate systemic risks.
Key Points
- Model structure: A five-state continuous-time dynamical system with intuitive, modular auxiliary functions:
- Variables: K (archive size), q (archive mean fidelity ∈[0,1]), θ (LLM capacity → accuracy via logistic σ(θ)), H (aggregate human skill), Q (query intensity).
- Inflows: human contributions α_H H and AI contributions α_A Q g(a) (g(a) is an admission gate function of answer accuracy a(θ,q)).
- Quality dynamics: new content nudges running quality toward human benchmark q_H or model answer accuracy a(θ,q).
- LLM learning: two modes — supervised corpus-driven training toward a scaling-law target θ*(K,q) and RLHF-driven gains proportional to Q when θ is below frontier.
- Human learning: via archive (Hmax(K,q), Hill/Michaelis–Menten form) and via LLM assistance S(Q)a(θ,q); forgetting/turnover included.
- Query demand: decreases with human skill; Q feeds back into LLM learning and AI contributions.
- Identified systemic risks are unified in one framework:
- Model collapse: poor-quality AI outputs admitted to archives lower q → weaker θ in future training → further poor outputs.
- Quality dilution: high α_A and lax gate g produce archive quality erosion.
- Human competence inversion: heavy reliance on low-quality archive or weak LLMs reduces H, lowering human contribution quality and amplifying degradation.
- Control levers (three paired levers emphasized):
- Content flow balance (α_H vs α_A): platform and market incentives that change relative rates of human vs AI content inflow.
- Learning pathways (β_K vs β_A): whether humans learn more from raw archives or LLM assistance.
- LLM training weighting (η_sup vs η_RLHF): investment balance between offline corpus scaling and interaction-driven RLHF.
- Numerical experiments: nine representative scenarios showing regimes such as pre-LLM steady growth, healthy co-growth with modest LLM involvement, inverted-flow regimes (AI additions dominate), inverted-learning (humans rely mostly on LLMs), oscillations, and collapse.
- Empirical calibration:
- Two case studies: PubMed (strong gate/moderation and slow growth) and GitHub+Copilot (fast growth, different moderation norms) map to distinct model regimes.
- Wikipedia calibration pre- vs post-ChatGPT: detected increased LLM additions and concurrent decline in human inflow post-ChatGPT, consistent with a model-predicted regime where AI inflow rises and human contributions fall.
Data & Methods
- Formal model: system of five ODEs (dK/dt, dq/dt, dθ/dt, dH/dt, dQ/dt). Auxiliary functions are modular and motivated by empirical/theoretical literature:
- σ(θ) = logistic skill curve.
- a(θ,q) = σ(θ) · q (answer accuracy depends on model skill and archive quality, capturing RAG dependence).
- Gate g(a) = logistic function (admission probability of AI outputs).
- θ*(K,q) = θ_max · [ln(1+K)/ln(1+K_max)] · q (scaling-law target: diminishing returns in K and proportional to quality).
- RLHF(θ,Q) = Q/(Q+Q1/2) · (θ_max − θ).
- H_max(K,q) = H_∞ · (qK)^β / (K_1/2^β + (qK)^β) (Hill form).
- S(Q) = Q/(Q+Q_sat) (saturating tutoring benefit).
- ξ(H) = ξ0 / (1 + T_difficulty · e^{−κ_H H}) (baseline query inflow decreasing with H).
- Numerical experiments: parameter sweeps and scenario simulations with fixed initial conditions; nine representative parameter configurations reported (appendix with exact settings for reproducibility).
- Case studies: parameter choices reflecting domain differences (e.g., gate strictness and growth rates) to place PubMed and GitHub/Copilot into different regimes.
- Calibration to Wikipedia: model fit separately on pre-ChatGPT and post-ChatGPT eras; inferred increase in α_A (LLM additions) and decrease in α_H (human inflow) post-ChatGPT.
- Model modularity allows substitution of alternative scaling laws and learning functions for sensitivity analysis.
Implications for AI Economics
- Externalities & public-good dynamics: Shared knowledge archives are public goods whose quality feeds both human productivity and commercial AI performance. Unpriced negative externalities (quality dilution, model collapse) can reduce social welfare across researchers, platforms, and firms that rely on these corpora.
- Incentives and market failure risks:
- Platform/firm incentives to maximize short-term engagement or content volume (raise α_A and Q) can unintentionally lower long-run archive quality q and human skill H, creating negative dynamic externalities.
- LLM providers face a trade-off: cheaper/fast gains from RLHF driven by user queries versus the potentially safer but costlier investment in supervised corpus curation and high-quality training (η_sup). Over-reliance on RLHF can amplify feedback loops that degrade corpora.
- Moderation/admission gates (g(a)) are private or public governance choices with real economic effects: stricter gates raise curation costs but protect long-term asset quality; lax gates lower near-term friction but risk long-term collapse.
- Labor and human-capital effects:
- Human competence inversion implies depreciation of aggregate skills H, which can reduce labor productivity in knowledge-intensive sectors. This has distributional consequences: novices vs experts, substitution/complementarity between human experts and LLMs, and potential erosion of demand for high-skills training if archives become dominated by low-quality AI content.
- Policy levers and welfare interventions:
- Subsidize or reward high-quality human contributions (raise α_H) and funding for moderation to internalize curation benefits.
- Mandate or incentivize transparency and provenance (labeling, provenance stamps) to strengthen gates g(a) and allow downstream consumers to discount low-quality content.
- Regulate training data disclosures or require datasets to meet quality thresholds (affects θ* via q and effective training corpora).
- Encourage balanced investment in supervised training (η_sup) and careful use of RLHF, possibly via audit standards for RLHF pipelines.
- Invest in public infrastructure for curation (e.g., community-moderated repositories, funding academic repositories) to maintain archive quality as a public good.
- Platform design and competitive dynamics:
- Platforms that internalize long-term quality (e.g., scholarly repositories, well-moderated ecosystems) create positive feedback that sustains both human skill and model performance and may enjoy durable comparative advantages.
- Fast-growing, lightly-moderated platforms (e.g., general code repos without strict quality gates) may extract short-term gains but risk regime shifts toward quality dilution that lower the long-run value of their knowledge assets.
- Measurement & monitoring for policy:
- Key measurable indicators to track: human vs AI content inflow rates (α_H, α_A proxies), archive mean fidelity proxies (q), fraction of AI-origin content admitted (g), rates of RLHF use and query volumes (Q), and human-skill proxies (H) via expertise/quality metrics.
- Early-warning signals: rising α_A alongside falling α_H and q, or rising reliance of human learning on LLM assistance (high β_A relative to β_K) should trigger mitigation actions.
- Open research and economic evaluation:
- Welfare quantification: the model provides a basis for formal cost–benefit or social-welfare analyses comparing policy options (gate strictness, subsidies for human curation, RLHF regulation).
- Market design: consider mechanisms (platform norms, verification markets, reputation rewards) that align individual incentives with collective archive quality.
- Distributional effects: evaluate who bears the costs of quality degradation (researchers, firms, consumers) and how policy can address inequities (e.g., by protecting expert knowledge production).
Overall, the paper offers a parsimonious but policy-relevant dynamical framework linking platform design, firm decisions, and human capital dynamics. For AI economics, it highlights dynamic externalities arising from feedback between model training and shared knowledge resources, identifies control levers that map to economic instruments, and provides a modeling and empirical toolkit for evaluating interventions to sustain long-run knowledge value.
Assessment
Claims (7)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Humans and large language models (LLMs) now co-produce and co-consume the web's shared knowledge archives. Other | mixed | co-production and co-consumption of web archives |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| Human-AI collective knowledge ecosystems contain feedback loops with both benefits (e.g., faster growth, easier learning) and systemic risks (e.g., quality dilution, skill reduction, model collapse). Output Quality | mixed | benefits (growth, learning) and risks (quality dilution, skill reduction, model collapse) |
Reading fidelity
high
Study strength
speculative
|
not reported
|
| We propose a minimal, interpretable dynamical model of the co-evolution of archive size, archive quality, model (LLM) skill, aggregate human skill, and query volume that captures: two content inflows (human, LLM) controlled by a gate on LLM-content admissions; two learning pathways for humans (archive study vs. LLM assistance); and two LLM-training modalities (corpus-driven scaling vs. learning from human feedback). Other | null_result | model state variables (archive size, archive quality, LLM skill, human skill, query volume) |
Reading fidelity
high
Study strength
high
|
not reported
|
| Numerical experiments with the model identify different growth regimes (e.g., healthy growth, inverted flow, inverted learning, oscillations). Organizational Efficiency | mixed | system growth regimes (archive growth patterns and dynamics) |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Platform and policy levers (gate strictness, LLM training, human learning pathways) shift the system across regime boundaries. Governance And Regulation | mixed | regime membership / qualitative system dynamics as a function of policy/platform parameters |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Two domain configurations (PubMed, GitHub and Copilot) illustrate contrasting steady states under different growth rates and moderation norms. Task Allocation | mixed | steady-state patterns of archive and actor variables under different domain parameterizations |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Fitting the model to Wikipedia's knowledge flow during pre-ChatGPT and post-ChatGPT eras separately shows a rise in LLM additions with a concurrent decline in human inflow, consistent with a regime identified by the model. Task Allocation | mixed | contribution flows to Wikipedia (LLM-origin additions vs human-origin additions) |
Reading fidelity
high
Study strength
medium
|
not reported
|