0 cumulative citations
View corpus contextGenerative AI copies the products of human thought but not the threaded activity that creates them, risking erosion of human capacities; the authors propose a seven-feature process standard and design/audit principles to make AI preserve rather than supplant human judgment.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Intelligence is constituted by \textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \textit{traces} (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define \textit{strong} equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.
Summary
Main Finding
Intelligence should be judged by the iterative process that produces outputs (the activity, practice, reframing, and feedback) rather than by final outputs or traces alone. Current generative AI (GenAI) systems are trained on traces of human cognition and therefore commonly achieve only weak equivalence (matching outputs). The authors propose a middle standard—strong (process) equivalence—defined by seven observable process features that can be assessed across human and machine substrates. Without architectures and training regimes that instantiate these process features, GenAI risks reproducing human outputs while eroding or outsourcing the underlying human capacities (deskilling), with important economic and policy implications.
Key Points
- Intelligence is process-constituted: the iterative activity (failed attempts, reframing, practice) is constitutive of cognitive capacity, not merely the final output or its trace.
- Weak vs hard vs strong (process) equivalence:
- Weak equivalence: two systems match input-output behavior (typical of current GenAI and what the Turing test certifies).
- Hard equivalence: systems match algorithms and architectures (Pylyshyn’s original “strong” equivalence; generally unrealistic between brains and transformers).
- Strong/process equivalence (proposed): systems match at the level of seven process features (substrate-neutral, coarser than algorithmic but finer than input–output).
- Seven candidate features of process-constituted intelligence (observable in solvers; realized differently across substrates):
- Generative trial and revision (many attempts, iterated refinement)
- Temporal extension (returning to and carrying state forward across occasions)
- Engagement with uncertainty (recognizing ill-posed problems, avoiding premature closure)
- Feedback with the medium (material resists and redirects activity; externalized memory)
- Value-laden framing (developing judgments about what problems/moves are valuable)
- Social and dialogical accountability (answerability to interlocutors, critique, peer review)
- Formative dimension (practice shapes who the practitioner becomes; tacit, developmental change)
- Current GenAI covers some features partially (e.g., branching search, self-refinement) but typically lacks sustained temporal extension, authentic engagement with uncertainty as a mode, value-laden framing generation, social accountability tied to reframing, and the long-term formative dimension.
- Apparent reasoning traces (chain-of-thought, multi-agent debate, tree-of-thoughts, self-refinement) often remain cosmetic: they output reasoning-shaped text without necessarily implementing the underlying cognitive process—unless task structure and training make the trace load-bearing.
- The process gap is not merely a data problem: training on more traces does not create the missing iterative human processes; architecture and training objectives must be designed to instantiate process features.
- Proposed remedies include process-preserving architecture and training commitments (e.g., dialogical-revision architectures that require challengers to attack framings, tasks that force externalized intermediate computation to be necessary, reward shaping toward reframing and sustained uncertainty handling).
- The authors propose "process audits"—empirical tasks and measures to test strong/process equivalence across matched human and machine tasks.
Data & Methods
- The paper is primarily conceptual and synthetic: it integrates theory and empirical findings from cognitive science, philosophy, behavioral science, and recent AI research.
- Methods used:
- Literature synthesis across domains (e.g., cognitive science accounts of practice-based skill acquisition, biological examples of distributed problem solving, empirical work on chain-of-thought and multi-agent debate).
- Conceptual analysis to identify seven process features with diagnostic signatures (detailed definitions and signatures indicated in supplementary material).
- Illustrative comparisons of contemporary GenAI techniques (chain-of-thought, tree-of-thoughts, self-refinement, multi-agent debate, Centaur fine-tuning) to show where they meet or fail the process criteria.
- Argumentation about how task structure and training objectives can make traces “load-bearing” (empirical citations showing limited faithfulness of reasoning traces and cases where faithfulness improves when the task demands intermediate computation).
- Empirical anchors cited include:
- Studies showing students reliant on unrestricted generative tools perform worse without them (cited as early evidence of capacity erosion).
- Evaluations showing reasoning traces often fail to track model computations, and debate frameworks sometimes collapse to voting or martingale-like dynamics.
- The Centaur model example: high predictive power for human choices but still weakly equivalent.
- The paper proposes empirical process audits (matched tasks across human and machine) but does not present new experimental data validating the seven-feature metric; rather it sets up testable criteria and experimental designs.
Implications for AI Economics
- Substitution vs. deskilling: GenAI that supplies only outputs (weak equivalence) can substitute for human generative work without supporting the formative, practice-based learning that builds durable human capital. This produces short-term productivity gains but long-term losses in human capability, potentially increasing dependence on AI and lowering resilience in the skilled labor force.
- Human capital measurement and policy:
- Standard productivity and labor-market metrics may undercount losses in cognitive capacity formation. Economists and policymakers should track not just output substitution but also changes in skill acquisition, retention, and quality over time.
- Education and training policy must consider task designs and assessments that require active process engagement (rather than passive verification), to preserve skill formation.
- Complementarity design and workplace adoption:
- AI adoption strategies that leave workers performing only verification tasks (instead of generative, formative tasks) risk converting complementary technologies into substitutes for skill development.
- Firms and product designers should prioritize "process-preserving" AI architectures and workflows—tools that scaffold, externalize, and co-train iterative processes rather than simply delivering final outputs.
- Procurement, regulation, and auditing:
- Procurement specifications for AI in workplaces (e.g., legal, medical, engineering domains) should include process-equivalence audits to assess whether tools preserve or erode worker capacities.
- Regulators could require documentation of how AI systems interact with human workflows, whether they externalize intermediate work, how they handle uncertainty, and whether they support social accountability and reframing.
- Metrics & empirical auditables for economics research:
- Develop "process-equivalence" metrics derived from the seven features (examples: counts of generative revisions per task, measures of persistent state across sessions, calibration/abstention statistics tied to task ambiguity, indicators of reframing activity, longitudinal skill trajectories).
- Use experimental designs comparing matched cohorts (with/without AI assistance) on immediate task outcomes and long-term capacity (retention, ability to perform unaided later, transfer to new tasks).
- Evaluate labor-market impacts not only via employment and wages but via changes in productivity elasticity with respect to human capital, concentration of skills, and resilience to AI downtime or adversarial inputs.
- Economic incentives and R&D:
- Funding and incentives (grants, procurement preferences) should favor AI research that targets process equivalence (training objectives that make intermediate work necessary, reward structures for reframing and dialectic challenge, multi-session persistent agents).
- Firms creating marketplace platforms should be mindful of incentives to optimize for short-term benchmark performance over long-term human-capacity preservation; standards and certifications could help align incentives.
- Distributional effects:
- If high-skill tasks become performable without the associated formation processes, entry barriers and credentialing dynamics may shift, with potential concentration of interpretative and oversight skills in fewer hands (those who can evaluate and correct AI outputs).
- Vulnerable groups relying on on-the-job learning may be disproportionately affected by process-obliterating automation.
Suggestions for applied economic research and policy action - Design and run longitudinal randomized controlled trials where workers learn with process-preserving AI tools vs. output-delivery AI vs. no AI, measuring both productivity and durable skill formation. - Develop standardized process-audit protocols for high-stakes domains (medicine, law, engineering, education) and pilot them in procurement. - Incorporate process features into machine-in-the-loop taxonomies used in labor market and productivity research (beyond simple task automatable/non-automatable labels).
Limitations - The paper is largely conceptual and calls for empirical validation; the seven features and proposed audits require operationalization and field testing. - Measuring long-term formative effects is costly and slow; short-term studies may miss important downstream impacts.
Overall, this paper reframes AI evaluation from output-matching to process-preservation, offering a concrete set of process features and a research agenda. For AI economics, it signals that analyses of automation should include process-level impacts on human capital formation, organizational design, and long-term productivity.
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Current generative AI is weakly equivalent to the human cognition it imitates: it can match human-like outputs while not reproducing the underlying cognitive processes that generated them. Ai Safety And Ethics | negative | Process-level equivalence between generative AI and human cognition |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Generative AI models are trained primarily on textual and visual traces of human cognitive activity rather than on the iterative processes that produced those traces. Ai Safety And Ethics | negative | Coverage of human cognitive processes in AI training data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Students given unrestricted access to a generative AI tool perform better when using the tool but worse without it than peers who never used such a tool. Skill Acquisition | mixed | Performance with and without generative AI assistance |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Unrestricted use of generative AI shifts users' cognitive effort away from generation and toward verification. Task Allocation | negative | Allocation of cognitive effort between generating and verifying work |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Reasoning traces produced by contemporary AI systems routinely fail to faithfully track the computation driving their answers. Ai Safety And Ethics | negative | Faithfulness of AI reasoning traces to the computation producing answers |
Reading fidelity
high
Study strength
high
|
not reported
|
| Contemporary reasoning systems reveal decision-shifting hints in fewer than one in five cases. Ai Safety And Ethics | negative | Frequency with which reasoning traces reveal decision-shifting information |
Reading fidelity
high
Study strength
medium
|
fewer than 1 in 5 cases
|
| Multi-agent debate does not reliably outperform single-agent self-consistency when compute is matched. Decision Quality | null_result | Task performance of multi-agent debate versus single-agent self-consistency |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Debate over belief trajectories forms a martingale, meaning that, on average, each round of debate leaves expected belief unchanged and adds little information beyond the eventual majority vote. Decision Quality | null_result | Information gain from intermediate rounds of AI debate |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Reasoning traces can become computationally load-bearing when a task is difficult enough that the answer cannot be reached without externalizing intermediate work. Ai Safety And Ethics | positive | Faithfulness and computational necessity of externalized reasoning |
Reading fidelity
high
Study strength
medium
|
not reported
|
| No existing generative-AI reasoning or agentic scaffold covers all seven proposed features of process-constituted intelligence at once. Ai Safety And Ethics | negative | Coverage of process features by current AI architectures |
Reading fidelity
high
Study strength
low
|
not reported
|
| Centaur, a model fine-tuned on more than ten million human choices across hundreds of experiments, predicts human behavior on held-out tasks better than bespoke models but remains only weakly equivalent to human cognition. Decision Quality | mixed | Prediction of human behavior on held-out tasks and process-level equivalence |
Reading fidelity
high
Study strength
medium
|
n=10000000
over ten million human choices
|