The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

AI agents can automate large swathes of the research pipeline—accelerating routine, codifiable tasks and methodological scaffolding—but theoretical creativity and tacit disciplinary knowledge remain hard to delegate, risking stratified research roles and pressures on training.

Vibe Researching as Wolf Coming: Can AI Agents with Skills Replace or Augment Social Scientists?
Zhang, Yongjun · February 25, 2026 · arXiv (Cornell University)
openalex theoretical n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Zhang, Yongjun provider ID

Semantic Scholar

Latest observation:

  1. Yongjun Zhang provider ID
The paper argues that AI agents can autonomously carry out many codifiable parts of the research pipeline—boosting speed and methodological scaffolding—but will struggle with theoretical originality and tacit field knowledge, creating augmentation benefits alongside risks of stratification and a pedagogical crisis.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

AI agents -- systems that execute multi-step reasoning workflows with persistent state, tool access, and specialist skills -- represent a qualitative shift from prior automation technologies in social science. Unlike chatbots that respond to isolated queries, AI agents can now read files, run code, query databases, search the web, and invoke domain-specific skills to execute entire research pipelines autonomously. This paper introduces the concept of vibe researching -- the AI-era parallel to vibe coding -- and uses scholar-skill, a 26-skill plugin for Claude Code covering the full research pipeline from idea to submission across 18 orchestrated phases with 53 quality gates, as an illustrative case. I develop a cognitive task framework that classifies research activities along two dimensions -- codifiability and tacit knowledge requirement -- to identify a delegation boundary that is cognitive, not sequential: it cuts through every stage of the research pipeline, not between stages. I argue that AI agents excel at speed, coverage, and methodological scaffolding but struggle with theoretical originality and tacit field knowledge. The paper concludes with an analysis of three implications for the profession -- augmentation with fragile conditions, stratification risk, and a pedagogical crisis -- and proposes five principles for responsible vibe researching.

Summary

Main Finding

AI agents that bundle multi-step reasoning, tool use, and specialist skills (exemplified by the scholar-skill Claude Code plugin) can autonomously execute large portions of the social‑science research pipeline—speeding work, ensuring methodical coverage, and producing publication‑calibrated outputs—yet a clear cognitive delegation boundary remains. Tasks that are highly codifiable and low in tacit knowledge can be delegated; tasks requiring tacit field expertise, deep theoretical originality, and contextual judgment remain distinctly human. The human–AI boundary therefore slices through every stage of the pipeline (not merely between stages), producing strong augmentation potential under fragile conditions and raising risks of stratification and pedagogical disruption.

Key Points

  • Wave framing: The paper situates current developments as a fourth wave of research automation (2024+), qualitatively distinct because it automates multi‑step reasoning across stages rather than single-stage execution.
  • Chatbot vs agent: AI agents are persistent, stateful, and tool‑using (file and code access, databases, sub‑agents, specialist skills), enabling "vibe researching" where a researcher specifies goals and the agent executes large parts of the pipeline.
  • Scholar-skill case: The author uses scholar-skill (a 26‑skill plugin) as an illustrative, operational example covering idea formation through submission, including internal QA and five hard stops (e.g., data safety, lit/theory verification, citation verification, ethics).
  • Cognitive task framework: Research tasks are classified along two dimensions—codifiability and tacit knowledge requirement. The delegation boundary is cognitive: tasks with high codifiability and low tacit requirement are delegable; tasks with low codifiability or high tacitness are not.
  • Agent strengths: speed, breadth/coverage, methodological scaffolding (including identification strategies, outcome‑type dispatch, analytic action dispatch), automated replication‑package construction, rigorous citation verification workflows, and peer‑review simulation.
  • Agent limitations: difficulties with theoretical originality, deep tacit field knowledge, recognizing subtle context‑specific tradeoffs, and some aspects of expert judgment; outcomes depend on data access, verification procedures, and institutional norms (i.e., fragile augmentation).
  • Professional implications: three main consequences identified—(1) augmentation is viable but fragile (depends on access, verification, and institutional rules), (2) risk of stratification (an "AI productivity premium" favoring better‑resourced actors), and (3) a pedagogical crisis for graduate training as routine research skills become automatable.
  • Responsible practice: the paper proposes five principles for responsible vibe researching (centered on transparency, verification/human oversight, replication/citation integrity, equitable access, and teaching reforms—see paper for exact wording).

Data & Methods

  • Approach: Conceptual + operational case study. The paper is not an empirical randomized evaluation; instead it develops a task framework and grounds it with an operational agentic system (scholar-skill) that the author used while preparing the manuscript. All outputs were reviewed and verified by the author.
  • Scholar-skill architecture (illustrative system details):
    • 26 specialist skills organized into 13 stage-groups (e.g., formulation, design, data, analysis, writing, ethics, submission, replication, QA, teaching, collaboration, extensions).
    • Orchestrator ("scholar-full-paper") coordinating 18 phases, 53 quality‑gate items, and five hard stops that block pipeline advancement until minimum standards are met.
    • Technical modules highlighted:
      • Idea formalization workflow with multi-agent stress‑testing panel (theorist, methodologist, domain expert, editor, devil’s advocate).
      • Literature synthesis integrated with hypothesis derivation and a six‑bin literature map.
      • Causal identification module producing DAGs, identification choice among 13 strategies, code in R and Stata, and diagnostic tests.
      • Outcome‑type dispatch for analysis across 11 outcome types and many estimation tools (multiple imputation, brms, SEM, latent class, etc.).
      • Asset‑driven writing that uses a three‑tier knowledge graph plus a Verified Citation Pool and post‑draft citation verification to prevent citation fabrication.
      • Peer‑review simulation spawning multiple reviewer agents and producing a triage dashboard and a Resolution Tracker.
      • Scholar-replication that constructs, tests, and audits replication packages with paper‑to‑code correspondence checks and artifact registries.
      • Analytic Action Dispatch (AAD): automated dispatch & execution of computational action items so that, in data‑available mode, the system executes analytic follow‑ups automatically.
    • Citation verification: a layered verification scheme querying local libraries, CrossRef, Semantic Scholar, OpenAlex, and finally web search; unverified references are flagged/removed.
  • Verification & limits: The system embeds multiple internal verification gates (citation verification, ethics/data safety, pre-draft verification), but the paper acknowledges these are brittle and depend on accurate access to local databases, APIs, and reliable rule sets.
  • Author role: The author reviewed, revised, and verified system outputs and claims full intellectual responsibility; no independent empirical benchmarking against human researchers is presented in this paper.

Implications for AI Economics

  • Labor supply and task reallocation:
    • Codifiable, routine research tasks (data munging, EDA, many robustness checks, replication‑package construction, formatted drafting) are subject to high substitution risk; demand for human time on those tasks will fall.
    • Demand will grow for non‑codifiable tasks: theoretical innovation, ethnographic/tacit knowledge work, high‑stakes judgment, interdisciplinary synthesis—raising returns to these skills.
  • Skill premium and stratification:
    • An "AI productivity premium" is likely: institutions and researchers with early and robust access to high‑quality agents (data, compute, curated knowledge bases, integration expertise) will produce more publishable outputs per unit of labor, amplifying inequality across institutions and countries.
    • Markets for agent tools, verified data connectors, and curated knowledge graphs will command rents; vendors and well‑resourced labs may internalize much of the productivity gains.
  • Wages, credentialing, and career paths:
    • Routine research assistant roles may shrink or transform (shift toward agent supervision, QA, and orchestration). Graduate training and early‑career pipelines may need to emphasize tacit skill acquisition, theory, and judgment rather than procedural execution.
    • Credential inflation could occur if publications become more numerous but individually cheaper to produce; alternative signals (replication, code quality, demonstrated theoretical contribution) may gain value.
  • Productivity, costs, and publication ecosystem:
    • Research production costs per paper may fall for actors using agents, increasing submission volumes and putting pressure on journals and peer review systems; this could shift equilibrium acceptance standards or spawn new gatekeeping mechanisms (e.g., stronger replication/verification requirements).
    • Increased supply of papers may reduce marginal returns to publication quantity and reweight incentives toward distinctiveness and theoretical novelty—areas where humans currently retain comparative advantage.
  • Externalities and market failures:
    • Quality assurance and public‑good provision (replication infrastructure, open verification services) are critical; absent public or interoperable verification, negative externalities (fabrication risks, replication failures, misinformation) may rise.
    • Concentration of research capacity may exacerbate knowledge monopolies, reduce pluralism, and alter the direction of research investment toward agent‑friendly, data‑heavy topics.
  • Policy and institutional responses:
    • Policies that promote equitable access to agentic tools, mandate disclosure of AI usage, and strengthen replication and citation verification could mitigate stratification and preserve incentive alignment.
    • Investment in training to preserve tacit and theoretical skills (e.g., curricula emphasizing conceptual work, field experience, and critical judgment) will be economically important to maintain human comparative advantages.
  • Research economics questions raised:
    • How will returns to different types of human capital (codifiable vs tacit) evolve?
    • What is the equilibrium effect on publication value, citation practices, and journal reputations as agentic outputs scale?
    • What market structure will emerge for agentic research tools—open interoperable ecosystems or proprietary platforms—and how will that affect distribution of gains?

Bottom line: Agentic AI tools like scholar-skill materially shift the production frontier of social‑science research, automating many codifiable tasks and reorganizing comparative advantage across researchers and institutions. The economic effects will depend heavily on access, verification institutions, and how the profession adapts training and norms—making policy and institutional design central to how benefits and harms are distributed.

Assessment

Paper Typetheoretical Evidence Strengthn/a — The paper is conceptual and argumentative rather than empirical: it develops a classification framework and uses an illustrative plugin (scholar-skill) but presents no systematic data, experiments, or causal tests. Methods Rigormedium — Theoretical contribution is structured and explicit (a two‑dimensional cognitive task framework, a staged pipeline, and a worked illustrative case), but it lacks empirical validation, robustness checks, user studies, or quantitative measures that would increase methodological rigor. SampleConceptual analysis with an illustrative case: 'scholar-skill', a 26-skill plugin for Claude Code covering an 18-phase research pipeline with 53 quality gates; no empirical dataset, randomized evaluation, or systematic user/field study is reported. Themeshuman_ai_collab skills_training productivity org_design labor_markets GeneralizabilityRelies on a single illustrative implementation (scholar-skill) and qualitative argumentation rather than cross-tool or cross-discipline evidence., AI agent capabilities, tool ecosystems, and access models are rapidly evolving, so conclusions may become outdated as systems improve., Focuses on academic research workflows; findings may not transfer to other professional or industrial contexts without adaptation., Assumes availability of high-quality models, plugins, and compute/resources that many researchers or institutions may not have., Cultural, disciplinary, and institutional variations in tacit knowledge and norms are not empirically accounted for.

Claims (12)

ClaimDirectionOutcomeConfidence & EvidenceDetails
AI agents represent a qualitative shift from prior automation technologies in social science. Research Productivity positive qualitative shift in automation capabilities
Reading fidelity high
Study strength speculative
not reported
0.02
Unlike chatbots that respond to isolated queries, AI agents can now read files, run code, query databases, search the web, and invoke domain-specific skills to execute entire research pipelines autonomously. Research Productivity positive ability to execute end-to-end research workflows autonomously
Reading fidelity high
Study strength medium
not reported
0.12
The paper introduces the concept of 'vibe researching' as the AI-era parallel to 'vibe coding.' Research Productivity mixed conceptual framing of research practice
Reading fidelity high
Study strength speculative
not reported
0.02
Scholar-skill is a 26-skill plugin for Claude Code covering the full research pipeline across 18 orchestrated phases with 53 quality gates. Research Productivity positive tool coverage of research pipeline (skills, phases, quality gates)
Reading fidelity high
Study strength medium
not reported
0.12
A cognitive task framework classifies research activities along two dimensions—codifiability and tacit knowledge requirement—to identify a delegation boundary that is cognitive, not sequential: it cuts through every stage of the research pipeline rather than between stages. Task Allocation mixed task delegation boundary based on codifiability and tacit knowledge
Reading fidelity high
Study strength speculative
not reported
0.02
AI agents excel at speed, coverage, and methodological scaffolding in research workflows. Task Completion Time positive speed of execution, coverage of tasks, and provision of methodological scaffolding
Reading fidelity high
Study strength medium
not reported
0.12
AI agents struggle with theoretical originality and tacit field knowledge. Output Quality negative ability to produce theoretical originality and apply tacit domain knowledge
Reading fidelity high
Study strength medium
not reported
0.12
The emergence of AI agents has three implications for the research profession: (1) augmentation with fragile conditions, (2) stratification risk, and (3) a pedagogical crisis. Inequality negative professional impacts (augmentation fragility, stratification, pedagogical disruption)
Reading fidelity high
Study strength speculative
not reported
0.02
Augmentation enabled by AI agents will be conditional and fragile (augmentation with fragile conditions). Worker Satisfaction mixed robustness of augmentation outcomes
Reading fidelity high
Study strength speculative
not reported
0.02
AI agents create or exacerbate stratification risk within the research profession (stratification risk). Inequality negative risk of stratification/inequality among researchers
Reading fidelity high
Study strength speculative
not reported
0.02
AI agents precipitate a pedagogical crisis for training researchers (pedagogical crisis). Training Effectiveness negative effect on pedagogy and training effectiveness
Reading fidelity high
Study strength speculative
not reported
0.02
The paper proposes five principles for responsible 'vibe researching.' Governance And Regulation positive guiding principles for responsible practice
Reading fidelity high
Study strength speculative
not reported
0.02

Notes