The Commonplace
Home Three-study pilot Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A field-scale mapping finds social-science research on large language models clusters into three stable domains—model 'minds', multi-agent 'societies', and human–LLM interactions—with human–LLM interaction work making up roughly 78% of topic mass, even as top conference citations tilt toward model-level and multi-agent studies.

Mapping the Emerging Social Science of Large Language Models
Yi Yang, Xiao Jia, Zeyun Dong, Chenzhang Wang, Zhanzhan Zhao · September 07, 2026
arxiv review_meta n/a evidence 7/10 relevance Full text usable extracted full text Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

Arxiv

Latest observation:

  1. Yi Yang unresolved corpus identity
  2. Xiao Jia unresolved corpus identity
  3. Zeyun Dong unresolved corpus identity
  4. Chenzhang Wang unresolved corpus identity
  5. Zhanzhan Zhao unresolved corpus identity
The authors produce a reproducible three-domain taxonomy of the social science of LLMs—LLM as Social Minds, LLM Societies, and LLM–Human Interactions—showing that human–LLM interaction research dominates volume while model-behavior and multi-agent work are more visible in top conferences.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

No provider observation is available for this paper.

Missing data, not a zero citation count.

Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.

Summary

Main Finding

The authors produce a reproducible, data-driven taxonomy of the social science of large language models (LLMs). From a curated corpus (N = 198) and a field-scale corpus (N = 47,719), they recover a stable three-domain structure—(1) LLM as Social Minds, (2) LLM Societies, and (3) LLM–Human Interactions—each with interpretable subcategories. The taxonomy is robust across embedding/clustering, topic modeling, author full-text labels, and LLM-based title/abstract classifications, and it reveals a literature dominated in volume by human–LLM interaction research but with influential subfields focused on model behavior and multi-agent systems.

Key Points

  • Three primary domains
    • LLM as Social Minds: socially interpretable model behavior (theory of mind, personality and demographic bias, political/moral judgment, persuasion/deception).
    • LLM Societies: interactions among LLM-based agents (behavioral games, collective intelligence, group decision-making, large-scale simulation).
    • LLM–Human Interactions: how people perceive, use, and are affected by LLMs (trust, support/empathy, generative work & collaboration, creativity, education & assessment).
  • Subtopics: within-cluster LDA produced 4 subcategories for the first two domains and 5 for the third (13 subcategories total mapped from the field).
  • Corpus and coverage
    • Curated corpus: 198 papers read in full (115 journals, 76 conferences, 7 preprints).
    • Field-scale corpus: 47,719 formally published papers aggregated from Semantic Scholar, OpenAlex, Scopus, PubMed, and Europe PMC.
  • Methodological triangulation
    • Document representations: MPNet sentence embeddings on title+abstract.
    • Clustering: K-means (K explored 2–9; K=3 selected via internal indices and resampling stability).
    • Topic models: within-cluster LDA (curated corpus) and structural topic modeling (field-scale).
    • Validation: comparisons to author full-text classifications and LLM-based classifications.
  • Robustness & correspondence statistics
    • K=3 clustering very stable under resampling (mean ARI = 0.952).
    • K-means assignments matched authors’ full-text classifications for 77.78% of curated papers.
    • LLM-based title/abstract classifications matched K-means ~75.25%; author vs. LLM agreement = 85.86%.
    • Field-scale mapping: 13 of 15 STM topics map to the taxonomy. K-means and STM domain assignments correspond ≈ 74–79% (reported ranges depending on overlap).
  • Domain prevalence and venue patterns
    • Overall (normalized over domain-mapped topics): LLM–Human Interactions accounts for ~78.02% of topic mass; LLM as Social Minds ~14.53%; LLM Societies ~7.44%.
    • Venue contrast: among highly cited papers in top conferences, Social Minds + Societies ≈ 66.37% (LLM–Human Interactions ≈ 33.63%); among top journals, LLM–Human Interactions ≈ 76.81%.

Data & Methods

  • Corpora
    • Curated corpus: 198 papers selected and read in full by the authors (broad interdisciplinary venues; rapid growth 2021–2025 with concentration in 2024–25).
    • Field-scale corpus: 47,719 formally published records aggregated from five bibliographic sources.
  • Representations and clustering
    • Sentence embeddings: MPNet on concatenated title + abstract.
    • Clustering: K-means across K = 2–9; model selection used cosine silhouette, Calinski–Harabasz, Davies–Bouldin, resampling stability (100 draws at 90% subsample), and random-initialization stability.
    • Chosen partition: K = 3 (high resampling stability; mean ARI = 0.952).
  • Topic modeling & interpretation
    • Within-cluster LDA used to discover lexical structure within each K-means cluster (to define subcategories).
    • Structural topic modeling (STM) applied to the field-scale corpus to identify 15 topics and estimate prevalence variation across publication period and venues.
  • Validation & cross-checks
    • Cross-comparison with human full-text labels (curated corpus) and LLM-based automated classifications on titles/abstracts.
    • Coverage checks: recovered 13 of 15 field-scale topics into the three-domain taxonomy; correspondence across methods in the 70–80% range.
  • Reproducibility emphasis: authors present procedural details and cross-method validation to demonstrate that the taxonomy is recoverable across evaluators and analytical choices.

Implications for AI Economics

The paper’s taxonomy and empirical mapping yield several actionable implications for economists studying AI adoption, labor markets, market design, and regulation.

  1. Clarifies where economic effects are likely to concentrate
  2. LLM–Human Interactions dominance suggests immediate economic impacts will center on task substitution/complementarity, productivity changes, workplace organization, human capital (education and assessment), and adoption/trust dynamics.
  3. LLM as Social Minds and LLM Societies, while smaller by volume, appear concentrated in high-citation technical venues—these literatures directly inform model-driven strategic behavior, automated persuasion, and multi-agent interactions with clear economic relevance (e.g., algorithmic collusion, market signaling, platform competition).

  4. Suggests priority research agendas for economics

  5. Labor and tasks: quantify complementarity vs. substitutability of LLMs across occupations, measure within-firm task reallocation, track wage and employment dynamics by skill and task type.
  6. Productivity & organization: field experiments or firm-level studies on generative tools’ effects on output, quality, measurement of human oversight costs, and changes in teamwork/role design.
  7. Market strategy & platform design: study multi-agent LLM societies to model strategic interactions among automated agents (e.g., pricing algorithms, bidding agents, recommender systems) and assess risks of coordination/collusion.
  8. Consumer behavior & regulation: preference alignment, persuasion, and deceptive behavior map to advertising, misinformation, consumer protection, and disclosure/regulation questions.
  9. Education & human capital: evaluate long-run effects of LLMs on skill acquisition, credential signaling, and inequality; redesign of assessment has macro labor supply implications.

  10. Methods and data implications for economists

  11. Use the taxonomy to structure literature reviews, construct focused search queries, or build labeled datasets for supervised classification of LLM-related studies.
  12. Adopt multi-method approaches mirrored in the paper: combine embeddings/clustering and topic models for large-scale literature mapping; triangulate with hand-curated validation sets.
  13. Employ LLMs and multi-agent simulations (LLM Societies) as experimental platforms to generate hypotheses about market dynamics and strategic interactions—then validate with field or lab experiments.
  14. Exploit venue contrasts: top conferences emphasizing Social Minds/Societies suggest that models of strategic behavior and agent-interaction are frontier topics with high technical traction; economists should bridge these findings to empirical market settings.

  15. Policy and regulatory considerations

  16. The taxonomy identifies specific problem clusters relevant for regulation: persuasion/deception (consumer protection), political/moral judgment (information ecosystems), and multi-agent coordination (antitrust and platform competition).
  17. Trust and adoption work points to informational frictions and behavioral responses that matter for designing disclosure, liability, and governance frameworks.

  18. Practical research design pointers

  19. When operationalizing LLM impacts, tag studies or datasets using the three-domain taxonomy to ensure comparability (e.g., separate measures for agent-level behavior vs. interaction-level systemic effects vs. human-facing adoption outcomes).
  20. Consider mixed designs: simulate LLM-agent markets to generate theoretical predictions (LLM Societies) and test them with firm/consumer data (LLM–Human Interactions).
  21. Use the taxonomy to prioritize data collection: labor force microdata, firm-level tool adoption logs, platform transaction records, and experimentally randomized interventions altering model outputs or disclosure.

Overall, the taxonomy provides a structured lens for economists to map where and how LLMs may generate economic change, to align empirical strategies with the types of social effects (individual behavior, multi-agent dynamics, human interaction) most relevant for specific policy or market questions, and to identify underexplored but economically consequential areas (e.g., large-scale multi-agent simulations, persuasion/deception impacts on markets).

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a systematic mapping and taxonomy paper synthesizing and organizing the literature on the social science of LLMs rather than presenting primary causal or experimental evidence; it does not make causal claims that require identification. Methods Rigorhigh — Uses a two-stage design with a carefully curated full-text corpus (N=198) and a much larger field-scale corpus (N=47,719) assembled from five bibliographic databases; applies modern embedding representations (MPNet), multiple clustering resolutions, stability/resampling checks, within-cluster LDA, structural topic modeling, and cross-method validation (author full-text labels and LLM classifications). Sensitivity analyses and venue/citation comparisons strengthen internal validity, though reliance on titles/abstracts for large-scale assignment and database coverage introduce limits. SampleTwo corpora: (1) a curated, hand-reviewed full-text corpus of 198 papers (115 journal articles, 76 conference papers, 7 preprints) concentrated in 2023-2025 and dispersed across many venues (Scientific Reports, Nature Human Behaviour, ACM CHI, EMNLP, etc.); (2) a field-scale corpus of 47,719 formally published papers retrieved from Semantic Scholar, OpenAlex, Scopus, PubMed, and Europe PMC, used for structural topic modeling and distributional analyses. Themeshuman_ai_collab productivity GeneralizabilityRelies on titles and abstracts for representation at field scale, which may miss nuance present in full texts and can misclassify interdisciplinary work., Coverage depends on selected bibliographic databases and search/eligibility criteria; non-indexed venues, non-English publications, or gray literature may be underrepresented., Taxonomy reflects literature up to the study's collection period (major growth 2023–2025) and may not capture rapid subsequent developments., Unsupervised clustering and topic models impose representational choices (MPNet embeddings, K-means, LDA/STM) which shape categories and boundaries; borderline papers can be ambiguously assigned., Citation- and venue-based comparisons reflect visibility rather than intrinsic quality and may bias interpretation toward well-cited subfields.

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
The study analyzed a curated corpus of 198 papers and a field-scale corpus of 47,719 formally published papers. Other positive Literature-corpus coverage
Reading fidelity high
Study strength medium
n=47719
198 curated papers; 47,719 field-scale papers
0.24
The literature is organized into three domains: LLM as Social Minds, LLM Societies, and LLM–Human Interactions. Other positive Recovered thematic domain structure
Reading fidelity high
Study strength medium
n=198
0.24
The three-cluster K-means solution was highly stable under resampling, with a mean adjusted Rand index of 0.952. Other positive Clustering stability
Reading fidelity high
Study strength high
n=198
mean ARI = 0.952
0.4
K-means assignments matched the authors’ full-text classifications for 77.78% of papers. Other positive Agreement between unsupervised clustering and author classification
Reading fidelity high
Study strength medium
n=198
77.78% agreement
0.24
The authors’ full-text classifications and the LLM’s primary classifications agreed for 85.86% of papers. Other positive Agreement between author and LLM domain classifications
Reading fidelity high
Study strength medium
n=198
85.86% agreement
0.24
At field scale, 13 of 15 structural-model topics mapped onto the three-domain taxonomy. Other positive Field-scale coverage of the taxonomy
Reading fidelity high
Study strength medium
n=47719
13 of 15 topics
0.24
LLM–Human Interactions accounted for 78.02% of domain-mapped topic mass, compared with 14.53% for LLM as Social Minds and 7.44% for LLM Societies. Other positive Relative thematic prevalence of the three domains
Reading fidelity high
Study strength medium
n=47719
78.02% versus 14.53% and 7.44% of topic mass
0.24
The domain composition differed by venue type: among highly cited papers in the highest-scoring conference venues, LLM as Social Minds and LLM Societies together accounted for 66.37%, whereas LLM–Human Interactions accounted for 76.81% in the corresponding journal subset. Other mixed Domain composition across publication venues
Reading fidelity high
Study strength medium
66.37% versus 76.81%
0.24
The curated corpus consisted primarily of formally disseminated scholarly publications: 115 journal articles, 76 conference papers, and 7 preprints. Other positive Composition of the curated evidence base by publication type
Reading fidelity high
Study strength high
n=198
96.5% journal or conference publications
0.4
The curated corpus grew rapidly from 1 paper in 2021 and 5 in 2022 to 39 in 2023, 97 in 2024, and 56 in 2025. Other positive Annual publication volume
Reading fidelity high
Study strength high
n=198
97 papers in 2024, representing 49.0% of the corpus
0.4

Notes