A field-scale mapping finds social-science research on large language models clusters into three stable domains—model 'minds', multi-agent 'societies', and human–LLM interactions—with human–LLM interaction work making up roughly 78% of topic mass, even as top conference citations tilt toward model-level and multi-agent studies.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
No provider observation is available for this paper.
Missing data, not a zero citation count.
Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.
Summary
Main Finding
The authors produce a reproducible, data-driven taxonomy of the social science of large language models (LLMs). From a curated corpus (N = 198) and a field-scale corpus (N = 47,719), they recover a stable three-domain structure—(1) LLM as Social Minds, (2) LLM Societies, and (3) LLM–Human Interactions—each with interpretable subcategories. The taxonomy is robust across embedding/clustering, topic modeling, author full-text labels, and LLM-based title/abstract classifications, and it reveals a literature dominated in volume by human–LLM interaction research but with influential subfields focused on model behavior and multi-agent systems.
Key Points
- Three primary domains
- LLM as Social Minds: socially interpretable model behavior (theory of mind, personality and demographic bias, political/moral judgment, persuasion/deception).
- LLM Societies: interactions among LLM-based agents (behavioral games, collective intelligence, group decision-making, large-scale simulation).
- LLM–Human Interactions: how people perceive, use, and are affected by LLMs (trust, support/empathy, generative work & collaboration, creativity, education & assessment).
- Subtopics: within-cluster LDA produced 4 subcategories for the first two domains and 5 for the third (13 subcategories total mapped from the field).
- Corpus and coverage
- Curated corpus: 198 papers read in full (115 journals, 76 conferences, 7 preprints).
- Field-scale corpus: 47,719 formally published papers aggregated from Semantic Scholar, OpenAlex, Scopus, PubMed, and Europe PMC.
- Methodological triangulation
- Document representations: MPNet sentence embeddings on title+abstract.
- Clustering: K-means (K explored 2–9; K=3 selected via internal indices and resampling stability).
- Topic models: within-cluster LDA (curated corpus) and structural topic modeling (field-scale).
- Validation: comparisons to author full-text classifications and LLM-based classifications.
- Robustness & correspondence statistics
- K=3 clustering very stable under resampling (mean ARI = 0.952).
- K-means assignments matched authors’ full-text classifications for 77.78% of curated papers.
- LLM-based title/abstract classifications matched K-means ~75.25%; author vs. LLM agreement = 85.86%.
- Field-scale mapping: 13 of 15 STM topics map to the taxonomy. K-means and STM domain assignments correspond ≈ 74–79% (reported ranges depending on overlap).
- Domain prevalence and venue patterns
- Overall (normalized over domain-mapped topics): LLM–Human Interactions accounts for ~78.02% of topic mass; LLM as Social Minds ~14.53%; LLM Societies ~7.44%.
- Venue contrast: among highly cited papers in top conferences, Social Minds + Societies ≈ 66.37% (LLM–Human Interactions ≈ 33.63%); among top journals, LLM–Human Interactions ≈ 76.81%.
Data & Methods
- Corpora
- Curated corpus: 198 papers selected and read in full by the authors (broad interdisciplinary venues; rapid growth 2021–2025 with concentration in 2024–25).
- Field-scale corpus: 47,719 formally published records aggregated from five bibliographic sources.
- Representations and clustering
- Sentence embeddings: MPNet on concatenated title + abstract.
- Clustering: K-means across K = 2–9; model selection used cosine silhouette, Calinski–Harabasz, Davies–Bouldin, resampling stability (100 draws at 90% subsample), and random-initialization stability.
- Chosen partition: K = 3 (high resampling stability; mean ARI = 0.952).
- Topic modeling & interpretation
- Within-cluster LDA used to discover lexical structure within each K-means cluster (to define subcategories).
- Structural topic modeling (STM) applied to the field-scale corpus to identify 15 topics and estimate prevalence variation across publication period and venues.
- Validation & cross-checks
- Cross-comparison with human full-text labels (curated corpus) and LLM-based automated classifications on titles/abstracts.
- Coverage checks: recovered 13 of 15 field-scale topics into the three-domain taxonomy; correspondence across methods in the 70–80% range.
- Reproducibility emphasis: authors present procedural details and cross-method validation to demonstrate that the taxonomy is recoverable across evaluators and analytical choices.
Implications for AI Economics
The paper’s taxonomy and empirical mapping yield several actionable implications for economists studying AI adoption, labor markets, market design, and regulation.
- Clarifies where economic effects are likely to concentrate
- LLM–Human Interactions dominance suggests immediate economic impacts will center on task substitution/complementarity, productivity changes, workplace organization, human capital (education and assessment), and adoption/trust dynamics.
-
LLM as Social Minds and LLM Societies, while smaller by volume, appear concentrated in high-citation technical venues—these literatures directly inform model-driven strategic behavior, automated persuasion, and multi-agent interactions with clear economic relevance (e.g., algorithmic collusion, market signaling, platform competition).
-
Suggests priority research agendas for economics
- Labor and tasks: quantify complementarity vs. substitutability of LLMs across occupations, measure within-firm task reallocation, track wage and employment dynamics by skill and task type.
- Productivity & organization: field experiments or firm-level studies on generative tools’ effects on output, quality, measurement of human oversight costs, and changes in teamwork/role design.
- Market strategy & platform design: study multi-agent LLM societies to model strategic interactions among automated agents (e.g., pricing algorithms, bidding agents, recommender systems) and assess risks of coordination/collusion.
- Consumer behavior & regulation: preference alignment, persuasion, and deceptive behavior map to advertising, misinformation, consumer protection, and disclosure/regulation questions.
-
Education & human capital: evaluate long-run effects of LLMs on skill acquisition, credential signaling, and inequality; redesign of assessment has macro labor supply implications.
-
Methods and data implications for economists
- Use the taxonomy to structure literature reviews, construct focused search queries, or build labeled datasets for supervised classification of LLM-related studies.
- Adopt multi-method approaches mirrored in the paper: combine embeddings/clustering and topic models for large-scale literature mapping; triangulate with hand-curated validation sets.
- Employ LLMs and multi-agent simulations (LLM Societies) as experimental platforms to generate hypotheses about market dynamics and strategic interactions—then validate with field or lab experiments.
-
Exploit venue contrasts: top conferences emphasizing Social Minds/Societies suggest that models of strategic behavior and agent-interaction are frontier topics with high technical traction; economists should bridge these findings to empirical market settings.
-
Policy and regulatory considerations
- The taxonomy identifies specific problem clusters relevant for regulation: persuasion/deception (consumer protection), political/moral judgment (information ecosystems), and multi-agent coordination (antitrust and platform competition).
-
Trust and adoption work points to informational frictions and behavioral responses that matter for designing disclosure, liability, and governance frameworks.
-
Practical research design pointers
- When operationalizing LLM impacts, tag studies or datasets using the three-domain taxonomy to ensure comparability (e.g., separate measures for agent-level behavior vs. interaction-level systemic effects vs. human-facing adoption outcomes).
- Consider mixed designs: simulate LLM-agent markets to generate theoretical predictions (LLM Societies) and test them with firm/consumer data (LLM–Human Interactions).
- Use the taxonomy to prioritize data collection: labor force microdata, firm-level tool adoption logs, platform transaction records, and experimentally randomized interventions altering model outputs or disclosure.
Overall, the taxonomy provides a structured lens for economists to map where and how LLMs may generate economic change, to align empirical strategies with the types of social effects (individual behavior, multi-agent dynamics, human interaction) most relevant for specific policy or market questions, and to identify underexplored but economically consequential areas (e.g., large-scale multi-agent simulations, persuasion/deception impacts on markets).
Assessment
Claims (10)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| The study analyzed a curated corpus of 198 papers and a field-scale corpus of 47,719 formally published papers. Other | positive | Literature-corpus coverage |
Reading fidelity
high
Study strength
medium
|
n=47719
198 curated papers; 47,719 field-scale papers
|
| The literature is organized into three domains: LLM as Social Minds, LLM Societies, and LLM–Human Interactions. Other | positive | Recovered thematic domain structure |
Reading fidelity
high
Study strength
medium
|
n=198
|
| The three-cluster K-means solution was highly stable under resampling, with a mean adjusted Rand index of 0.952. Other | positive | Clustering stability |
Reading fidelity
high
Study strength
high
|
n=198
mean ARI = 0.952
|
| K-means assignments matched the authors’ full-text classifications for 77.78% of papers. Other | positive | Agreement between unsupervised clustering and author classification |
Reading fidelity
high
Study strength
medium
|
n=198
77.78% agreement
|
| The authors’ full-text classifications and the LLM’s primary classifications agreed for 85.86% of papers. Other | positive | Agreement between author and LLM domain classifications |
Reading fidelity
high
Study strength
medium
|
n=198
85.86% agreement
|
| At field scale, 13 of 15 structural-model topics mapped onto the three-domain taxonomy. Other | positive | Field-scale coverage of the taxonomy |
Reading fidelity
high
Study strength
medium
|
n=47719
13 of 15 topics
|
| LLM–Human Interactions accounted for 78.02% of domain-mapped topic mass, compared with 14.53% for LLM as Social Minds and 7.44% for LLM Societies. Other | positive | Relative thematic prevalence of the three domains |
Reading fidelity
high
Study strength
medium
|
n=47719
78.02% versus 14.53% and 7.44% of topic mass
|
| The domain composition differed by venue type: among highly cited papers in the highest-scoring conference venues, LLM as Social Minds and LLM Societies together accounted for 66.37%, whereas LLM–Human Interactions accounted for 76.81% in the corresponding journal subset. Other | mixed | Domain composition across publication venues |
Reading fidelity
high
Study strength
medium
|
66.37% versus 76.81%
|
| The curated corpus consisted primarily of formally disseminated scholarly publications: 115 journal articles, 76 conference papers, and 7 preprints. Other | positive | Composition of the curated evidence base by publication type |
Reading fidelity
high
Study strength
high
|
n=198
96.5% journal or conference publications
|
| The curated corpus grew rapidly from 1 paper in 2021 and 5 in 2022 to 39 in 2023, 97 in 2024, and 56 in 2025. Other | positive | Annual publication volume |
Reading fidelity
high
Study strength
high
|
n=198
97 papers in 2024, representing 49.0% of the corpus
|