The Commonplace
Home Papers Evidence Explore Trends Syntheses Digests References Docs 🎲 Workforce Futures
← Papers
Direction, evidence grade, and study type are AI-generated labels (gpt-5-mini), not human-verified. Syntheses are LLM-written. "Tensions" are machine-detected candidates, not confirmed contradictions. A research-acceleration tool, not peer review. How this is built →

A comprehensive survey finds that the promise of large foundation models for human-AI collaboration depends less on model scale and more on human-centered design, preference-driven objective shaping, and governance; the literature is expanding quickly but remains non-systematic and highlights open challenges in bias, evaluation, and socio-economic effects.

A Survey on Human-AI Collaboration with Large Foundation Models
Vanshika Vats, Marzia Binta Nizam, Minghao Liu, Ziyuan Wang, Richard Ho, Mohnish Sai Prasad, Vincent Titterton, Sai Venkat Malreddy, Riya Aggarwal, Yanwen Xu, Lei Ding, Jay Mehta, Nathan Grinnell, Li Liu, Sijia Zhong, Devanathan Nallur Gandamani, Xinyi Tang, Rohan Ghosalkar, Celeste Shen, Rachel Shen, Nafisa Hussain, Kesav Ravichandran, James Davis · August 22, 2026 · ACM Transactions on Intelligent Systems and Technology
openalex review_meta n/a evidence 7/10 relevance Full text usable extracted full text DOI Source PDF

Structured author observations

Linked only from stored provider relations; the raw author line above is never matched by name.

OpenAlex

Latest observation:

  1. Vanshika Vats provider ID
  2. Marzia Binta Nizam provider ID
  3. Minghao Liu provider ID
  4. Ziyuan Wang provider ID
  5. Richard Ho provider ID
  6. Mohnish Sai Prasad provider ID
  7. Vincent Titterton provider ID
  8. Sai Venkat Malreddy provider ID
  9. Riya Aggarwal provider ID
  10. Yanwen Xu provider ID
  11. Lei Ding provider ID
  12. Jay Mehta provider ID
  13. Nathan Grinnell provider ID
  14. Li Liu provider ID
  15. Sijia Zhong provider ID
  16. Devanathan Nallur Gandamani provider ID
  17. Xinyi Tang provider ID
  18. Rohan Ghosalkar provider ID
  19. Celeste Shen provider ID
  20. Rachel Shen provider ID
  21. Nafisa Hussain provider ID
  22. Kesav Ravichandran provider ID
  23. James Davis provider ID

Semantic Scholar

Latest observation:

  1. Vanshika Vats provider ID
  2. Marzia Binta Nizam provider ID
  3. Minghao Liu provider ID
  4. Ziyuan Wang provider ID
  5. Richard Ho provider ID
  6. M.Sai Prasad provider ID
  7. Vincent Titterton provider ID
  8. Sai Venkat Malreddy provider ID
  9. Riya Aggarwal provider ID
  10. Yanwen Xu provider ID
  11. Lei Ding provider ID
  12. Jay Mehta provider ID
  13. N. Grinnell provider ID
  14. Li Liu provider ID
  15. Sijia Zhong provider ID
  16. Devanathan Nallur Gandamani provider ID
  17. Xinyi Tang provider ID
  18. Rohan Ghosalkar provider ID
  19. Ce Shen provider ID
  20. Rachel Shen provider ID
  21. N. Hussain provider ID
  22. Kesav Ravichandran provider ID
  23. James Davis provider ID
This narrative survey (2014–2024) reviews how large foundation models reshape human-AI collaboration across human-guided model development, design principles, ethics/governance, and applications, arguing that successful HAI depends on human-centered design, preference-based objective shaping (e.g., RLHF), and careful evaluation rather than model scale alone.

Citation observations

Cumulative provider counts captured on specific dates; providers are never combined.

As the capabilities of artificial intelligence (AI) continue to expand rapidly, Human-AI (HAI) Collaboration, combining human intellect and AI systems, has become pivotal for advancing problem-solving and decision-making processes. The advent of Large Foundation Models (LFMs) has greatly expanded its potential, offering unprecedented capabilities by leveraging vast amounts of data to understand and predict complex patterns. At the same time, realizing this potential responsibly requires addressing persistent challenges related to safety, fairness, and control. This paper reviews the crucial integration of LFMs with HAI, highlighting both opportunities and risks. We structure our analysis around: human-guided model development, collaborative design principles, ethical and governance frameworks, and applications in high-stakes domains. Our review shows that successful HAI systems are not the automatic result of stronger models but the product of careful, human-centered design. By identifying key open challenges, this survey aims to give insight into current and future research that turns the raw power of LFMs into partnerships that are reliable, trustworthy, and beneficial to society.

Summary

Main Finding

Large Foundation Models (LFMs) substantially expand the potential of Human-AI (HAI) collaboration, but better models alone do not guarantee effective, safe, or equitable partnerships. Real-world HAI systems require careful human-centered design across data curation, objective shaping, interaction design, evaluation, and governance. Human roles are shifting—from mass annotators toward targeted teachers, evaluators, and overseers—while methods like RLHF/DPO and HITL active learning are central to aligning LFMs with human values. The survey emphasizes opportunities (productivity, richer multimodal collaboration, new applications) and risks (bias reinforcement, misaligned incentives, labor impacts, privacy, governance gaps), and identifies open technical and socio-economic research challenges.

Key Points

  • Scope and approach

    • Survey of literature (2014–2024) on HAI collaboration, with attention to LFMs (LLMs, LVMs).
    • Non-systematic Google Scholar search using targeted keywords; includes published papers and arXiv preprints.
    • Authors used an LLM (ChatGPT-4) to assist in drafting and revision.
  • Human-guided model development

    • Human-in-the-loop (HITL) and active learning: humans now act as targeted teachers for LFMs, focusing labeling and correction where models are uncertain or biased.
    • Objective shaping: preference-based methods (RLHF) and gradient-based alternatives (DPO, Constitutional AI) convert human judgments into model objectives; multimodal feedback is emerging.
    • Evaluation: human experts and end-users remain essential for measuring alignment, safety, trustworthiness, and usability.
  • Design principles for HAI systems

    • Interaction and interface design: UI/UX, explainability, and feedback channels determine effective human-AI teaming.
    • Collaboration patterns & role allocation: research explores task decomposition, dynamic role switching, and allocation of responsibilities between humans and LFMs.
    • System integration and evaluation: end-to-end evaluation metrics and deployment strategies are needed to validate HAI effectiveness.
  • Ethics, governance, and societal concerns

    • Fairness and bias: LFMs can perpetuate or amplify biases; bias-aware active learning and evaluation methods are open needs.
    • Labor & autonomy: HAI affects job content, autonomy, and worker well-being (augmentation vs. displacement trade-offs).
    • Privacy, security, and trust: data governance, adversarial risks, and accountability frameworks are required.
    • Policy and regulation: legal, regulatory, and governance frameworks lag technical advances.
  • Applications

    • Broad sectors covered: healthcare, autonomous vehicles, surveillance/security, games, education, accessibility.
    • Domain-specific challenges (safety, liability, interpretability) are particularly acute in high-stakes settings.
  • Open technical challenges

    • Scalable, faithful preference elicitation for RLHF; multimodal human feedback; bias-aware sampling in active learning; standardization of HAI evaluation metrics.
    • Socio-technical research: measuring real-world impacts, adoption barriers, and long-term equilibrium effects.
  • Limitations of the survey

    • Non-systematic selection and reliance on visible literature; includes preprints; rapidly evolving field—findings are provisional.

Data & Methods

  • Type of paper: literature survey (30 pages; arXiv preprint).
  • Search procedure: Google Scholar searches covering 2014–2024 with keywords including “Human-AI”, “Human-AI collaboration”, “large models”, “foundation models”, “large language models”, and “large vision models”.
  • Organization: papers were filtered and grouped into four thematic areas—(1) Human-guided AI development, (2) Design principles for HAI collaboration, (3) Ethics and governance, (4) Applications.
  • Analytical framing:
    • Adopted a three-phase model (adapted from Maadi et al.): data control → model optimization → evaluation.
    • Synthesizes methods (HITL, active learning, RLHF/DPO, constitutional approaches), interface/role design literature, ethics and governance work, and applied case studies.
  • Empirical content: survey compiles and synthesizes prior empirical and methodological work rather than presenting new field experiments or econometric analysis.
  • Note on reproducibility: selection was not systematic; the survey maps visible research trends rather than providing a comprehensive meta-analysis.

Implications for AI Economics

  • Task reallocation and labor market effects

    • Complementarity vs substitution: LFMs plus HAI design change task boundaries—routine/automatable subtasks are likely to be automated, while demand rises for tasks requiring oversight, judgment, labeling, and model steering (AI-literate, higher-skill roles).
    • Role shift: human workers move from bulk annotation to higher-value activities (curation, feedback, evaluation), altering skill demands and wage premiums.
    • Crowdwork & compensation design: platforms that monetize human feedback (for RLHF, evaluation) will shape bargaining power, compensation norms, and the microeconomics of data labor.
  • Productivity, value capture, and industry structure

    • Productivity gains: properly designed HAI can raise productivity in many sectors (medical diagnostics, design, coding), but realized gains depend on integration, retraining, and complementary investments.
    • Value capture: platform owners and model proprietors may capture disproportionate returns (control over LFMs, fine-tuning datasets, deployment pipelines), potentially increasing market concentration.
    • Measurement challenges: traditional productivity statistics may understate or mis-attribute HAI-driven gains—new metrics needed to capture human-AI joint output and reallocation effects.
  • Wages, inequality, and distributional outcomes

    • Skill-biased demand: increased returns to AI oversight and model engineering skills could widen wage dispersion absent compensating policies.
    • Displacement risk: some intermediate-skilled roles could shrink, pushing workers into lower-paid tasks or requiring transition support.
  • Markets for data, feedback, and governance

    • Markets and incentives for human feedback: RLHF and other preference-data markets create new economic goods (quality-rated human judgments); designing incentives (monetary, reputation, regulatory) affects data quality and fairness.
    • Data ownership and privacy externalities: governance of training/evaluation data affects firms’ costs, consumer welfare, and potential liabilities—policy choices alter competitive dynamics.
  • Policy and regulation implications

    • Labor policy: policies for retraining, bargaining rights in platform-mediated feedback markets, and safety nets will shape transition costs.
    • Competition policy: concentration around LFMs and feedback/data ecosystems calls for antitrust scrutiny and possibly data-portability rules.
    • Standards and certification: sector-specific certification for HAI systems (e.g., healthcare, autonomy) can influence adoption rates and liability allocation.
    • Measurement & monitoring: regulators need new tools (audits, disclosure regimes, standardized HAI evaluation) to oversee economic and societal impacts.
  • Research agenda for economists

    • Causal studies: randomized trials or natural experiments evaluating productivity and employment effects of HAI deployments.
    • Microdata on feedback labor: datasets on who provides RLHF/evaluation labels, compensation structures, and quality outcomes.
    • Equilibrium models: task-allocation models that incorporate human feedback costs, model-improvement dynamics, and monopoly rents.
    • Metrics development: frameworks to measure joint human-AI output, complementarity indices, and welfare impacts.

Summary implication: LFMs magnify opportunities for productivity and new services, but they also reshape labor demand, value capture, and regulatory needs. Economic outcomes will hinge on how HAI systems are designed, who supplies and gets paid for human feedback, and the policy frameworks that govern data, labor, and market power.

Assessment

Paper Typereview_meta Evidence Strengthn/a — This is a narrative survey synthesizing prior work rather than presenting original causal or empirical identification; it does not attempt causal inference or provide primary empirical estimates. Methods Rigormedium — The authors conducted a structured but non-systematic literature search (Google Scholar, keyword-based, 2014–2024) and organized papers into thematic sections; however, they concede the sample is not systematic, rely on a single search engine, include preprints, and do not report formal inclusion/exclusion criteria or reproducible search procedures. SampleA curated (non-systematic) set of literature identified via Google Scholar searches for 2014–2024 using keywords like "Human-AI", "Human-AI collaboration", and LFM-related terms; includes peer-reviewed papers and arXiv preprints across HCI, ML, ethics, and application domains; authors also used ChatGPT-4 to help draft and revise sections. Themeshuman_ai_collab productivity adoption governance skills_training GeneralizabilityNot a systematic or reproducible literature review—selection and reporting bias possible, Search limited to Google Scholar and English-language (implied) sources, Includes preprints that may be unreviewed or later revised, Findings are a thematic synthesis and not quantitative/meta-analytic, so magnitude statements are not provided, Rapidly evolving field: coverage up to 2024 may miss important 2025+ developments

Claims (10)

ClaimDirectionOutcomeConfidence & EvidenceDetails
Large Foundation Models can support human-AI collaborative problem-solving by processing data at large scales, identifying patterns, and augmenting human decision-making. Decision Quality positive Human-AI collaborative problem-solving and decision-making capability
Reading fidelity high
Study strength low
not reported
0.12
Active learning enables models to selectively identify data that requires human labeling, thereby improving model performance while using less training data or annotation effort. Training Effectiveness positive Model performance relative to annotation or training-data requirements
Reading fidelity high
Study strength low
not reported
0.12
Active learning can use human expertise more efficiently by directing experts toward uncertain cases and reducing annotation effort. Training Effectiveness positive Human annotation effort required for model improvement
Reading fidelity high
Study strength low
not reported
0.12
Human-guided active-learning datasets must be sufficiently diverse because small or carefully selected datasets can otherwise reinforce biases already present in pretrained models. Ai Safety And Ethics negative Bias in model training data and resulting model behavior
Reading fidelity high
Study strength medium
not reported
0.24
Reinforcement Learning from Human Feedback uses human preference judgments to train a reward model and fine-tune a foundation model, producing assistants described as more helpful and safer. Ai Safety And Ethics positive Assistant helpfulness and safety/alignment
Reading fidelity high
Study strength medium
not reported
0.24
Human evaluation in RLHF can help identify and mitigate harmful biases and unsafe outputs, resulting in fairer and more trustworthy AI systems. Ai Safety And Ethics positive Fairness, trustworthiness, harmful bias, and unsafe-output rates
Reading fidelity high
Study strength low
not reported
0.12
Direct Preference Optimization simplifies the training process by directly mapping human feedback into policy improvements and is reported in the survey as enhancing output quality. Output Quality positive Model output quality and training complexity
Reading fidelity high
Study strength medium
not reported
0.24
Web-browser-assisted agents trained with RLHF have shown significant improvements in real-world usability. Organizational Efficiency positive Real-world usability of AI agents
Reading fidelity high
Study strength medium
not reported
0.24
Human-generated feedback and reward-model insights can improve transparency by providing visibility into the rationale behind model decision-making. Ai Safety And Ethics positive Transparency of model decision-making
Reading fidelity high
Study strength low
not reported
0.12
The survey argues that effective human-AI systems result from careful human-centered design rather than simply from using more capable models. Ai Safety And Ethics positive Reliability, trustworthiness, and societal benefit of human-AI systems
Reading fidelity high
Study strength low
not reported
0.12

Notes