1 cumulative citations
View corpus contextMultimodal perception has raced ahead, but conversational AI still cannot reliably sustain coherent multi-turn sessions: memory, cross-turn grounding, full‑duplex timing and cultural alignment remain major shortcomings while datasets and benchmarks stay English‑ and text‑heavy.
Citation observations
Cumulative provider counts captured on specific dates; providers are never combined.
Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)
Summary
Main Finding
Multi-turn conversational AI is rapidly expanding from text-only interactions to speech-native, multimodal, and omni-modal systems, but advances in perception and generation across modalities have outpaced progress on sustaining coherent, context-aware interaction across sessions. Persistent memory, cross-turn grounding, timing (full‑duplex) behavior, robust session-level evaluation, and cultural/linguistic alignment remain key unsolved problems.
Key Points
- Scope and contribution
- Unified survey of multi-turn conversational AI across text, speech (AudioLLMs), vision/video, omni-modal systems, tool-augmented agents, and culturally grounded resources.
- Organizes literature by interaction depth (multi-turn sessions), modality complexity, and cultural/linguistic diversity; focuses on the session as the primary unit of analysis.
- Principal gaps identified
- Capability gap: large context windows and stronger base LLMs do not guarantee session-level competence (memory, cross-turn grounding, assumption revision, tool-state preservation, spoken timing).
- Resource gap: heavy English and text bias (>80%); relatively few multi-turn spoken, video, omni-modal, or culturally grounded datasets.
- Evaluation gap: benchmarks often use LLM-as-judge or turn-level scoring; session-level metrics, human validation, reproducible scoring underdeveloped.
- Integration gap: most resources evaluate a single dimension (e.g., memory OR multilinguality OR tool use); few benchmarks combine long-horizon memory, speech, vision, tool use, safety, and cultural alignment.
- Modeling evolution and limitations
- Shift from modular dialogue pipelines → instruction‑tuned LLMs → AudioLLMs → omni‑modal unified models and tool-augmented agents.
- Common shortcoming: dialogue history typically treated as a flat context; limited explicit mechanisms for selective memory, state updates, or efficient long-session reasoning.
- AudioLLMs and omni‑modal models add real-time, paralinguistic, and visual grounding challenges (ASR/acoustic-token errors, decay of visual grounding over turns, duplex interaction).
- Datasets & benchmarks
- The survey catalogs many datasets and benchmarks (organized in tables): numerous text-image multi-turn resources (~18), fewer spoken (≈8) and video (~6), many single-turn cultural datasets but few multi-turn culturally grounded resources.
- Benchmark practices: mixed evaluation types (LLM-judge, human, mixed), but scaling these requires stronger human evaluation and reproducibility.
- Open technical challenges
- Long-horizon memory (within and across sessions), cross-turn multimodal grounding, robust full‑duplex spoken interaction, paralinguistic understanding, evaluation methodology for sessions, and cultural/dialectal adaptation.
Data & Methods
- Review protocol
- PRISMA-ScR framework for systematic, reproducible literature review.
- Searches across Google Scholar, Semantic Scholar, IEEE Xplore, ACM DL, Elsevier, DBLP, ACL Anthology, and arXiv; venues included ACL, NeurIPS, ICLR, ICML, Interspeech, SIGDIAL, CVPR, ICCV, AAAI, etc.
- Initial pool ≈ 4,000 papers; screening reduced to ~200 papers for detailed review.
- Organizing framework
- Three axes: interaction depth (multi-turn session focus), modality complexity (text → speech → vision/video → omni-modal), and cultural/linguistic diversity (multilingual, dialectal, culturally grounded).
- Formal session model: a session D = {(u_t, y_t, c_t, m_t, z_t)}_t with y_t = f_θ(u_t, c_t, m_t, z_t) where c_t = accumulated dialogue context, m_t = multimodal context, z_t = external state (memory, retrievals, tool outputs, user profile).
- Dataset & benchmark assembly
- Compiled modality-specific dataset table (text-only, spoken, multimodal, video, cultural) and benchmark table (grouped by evaluation focus: instruction-following, multimodal, spoken, multilingual, robustness/safety).
- Annotated dataset/benchmark metadata: task types, dialog turns per session, size, languages, curation method (human vs synthetic), evaluation style (LLM-judge, mixed, human).
- Limitations noted by authors
- English and text bias in available resources.
- Most cultural datasets are single-turn; few multi-session or long-horizon evaluations.
- Scalability of LLM-as-judge evaluations depends on better human validation and reproducible scoring.
Implications for AI Economics
- Market and product implications
- New product opportunities: genuinely multimodal, session-aware assistants (voice-first customer support, multimodal help desks, multi-turn tutoring, real-time collaborative assistants).
- Productization requires solving session-level memory and grounding; until then, user trust and commercial adoption may be constrained, reducing realized ROI from perceptual improvements alone.
- Investment and R&D priorities
- High economic value in funding multi-turn, multimodal, multilingual datasets and costly human-evaluation infrastructure (for session-level validation). These are public‑good style inputs likely to increase downstream innovation.
- Engineered memory, efficient long-context mechanisms, and tool-state integrations are strategic R&D areas with strong leverage for enterprise deployments (CRM, healthcare, legal assistants).
- Labor, productivity, and task shifting
- Better multi-turn multimodal assistants could automate higher-complexity customer interactions and knowledge work, accelerating task reallocation—especially in service and support industries—but risks of errors on multi-step tasks imply phased augmentation rather than immediate substitution.
- Demand shifts toward skilled roles: prompt engineering for session design, human evaluators for long-horizon behavior, cultural/dialectal annotators, and system integrators for tool-augmented agents.
- Measurement and competition
- Current evaluation gaps (reliance on turn-level metrics and LLM judges) can misstate system quality, affecting buyer decisions and market competition; robust, session-level benchmarks would improve procurement and regulatory assessment.
- Firms that invest early in session-level evaluation and culturally diverse resources can obtain competitive advantage in global markets.
- Policy, externalities, and regulation
- Cultural and language biases concentrated in English-text datasets imply distributional harms and uneven access—economic policy should incentivize creation of multilingual, culturally grounded resources.
- Misalignment risks (e.g., persistent memory mistakes, cross-turn hallucinations) have regulatory relevance in high-stakes domains (healthcare, finance). Regulators may require session-level reliability audits and disclosure of memory/grounding limits.
- Practical recommendations for economists, investors, and policymakers
- Fund and subsidize creation of multi-turn multimodal and multilingual datasets, plus human-evaluation labs for session-level metrics.
- Prioritize benchmarks and procurement standards that measure session-level correctness, memory fidelity, and cultural alignment.
- Monitor labor impacts in sectors where multi-turn multimodal assistants are deployed; support retraining programs focused on evaluator, curator, and integrator roles.
- Encourage transparency around model memory capabilities and failure modes in commercial deployments to mitigate downstream economic risks.
If you want, I can (a) extract the key dataset and benchmark tables into a compact spreadsheet-style summary prioritized for investment decisions, or (b) produce a 1‑page policy brief targeted at regulators summarizing recommended auditing and procurement standards for multi-turn conversational systems. Which would be most useful?
Assessment
Claims (11)
| Claim | Direction | Outcome | Confidence & Evidence | Details |
|---|---|---|---|---|
| Current conversational AI systems support more modalities, but their ability to sustain coherent interaction across a session has advanced more slowly. Other | mixed | Relative progress in multimodal capability versus session-level conversational coherence |
Reading fidelity
high
Study strength
medium
|
n=200
|
| Frontier models still perform below human-level reliability on realistic multi-turn instruction following. Error Rate | negative | Reliability of multi-turn instruction following |
Reading fidelity
high
Study strength
medium
|
n=273
below human-level reliability
|
| Models can fail to use relevant information even when that information remains inside the context window. Decision Quality | negative | Use of relevant information from the conversation context |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Models degrade when a task is distributed across several turns rather than stated as a complete single-turn instruction. Task Completion Time | negative | Task performance under multi-turn versus single-turn instruction presentation |
Reading fidelity
high
Study strength
medium
|
not reported
|
| The review identified approximately 4,000 papers initially and selected 200 papers for detailed review after PRISMA screening. Other | null_result | Literature-review coverage and selection |
Reading fidelity
high
Study strength
medium
|
n=200
approximately 4K papers initially; 200 papers selected
|
| English-only datasets dominate the reviewed multi-turn dialogue resources: 43 of 52 dataset entries outside the cultural and cross-lingual block are English-only. Skill Acquisition | negative | Language diversity of multi-turn dialogue datasets |
Reading fidelity
high
Study strength
medium
|
n=52
43 of 52 dataset entries
|
| Text-only datasets dominate the benchmark landscape: 32 of 53 benchmarks focus on text-only datasets. Other | negative | Modality coverage of multi-turn dialogue benchmarks |
Reading fidelity
high
Study strength
medium
|
n=53
32 of 53 benchmarks
|
| Modality coverage is expanding but remains uneven: the review lists 18 text-image resources, compared with only 8 spoken and 6 video-oriented datasets. Other | mixed | Availability of multi-turn datasets across modalities |
Reading fidelity
high
Study strength
medium
|
n=32
18 text-image resources; 8 spoken resources; 6 video-oriented datasets
|
| Long-horizon and multi-session dialogue resources remain rare. Other | negative | Availability of long-horizon and multi-session dialogue data |
Reading fidelity
high
Study strength
medium
|
not reported
|
| Cultural coverage is increasing, but most resources in the cultural block are single-turn: 4 of the 5 listed entries do not test sustained interaction. Other | mixed | Cultural coverage and sustained-interaction coverage of dialogue resources |
Reading fidelity
high
Study strength
medium
|
n=5
4 of 5 entries are single-turn
|
| Session-level metrics, human validation, and reproducible scoring remain underdeveloped in current evaluation benchmarks. Ai Safety And Ethics | negative | Maturity and reliability of multi-turn evaluation methodology |
Reading fidelity
high
Study strength
medium
|
n=200
|