Physical Sciences › Computer Science › Artificial Intelligence
Speech and dialogue systems
298 papers indexed
AI speech and dialogue systems explore how models process vocal or textual exchanges, whether between humans or with artificial agents. This research addresses challenges such as understanding spatial sounds, managing real-time conversations with low latency, or detecting misunderstandings and disagreements in interactions. It also examines mechanisms like in-context learning, goal-oriented dialogue optimization, or modeling preferences and constraints in exchanges to enhance system fluency and coherence.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States38% · 53 papers
- China34% · 48 papers
- United Kingdom9.9% · 14 papers
- Germany9.9% · 14 papers
- Japan7.1% · 10 papers
- Singapore5.7% · 8 papers
- India5.7% · 8 papers
- South Korea4.3% · 6 papers
Across 141 papers on this subject with at least one lab located. 34 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
Yuxiang Wang, Kunyu Feng, Yuancheng Wang, Zihang Liu, Shengbo Cai, Qinke Ni, Wan Lin, Tao Feng, Yingda shen, Ming-Hao Hsu, Zhixian Zhao, Liqiang Zhang, Teddy Sun, Steve Yves, Zhizheng Wu · 2 October 2026
Model intelligence and fast response jointly shape the quality of interaction with speech language models, yet remain difficult to achieve together. Explicit chain-of-thought (CoT) improves reasoning and audio understanding, but generating intermediate reasoning tokens delays responses. Describing f…
- Mixture of Decoders for Diverse Dialog Response Generation
Wenchao Du · 2 October 2026
Mixture modeling is a long established machine learning technique for learning large sets of multi-modal data. While it is known that sequence-to-sequence models for dialog response generation suffer from the problem of low diversity, we hypothesize that it is because sequence-to-sequence models ten…
- Multi-Party Backchannel Prediction: a Diagnosis, a Benchmark, and a Ceiling
Mohammed Hafsati, Ahmed Loughzali · 2 October 2026
Backchannel prediction has been studied almost entirely in dyadic conversation. We introduce a multi-party benchmark based on the AMI corpus, comprising 682 masked-listener views from 171 meetings, 190 speakers, and 18,697 backchannel events, with a person-disjoint held-out split. A state-of-the-art…
- DuplexSpeechBench-Document Grounding: Benchmarking Document Grounding and Hallucinations in Voice Agents
Puneet Mathur, Nedim Lipka, Zeyu Jin, Dinesh Manocha · 2 October 2026
Voice agents enable low-latency, natural interaction, yet their ability to faithfully ground responses in external documents remains underexplored. We introduce DuplexSpeechBench-Document Grounding (DSB-DG), a benchmark for evaluating document grounding in voice agents across five professional domai…
- Automated Evaluation of Multi-Turn Dialogues in In-Car Conversational Assistants
Vaishnav Negi, Lev Sorokin, Soroosh Tayebi Arasteh, Andrea Stocco · 1 October 2026
In-car conversational assistants (ICAs) are increasingly integrated into vehicles to support route planning, vehicle control, and information access. Ensuring their reliability is challenging due to multi-turn interactions, the absence of explicit ground truth, and strict safety constraints. Existin…
- FD-VAD: Semantic Endpoint Detection for Streaming Full-Duplex Speech
Puneet Mathur, Dinesh Manocha · 1 October 2026
Natural turn-taking in full-duplex voice interaction requires determining from partial speech whether a pause reflects hesitation or a completed conversational intent. Acoustic voice activity detection lacks this semantic information, while cascaded ASR-based endpointing introduces transcription dep…
- SCB: SpeechConversationBench for Evaluating Multi-Turn Reasoning in Speech-to-Speech Models
Kanpat Vesessook, Saksorn Ruangtanusak · 1 October 2026
Speech-to-speech systems must solve tasks whose requirements emerge across conversational turns. We introduce SpeechConversationBench (SCB), a focused evaluation of spoken mathematical reasoning using 103 sharded GSM8K problems. The framework compares the original problem delivered in one turn (full…
- Does Gradient Conflict Predict the Understanding--Generation Trade-off? A Controlled Audit of Conflict-Metric Validity in Unified Multimodal Models
Shuyang Jiang, Fucheng Deng, Yuchuan Luo, Zhenyu Wu · 1 October 2026
Unified multimodal models (UMMs) are increasingly designed around gradient conflict between understanding and generation objectives. The premise that reducing these metrics improves the downstream understanding-generation trade-off has never been tested directly. We audit it in a controlled testbed,…
- Talk2Agent: Benchmarking Voice Interfaces for Text Agents
Terumi Chiba, Guangzhi Sun, Zheqi Yuan, Chao Zhang · 1 October 2026
Large language model (LLM) computer-use agents are typically evaluated with clean written instructions, despite speech being an increasingly popular interface for interacting with such systems. Speech input introduces an additional failure point: transcription errors can alter task-critical entities…
- Beyond Mode Collapse: Generating Diverse Synthetic Expert Conversations via Generative Flow Networks
Sumit Asthana, Michael Ion, Kevyn Collins Thompson · 1 October 2026
High quality synthetic data is central to post training LLMs for adaptive AI applications that represent the diverse expert strategies and decisions in conversations. Prompting LLMs directly or conditioning them on end use scenarios yields low diversity data that collapses onto dominant modes. We pr…
- LLMs are General Asynchronous Agents
George Yakushev, Denis Mazur, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko · 30 September 2026
Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform a…
- Distributional Metrics for Evaluating Spoken Conversational Systems
Shree Harsha Bokkahalli Satish, Erica Cooper, Patr\'icia Schmidtov\'a, Maike Z\"ufle, \'Eva Sz\'ekely, Nicholas Sanders, Ond\v{r}ej Klejch · 30 September 2026
Evaluating conversational systems is a difficult and unresolved problem. We introduce the Conversational Distribution Score (CDS), which compares distributions of conversational behaviour using human conversations as a reference. CDS describes speech rate, syllabic rhythm, and turn interaction throu…
- Towards Communication-Efficient Social Intelligence in Language Agents
Linxiao Gong, Yijie Xu, Tianfu Wang, Yin Wu, Yili Wang, Xingbo Yao, Huizai Yao, Xilin Xia, Haowen Yang, Hui Xiong · 30 September 2026
Socially intelligent language agents must negotiate, coordinate, and resolve conflicting preferences while respecting the time and attention of both participants. Balancing these demands is challenging because agents must convey enough to address a partner's constraints and advance their goals witho…
- SALMONN-duo: Adaptive Dual-System Coordination for Full-Duplex Voice Agents
Wenyi Yu, Siyin Wang, Terumi Chiba, Xianzhao Chen, Xiaohai Tian, Jun Zhang, Lu Lu, Chao Zhang · 30 September 2026
Full-duplex speech large language models (LLMs) enable low-latency, natural voice interaction. However, real-world agents must also use tools and perform deliberative reasoning-operations whose variable latency and computational cost conflict with the stringent timing requirements of real-time conve…
- Understanding Clinical Cognitive Dialogues Using Large Language Models
Vishalakshi Arumugam, Dan Schumacher, Veronica Rammouz, Enrique Gonzalez Guerrero, Jeremy Davis, Anthony Rios · 30 September 2026
In-person cognitive assessment is both a test and an interaction. Clinicians explain tasks, repair misunderstandings, and adapt to patient responses, while patients may hesitate, seek clarification, or disengage. Yet clinical dialogue resources rarely label the interaction structure needed to study …
- Lost with a Map: Conversational State and Behavioral Reliability in Language Models
Atahan Dokme, Larry Heck · 30 September 2026
Task-oriented dialogue requires maintaining and updating information across turns, yet language models expose no explicit belief-state object. We study how conversational state is represented, updated, and used inside eight instruction-tuned language models from four families on MultiWOZ and SGD. St…
- Context Spanning: A Communication Framework for Full-Duplex Speech Models and External LLM Backends
Seonghyeon Go, Yongwoo Kim, Hyeonjin Cha, Jaeho Shin · 30 September 2026
Full-duplex spoken dialogue models can listen and speak simultaneously like the real-time dynamics of human conversation. For natural dialogue, the ability to search for external information in real-time is also an important capability. Many models remain trapped in parametric knowledge, leaving the…
- A model of rational interlocutors: Unification of comprehension and production
Hanlin Wu, Zhenguang G. Cai · 30 September 2026
Who we communicate with influences both our interpretation of their utterances and the design of our own. Such adjustment to the conversational partner is studied as speaker modeling in comprehension and as audience design in production, with the two literatures having developed largely separately. …
- IndicFDB: Benchmarking Full-Duplex Voice Agents across Indian Languages
Rajarshi Roy, Shobhit Banga, Jonathan Raiman, Supriya Paul, Bhaskar Singh, Manmeet Kaur, Sagar Jain, Hanuman Sidh, Pranav Sharma, Aditya Singh, Aaditya Pareek, Manas Dhir, Adi Margolin, Niket Agarwal, Bryan Catanzaro · 30 September 2026
Full-duplex voice agents must handle pauses, take turns, backchannel, and respond to user interruptions in real time. Full-Duplex-Bench evaluates these behaviors, but its English-only corpus and reliance on word-timestamped ASR and an English-prompted LLM judge make it difficult to extend to Indian …
- Thinking in Depth, Speaking Directly: Recurrent Latent Reasoning for Paralinguistically Grounded Spoken Dialogue
Shengbo Cai, Yuxiang Wang, Jingran Xie, Zhisheng Zhang, Shun Lei, Di Cao, Teddy Sun, Zhiyong Wu · 30 September 2026
Empathetic spoken dialogue requires models to use both what is said and how it is said to decide how to respond. Explicit CoT can improve paralinguistic perception and make acoustic cues more explicit in replies, yet does not ensure their effective use in response planning. We call this mismatch the…
- MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
Ke Wang, Houxing Ren, Zimu Lu, Yunqiao Yang, Zhuofan Zong, Mingjie Zhan, Hongsheng Li · 30 September 2026
End-to-end full-duplex speech models have brought open-source machine conversation closer to human-like interaction, yet existing systems remain limited in two intertwined dimensions: long-context robustness and multi-party interaction. Real-world scenarios such as meetings, group lessons, and socia…
- When Users Change Their Minds: Measuring and Repairing Intent Drift in LLM Agents
Yanjie Zhang, Bowen Cao, Zixin Chen, Yushi Sun · 29 September 2026
LLM agents often operate over multi-turn interactions in which user intent changes before execution. We study intent drift: the failure mode in which superseded parts of the user's intent continue to influence the final answer or tool action. We introduce IntentFlux, an executable benchmark that con…
- Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game
Vivek Anand, Muthu Chandrasekaran, Shiva Chaitanya · 29 September 2026
We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referential game: one sees an object and describes it in a short fixed-length message over a small alphabet; the other must pick t…
- DuplexCadence: Exact State and Execution from a Speech Model's Declared Timelines
Haixiao Gao, Yimin Zheng, Linyou Xiao, Zeke Xie · 29 September 2026
Full-duplex speech models support streaming interaction that listens and speaks at the same time. Serving them is governed by a strict, repeating deadline: conversation advances on a one-second cadence, and every second of input must be turned into a second of speech before the next second arrives. …
- APEX-Voice: Can Voice Agents Complete Professional Workflows Through Full-Duplex Interaction
Puneet Mathur, Dinesh Manocha · 29 September 2026
Full-duplex voice agents can now listen, speak, use tools, and act during spoken interactions, but fluent dialogue does not guarantee correct completion of delegated professional workflows. We introduce APEX-Voice, a benchmark of 120 interactive professional workflows spanning ten work archetypes su…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
