Physical Sciences › Computer Science › Artificial Intelligence
Speech Recognition and Synthesis
567 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Unifying Speech Editing Detection and Content Localization via Prior-Enhanced Audio LLMs
Jun Xue, Yi Chai, Yanzhen Ren, Jinshen He, Zhiqiang Tang, Zhuolin Yi, Yihuan Huang, Yuankun Xie, Yujie Chen · 30 January 2026
Speech editing achieves semantic inversion by performing fine-grained segment-level manipulation on original utterances, while preserving global perceptual naturalness. Existing detection studies mainly focus on manually edited speech with explicit splicing artifacts, and therefore struggle to cope …
- VoxMorph: Scalable Zero-shot Voice Identity Morphing via Disentangled Embeddings
Bharath Krishnamurthy, Ajita Rattani · 30 January 2026
Morphing techniques generate artificial biometric samples that combine features from multiple individuals, allowing each contributor to be verified against a single enrolled template. While extensively studied in face recognition, this vulnerability remains largely unexplored in voice biometrics. Pr…
- Reducing Prompt Sensitivity in LLM-based Speech Recognition Through Learnable Projection
Sergio Burdisso, Esa\'u Villatoro-Tello, Shashi Kumar, Srikanth Madikeri, Andr\'es Carofilis, Pradeep Rangappa, Manjunath K E, Kadri Hacioglu, Petr Motlicek, Andreas Stolcke · 30 January 2026
LLM-based automatic speech recognition (ASR), a well-established approach, connects speech foundation models to large language models (LLMs) through a speech-to-LLM projector, yielding promising results. A common design choice in these architectures is the use of a fixed, manually defined prompt dur…
- Text-only adaptation in LLM-based ASR through text denoising
Sergio Burdisso, Esa\'u Villatoro-Tello, Andr\'es Carofilis, Shashi Kumar, Kadri Hacioglu, Srikanth Madikeri, Pradeep Rangappa, Manjunath K E, Petr Motlicek, Shankar Venkatesan, Andreas Stolcke · 30 January 2026
Adapting automatic speech recognition (ASR) systems based on large language models (LLMs) to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on target-domain text often disrupts the critical alignment between speech and text modalities l…
- MK-SGC-SC: Multiple Kernel Guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
Nikhil Raghav, Avisek Gupta, Swagatam Das, Md Sahidullah · 30 January 2026
Speaker diarization aims to segment audio recordings into regions corresponding to individual speakers. Although unsupervised speaker diarization is inherently challenging, the prospect of identifying speaker regions without pretraining or weak supervision motivates research on clustering techniques…
- EuleroDec: A Complex-Valued RVQ-VAE for Efficient and Robust Audio Coding
Luca Cerovaz, Michele Mancusi, Emanuele Rodol\`a · 29 January 2026
Audio codecs power discrete music generative modelling, music streaming and immersive media by shrinking PCM audio to bandwidth-friendly bit-rates. Recent works have gravitated towards processing in the spectral domain; however, spectrogram-domains typically struggle with phase modeling which is nat…
- Do we really need Self-Attention for Streaming Automatic Speech Recognition?
Youness Dkhissi (LIUM), Valentin Vielzeuf (LIUM), Elys Allesiardo (LIUM), Anthony Larcher (LIUM) · 29 January 2026
Transformer-based architectures are the most used architectures in many deep learning fields like Natural Language Processing, Computer Vision or Speech processing. It may encourage the direct use of Transformers in the constrained tasks, without questioning whether it will yield the same benefits a…
- CTC-DRO: Robust Optimization for Reducing Language Disparities in Speech Recognition
Martijn Bartelds, Ananjan Nandi, Moussa Koulako Bala Doumbouya, Dan Jurafsky, Tatsunori Hashimoto, Karen Livescu · 29 January 2026
Modern deep learning models often achieve high overall performance, but consistently fail on specific subgroups. Group distributionally robust optimization (group DRO) addresses this problem by minimizing the worst-group loss, but it fails when group losses misrepresent performance differences betwe…
- In-context Language Learning for Endangered Languages in Speech Recognition
Zhaolin Li, Jan Niehues · 29 January 2026
With approximately 7,000 languages spoken worldwide, current large language models (LLMs) support only a small subset. Prior research indicates LLMs can learn new languages for certain tasks without supervised data. We extend this investigation to speech recognition, investigating whether LLMs can l…
- LTS-VoiceAgent: A Listen-Think-Speak Framework for Efficient Streaming Voice Interaction via Semantic Triggering and Incremental Reasoning
Wenhao Zou, Yuwei Miao, Zhanyu Ma, Jun Xu, Jiuchong Gao, Jinghua Hao, Renqing He, Jingwen Xu · 29 January 2026
Real-time voice agents face a dilemma: end-to-end models often lack deep reasoning, while cascaded pipelines incur high latency by executing ASR, LLM reasoning, and TTS strictly in sequence, unlike human conversation where listeners often start thinking before the speaker finishes. Since cascaded ar…
- Addressing Gradient Misalignment in Data-Augmented Training for Robust Speech Deepfake Detection
Duc-Tuan Truong, Tianchi Liu, Junjie Li, Ruijie Tao, Kong Aik Lee, Eng Siong Chng · 29 January 2026
In speech deepfake detection (SDD), data augmentation (DA) is commonly used to improve model generalization across varied speech conditions and spoofing attacks. However, during training, the backpropagated gradients from original and augmented inputs may misalign, which can result in conflicting pa…
- FastWhisper: Adaptive Self-knowledge Distillation for Real-time Automatic Speech Recognition
Junseok Lee, Nahoon Kim, Sangyong Lee, Chang-Jae Chun · 29 January 2026
Knowledge distillation is one of the most effective methods for model compression. Previous studies have focused on the student model effectively training the predictive distribution of the teacher model. However, during training, the student model may inherit the shortcomings of the teacher model, …
- MK-SGC-SC: Multiple Kernel guided Sparse Graph Construction in Spectral Clustering for Unsupervised Speaker Diarization
Nikhil Raghav, Avisek Gupta, Swagatam Das, Md Sahidullah · 29 January 2026
Speaker diarization aims to segment audio recordings into regions corresponding to individual speakers. Although unsupervised speaker diarization is inherently challenging, the prospect of identifying speaker regions without pretraining or weak supervision motivates research on clustering techniques…
- Hyperbolic Additive Margin Softmax with Hierarchical Information for Speaker Verification
Zhihua Fang, Liang He · 28 January 2026
Speaker embedding learning based on Euclidean space has achieved significant progress, but it is still insufficient in modeling hierarchical information within speaker features. Hyperbolic space, with its negative curvature geometric properties, can efficiently represent hierarchical information wit…
- SE-DiCoW: Self-Enrolled Diarization-Conditioned Whisper
Alexander Polok, Dominik Klement, Samuele Cornell, Matthew Wiesner, Jan \v{C}ernock\'y, Sanjeev Khudanpur, Luk\'a\v{s} Burget · 28 January 2026
Speaker-attributed automatic speech recognition (ASR) in multi-speaker environments remains a major challenge. While some approaches achieve strong performance when fine-tuned on specific domains, few systems generalize well across out-of-domain datasets. Our prior work, Diarization-Conditioned Whis…
- A Benchmark for Audio Reasoning Capabilities of Multimodal Large Language Models
Iwona Christop (Adam Mickiewicz University), Mateusz Czy\.znikiewicz (Samsung R&D Institute Poland), Pawe{\l} Sk\'orzewski (Adam Mickiewicz University), {\L}ukasz Bondaruk (Samsung R&D Institute Poland), Jakub Kubiak (Samsung R&D Institute Poland), Marcin Lewandowski (Samsung R&D Institute Poland), Marek Kubis (Adam Mickiewicz University) · 28 January 2026
The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer the questions that require reasoning skills to combine audio t…
- Language Family Matters: Evaluating LLM-Based ASR Across Linguistic Boundaries
Yuchen Zhang, Ravi Shekhar, Haralambos Mouratidis · 28 January 2026
Large Language Model (LLM)-powered Automatic Speech Recognition (ASR) systems achieve strong performance with limited resources by linking a frozen speech encoder to a pretrained LLM via a lightweight connector. Prior work trains a separate connector per language, overlooking linguistic relatedness.…
- SICL-AT: Another way to adapt Auditory LLM to low-resource task
Haolong Zheng, Siyin Wang, Zengrui Jin, Mark Hasegawa-Johnson · 28 January 2026
Auditory Large Language Models (LLMs) have demonstrated strong performance across a wide range of speech and audio understanding tasks. Nevertheless, they often struggle when applied to low-resource or unfamiliar tasks. In case of labeled in-domain data is scarce or mismatched to the true test distr…
- Residual Tokens Enhance Masked Autoencoders for Speech Modeling
Samir Sadok, St\'ephane Lathuili\`ere, Xavier Alameda-Pineda · 28 January 2026
Recent speech modeling relies on explicit attributes such as pitch, content, and speaker identity, but these alone cannot capture the full richness of natural speech. We introduce RT-MAE, a novel masked autoencoder framework that augments the supervised attributes-based modeling with unsupervised re…
- SonoEdit: Null-Space Constrained Knowledge Editing for Pronunciation Correction in LLM-Based TTS
Ayush Pratap Singh, Harshit Singh, Nityanand Mathur, Akshat Mandloi, Sudarshan Kamath · 27 January 2026
Neural text-to-speech (TTS) systems systematically mispronounce low-resource proper nouns, particularly non-English names, brands, and geographic locations, due to their underrepresentation in predominantly English training corpora. Existing solutions typically rely on expensive multilingual data co…
- Shortcut Learning in Binary Classifier Black Boxes: Applications to Voice Anti-Spoofing and Biometrics
Md Sahidullah, Hye-jin Shim, Rosa Gonzalez Hautam\"aki, Tomi H. Kinnunen · 27 January 2026
The widespread adoption of deep-learning models in data-driven applications has drawn attention to the potential risks associated with biased datasets and models. Neglected or hidden biases within datasets and models can lead to unexpected results. This study addresses the challenges of dataset bias…
- VIBEVOICE-ASR Technical Report
Zhiliang Peng, Jianwei Yu, Yaoyao Chang, Zilong Wang, Li Dong, Yingbo Hao, Yujie Tu, Chenyu Yang, Wenhui Wang, Songchen Xu, Yutao Sun, Hangbo Bao, Weijiang Xu, Yi Zhu, Zehua Wang, Ting Song, Yan Xia, Zewen Chi, Shaohan Huang, Liang Wang, Chuang Ding, Shuai Wang, Xie Chen, Furu Wei · 27 January 2026
This report presents VibeVoice-ASR, a general-purpose speech understanding framework built upon VibeVoice, designed to address the persistent challenges of context fragmentation and multi-speaker complexity in long-form audio (e.g., meetings, podcasts) that remain despite recent advancements in shor…
- Neural Multi-Speaker Voice Cloning for Nepali in Low-Resource Settings
Aayush M. Shrestha, Aditya Bajracharya, Projan Shakya, Dinesh B. Kshatri · 27 January 2026
This research presents a few-shot voice cloning system for Nepali speakers, designed to synthesize speech in a specific speaker's voice from Devanagari text using minimal data. Voice cloning in Nepali remains largely unexplored due to its low-resource nature. To address this, we constructed separate…
- ARTI-6: Towards Six-dimensional Articulatory Speech Encoding
Jihwan Lee, Sean Foley, Thanathai Lertpetchpun, Kevin Huang, Yoonjeong Lee, Tiantian Feng, Louis Goldstein, Dani Byrd, Shrikanth Narayanan · 27 January 2026
We propose ARTI-6, a compact six-dimensional articulatory speech encoding framework derived from real-time MRI data that captures crucial vocal tract regions including the velum, tongue root, and larynx. ARTI-6 consists of three components: (1) a six-dimensional articulatory feature set representing…
- Probing the Hidden Talent of ASR Foundation Models for L2 English Oral Assessment
Fu-An Chao, Bi-Cheng Yan, Berlin Chen · 27 January 2026
In this paper, we explore the untapped potential of Whisper, a well-established automatic speech recognition (ASR) foundation model, in the context of L2 spoken language assessment (SLA). Unlike prior studies that extrinsically analyze transcriptions produced by Whisper, our approach goes a step fur…
