Physical Sciences › Computer Science › Artificial Intelligence
Speech Recognition and Synthesis
960 papers indexed
Speech recognition and synthesis explore methods enabling machines to understand and generate vocal signals. This work addresses challenges such as detecting manipulated audio content, adapting multilingual models to low-resource languages, or controlling vocal characteristics in artificial voice generation. Research also focuses on semantic biases, model robustness to variations in audio quality, or the integration of multiple modalities for applications like dialogue analysis or understanding specific contexts.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States34% · 204 papers
- China26% · 154 papers
- France8% · 48 papers
- United Kingdom7.1% · 43 papers
- India6.5% · 39 papers
- Germany6.3% · 38 papers
- South Korea6% · 36 papers
- Japan5.1% · 31 papers
Across 602 papers on this subject with at least one lab located. 75 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- MGhana-ST: A Low-Resource Speech Translation Dataset for Ghanaian Languages and an Analysis of Multilingual Training Trade-offs
Frank Lawrence Nii Adoquaye Acquaye, Eric George Parakal, Jesse Johnson, Kishankumar Bhimani, Jochebed Afua Basil · 5 October 2026
We present MGhana-ST, a speech translation dataset for four low-resource Ghanaian language varieties: Ga, Twi (Akuapem and Asante), Ewe, and Fante. MGhana-ST is an ongoing annotation effort; the experiments here use a fixed subset of about 16.1 hours of paired speech and English translations. The au…
- Learning When to Commit from Partial Speech for End-to-End Simultaneous Speech Translation
Hieu Hoang, Amittai Axelrod · 5 October 2026
Simultaneous speech translation must emit useful target text before the source is complete while preserving every committed token. We adapt a full-utterance speech language model using prefix supervision derived from its own complete- and partial-waveform translations, requiring neither transcripts …
- DriftTTS: Few-Step Text-to-Speech Without Distillation via Distribution-Matching Drift
Mohammad Nur Hossain Khan, Subrata Biswas, Bashima Islam · 5 October 2026
Few-step neural text-to-speech models often rely on short- ened diffusion or flow-matching schedules, or on distillation from pretrained multi-step teachers. To avoid these depen- dencies, we present DriftTTS, a few-step mel-spectrogram generator trained without a generative teacher, distillation, o…
- Refinement Buys Intelligibility, Search Buys Identity: What Test-Time Compute Buys in Masked-Diffusion TTS
Nityanand Mathur, Hamees Sayed, Ayush Pratap Singh · 5 October 2026
Diffusion language models for text-to-speech combine two forms of computation: model depth (parameters) and refinement steps (inference budget). We ask whether they scale equally across capabilities. We train 15 masked-diffusion codec TTS models varying depth (19-133M parameters, 3 seeds) on 2,000 h…
- Q-SPT: Learnable Query-Based Compression for Low-Frame-Rate Speech Tokenization
Jeeyoung Yun, Seohwan Yun, Sungwoong Kim · 2 October 2026
Neural speech codecs increasingly serve as tokenizers for speech language models (SLMs). Lowering the frame rate reduces the computational and memory costs of SLMs, but makes it difficult to preserve both linguistic information and acoustic detail. Existing approaches rely on rule-based compression:…
- Code-Switching Spoken Language Identification as Multi-Label Set Prediction
Shunsuke Mitsumori, Matthew Wiesner, Shigeo Morishima, Shinji Watanabe · 2 October 2026
Code-switched (CS) speech leaks through the monolingual language identification (LID) filters used to curate massive speech corpora, calling for CS-aware LID (CS-LID). We formulate utterance-level CS-LID as multi-label language-set prediction and propose a set generator that directly outputs the lan…
- How Robust Are Neural Audio Codecs for African Speech? A Multi-Task Benchmark and the Limits of Perceptual Quality
Chibuzor Okocha, Christan Earl Grant · 2 October 2026
Neural audio codecs enable low-bitrate speech compression and tokenization, yet their robustness on accented and multilingual speech remains under-evaluated. We benchmark seven open-source neural codecs (DAC, EnCodec, FocalCodec, LanguageCodec, SemantiCodec, UniCodec, WavTokenizer) on three African …
- High-Value Synthetic Supervision for Parameter-Efficient Adaptation of a Compact Japanese Speech Model
Sidi Chang, Peiying Zhu · 2 October 2026
Private domain speech is difficult to collect and redistribute, while compact models need task-specific supervision. We study an auditable synthetic pipeline that maps Japanese care handoffs directly to six-field structured notes. Using 182 synthetic training and development clips, we adapt a 1.47B …
- SHAMS: An Audio-Grounded Pronunciation Benchmark for Levantine Arabic
Ben Sapirstein, Roy Mattar, Guy Mor-Lan, Ahlam Mohamed, Letizia Cerqueglini, Morris Alper · 2 October 2026
Levantine Arabic (LA) is spoken by tens of millions of people, creating a pressing need for shared benchmarks to evaluate LA speech-language technologies. Evaluating such technology is particularly challenging given LA's internal diversity and its opaque and non-standardized orthography. We present …
- Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar · 1 October 2026
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio …
- Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
Ki Woong Moon, Daniel Brenner · 1 October 2026
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBER…
- Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian · 1 October 2026
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ev…
- Benchmarking Automatic Speech Recognition Tools for Iberian Languages
Fernando L\'opez, Pablo G\'omez, David Solans, Paulo Villegas, Jordi Luque · 1 October 2026
Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Ca…
- MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo · 1 October 2026
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationall…
- Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Shantanu Vispute, Aditya Mishra, Siddhartha Saxena · 1 October 2026
Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across me…
- SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS
Longyu Lu, Zongwei Du, Mengtao Xing, Zhuoqun Liu, Zifan Guan, Meiguang Jin, Junfeng Ma · 1 October 2026
Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline P…
- Audio Token Attention Is Predictable Before the Language Model Runs
Kyoungjun Park, Yunzhe Li, Lili Qiu · 1 October 2026
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far fro…
- VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
Joshua Meyer, Sahar Shayegan, Ritiz Tambi, Ali Khan, Sun Kim, Victor Shih, Mehdi Jamei, Andi Partovi · 1 October 2026
Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen …
- CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols · 30 September 2026
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how …
- Controlling Speaking Rate in Autoregressive TTS via Activation Steering
Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet, Juan Pablo Zuluaga Gomez · 30 September 2026
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-…
- DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu, Yuanyuan Wang, Weidong Chen, Helen M. Meng, Shujie Liu, Xixin Wu · 30 September 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explic…
- Acoustic Progress Propagation for Long-Horizon Speculative Decoding in ASR
Yuanyuan Jia, Qianqian Yang · 30 September 2026
Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates an acoustic progress state across draft s…
- Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, Iheb Zouari, Essa Ahmed Abou Jabal, Ahmed Ezzat, Hind AL-Merekhi, Aisha Hamad M A Al-Naimi, Hadi Wazni, Bushra Alnajjar, Omar Amin, Haya Al-Thani, Houssam Eddine-Othman Lachemat, Marwa Elwakedy, Sundus Abdulmalik Al Nahari, Elahe Zahiri, Osamah Sarraj, Raghad Mousa, Mckeen Assi, Ahd Al Jumah, Heyam Salman, Alhanouf Abdulraqib, Sara Benoumhani, Alia Hamwi, Ayaat Al-Yasseri, Rim Ibrahim Ghazal, Lamia Ben hiba, Mohamed Eltabakh, Fatima Al-Raisi, Yassine El Kheir, Mohammed Abdulrahman, Hamdy Mubarak, Ayah Hashem, Lefkir Meriem, Ehsaneddin Asgari · 30 September 2026
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unsee…
- Evaluating Machine Unlearning in ASR
Diogo Dinis, Francisco Teixeira, Bhiksha Raj, Alberto Abad, Isabel Trancoso · 30 September 2026
Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation …
- NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
Ziwei Chen · 30 September 2026
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final t…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
