Physical Sciences › Computer Science › Artificial Intelligence
Speech Recognition and Synthesis
960 indexierte Paper
Die Erkennung und Synthese von Sprache untersuchen Methoden, mit denen Maschinen Sprachsignale verstehen und generieren können. Diese Arbeiten behandeln Herausforderungen wie die Erkennung manipulierter Audioinhalte, die Anpassung mehrsprachiger Modelle an ressourcenarme Sprachen oder die Steuerung von Stimmmerkmalen bei der Erzeugung künstlicher Stimmen. Die Forschung befasst sich auch mit semantischen Verzerrungen, der Robustheit von Modellen gegenüber Schwankungen der Audioqualität oder der Integration multipler Modalitäten für Anwendungen wie Dialoganalyse oder das Verständnis spezifischer Kontexte.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten34 % · 204 Artikel
- China26 % · 154 Artikel
- Frankreich8 % · 48 Artikel
- Vereinigtes Königreich7,1 % · 43 Artikel
- Indien6,5 % · 39 Artikel
- Deutschland6,3 % · 38 Artikel
- Südkorea6 % · 36 Artikel
- Japan5,1 % · 31 Artikel
Über 602 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 75 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Pruning for Efficiency, Paying in Fairness: Demographic Disparities in Pruned Speech-LLMs
Ganesh Pavan Kartikeya Bharadwaj Kolluri, Michael Kampouridis, Ravi Shekhar · 1. Oktober 2026
Speech-LLMs are expensive to run, making compression important for real-world deployment. However, compressed models are usually selected using aggregate word error rate (WER), which can hide how pruning affects different demographic groups. In this work, we systematically study the effect of audio …
- Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
Ki Woong Moon, Daniel Brenner · 1. Oktober 2026
Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBER…
- Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
Kirill Borodin, Vasilii Kudryavtsev, Maxim Maslov, Grach Mkrtchian · 1. Oktober 2026
Text-to-speech models loop, truncate and lose count on text that repeats a phrase many times. We show that repetition itself is what breaks them, not the length that comes with it. Every repeated sentence in our test set is paired with a control of matched sentence and word count in which no word ev…
- Benchmarking Automatic Speech Recognition Tools for Iberian Languages
Fernando L\'opez, Pablo G\'omez, David Solans, Paulo Villegas, Jordi Luque · 1. Oktober 2026
Comprehensive evaluations of automatic speech recognition (ASR) for Iberian languages remain limited, and low-resource languages, biases, and efficiency trade-offs are underexplored. We benchmark eleven systems, ten open-weight models and one commercial API, across five Iberian languages (Basque, Ca…
- MeanVoiceFlow2: Joint Optimization of Mean Flow and Content Encoder for Fast One-Step Zero-Shot Voice Conversion
Takuhiro Kaneko, Hirokazu Kameoka, Kou Tanaka, Yuto Kondo · 1. Oktober 2026
Flow-matching approaches to voice conversion (VC) have gained attention owing to their high speech quality and strong speaker similarity. Among them, one-step models such as MeanVoiceFlow are particularly attractive because they enable efficient inference; however, their reliance on a computationall…
- Who Said What, and Will It Be Remembered? Evaluating Persistent Speaker Attribution Across Meetings
Shantanu Vispute, Aditya Mishra, Siddhartha Saxena · 1. Oktober 2026
Speech transcripts used as long-term memory must preserve both words and stable speaker identities. Existing meeting-transcription metrics either ignore speakers or remap anonymous speakers independently in each recording, so they cannot measure whether the same person retains one identity across me…
- SCIC: Scope- and Codebook-Aware Instruction Conditioning for Speaker-Adapted Expressive TTS
Longyu Lu, Zongwei Du, Mengtao Xing, Zhuoqun Liu, Zifan Guan, Meiguang Jin, Junfeng Ma · 1. Oktober 2026
Long-form live-streaming TTS requires context-dependent prosody and paragraph-level coherence. However, many existing instruction-based TTS systems use global or uniform conditions, providing limited explicit control over clause-level relative prosodic changes. We introduce Speaker-Relative Inline P…
- Audio Token Attention Is Predictable Before the Language Model Runs
Kyoungjun Park, Yunzhe Li, Lili Qiu · 1. Oktober 2026
A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language model's first layers, where image tokens draw little attention. Audio tokens draw much more attention there, and their ranking is still far fro…
- VAmoS Part Deux: Harder, More Realistic Voice-Agent Simulation
Joshua Meyer, Sahar Shayegan, Ritiz Tambi, Ali Khan, Sun Kim, Victor Shih, Mehdi Jamei, Andi Partovi · 1. Oktober 2026
Voice agents in production must handle several requests, background speech, and customers who lose patience. We introduce VAmoS Energy, a benchmark that combines these challenges in 100 calls about utility billing and payment assistance. Each caller makes two to four requests. The agent has sixteen …
- CoSE-E: A Benchmark for Code-switched Speech Evaluation in Enterprise Settings
Shama Gupta, Hoang H Nguyen, Chelsea Huang, Lindsay Devon Brin, Fanny Riols · 30. September 2026
Code-switching (CS), a seamless alternation between languages within a single utterance, remains a critical challenge in automatic speech recognition (ASR). While prior works focus on conversational CS-ASR, enterprise settings demand evaluation of operational impact beyond edit-distance errors: how …
- Controlling Speaking Rate in Autoregressive TTS via Activation Steering
Francesco Verdini, Antonis Asonitis, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet, Juan Pablo Zuluaga Gomez · 30. September 2026
Autoregressive text-to-speech (TTS) systems synthesize natural speech but, once trained, offer little control over speaking rate. We show that speaking rate can be steered at inference time, without retraining, by clamping a single decoder block's activation along a discovered speed axis. A decoder-…
- DuraS2ST: Chain-of-Thought and Reinforcement Learning for Duration-Aligned Speech-to-Speech Translation
Yayue Deng, Dingdong Wang, Yuxuan Hu, Jinyu Li, Yanqing Liu, Yuanyuan Wang, Weidong Chen, Helen M. Meng, Shujie Liu, Xixin Wu · 30. September 2026
Speech-to-speech translation (S2ST) in time-sensitive applications such as video dubbing requires not only semantic fidelity and speaker preservation, but also strict duration consistency to avoid audio-visual misalignment. However, existing S2ST systems largely generate target speech without explic…
- Acoustic Progress Propagation for Long-Horizon Speculative Decoding in ASR
Yuanyuan Jia, Qianqian Yang · 30. September 2026
Speculative decoding accelerates autoregressive automatic speech recognition (ASR), but the acceptance length of alignment-aware drafters can saturate as the draft horizon increases. We propose a progress-aware speculative drafter that recurrently propagates an acoustic progress state across draft s…
- Almieyar: A Culturally Grounded Benchmark for Multi-Dialect Arabic Speech Recognition
Omid Ghahroodi, Anas Madkoor, Dima Faris Al Saudi, Fagr Tahir, Malak Annan, Talha shahid javad allah rakha, Omar Al-Busaidi, Zineb El Kahla, Iheb Zouari, Essa Ahmed Abou Jabal, Ahmed Ezzat, Hind AL-Merekhi, Aisha Hamad M A Al-Naimi, Hadi Wazni, Bushra Alnajjar, Omar Amin, Haya Al-Thani, Houssam Eddine-Othman Lachemat, Marwa Elwakedy, Sundus Abdulmalik Al Nahari, Elahe Zahiri, Osamah Sarraj, Raghad Mousa, Mckeen Assi, Ahd Al Jumah, Heyam Salman, Alhanouf Abdulraqib, Sara Benoumhani, Alia Hamwi, Ayaat Al-Yasseri, Rim Ibrahim Ghazal, Lamia Ben hiba, Mohamed Eltabakh, Fatima Al-Raisi, Yassine El Kheir, Mohammed Abdulrahman, Hamdy Mubarak, Ayah Hashem, Lefkir Meriem, Ehsaneddin Asgari · 30. September 2026
Arabic speech technology has largely focused on Modern Standard Arabic, leaving the living dialects spoken by hundreds of millions under-served. We introduce ALMIEYAR, a culturally grounded ASR benchmark covering 17 Arabic dialects across six families, built entirely from newly recorded speech unsee…
- Evaluating Machine Unlearning in ASR
Diogo Dinis, Francisco Teixeira, Bhiksha Raj, Alberto Abad, Isabel Trancoso · 30. September 2026
Machine unlearning (MU) offers a path to compliance with "right to be forgotten" regulations. While MU has received increasing attention for speech tasks, it remains largely unexplored for Automatic Speech Recognition (ASR). In this work, we investigate whether existing MU algorithms and evaluation …
- NSV-Shift: A Contrastive Benchmark for Non-Speech Vocalization Understanding and Response Adaptation in Speech-to-Speech Models
Ziwei Chen · 30. September 2026
We introduce NSV-Shift, a contrastive benchmark for evaluating whether speech-to-speech models can understand non-speech vocalizations (NSVs) and adapt their responses accordingly. Each pair contains two conversations with identical lexical content that differ only in the NSV embedded in the final t…
- In-Context Adaptation of Encoder-Decoder Models in Speech Recognition
Yen Meng, Sharon Goldwater, Hao Tang · 30. September 2026
In-context learning offers an appealing approach to adapt automatic speech recognition (ASR) models to new speakers, accents, and domains by providing speech-text pairs as demonstrations at inference time. Recent work shows that some LLM-based speech models are capable of ASR in-context adaptation, …
- Jev Matches 7B Language Models for Speech-Neuroprosthesis Rescoring
Gabriele Cin\`a · 30. September 2026
A speech neuroprosthesis decodes attempted speech from brain activity and ends by rescoring the decoder's candidate sentences with a language model of several billion parameters, the only component that needs a GPU. Replacing that model with a cheaper one is hard: general language models asked to pi…
- Language Discrimination Improves Linguistic Learning in Multilingual Speech Models
Maureen de Seyssel, Jie Chi, Zakaria Aldeneh · 30. September 2026
Multilingual self-supervised speech models can benefit from sharing information across languages, but under a matched total pretraining data budget they still fall short of monolingual models. We show that strengthening the model's ability to discriminate languages during pretraining reduces and, on…
- CoLMbo-SV: A Grounded Language Model for Explainable Speaker Verification
Massa Baali, Sarthak Bisht, Ziyue Qiu, Joseph Konan, Rita Singh, Bhiksha Raj · 30. September 2026
Speaker verification systems achieve high accuracy but provide little account of the acoustic evidence behind their judgments. Making these systems inspectable requires exposing interpretable evidence while retaining the richer information on which their decisions depend. We present \textbf{CoLMbo-S…
- Turning Speech Language Models into Multilingual Listeners
Tol\'{u}lop\'{e} \`{O}g\'{u}nr\`{e}m\'{i}, Dan Jurafsky, Chris Manning, Ahmet \"Ust\"un, Martijn Bartelds · 30. September 2026
Speech Language Models (SLMs) that understand spoken language questions support only a few high-resource languages, limiting access to millions of people worldwide. This gap stems from the scarcity of multilingual speech instruction-tuning datasets. We present MULTISPEECHQA, a large-scale, synthetic…
- How to Reduce Whisper Hallucination
Husein Zolkepli · 30. September 2026
Whisper is still what runs in production: one permissively licensed checkpoint, 99 languages, no per-language tuning. But it writes sentences nobody said. On 42 clips of pure room tone, whisper-large-v3 emits words on 61.9% of them and emits something on 100%. The usual response is to distil a stude…
- FA-Bench: A Benchmark for Word-Level and Phone-Level Forced-Alignment and ASR Timestamps Under Clean and Noisy Conditions
Wei Chu, Yuanzhe Dong, Ke Tan, Dong Han, Yichao Zhou, Ruchao Fan, Bingshen Mu, Jingbei Li, Vishwas Shetty, Sarthak Bisht, Ziyue Qiu, Massa Baali, Rita Singh, Bhisha Raj · 30. September 2026
Forced alignment aligns speech audio with a text transcript to generate word and phone timestamps. Published comparisons normalize transcripts, split the data and match boundaries differently, so their numbers cannot be read together. We present FA-Bench, an open framework that fixes those choices o…
- mu-bench: A Multilingual Utterance Transcription Benchmark
Andrea Li (UC Berkeley), Soham Ray (Sierra AI) · 30. September 2026
Voice agents depend on accurate automatic speech recognition (ASR) to act on what callers say, yet ASR is evaluated on read, English-centric speech with word error rate (WER), which penalizes surface rather than semantic differences. We introduce mu-bench, a dataset of 4,270 caller utterances from 2…
- An Evaluation of AI-Supported Evidence-Based Learning for Public Speaking Skill Development
Sashini Hettiarachchi, Shahbaz Siddeeq, Mika Saari, Pekka Abrahamsson · 30. September 2026
Public speaking is an essential skill in academic and professional contexts, but it often causes anxiety. Although several AI-based speech coaching tools exist, they typically provide generic feedback and lack model speeches for learning. This study addresses these gaps by developing a system that i…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
