Physical Sciences › Computer Science › Signal Processing
Speech and Audio Processing
339 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- China41% · 85 papers
- United States30% · 62 papers
- United Kingdom9.3% · 19 papers
- South Korea8.3% · 17 papers
- Germany7.3% · 15 papers
- Israel5.4% · 11 papers
- Japan5.4% · 11 papers
- France2.9% · 6 papers
Across 205 papers on this subject with at least one lab located. 44 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Structured-Noise Masked Modeling for Video, Audio and Beyond
Aritra Bhowmik, Carlos Hinojosa, Fida Mohammad Thoker, Bernard Ghanem, Cees G. M. Snoek · 2 October 2026
Masked modeling has emerged as a robust self-supervised learning framework. However, most methods rely on random masking, which disregards the structural properties of different data modalities. To align with the spatiotemporal and spectral characteristics of video and audio data, we introduce a str…
- Supervising Sound Localization by In-the-wild Egomotion
Anna Min, Ziyang Chen, Hang Zhao, Andrew Owens · 2 October 2026
We present a method for learning binaural sound localization using egomotion as a supervisory signal. Over the course of a video, the cameras direction to a sound source will change as the camera moves. We train an audio model to predict sound directions that are consistent with visual estimates of …
- Multi-agent Auditory Scene Analysis: Improved Localization Speed and Robustness by Multi-beamformed Speech Quality Feedback
Caleb Rascon · 2 October 2026
A real-time auditory scene analyzer (ASA) aims to carry out the tasks of locating, separating and classifying the sound sources present in a given acoustic environment. Recently, an effort has been made into modelling an ASA as a multi-agent system, with each one of its agents performing one of the …
- OP-CAD: On-Policy Clean-Audio Distillation for Robust Audio-Visual Reasoning
Xingming Shui, Dapeng Chen, Bowei Liu, Jingqi Tian, Minfu Li, Kun Yi, Jiapeng Hong, Yansong Tang · 1 October 2026
Omni-modal large language models deployed in real-world environments encounter external noise that can interfere with their perception and understanding of multimodal inputs. We study their robustness in audio-visual understanding, focusing on question answering under environmental noise and competi…
- Here the World in Stereo: Learning Dynamic Spatial Correspondence for Immersive Joint Video-Audio Generation
Hanmo Chen, Chengcheng Liu, Tianxiao Chen, Zheyu Zhang, Siming Zheng, Jinwei Chen, Xu Yang, Cheng Deng, Bo Li, Peng-tao Jiang · 1 October 2026
Recent joint video-audio generation models have achieved strong semantic correspondence and temporal synchronization. However, applications such as AR/VR and interactive gaming further require stereo audio to provide an immersive sense, which remains largely overlooked. Effective stereo audio requir…
- Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
Eric Ming Chen, Jin Woo Lee, Vincent Sitzmann · 1 October 2026
The vocal tract is the region of the human body responsible for filtering one's voice to create speech. In this paper, we present a differentiable and GPU accelerated acoustic simulator for the vocal tract. The differentiable simulator synthesizes speech by propagating sound along an acoustic tube m…
- FFASR: Benchmarking Far-Field Automatic Speech Recognition using High-Fidelity Simulated RIRs
Shivam Saini, Eric Bezzam, Georg G\"otz, Alessia Milo, Steinar Gu{\dh}j\'onsson, Konstantinos Gkanos, Finnur Pind, Daniel Gert Nielsen · 1 October 2026
Far-field automatic speech recognition(ASR) degrades under reverberation, noise, and talker motion, yet the benchmarks that drive model selection emphasize close-microphone speech. We present FFASR, a held-out corpus of 15,637 utterances and an open leaderboard spanning nine conditions, each varying…
- Improving Audiovisual Speech Recognition through Synthetic Visual Data Augmentation
Pol Buitrago, Pol G\`alvez, Javier Hernando · 30 September 2026
Audiovisual Speech Recognition (AVSR) is a multimodal approach to speech recognition that incorporates visual information from lip movements to enhance model performance. Despite its advantages, its development remains constrained by the limited availability of labeled audiovisual (AV) datasets. Thi…
- Multichannel Audio Quality Assessment: Extending Pretrained Perceptual Models to Spatial Audio
Gouthaman KV, Shiv Gehlot, Vishnu Raj, Lars Villemoes, Arijit Biswas · 30 September 2026
Accurate perceptual quality assessment is essential for evaluating and optimizing spatial audio, where perceived quality depends on both signal fidelity and inter-channel spatial relationships. However, subjective evaluation is costly, while existing perceptual models are often trained for limited c…
- Long-Term Memory-Guided Enhancement for Target Perception in Audio-Language Models
Zhenhong Zhou, Xuanyue Zhao, Youji Liu, Yuanhe Zhang, Xiaoyu Ma, Lianyu Hu, Yang Liu · 30 September 2026
Audio large language models (ALLMs) can reason about the content of audio recordings to perform complex tasks. However, these capabilities usually collapse in real-world environments when background noise and competing sources mix the target sound. Inspired by long-term memory in human listening, we…
- Estimation of Room Impulse Responses from Handclaps
Shih-Yu Lai, Kyung Yun Lee, Nils Meyer-Kahlen, Eloi Moliner, Bing-Yu Chen, Vesa V\"alim\"aki · 30 September 2026
Handclaps provide an equipment-free excitation for room acoustics, but their unknown and variable source waveform makes room impulse response (RIR) estimation challenging. In this work, we investigate whether RIRs can be estimated directly from handclaps. To this end, we introduce an anechoic handcl…
- AVIO: Learning to Add and Remove Sounding Objects in Audiovisual Scenes
Weihan Xu, Kan Jen Cheng, Koichi Saito, Jingyu Shi, Tingle Li, Yisi Liu, Liming Wang, Masato Ishii, Takashi Shibuya, Gopala Anumanchipalli, Paul Pu Liang · 30 September 2026
Adding or removing a sounding object requires coordinated changes to visual content and sound while preserving the surrounding scene. Yet paired supervision for localized non-speech audiovisual editing remains limited, as visual and acoustic edits must target the same object and isolate its sound fr…
- Transformer-based Neural Beamforming for Real-Time Speech Enhancement on Smart Low-Power Hearable Devices
Luca Bompani, Marco Fariselli, Giovanni Oltrecolli, Francesco Conti · 29 September 2026
Accurate, efficient, and low-latency spatial beamforming is a key component in emerging smart hearable devices, enhancing speech while suppressing noise and interference. However, handling multiple input sources under strict real-time constraints poses significant challenges for the low-power, resou…
- What Survives the Codec Shift: Pooled No-Vocals Residuals for Speech Deepfake Detection
Jiajun Xu, Menglu Li, Xiao-Ping Zhang · 29 September 2026
The transition from vocoder-based to neural-codec speech synthesis makes generalization more difficult for speech deepfake detectors, particularly those relying on speech-oriented representations. It remains unclear which acoustic representations retain discriminative information when the generation…
- PRIME-ANC: Path-Ratio-Informed Modeling for Efficient Neural Filter Synthesis in Active Noise Control
Yaokun Huang, Chunyang Xu, Haowen Hua, Sen Lin, Shichao Hu, Mengyao Zhu · 29 September 2026
Changes in listener acoustics require active noise control (ANC) filters to be redesigned for new acoustic paths. We introduce PRIME-ANC, a shared neural synthesizer that learns a bounded, path-dependent logmagnitude correction to a regularized path-ratio base. Minimum-phase reconstruction and trunc…
- Adapting Personalized Speech Enhancement for Low-Latency Audio-Visual Target-Speaker Extraction
Rayhan Rashed, Senja Filipi, Ross Cutler · 28 September 2026
Online audio-visual target-speaker extraction aims to remove competing voices while preserving speech quality and bounding lookahead. Existing extractors are built and evaluated for separation on synthetic mixtures, leaving listening quality and meeting behavior largely untested. We introduce Audio-…
- Who Says What: Symbolic Trimodal Binding Mechanisms in Audio-Visual LLMs
Jihoo Jung, Youngjoon Jang, Joon Son Chung · 28 September 2026
Current Audio-Visual LLMs (AVLLMs) struggle with reasoning over videos featuring multi-speaker dialogues. In such videos, resolving "who says what" is crucial, which necessitates trimodal (text-audio-visual) binding. Motivated by these challenges, we systematically investigate how this trimodal bind…
- BAT-CLIP: Trimodal Alignment of Brain, Audio and Text
Suhyun Kim, Jinmo Han, Danny Dongyeop Han, Ahhyun Lucy Lee, Jewoon Lee, Yonghyeon Gwon, Zach Paris, Chun Kee Chung, Saewoong Bahk, Nam Soo Kim, Seong Jae Hwang, Jiook Cha · 28 September 2026
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently mu…
- Does per-frame early exit pay? A compute-matched study of dynamic depth for on-device speech enhancement
Cl\'ement Laroche, Riccardo Miccini · 25 September 2026
Deep learning-based speech enhancement is increasingly deployed on-device in hearing aids, headsets, and earbuds. Most of these devices, however, can only accelerate static int8 graphs, so a depth-varying network must be implemented as several graphs, orchestrated by a policy. In this paper, we supe…
- Beyond Model Size: Redesigning LiSenNet for embedded speech enhancement
Cl\'ement Laroche, Rasmus Kongsgaard Olsson · 25 September 2026
Deploying real-time speech enhancement on resource-constrained devices requires meeting strict latency, memory, and energy constraints. Microcontroller NPUs can accelerate neural inference under these constraints, but only through a restricted set of operators in static, integer-quantized graphs. Re…
- Transcript-Supervised Post-Training of Generative Speech Enhancement on Real Recordings via Reinforce Adjoint Matching
Julius Richter, Christoph Boeddeker, Yoshiki Masuyama, Kohei Saijo, Dominik Klement, Gordon Wichern, Jonathan Le Roux · 25 September 2026
We adapt Reinforce Adjoint Matching (RAM), a reward-based post-training method, to generative speech enhancement (SE). Starting from a pretrained SE model, RAM tilts the model's conditional distribution toward outputs with higher reward. During training, the current model generates enhanced speech o…
- ROAM-ASD: Robust Open-World Active Speaker Detection with Flexible Multimodal Fusion
Pu Wang, Hugo Van hamme · 25 September 2026
Active speaker detection (ASD) requires reliable association between visible faces and acoustic speech, yet existing systems often degrade under challenging domains or incomplete observations. We introduce ROAM-ASD, a robust audiovisual framework that jointly models audio, full-face, and fine-graine…
- AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation
Zhiyu Xu, Weilong Yan, Yufei Shi, Shiyang Li, Yihao Liu, Kin-Man Lam, Yuewen Cao · 25 September 2026
Recent years have witnessed major progress in joint audio-video generation. Existing models still suffer from limited per-modality fidelity, insufficient text-modality alignment and weak cross-modal synchronization. While reinforcement-learning post-training offers a promising remedy, directly adapt…
- ComplexSync: High-Fidelity and Real-Time Lip Sync in Complex Scenarios
Jiaran Cai, Xingpei Ma, Shenneng Huang · 25 September 2026
Lip synchronization aims to generate visual lip dynamics that align precisely with speech audio. Despite the high generation quality of diffusion models, they often struggle in complex scenarios and suffer from prohibitive inference latency, limiting real-world deployment. We present ComplexSync, a …
- Personalized Korean Lipreading as Visual Speech Recognition: Transfer, Census and Adaptation on OLKAVS
Se Un Park, Hakjun Kim, Taehoon Roh, Junyoung Park · 25 September 2026
We present a personalized Korean visual speech recognition (VSR) system and quantify, on the nine-camera OLKAVS corpus, the gap between the population-level benchmark score and an individual user's error. A video-only Conformer initialized from English-trained weights attains 9.95 - 12.19% character…
