Physical Sciences › Computer Science › Signal Processing
Music and Audio Processing
348 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos36 % · 74 artículos
- China34 % · 70 artículos
- Reino Unido12 % · 24 artículos
- Corea del Sur7,8 % · 16 artículos
- Taiwán6,3 % · 13 artículos
- Francia6,3 % · 13 artículos
- Italia5,4 % · 11 artículos
- Alemania4,9 % · 10 artículos
Sobre 205 artículos de este tema con al menos un laboratorio localizado. 40 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- From Isolated Feature to Orbits: Discovering Music Concepts via Multi-SAE Alignment
Liwei Lin, Gus Xia · 2 de octubre de 2026
How can we understand what a music foundation model has learned \textit{internally}? Most interpretability approaches, such as probing and Sparse Autoencoders (SAEs), focus on identifying individual features with minimal structural assumptions. We argue that many concepts are better understood as \t…
- Audible World Models: Spatially Aware Sound Generation for 3D Worlds
Duowen Chen, Jinjin He, Gouthaman KV, Sandeep Bangalore Venkatesh, Bo Zhu · 1 de octubre de 2026
Text- and image-conditioned world generators can create visually rich 3D environments, yet these worlds often remain silent or rely on soundtracks synthesized solely from text or rendered video. Although such audio can convey what should be heard, it lacks an explicit representation of where sound s…
- UniAE-MoE: A Unified Audio Encoder via Mixture of Experts
Shengbo Cai, Zhisheng Zhang, Zichao Nie, Jing Peng, Jingran Xie, Zhiyong Wu · 1 de octubre de 2026
Large Audio Language Models (LALMs) rely on effective audio encoders for multi-task performance. We introduce UniAE-MoE, a unified audio encoder designed to model cross-domain audio representations and achieve outstanding downstream understanding performance via a Mixture-of-Experts (MoE) architectu…
- CORA: A Protocol for Diagnosing Boundary Robustness in Text-to-Audio Retrieval under Query Reformulations
Jae Min Woo, Kyongmin Kong, Bogyung Jeong, Minjeong Kim, HaeJun Yoo, Du-Seong Chang · 30 de septiembre de 2026
Text-to-Audio (T2A) retrievers are typically evaluated with caption style queries, but the same user intent can be expressed in many forms. We introduce CORA (Caption-Offset Retrieval for Audio), a caption anchored diagnostic protocol that rewrites each source caption into five intent preserving for…
- Uncovering shortcut learning in audio classifiers by discovering recurring concepts in temporal explanations
Cecilia Bola\~nos, Luciana Ferrer, Magdalena Fuentes · 29 de septiembre de 2026
Correlations between events in machine learning datasets may result in shortcut learning, where models learn to predict the target event based on the presence of a correlated event. When these correlations are spurious -- arising from data collection artifacts -- models are likely to perform poorly …
- SCISSOR: Score-Conditioned Instrument Source Separation for Orchestral Recordings
Yiheng Lu, Hao-Wen Dong · 29 de septiembre de 2026
Orchestral separation recovers instrument sections from mixtures in which shared pitches, harmonics, and timbres obscure source identity. An aligned score provides instrument labels, note pitches, and activity times. A score-informed approach appends piano rolls to audio features before mask predict…
- Video-to-Music Generation for Gameplay Videos
Felipe Marra, Lucas N. Ferreira · 29 de septiembre de 2026
Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we investigate this problem in the video game domain, which introduces new challenges for these models: video frames are rendered graphics, music is mostly synthe…
- Synth-JEPA: Joint Embedding Prediction for Renderer-Free Synthesizer Parameter Search
Ben Hayes, Haokun Tian, Stefan Lattner · 28 de septiembre de 2026
Sound matching can be formulated as optimizing synthesizer parameters against an audio-domain objective. However, objectives derived from generic audio representations are often difficult to optimize, while direct search requires rendering every candidate. We introduce Synth-JEPA, which learns mutua…
- Symbiotic Architecture for Post-Hoc Audio Extension of Frozen Language Models
Yotaro Kubo, Qi Sun, Yujin Tang · 28 de septiembre de 2026
This paper proposes an architecture for equipping large language models (LLMs) with audio-understanding capabilities without fine-tuning their weights. The proposed symbiotic architecture employs an injector module that writes audio-conditioned vectors directly into the target LLM's short-term memor…
- Don't CLAP: Are Music-Text Models Bag-of-Words?
Yuan-Chiao Cheng, Alexander Lerch · 28 de septiembre de 2026
Text-to-music systems are assessed on audio quality and on how faithfully the music follows its prompt, and the CLAP score, the cosine similarity between a music-text model's audio and text embeddings, is the standard objective metric of faithfulness. We ask how accurately that score reflects the te…
- AcoustiClaim: A Numeric Claim Benchmark with Instrument Ground Truth
Sheng-Tse Lin, Siyuan Zhai, Chien-Liang Kuo, Massa Baali, Bhiksha Raj · 28 de septiembre de 2026
Audio language models state numbers for acoustic quantities, and neither human opinion nor a judge model says whether such a number is true of the signal. AcoustiClaim extracts each numeric claim from free text, scores it against the instrument that defines the quantity, and classes each quantity by…
- EvoAudio: Recursive Self-Improvement for Audio Understanding
Yuxiang Wang, Shengbo Cai, Yingda Shen, Ming-Hao Hsu, Qinke Ni, Liqiang Zhang, Teddy Sun, Steve Yevs, Zhizheng Wu · 24 de septiembre de 2026
Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, labels from stronger models inherit their errors and limits, and fixed data cannot adapt as the learner improves. We therefore propose EvoAudio,…
- Pose-Aware Multimodal Automatic Tagging for Greek Traditional Music
Alexandros Alexiou, Charilaos Papaioannou, Alexandros Potamianos · 24 de septiembre de 2026
Automatic tagging is a core task in Music Information Retrieval (MIR), yet most tagging systems exploit only audio. Live music performance is inherently multimodal, as semantic labels such as instruments, regional styles, and dance forms are encoded simultaneously across acoustic, visual, and embodi…
- Mizar: A 159M-Parameter Audio-Language Model for Audio Understanding
Kaiyang Li, Shaobo Han, Yue Tian, Shihao Ji · 24 de septiembre de 2026
Audio-language models (ALMs) integrate acoustic perception with the knowledge encoded in language models, enabling contextual understanding of auditory events. Making these capabilities practical on devices with limited memory and computation motivates our focus on small ALMs with fewer than 200M pa…
- Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs
Jing Peng, Zichao Nie, Zhisheng Zhang, Jingran Xie, Zhiyong Wu · 23 de septiembre de 2026
Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To address these gaps, we …
- REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States
Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu, Guowu Tan, Xiang Xie · 23 de septiembre de 2026
Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight met…
- ParA-LLM: A Unified Approach to Paralinguistic and Acoustic Speech Understanding
Nishit Anand, Jiaqi Su, Ke Chen, Yunyun Wang, Dinesh Manocha, Ramani Duraiswami, Rithesh Kumar, Zeyu Jin · 22 de septiembre de 2026
Recent advances in Audio LLMs have achieved human-level speech recognition, yet existing systems struggle to capture paralinguistic aspects such as speaker traits, expressive variations, and environmental acoustic conditions. To address this, we design a framework of 22 paralinguistic characteristic…
- Hi-Singers: A Comprehensive High-Quality Dataset for Expressive Audio-Driven Singing Head Synthesis
Yichi Zhang, Hui Zhang, Guanjun Liu, Yuefeng Zou, Fengzhao Sun, Jun Yu · 22 de septiembre de 2026
State-of-the-art models for audio-driven digital human generation have achieved photo-realistic results in talking-head synthesis. However, extending these models to singing-head synthesis remains challenging due to a significant Domain Gap: singing requires more exaggerated expressions, vivid jaw o…
- MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing
Zeyu Yang, Xinyu Zhang, Zibo Bi, Pei Zhang, Xize Cheng, Jin Xu, Baosong Yang, Satoshi Nakamura · 22 de septiembre de 2026
Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight…
- I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance
Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain, Pinar Donmez, Babak Damavandi · 21 de septiembre de 2026
Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for …
- Enhancing Audio Reasoning via Semantic Summary Prediction
Francesco Bonzi, Pooneh Mousavi, Cem Subakan, Mirco Ravanelli · 21 de septiembre de 2026
Large Audio Language Models (LALMs) perform well on complex question answering but often show a reasoning gap, where explicit Chain-of-Thought (CoT) reduces accuracy compared to direct answers. We hypothesize that long reasoning sequences shift attention away from the audio input. To address this, w…
- Enabling automatic transcription of child-centered audio recordings from real-world environments
Daniil Kocharov, Azarias Galama, Okko R\"as\"anen · 18 de septiembre de 2026
Longform audio recordings obtained with microphones worn by children-also known as child-centered daylong recordings-have become a standard method for studying children's language experiences and their impact on subsequent language development. Transcripts of longform speech audio would enable rich …
- FRAUDSkill: Structured Frozen-Weight Skill Optimization for Audio Anti-Fraud Detection
Chengxian Hu, Zhiming Ma, Mingjun Pan, Yifan Wang, Shun Zhang, Qifan Wang, Zhilei Zhao, Yijin Zhou, Yuxi Zhao, Huiyuan Liu, Peidong Wang, Peng Chen · 18 de septiembre de 2026
Large audio-language models have shown promise for anti-fraud detection by directly processing speech and reasoning over fraud-related evidence. Their deployment, however, requires predictions to follow a predefined label space and a structured decision protocol consisting of service-scenario identi…
- Music Hallucination in Audio-Language Models: A Hierarchical Formulation and Empirical Study
Yu Liu, Jiahui Liu, Zhilin Liu, Cong Cao, Fangfang Yuan, Yuling Yang, Pin Xu, Yanbing Liu · 18 de septiembre de 2026
Audio-language models increasingly generate confident music descriptions that are unsupported by the input audio. We present, to our knowledge, the first music-specific, layer-wise, multi-paradigm empirical study of hallucination in audio-language models and formulate it as a hierarchical perceptual…
- Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception
Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo, Haoyu Song, Ian McLoughlin · 15 de septiembre de 2026
Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. Howeve…
