Physical Sciences › Computer Science › Computer Vision and Pattern Recognition
Music Technology and Sound Studies
132 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- China43 % · 34 Artikel
- Vereinigte Staaten39 % · 31 Artikel
- Vereinigtes Königreich8,8 % · 7 Artikel
- Taiwan6,3 % · 5 Artikel
- Sonderverwaltungsregion Hongkong6,3 % · 5 Artikel
- Spanien5 % · 4 Artikel
- Griechenland3,8 % · 3 Artikel
- Indien3,8 % · 3 Artikel
Über 80 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 27 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Emergent Tonal Structure in Learned Chord Embeddings and Its Relation to Tonal Tension
Maral Ebrahimzadeh, Gilberto Bernardes, Sebastian Stober · 30. September 2026
Several tonal pitch spaces and computational models have been proposed to analyze tonal structure in Western tonal music, many of them grounded in principles from music theory and used to support tonal analysis with important implications for tonal tension. In parallel, data-driven methods such as s…
- Getting Motif-ated: Controllable AI Compositions from Injected Motif Prompts
Chao Peter Yang, Cynthia Rudin, Yue Jiang, Simon Mak, Stephen Ni-Hahn · 29. September 2026
Deep learning has transformed symbolic music generation by borrowing the training paradigms of large language models, with systems such as NotaGen now producing complete, stylistically convincing classical scores from a short prompt. These systems could become powerful creative partners, helping mus…
- Open-Qwen-Music: An Auditable Framework for LLM-Based Music Composition and Diffusion Rendering
Yangbin Yu, Mingyu Yang · 29. September 2026
We present Open-Qwen-Music, an open reconstruction of Qwen-Music and a fully specified research system for text-to-music generation that couples LLM-based semantic composition with diffusion-based acoustic rendering. The system comprises a 25 Hz single-codebook music tokenizer, a 3B-parameter autore…
- From One-Shot Generation to Incremental Music Composition: Adapting a General-Purpose Instruction LLM for Persistent Symbolic Editing
Andr\'e Ricardo Ducca Fernandes, Jean-Pierre Briot, Simone Diniz Junqueira Barbosa1, H\'elio C\^ortes Vieira Lopes · 29. September 2026
Most music-generation systems are still framed and evaluated primarily as producers of complete outputs, whereas composition often proceeds through successive revisions to a shared musical artifact. This paper studies a different use of a general-purpose instruction-following large language model: n…
- From Human Narrative to Harmonic Structure: A Human-Centered Investigation of Algorithmic Music Generation through the Chord Wheel Diagram
Josef Pavl\'i\v{c}ek, Petra Pavl\'i\v{c}kov\'a, Irena \v{S}trausov\'a · 29. September 2026
Contemporary AI-based music generation can produce compositions that satisfy formal requirements of tonality and musical coherence. However, whether musical expression can be described by mathematical properties alone remains a fundamental question. Human composers operate within personal and cultur…
- Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning
Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos · 23. September 2026
Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantise…
- CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling
Chong Jing, Junan Zhang, Zhizheng Wu · 17. September 2026
Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal te…
- CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation
Jinting Wang, Chenxing Li, Dong Yu, Li Liu · 14. September 2026
Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including struct…
- Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach
Elena Badillo-Goicoechea, Fengfeng He · 10. September 2026
Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation …
- Full-Page Optical Music Recognition of Handwritten Monophonic Scores
Adrian Rosello, Antonio R\'ios-Vila, David Rizo, Jorge Calvo-Zaragoza · 9. September 2026
Full-page end-to-end Optical Music Recognition seeks to transcribe entire music pages directly into symbolic notation, avoiding the limitations of traditional pipelines that rely on accurate staff segmentation. Recent Transformer-based architectures have achieved strong performance on typeset scores…
- InterSing: Explicit Interaction Dynamics for 3D Duet Singing Animation and Beyond
Yihan Zhou, Zikai Huang, Yuyang Yu, Xuemiao Xu, Cheng Xu, Shengfeng He · 7. September 2026
We present InterSing, a framework for generating realistic 3D head animations for duet singing performances. Unlike solo singing, duet performance requires each singer to balance individual expressiveness with intermittent interaction at musically salient moments, such as phrase boundaries, synchron…
- Memory as transformation: LETHE, a self-referential gan-inspired architecture
Francesco Vitucci, Anthony Di Furia, Francesco Scagliola · 7. September 2026
LETHE (Latent-parameter Evolution with Temporal Hierarchical quasi-Equilibrium) is a self-referential sonic-oblivion system implemented in SuperCollider. It adopts the formal vocabulary of Generative Adversarial Networks in a closed configuration without external datasets or supervision after initia…
- Pitch-class Steering for Diffusion-based Music Generation via Latent-space Probes
Yushi Ye, Wilson Zheng, Yongyi Zang · 7. September 2026
Recent work on controllable music generation has focused on autoregressive models, leaving diffusion-based systems comparatively underexplored. We present a lightweight method for steering the pitch content of audio produced by Stable Audio Open, a latent diffusion model for music synthesis. A small…
- Artificial Rosetta Stone: Constrained Maximum A Posteriori (MAP) Reconstruction of Symbolic Raga Sequences via Order-k Markov Models
Saanvi Raghavendran (Abstract Math Institute), Abhishek Bhattacharjee (Abstract Math Institute) · 2. September 2026
Reconstructing a damaged musical fragment is an inverse problem: the observed sequence contains partial information, while a raga encodes constraints limiting allowable completions. This paper formalizes a mathematical framework for this, proposing the Artificial Rosetta Stone (ARS). We separate thr…
- MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI
Laura Ib\'a\~nez-Mart\'inez, Roser Batlle-Roca, Xavier Serra, Mart\'in Rocamora · 1. September 2026
Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains underexplored. Prior work includes openness-focused evaluation frameworks, such as MusGO (Music-Generative Open AI), as well as qualitative studies of m…
- Arbitrary Polygon Oscillator: Generalizing Polygonal Synthesis to Arbitrary Shapes, Morphing, and Three-Dimensional Polyhedra
Antonio Argentieri, Francesco Scagliola · 26. August 2026
Polygonal synthesis generates audio by traversing the perimeter of a polygon with a phasor; prior work uses a constant angular velocity, whereas the proposed system adopts constant arc-length (perimeter) velocity. Existing formulations operate on regular, parametrically defined polygons, producing s…
- One Timeline, Many Renderings: A Wolfram Language Paclet for heterogeneous musical output
Francesco Vitucci, Michele Lorusso, Francesco Scagliola · 26. August 2026
One algorithmic composition may require a Csound score, engraved notation, real-time control, and a rehearsal click. Authored separately, their timelines drift. Temporal System is a Wolfram Language paclet that instead compiles one immutable store of typed entities on a rational beat timeline throug…
- From local kernels to global form: modeling the emergence of musical content
Francesco Vitucci, Michele Lorusso, Francesco Scagliola · 26. August 2026
Markov models are established tools for symbolic music, including non-homogeneous formulations. The narrower contribution examined here is an observation-driven estimation mechanism: overlapping sliding windows derive a trajectory of local transition kernels from one symbolic sequence rather than fr…
- Jiuge-Tuiqiao: An Interpretable Human-AI System for Classical Chinese Poetry Refinement
Yufeng Han, Lifan Deng, Cunliang Kong, Wenhao Li, Xin Cong, Yuzhuo Bai, Kangyang Luo, Maosong Sun · 25. August 2026
Classical Chinese poetry composition has long valued Tuiqiao, the iterative refinement of words, imagery, and prosody. However, many current AI poetry systems follow a one-shot generation paradigm, which reduces users to prompt providers and weakens their creative agency. We present Jiuge-Tuiqiao, a…
- Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
Yi Wang · 19. August 2026
GPT-style models achieve strong performance by representing language with finite vocabularies of reusable discrete tokens. This success has motivated symbolic music tokenizations to treat recurring musical structures, such as chords, motifs, and phrases, as reusable units analogous to linguistic tok…
- Architecture and Affordances of PLAUD: Performative Latents and Unsupervised DDSP
B{\l}a\.zej Kotowski, Frederic Font · 17. August 2026
PLAUD (Performative Latents and Unsupervised DDSP) is a neural synthesizer and Max for Live instrument for live electronic music, built on NoiseBandNet and trained on small personal sound corpora. We present its architecture, combining a variational DDSP synthesis model, latent smoothing, multi-scal…
- Musical Agent Systems: MACAT and MACataRT
Keon Ju M. Lee, Philippe Pasquier · 17. August 2026
Our research explores the development and application of musical agents, human-in-the-loop generative AI systems designed to support music performance and improvisation within co-creative spaces. We introduce MACAT and MACataRT, two distinct musical agent systems crafted to enhance interactive music…
- Musical Mirrors: The LLM as Sounding Board in Songwriting
Xiao Xiao · 17. August 2026
This paper examines a use of AI in creative practice as an interpretive sounding board for human-generated material, rather than the more familiar pattern of AI generation followed by human curation. Through the lens of resonance as theorized by Hartmut Rosa, I present a first-person case study of s…
- Drive-to-Music: Context-Aware Generative Audio for In-Vehicle Experiences
Cosmin Dragoiu, Nooshin Nabizadeh · 14. August 2026
In-vehicle music can serve as an adaptive interface to enhance driver experience, attention, and well-being. We present Drive-to-Music, a context-aware system that generates music in real time from multimodal driving signals. Using dashcam imagery and vehicle telemetry, the system extracts scene sem…
- MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques
Jiabao Zhuang, Changhao Jiang, Hanchen Wang, Jiahao Chen, Zhixiong Yang, Zhenghao Xiang, Yifei Cao, Jiajun Sun, Hui Li, Ming Zhang, Tao Ji, Tao Gui, Qi Zhang, Xuanjing Huang · 13. August 2026
Long-form song generation models continue to improve in duration, structural integrity, and acoustic complexity, making reliable aesthetic rewards increasingly important for aligning these models with human preferences. However, reward models for complete songs remain limited, and existing evaluator…
Weitere Unterthemen aus Bildverarbeitung und Mustererkennung
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Multimodal Machine Learning Applications8.069 Papiere / 12 Monate+191 %
- Generative Adversarial Networks and Image Synthesis4.992 Papiere / 12 Monate+39 %
- Advanced Neural Network Applications2.354 Papiere / 12 Monate+48 %
- Advanced Vision and Imaging841 Papiere / 12 Monate+78 %
- Human Pose and Action Recognition836 Papiere / 12 Monate+457 %
- Face recognition and analysis482 Papiere / 12 Monate+88 %
