This week, sound and speech processing dominate the trends. Two topics show similar growth: 51 papers on music and audio processing (compared to 23 the previous month) and 97 on speech recognition and synthesis (compared to 44). Practical applications are multiplying, particularly for music analysis and voice synthesis.
A few recent works:
- Listening Like a Judge: A Music-Aware Framework for Automatic Singing Performance Evaluation
- Velocity Prediction in Automatic Guitar Transcription
- SLEEPING-DISCO 9M: A large-scale pre-training dataset for generative music modeling
- Neural Speaker Diarization via Multilingual Training: Evaluation on Low-Resource Nepali-Hindi Speech
- VoiceTTA: Enhancing Zero-Shot Text-to-Speech via Reinforcement Learning-Based Test-Time Adaptation
