This week, speech synthesis and recognition clearly dominate the trends. The topic Speech Recognition and Synthesis has risen from 36 to 91 articles in four weeks, marking the strongest growth at the moment.
Three key themes emerge from recent publications:
- Fine-grained evaluation of speech quality, with models specialized in accent or prosody errors.
- Adapting models to atypical voices, such as dysarthric speech.
- Repurposing speech classifiers to generate speech via guided diffusion.
A few representative papers: PASQA: Pitch-Accent-Focused Speech Quality Assessment Model Trained on Synthetic Speech with Accent Errors Repurposing a Speech Classifier for Guided Diffusion-Based Speech Generation Improving End-to-End Speech Recognition for Dysarthric Speech through In-Domain Data Augmentation
