Social Sciences › Psychology › Experimental and Cognitive Psychology
Phonetics and Phonology Research
82 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos36 % · 17 artículos
- Alemania19 % · 9 artículos
- Japón13 % · 6 artículos
- Países Bajos11 % · 5 artículos
- China8,5 % · 4 artículos
- Francia6,4 % · 3 artículos
- India6,4 % · 3 artículos
- Reino Unido4,3 % · 2 artículos
Sobre 47 artículos de este tema con al menos un laboratorio localizado. 19 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- The Shape of Speech: A Geometric Measure of Coarticulation for Speech-Driven 3D Facial Animation
Danzel Serrano, Przemyslaw Musialski · 5 de octubre de 2026
Speech-driven 3D facial animation can reproduce recognizable mouth poses. However, it can simplify the motion between them, and that motion carries coarticulation, the way the sounds around each sound shape its articulation. We introduce a geometric measure of this trajectory shaping: lip-path lengt…
- Child-Adapted Structured Phonological Representations for Interpretable Speech Sound Analysis
Abner Hernandez, Tom\'as Arias Vergara, Andreas Maier, Paula Andrea P\'erez-Toro · 2 de octubre de 2026
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, a…
- Do Audio Language Models Hear and Read Distinctive Features Alike?
Yuanhao Chen, Peter Chin · 25 de septiembre de 2026
Audio language models pass speech and text through a single decoder. We ask whether that decoder represents a distinctive feature in the same direction when a phoneme is heard and when it is read. For minimal pairs of phonemes differing in one feature, we take the offset between the two members' mea…
- A Native-Reference Phone-Class Geometry for Second-Language Pronunciation Analysis
Tina Raissi, Nhan Phan, Chenxiao Wang, Mikko Kurimo · 25 de septiembre de 2026
Automatic speaking assessment systems can provide holistic proficiency scores, but often lack interpretable measures that characterize pronunciation quality. We propose a native-reference phone-class geometry for measuring second language (L2) pronunciation deviation without requiring pronunciation …
- Anatomy-aware cross-speaker adaptation of complete vocal-tract acoustic-to-articulatory inversion
Nhat-Nam Nguyen, Pierre-Andre Vuissoz, Yves Laprie · 25 de septiembre de 2026
Cross-speaker acoustic-to-articulatory inversion requires accounting for anatomical differences between speakers. We propose a geometric adaptation framework that uses anatomical landmarks, primarily on vertebrae and dental structures,to transfer predictions from a fixed inversion model to unseen sp…
- A Native-Reference Coordinate Geometry for L2 Pronunciation Deviation Using Self-Supervised Speech Models
Tina Raissi, Nhan Phan, Mikko Kurimo · 24 de septiembre de 2026
Self-supervised speech models encode rich phonetic information, but it remains unclear how to transform this information into interpretable metrics for second-language (L2) pronunciation assessment in spontaneous speech. We propose a native-reference coordinate geometry in which phone-class averages…
- Morpho-VITS: Variational Inference with Morphological Modeling for End-to-End Speech Synthesis of a Tonal Bantu Language
Antoine Nzeyimana · 22 de septiembre de 2026
Text-to-speech models for Bantu tonal languages are challenged by a tonal system that is rooted in both the lexis (i.e., the inventory of words, stems, and affixes) and the grammar (i.e., morpho-syntax). To complicate matters, the standard writing systems of these languages often omit tone markings …
- ToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification
Qisheng Liao, Youngah Do · 22 de septiembre de 2026
Tone languages constitute over 50-70% of the world's languages, but the vast majority are low-resource, lacking the large transcribed corpora needed for automatic tone classification. Existing datasets are typically collected at the sentence level, whereas field linguists require fine-grained syllab…
- Beyond Atomic Tokens: Factorizing Syllables for Language Model Pretraining
Nghia Hieu Nguyen, Thai Bao Huynh, Binh-An Dinh-Le, Phu Gia Hoang, Dat Tien Nguyen, Kiet Van Nguyen, Ngan Luu-Thuy Nguyen · 21 de septiembre de 2026
Conventional tokenizers represent text as characters or statistically derived subwords, overlooking the internal phonological structure of syllables and often requiring large vocabularies. We introduce \textbf{Phonemic Tokenizer}, a linguistically motivated tokenizer for Vietnamese and Chinese that …
- A Phonemically Comprehensive, ASCII-Only Romanization Scheme for Thai and Lao: Systematic Cross-Lingual Correspondence and Chinese-User-Friendly Design
Zijie Zhang, Tan Lee · 18 de septiembre de 2026
This paper proposes a phonemically comprehensive, ASCII-only romanization scheme for Thai and Lao, treating the two closely related languages as a unified cross-lingual design problem. The scheme represents segmental contrasts, vowel length, and lexical tone while maintaining one-symbol-one-phoneme …
- CLASH: Counterfactual Auditing of Lexical and Prosodic Reliance in Spoken Sarcasm Detection
Qiyang Sun, Xudong Li, Yupei Li, Jiabin Xue, Yuhang Dai, Jiaming Li, Bjorn W. Schuller · 16 de septiembre de 2026
Spoken sarcasm detectors may exploit lexical content, prosody, or their interaction, yet conventional evaluation cannot reveal which cues drive their predictions. We introduce CLASH (Controlled Lexical-Acoustic Separation Harness), a bilingual counterfactual diagnostic framework that evaluates each …
- Speaker or Language? Explaining Variance in Charismatic Prosody Across Luxembourgish and French
Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr · 16 de septiembre de 2026
Charismatic speech is shaped by language and speaking style, yet their relative contribution in bilingual public speaking remains unclear. We analyzed spontaneous speeches of 10 politicians who address audiences in Luxembourgish and French, in highly comparable communicative contexts across language…
- Speaker-Specific and Language-Dependent Temporal Organization in Bilingual Political Speech
Nina Hosseini-Kivanani, Nafiseh Taghva, Peter Gilles, Oliver Niebuhr · 16 de septiembre de 2026
Speech rhythm helps structure persuasive speech, but most empirical work examines monolingual English. This study asks how politicians organize timing when speaking Luxembourgish and French. We analyze 400 sentences from ten politicians, annotated for segments and pauses. We compute rhythm metrics, …
- Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking
Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath · 14 de septiembre de 2026
Consonants contribute unequally to whether a word is understood. Given the limited time available for therapy, ranking consonants by contribution to intelligibility helps prioritize intervention targets in motor speech disorders. However, measuring this contribution relies on perceptual studies that…
- Evaluation of Phonetic Encoding Algorithms on Transcription Datasets
Can \"Ozbey, Emre Kaplan, Berkin Deniz Kahya · 7 de septiembre de 2026
In this work, a novel evaluation scheme built on a generalized variant of the Rand Index measure, namely, the H\"ullermeier-Rifqi Index, is proposed in order to assess how well phonetic encoding algorithms conform to word-based transcriptions in IPA (International Phonetic Alphabet) notation. For th…
- Opinionated, Hesitant and Stressed: Three Studies of How Politicians Speak in Four Slavic Parliaments
Ivan Porupski, Nikola Ljube\v{s}i\'c · 1 de septiembre de 2026
We present three large-scale studies of spoken parliamentary speech across four Slavic languages (Croatian, Czech, Polish, Serbian), drawing on over 6,000 hours from the ParlaSpeech 3.0 corpus. The first study examines how utterance-level sentiment shapes acoustic realisation: negative speech is con…
- Using Prosody to Predict Syntactic Structure
Junghyun Min, Alex Warstadt, Tamar I. Regev, Tiago Pimentel, Ethan Gotlieb Wilcox · 1 de septiembre de 2026
While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two domains remains contested. We investigate the syntax-prosody interface through an information-theoretic lens, quantifying the interaction between prosodic…
- Toward a Cross-Lingual Romanization Ecosystem for Sinitic Languages: A Paired Mandarin-Cantonese Case Study
Zijie Zhang, Tan Lee, Yong Cao, Benyou Wang · 1 de septiembre de 2026
This paper proposes the Sinitic Romanization Ecosystem, a cross-lingual Sinitic romanization design framework with supporting digital infrastructure and a community-driven open-source workflow. The design framework addresses the lack of systematic cross-lingual romanization alignment among Sinitic l…
- Is Prosody Lost in Translation? Fine-Grained Cross-Lingual Prosody Similarity Across Languages
Haopeng Xie, Ismail Rasim Ulgen, Sofia Son, Berrak Sisman, Philipp Koehn · 31 de agosto de 2026
Prosody plays an important role in speech translation, conveying information such as emphasis, emotion, and intent beyond lexical content. However, despite recent progress in expressive speech-to-speech translation (S2ST), little is known about how prosodic patterns are similar/different across lang…
- Letters hide the truth from our eyes: English homophones have meaningfully different phonetic realizations
Yu-Hsiang Tseng, Mirjam T. C. Ernestus, Louis F. M. ten Bosch, R. Harald Baayen · 28 de agosto de 2026
The distribution of spoken word duration of English homophones is known to co-vary with frequency of use. This study investigates whether other aspects of the phonetic realization of homophones also differ. A series of quantitative investigations of 14,000 homophone tokens in American television new…
- A Simulator-Grounded Framework For Constructing Verifiable Muscle-Grounded QA From 3D Tongue Meshes (extended version)
Seungho Eum, Unsang Park · 27 de agosto de 2026
Existing articulatory corpora based on real-time MRI and electromagnetic articulography capture tongue shape and motion but do not provide traceable labels for the muscle-driven process that generated an observed configuration. We introduce a simulator-grounded data-construction framework and instan…
- Preference Optimization for Non-Verbal Vocalization Synthesis
Haoyang Li, Chenglin Xu, Junchuan Zhao, Yuang Cao, Liumeng Xue, Yiwen Guo, Eng Siong Chng · 26 de agosto de 2026
Non-verbal vocalizations (NVs), such as laughter, coughs, and sighs, are essential for expressive TTS, but the effectiveness of preference optimization for NV generation remains poorly understood. We systematically study preference optimization for NV-capable TTS, focusing on preference signals, pre…
- Represented but Ignored: A Causal Account of Prosodic Underuse in Audio-Language Models
Linkai Peng, Baorian Nuchged · 21 de agosto de 2026
Human speech is richly expressive, with prosody carrying linguistic and emotional information beyond the lexical content. A capable large audio-language model (audio-LLM) should therefore support expressive speech understanding, not only transcribing what was said but also interpreting how it was sa…
- Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
Abner Hernandez, Tomás Arias Vergara, Daiqi Liu, Andreas Maier, Paula Andrea Pérez-Toro · 11 de agosto de 2026
Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides …
- Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols · 11 de agosto de 2026
Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive. We deconstruct…
