Social Sciences › Social Sciences › Linguistics and Language
Linguistic Variation and Morphology
48 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- BanglaDial-Abuse: A Corpus-Grounded Dataset for Regional Dialect Identification in Abusive Bangla Text
Hasin Almas Sifat · 2 October 2026
Regional linguistic variation remains an important challenge for Bangla natural language processing, particularly in informal and non-standard text. This paper introduces BanglaDial-Abuse, a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla t…
- Precision over Scale: A Polish-Silesian Benchmark and a Translation System Outperforming Open-Source and Commercial Models
Grzegorz Kulik, Miko{\l}aj Pokrywka, Adam Jatowt, Wojciech Kusa · 2 October 2026
Dialectal machine translation remains challenging due to limited data and strong linguistic variation not captured by standard benchmarks, which often assume standardized and well-edited text. We study Polish-Silesian MT using neural and rule-based systems, evaluating on SiLTT - a new Pol-Szl testse…
- Metric-Construction Coupling Inflates Measured Synthetic Dialect Recovery
Hyojung Han (ThakiCloud) · 2 October 2026
Korean dialect corpora are available but not redistributable: weights may be released, while reproducing training and evaluation from the underlying data cannot be. We ask how much of that supervision synthetic data recovers, and whether that recovery can be measured independently of the synthesis p…
- CARDAMOM: A Micro-Dialectal Arabic Speech Dataset for ASR
Bashar Talafha, Samar M. Magdy, Aisha Alansari, Alaa Alkhawaldeh, Abdurrahman Juma, Sharaf Makahleh, Nour Gamal, Omar Attia, Hanaa Kurdi, Najwa Rizk, Maysa Anaya, Hessah Altimyat, Layal Alhazmi, Shumukh Alotaibi, Hajar Alhadaris, Rayan Alomari, Rahaf Almalaq, Malak Alkhorasani, Sara alghamdi, Rahaf Alshamrani, Nsrin Ashraf, Ibrahim Jaradat, Nada Qardahji, Yasmin Zaraket, Elmoukhtar Brahim, Sidi Ebeidy, Oumoulmouminin Mahmoud, Yahjeb Bouha Khatraty, Meya Haroune, Mohammad Ghaddar, Mohamad Eldirany, Rashed Alamoush, Tala Chhaytle, Nuha Albadi, Yahya El Hadj, Hamzah Luqman, Fadi A. Zaraket, Mustafa Jarrar, Muhammad Abdul-Mageed · 30 September 2026
We present Cardamom, a micro-dialectal Arabic speech dataset designed to support fine-grained evaluation and adaptation of automatic speech recognition (ASR) systems. Community-curated by native speakers familiar with the represented varieties, Cardamom contains approximately 40 hours of transcribed…
- NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task
Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Est\`eve, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed · 24 September 2026
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks spanning Automatic Speech Recognition (ASR), Spoken Dialect Identification…
- Some Dialects Are More Equal Than Others: Non-Prestigious Arabic Dialectal Bias in LLMs
Mai Mohamed Eida, Ryan Dolan, Paul de Nijs, Jonathan Dunn · 22 September 2026
Previous work on Egyptian Arabic in NLP has focused largely on the prestigious Cairene Egyptian Arabic (CEA) dialect, resulting in a lack of representation for the less prestigious Sa'idi Egyptian Arabic (SEA) dialect both in LLM and resource development. Does this lack of representation influence a…
- AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation
Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy, Saad Ezzini, Mo El-Haj, Mustafa Jarrar, Zaid Alyafeai, Bernard Ghanem, Muhammad Abdul-Mageed · 22 September 2026
Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective translation requires modeling not only semantic content but also dialectal variation, conversational context, speaker and addressee characteristics, a…
- Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI
Shamim Rahim Refat, Faika Fairuj Preotee, Shuvashis Sarker, Shifat Islam, Bidyarthi Paul, Mohammad Ashraful Hoque · 22 September 2026
Bangla, the seventh most spoken language in the world, exhibits significant regional dialectal diversity, with dialects such as Barishal, Chattogram, Sylhet, Noakhali, and Mymensingh differing in lexical, morphological, and syntactic characteristics. These variations pose substantial challenges for …
- From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale
Chowdhury Mohammad Abdullah, Rita Orji · 17 September 2026
Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions untouched. Adapting the matched-…
- Rosetta at AlexandriaX-2026: LoRA-Adapted NileChat for Context-Aware Dialectal Arabic Dialogue Translation
Nada Esmaeil, Fathima Rena, Sibi Subhash, Osama Elgendy, Mina Naguib, Salma Omar, Muhammad Arif · 10 September 2026
This paper describes the Rosetta system for Subtask 1 (Context-Aware English-to-Dialectal Arabic Dialogue Translation) of the AlexandriaX shared task, participating in both constrained and unconstrained tracks. The approach fine-tunes a LoRA adapter on NileChat-3B using structured system/user prompt…
- 5-Dialects-BN: Unmasking the Impact of Transliteration on Bangla Dialectal LLMs
Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed, Mir Sazzat Hossain, Md Fahim, Md Farhad Alam Bhuiyan · 10 September 2026
Large Language Models (LLMs) have achieved remarkable progress across natural language processing (NLP) tasks, yet their capabilities degrade sharply for low-resource languages and dialectally diverse settings. Bangla, the world's sixth most spoken language, exemplifies this gap: existing resources …
- EDRAC: Benchmarking Arabic Dialect Reading Comprehension
Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji · 2 September 2026
Dialectal Arabic (DA) remains under-resourced compared to Modern Standard Arabic (MSA), particularly for machine reading comprehension (MRC) and question answering (QA). Existing Arabic QA benchmarks primarily focus on formal written MSA or multiple-choice QA, with limited coverage of naturally spok…
- Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models
Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg · 1 September 2026
Large language models are increasingly deployed in Arabic-speaking markets, yet standard benchmarks overwhelmingly reward Modern Standard Arabic (MSA) fluency while leaving dialectal and culturally grounded competence unmeasured. This gap is consequential: everyday Arabic is largely dialectal, and d…
- The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline
Elle · 27 August 2026
Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this "dialect tax" across the natural language processing pipeline. Using parallel English dialect co…
- Wiktionary as a Crowdsourced Lexicon for English Dialects
Sidney Wong · 18 August 2026
This paper evaluates Wiktionary as an ethically crowdsourced lexicon for English dialects. We took a two-phase approach, providing an in-depth descriptive analysis of the crowdsourced lexicon for 12 national varieties of English before applying the lexicon to geo-referenced, country-level social med…
- Poly-Dialectal Neural Machine Translation System for Bangla Regional Dialects
Rakib Ullah, Ruhul Islam Rahul, Tanbir Ahmed · 13 August 2026
Regional dialectal variation poses a fundamental challenge to natural language processing (NLP) in Bangla, where over 240 million speakers communicate across diverse regional variants that diverge significantly from Standard Colloquial Bangla (SCB) in phonology, morphology, and lexicon. Contemporary…
- How Robust Are LLMs to Vietnamese Dialects?
Minh Tran, Trinh Chau, Thanh-Nhan Le, Nam Tran, Luan Thanh Nguyen, Cuong Dang, Duc Hoang · 12 August 2026
Large Language Models (LLMs) are typically evaluated on standard written Vietnamese, yet everyday communication frequently involves regional dialects that preserve meaning but differ in surface form. Existing Vietnamese dialect work largely addresses this issue through dialect-to-standard normalizat…
- Different Perturbations, Different Mechanisms: Understanding Continued Pre-training for Zero-Shot Dialect Robustness
Aarohi Srivastava, David Chiang · 7 August 2026
Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insigh…
- Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
Mohamed Aziz Khadraoui, Adel Ammar, Bilel Benjdira, Zahid Khan, Skander Turki, Wadii Boulila · 23 July 2026
We present a regression-based approach to Arabic dialect geolocation that models dialectal variation as a continuous geographic space rather than discrete categories. Speaker origin is predicted as continuous latitude-longitude coordinates using a hierarchical neural architecture that fuses frame-le…
- LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Huan Wu, Ali Emami, Muhammad Furquan Hassan, Osaretin Igbinoba, Osakpolor Idusuyi, Osamede Igbinoba, Faiza Khan Khattak, Laleh Seyyed-Kalantari · 9 July 2026
African American English (AAE), a rule-governed dialect spoken by over 30 million people, is routinely misinterpreted and "corrected" by large language models (LLMs). Across six instruction-tuned LLMs (14B to 70B), we show that state-of-the-art models systematically prefer Standard American English …
- DiaLLM: An Investigation into the Robustness-Generation Gap in English Dialect Adaptation
Jordan Painter, Dipankar Srirag, Adarsh Kappiyath, Diptesh Kanojia, Aditya Joshi, Lu Yin · 9 July 2026
Large language models increasingly \emph{understand} dialectal English, yet still \emph{produce} only standard, US-leaning English, leaving dialectal generation, the harder half of the problem, largely unaddressed. We introduce \textbf{DiaLLM}, which continually pretrains three open-weight language …
- Jointly Improving Dialect Identification and ASR in Indian Languages using Multimodal Feature Fusion
Saurabh Kumar, Amartyaveer, Prasanta Kumar Ghosh · 6 July 2026
Automatic Speech Recognition (ASR) and Dialect Identification (DID) are crucial for Indian languages, many of which are low-resource and exhibit significant dialectal differences. Existing methods often optimize ASR or DID individually, resulting in performance trade-offs. In this work, we propose a…
- Benchmarking Frontier LLMs on Arabic Cultural and Sociolinguistic Knowledge: A Cross-Evaluation Framework with Human SME Ground Truth
Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh, Clayton W. Taylor, Ahmed Rashad · 2 July 2026
The cost of human expert evaluation is a principal bottleneck to deploying language models in specialized, high-stakes domains. This is particularly acute for Arabic sociolinguistic knowledge: credible grading requires not only linguistic fluency but deep cultural familiarity that cannot be approxim…
- TMASC: Transmasculine Attitude and Speech Corpus
Sidney Wong · 16 June 2026
We introduce the Transmasculine Attitudes and Speech Corpus (TMASC), a multimodal corpus of 196 transmasculine individuals, including questionnaire responses and 66 audio recordings. The questionnaire includes items exploring the vocal health of transmasculine individuals. The audio recordings inclu…
- P3B3: A Multi-Turn Conversational Benchmark for Measuring European and Brazilian Portuguese Variety Bias in LLMs
Rafael Ferreira, In\^es Vieira, In\^es Calvo, James Furtado, Iago Paulo, Diogo Tavares, Diogo Gl\'oria-Silva, David Semedo, Jo\~ao Magalh\~aes · 16 June 2026
As Large Language Models (LLMs) become embedded in everyday communication, capturing regional linguistic variation is essential for reliable and equitable language use. In Portuguese, European (pt-PT) and Brazilian (pt-BR) varieties remain unevenly represented, with pt-BR dominating in data quantity…
