Physical Sciences › Computer Science › Artificial Intelligence
Natural Language Processing Techniques
1,595 papers indexed
Natural language processing techniques explore how language models, such as Large Language Models, analyze, generate, or adapt text across different languages and contexts. This research addresses methods to enhance their performance in multilingual settings, optimize their training with targeted data, or refine their behavior without relying on extensive weight adjustments. It also examines approaches to assess their effectiveness, structure their internal memory, or accelerate their decoding, leveraging architectures like transformers or strategies such as zero-shot learning.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States34% · 308 papers
- China25% · 221 papers
- Germany7.8% · 70 papers
- United Kingdom7.6% · 68 papers
- India6.2% · 56 papers
- Canada4.8% · 43 papers
- France4.7% · 42 papers
- Japan4.3% · 39 papers
Across 897 papers on this subject with at least one lab located. 88 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abhinav Bohra, Anuj Bohra · 2 October 2026
Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spani…
- The Asymptotics of Language Model Alignment with Memory
Haricharan Balasundaram, V. Arvind Rameshwar · 2 October 2026
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowle…
- Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yor\`ub\'a
Ahmad Samuel Gali (University of Lagos), Shamsuddeen Hassan Muhammad (Bayero University Kano, Imperial College London) · 2 October 2026
Yor\`ub\'a is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Resto…
- Evaluating the Robustness of Japanese LLMs to IME-Related and Typographical Errors
Ryota Mibayashi, Hiroaki Ohshima · 2 October 2026
Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this stud…
- Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models
Zhanyu Chen, Jaap Jumelet · 2 October 2026
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from…
- FACET at WMT 2026 Automated Translation Quality Evaluation Task
Ahrii Kim, Chanjun Park, Seong-heum Kim · 2 October 2026
Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free…
- Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng · 2 October 2026
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, high…
- Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song · 2 October 2026
Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (Mi…
- Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning
Gunmay Jhingran · 2 October 2026
In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as wo…
- Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs
Dmitrij \.Zatuchin · 2 October 2026
We evaluate six language-model configurations on requests for six distinct random integers from 1-49. Across 1,200 attempted calls using four English prompt variants, 1,184 responses yielded valid tickets. Effective diversity of number frequencies ranged from 9.9 to 18.0, compared with simulated fif…
- Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
Scott Geng, Yufei Zhang, Joseph Lee, Jerry Li, Marjan Ghazvininejad, Pang Wei Koh · 2 October 2026
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existi…
- Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva · 1 October 2026
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discar…
- Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping
Lalita Lowphansirikul, Attapol Rutherford, Jian Gang Ngui, Sarana Nutanong, Peerat Limkonchotiwat · 1 October 2026
Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To …
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Zexiao Wang, Zihao Zhang, Xudong Wang, Pan Wang, Ziyi Ye, Haoyu Zhao, Zuxuan Wu, Shuicheng Yan · 1 October 2026
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized set…
- Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models
Yusheng Zhou, Eleanor Lin, David Jurgens · 1 October 2026
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is i…
- A Character-Level Neural Approach to Sinhala Sandhi Splitting
Yasas Ekanayaka, Deshan Sumanathilaka · 1 October 2026
Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We p…
- Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav · 1 October 2026
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neithe…
- Synthetic Pre-pretraining Survives Scale, but Not as a Grammatical Prior
Atsuki Yamaguchi, Tatsuro Inaba, Joel Niklaus, Michal \v{S}tef\'anik, Aline Villavicencio, Nikolaos Aletras · 1 October 2026
Pre-pretraining (PPT) on synthetic non-natural language data improves token efficiency during language model pre-training (PT). Prior work attributes this gain to a grammatical prior, i.e., a structural inductive bias learned during PPT that transfers to natural language grammar. However, PPT has on…
- Zero-Compute Cross-Lingual Transferability Estimation Using Typological Feature Proxies
Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zar\`e Palanciyan, Joaquin Vanschoren · 1 October 2026
Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingual pre-training, as prior work has done with cross-lingual transfer matrices. We ask whether transfer is predictable from freely available typological …
- The Invisible Language Tax: Token Premiums of French and Regional Languages in 2026 LLM Tokenizers, and a French-Optimized Prototype
Thomas Serval · 1 October 2026
LLM services are billed per token and context windows are measured in tokens, yet the number of tokens needed for the same content varies across languages. We measure this token premium on seven tokenizers of widely used 2026 models (OpenAI o200k, Llama 3, Qwen3, DeepSeek V3/V4, Gemma 3, Mistral Tek…
- Breaking Babel: A Self-Evolving Multi-Agent System for Long-Form Subtitle Translation
Haibo Jin, Xinjie Li, Najmeh Sadoughi, Yang Liu, Yibo Wang, Zhu Liu, Yuzong Liu · 1 October 2026
Long-form subtitle translation requires reasoning over discourse and cultural context spanning episodes or entire series, while maintaining consistent terminology and style. Existing single-LLM methods are largely sentence-level, and multi-agent systems often use static workflows that do not adapt t…
- NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space
Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri · 1 October 2026
In this paper we propose NinaXander, a series of composed language models obtained by connecting layers of frozen language models from different architecture families with a single trained shared-latent adapter. A composed model runs the first layers of one model, converts the resulting intermediate…
- GrammarRL: Effective Grammar-Constrained Decoding via Reinforcement Learning
Gabriele Tuccio, Antonino Furnari, Aldo Gangemi, Misael Mongiov\`{\i} · 1 October 2026
Grammar-constrained generation guarantees syntactic validity, but can substantially degrade semantic quality when the model's preferred outputs are poorly aligned with the imposed grammar. This trade-off is particularly severe when the prompt is underspecified or the model has limited instruction-fo…
- A helps B while B hurts A: directed transfer in instruction-tuning mixture
Nima H. Siboni, Vahid Rostami · 1 October 2026
Adapting a language model to a specialized corpus means choosing which instruction-tuning tasks to train on under a fixed budget, and testing one choice costs a fine-tuning run. Common heuristics add more source tasks or pick sources similar to the target. The first assumes transfer is never negativ…
- R2T: Rule-Encoded Loss Functions for Sequence Tagging in Low-Resource Languages
Mamadou K. Keita, Christopher Homan, Sebastien Diarra · 30 September 2026
We introduce Rule-to-Tag (R2T), a framework that turns linguistic rules into the training signal for neural sequence taggers in low-resource languages. R2T encodes lexical, morphological, and syntactic rules as differentiable loss terms, so that a tagger learns from rules and unlabeled text, with no…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Reinforcement Learning in Robotics2,519 papers / 12 months+117%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
