Physical Sciences › Computer Science › Artificial Intelligence
Natural Language Processing Techniques
1 595 papiers indexés
Les techniques de traitement automatique du langage naturel explorent comment les modèles de langage, comme les Large Language Models, analysent, génèrent ou adaptent du texte dans différentes langues et contextes. Ces travaux abordent des méthodes pour améliorer leur performance dans des environnements multilingues, optimiser leur entraînement avec des données ciblées, ou affiner leur comportement sans recourir à des ajustements lourds des poids du modèle. On y étudie aussi des approches pour évaluer leur efficacité, structurer leur mémoire interne, ou accélérer leur décodage, en s’appuyant sur des architectures comme les transformers ou des stratégies comme le zero-shot learning.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis34 % · 308 articles
- Chine25 % · 221 articles
- Allemagne7,8 % · 70 articles
- Royaume-Uni7,6 % · 68 articles
- Inde6,2 % · 56 articles
- Canada4,8 % · 43 articles
- France4,7 % · 42 articles
- Japon4,3 % · 39 articles
Sur 897 articles de ce sujet dont au moins un laboratoire est situé. 88 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Context-Tower Conversion Preserves Generation While Freezing Retains Knowledge: Low-Budget AR-to-Diffusion Conversion of MoE LLMs
Wentao Lu, Jesse Clark, Tianyu Zhu · 5 octobre 2026
Converting a pretrained autoregressive (AR) model to a diffusion language model (dLLM) enables parallel generation without pretraining a new model. Published conversion methods differ by roughly three orders of magnitude in training data and have not been compared under a common protocol. We compare…
- TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation
Prasanth Bathala, Anubhav Shrimal, Sukhdeep Singh Kharbhanda, Pradyumna Lanka, Rohit Dhaipule · 5 octobre 2026
Large-scale machine-translation (MT) systems are typically evaluated on random samples from a corpus whose distributional composition is an artifact of how it was assembled. Such a sample inherits the phenomena the collection happens to contain rather than the full space a system must handle, spanni…
- The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla · 5 octobre 2026
Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. D…
- Automatic register identification for the open web using multilingual deep learning
Erik Henriksson, Amanda Myntti, Saara Hellstr\"om, Anni Eskelinen, Selcen Erten-Johansson, Veronika Laippala · 5 octobre 2026
This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which contain over 72,000 documents annotated with a hierarchical taxonomy of 25 regi…
- Trained Agentic Context Management
Bryce Sandlund · 5 octobre 2026
We study long context language models. Instead of training long context natively, or designing a long context harness, we train a model over the simplest possible harness: a tool to call itself with any specified prompt and a tool to read tokens in a range from the input context. We finetune Qwen3.6…
- A generative-informed neuro-symbolic framework for syntactic ambiguity resolution: Evidence from Arabic DPs
Mohammed Damom, Muneef Y. Alshawsh, Ashraf A. Naji, Mustafa Ali Alhamzi, Fawwaz An-Nashef, Jameel Ahmed Elayah, Mohammed Q. Shormani, Noman AL-Sayadi · 5 octobre 2026
Syntactic ambiguity poses a persistent challenge for Arabic NLP, particularly in morphologically rich nominal constructions where multiple structu6ral interpretations may be compatible with the same surface sequence. This study proposes a generatively informed neuro-symbolic framework for resolving …
- SyntaxBench: A Statistical Diagnostic Framework for Character-Level Reasoning in Large Language Models
Mohsen Larni (Department of Computer Science, University of Nevada, Las Vegas), Sobhan Ebrahimi Azar (Department of Computer Science, University of Nevada, Las Vegas), Pouyan Nahed (Department of Computer Science, University of Nevada, Las Vegas), Kazem Taghva (Department of Computer Science, University of Nevada, Las Vegas) · 5 octobre 2026
Large language models are increasingly used where small syntactic errors matter, yet character-level reasoning is still evaluated mostly through isolated probes and aggregate accuracy. We introduce SyntaxBench, a diagnostic benchmark and statistical evaluation framework for character-level reasoning…
- Multilingual GSM-Symbolic: What determines capability transfer across languages?
Kenneth Enevoldsen, Riley Herchert, Sofie Mosegaard, Dan Saattrup Smart, Simon Enni, Isaac Chung, Sofie Bruun, Ayush Sunil Munot, Max M\"uller-Eberstein, Adnan El-Assadi, Elisa Bassignana, Gianluca Barmina, Hafsteinn Einarsson, Iben Nyholm Debess, Linda Freienthal, Lukas Galke Poech, Mike Zhang, Nicolas Legrand, Vladimir Salnikov, Yevhen Kostiuk, Zafar Hussain, Sagandeep Kaur, Agnes Toftg{\aa}rd, Marie Mattson, Kristoffer Nielbo · 5 octobre 2026
We understand little about how capabilities acquired in one language carry over to another, or what governs this transfer: evaluations rely on incomparable, saturation-prone datasets and rarely examine its determinants jointly. Identifying what predicts transfer would let us avoid exhaustive evaluat…
- Do Multilingual Encoders Produce Language-Consistent Semantic IDs?
Abhinav Bohra, Anuj Bohra · 2 octobre 2026
Semantic IDs (SIDs) compress item embeddings into discrete code sequences used in generative retrieval. We ask whether a multilingual encoder is sufficient for different-language renderings of the same product to receive language-consistent SIDs. Using Amazon ESCI listings rendered in English, Spani…
- The Asymptotics of Language Model Alignment with Memory
Haricharan Balasundaram, V. Arvind Rameshwar · 2 octobre 2026
Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in probability, ii) $q$ has a higher expected reward than $Q$. Two common techniques for LM alignment are: KL-constrained RL, which requires knowle…
- Yo-ByT5: Efficient and High-Fidelity Diacritic Restoration for Yor\`ub\'a
Ahmad Samuel Gali (University of Lagos), Shamsuddeen Hassan Muhammad (Bayero University Kano, Imperial College London) · 2 octobre 2026
Yor\`ub\'a is a widely spoken tonal language that depends on diacritics to avoid lexical ambiguity. However, it is often written without these diacritics, thereby hindering downstream Natural Language Processing (NLP) tasks. In this paper, we introduce Yo-ByT5, a byte-level Automatic Diacritic Resto…
- Evaluating the Robustness of Japanese LLMs to IME-Related and Typographical Errors
Ryota Mibayashi, Hiroaki Ohshima · 2 octobre 2026
Large language models (LLMs) have achieved strong performance across various natural language processing tasks. However, their robustness to typographical errors remains underexplored, particularly in Japanese, where text input involves multiple writing systems and IME-based conversion. In this stud…
- Assessing the Impact of Language Disparity on Multilingual Linguistic Ability in Large Language Models
Zhanyu Chen, Jaap Jumelet · 2 octobre 2026
Claims about the grammatical competence of multilingual language models vary sharply with how competence is measured, yet the interaction between evaluation paradigm, post-training, and language resource availability has not been systematically examined. We evaluate base and post-trained models from…
- FACET at WMT 2026 Automated Translation Quality Evaluation Task
Ahrii Kim, Chanjun Park, Seong-heum Kim · 2 octobre 2026
Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free…
- Cross-Lingual Alignment for Decoder-Only Models using MoE Routers
Lucas Bandarkar, Clark Peng, Ahmed Haj Ahmed, Aditi Khandelwal, Nanyun Peng · 2 octobre 2026
Cross-lingual contrastive learning has been a core component of multilingual encoder training, but the ability to explicitly align representations is not possible in decoder-only LLMs because of varying multilingual tokenization. However, a growing amount of research suggests that even in LLMs, high…
- Match the Distribution, Not the Compute: Post-Training Multi-Token Prediction Heads
Prachi Badarayani, Aidan Jay, Chenghui Zhou, Dayquan Julienne, Yuan Gao, Tianwei Chen, George Zerveas, Ishmam Zabir, Xiren Zhou, Chris Quirk, Xia Song · 2 octobre 2026
Multi-token prediction (MTP) improves the throughput of autoregressive generation by enabling the language model to draft multiple next tokens per forward pass, while a verification step over draft tokens ensures that token distribution of the backbone is preserved. Every open MTP-family release (Mi…
- Rules Amortize, Pairings Don't: Linguistic Structure Determines What Latent Task Representations Can Replace In-Context Learning
Gunmay Jhingran · 2 octobre 2026
In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at zero-shot inference cost, but recent theory shows a static vector acts as a single synthetic demonstration and must fail on high-rank mappings such as wo…
- Ask a Language Model for Lottery Numbers: Concentration in Repeated Six-of-49 Outputs
Dmitrij \.Zatuchin · 2 octobre 2026
We evaluate six language-model configurations on requests for six distinct random integers from 1-49. Across 1,200 attempted calls using four English prompt variants, 1,184 responses yielded valid tickets. Effective diversity of number frequencies ranged from 9.9 to 18.0, compared with simulated fif…
- Gacha Decoding: Eliciting Diverse Generations Through Instruction Following
Scott Geng, Yufei Zhang, Joseph Lee, Jerry Li, Marjan Ghazvininejad, Pang Wei Koh · 2 octobre 2026
We introduce Gacha Decoding, an inference-time method for eliciting diverse language model generations that scales with model capability. Across open-ended domains (in-the-wild chat, creative writing, planning for image generation, and protein design), Gacha Decoding significantly outperforms existi…
- Pretraining Latent Information Feedback Transformers with Teacher Supervision
Dor Tirosh, Ido Amos, Mor Geva · 1 octobre 2026
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discar…
- Zero-shot Dependency Parsing with Unsupervised Cross-Lingual Bootstrapping
Lalita Lowphansirikul, Attapol Rutherford, Jian Gang Ngui, Sarana Nutanong, Peerat Limkonchotiwat · 1 octobre 2026
Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various language understanding tasks. However, applying this technique to dependency parsing remains a significant challenge due to its syntactic nature. To …
- Chinese-Jev: Bringing System One Model to Chinese-Language Tasks
Zexiao Wang, Zihao Zhang, Xudong Wang, Pan Wang, Ziyi Ye, Haoyu Zhao, Zuxuan Wu, Shuicheng Yan · 1 octobre 2026
System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended responses. However, existing Jev models exhibit limited Chinese-language decision accuracy, restricting their utility in both general and specialized set…
- Lost in Translation: Measuring the Effect of Non-Native English on End User Performance of Large Language Models
Yusheng Zhou, Eleanor Lin, David Jurgens · 1 octobre 2026
Large language models (LLMs) are increasingly used by people whose first language is not English, yet these users have been shown to receive systematically lower-quality responses than fluent speakers. Which specific features of non-native English drive this gap remains unclear, because fluency is i…
- A Character-Level Neural Approach to Sinhala Sandhi Splitting
Yasas Ekanayaka, Deshan Sumanathilaka · 1 octobre 2026
Sinhala Sandhi splitting recovers the constituent words or morphemes hidden inside a phonologically merged surface form. The task is important for Sinhala NLP because Sandhi obscures lexical boundaries, but no prior published work has established a neural benchmark for Sinhala Sandhi splitting. We p…
- Linguistic Loopholes in LLM Unlearning: From a 174-Language Benchmark to Coverage-Aware Unlearning
Tyler Skow, Shravan Chaudhari, Rama Chellappa, Abhay Yadav · 1 octobre 2026
Unlearning a fact in one language does not guarantee its removal in others as changing the query or even the requested answer language can reopen seemingly forgotten knowledge -- a cross-lingual loophole. The most straightforward solution to this challenge -- unlearning in all languages -- is neithe…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Large Language Models7 407 papiers / 12 mois+247 %
- Adversarial Robustness in Machine Learning3 552 papiers / 12 mois+118 %
- Reinforcement Learning in Robotics2 519 papiers / 12 mois+117 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
