Physical Sciences › Computer Science › Artificial Intelligence
Large Language Models
7 407 papiers indexés
Les travaux rassemblés ici explorent les mécanismes internes et les performances des Large Language Models, en cherchant à optimiser leur fonctionnement sans modifier leur architecture de base. Certains se concentrent sur des méthodes pour améliorer la précision des réponses, comme le contrôle de l’attention linéaire ou la compression adaptative des prompts, tandis que d’autres analysent comment ces modèles gèrent l’abstention, l’interprétation des données ou la prise de décision en contexte incertain. D’autres encore proposent des cadres pour évaluer leur fiabilité, diagnostiquer leurs erreurs ou structurer leur raisonnement, notamment dans des tâches complexes comme la réponse à des questions multi-étapes.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis41 % · 2 033 articles
- Chine37 % · 1 824 articles
- Royaume-Uni6,4 % · 319 articles
- Allemagne5,8 % · 291 articles
- Inde4,9 % · 247 articles
- Canada4 % · 199 articles
- Corée du Sud3,4 % · 171 articles
- France3,4 % · 168 articles
Sur 4 994 articles de ce sujet dont au moins un laboratoire est situé. 100 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
Yinheng Li, Justin Wagle · 2 octobre 2026
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of t…
- Which LLM to pick? Online Active Model Selection for Large Language Models
Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel · 2 octobre 2026
Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult …
- Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias · 2 octobre 2026
Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicat…
- Role-aware Heuristic Episodic Attention for Conversational LLMs
Wanyang Hong, Zhaoning Zhang, Yi Chen, Libo Zhang, Baihui Liu, Linbo Qiao, Zhiliang Tian, Dongsheng Li · 2 octobre 2026
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attentio…
- Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations
Jungwon Park, Changin Choi, Jimyeong Kim, Nojun Kwak, Wonjong Rhee · 2 octobre 2026
As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-inje…
- Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow
Polina Tsvilodub, Andreas Waldis, Linlu Qiu, Tal Linzen, Michael Franke · 2 octobre 2026
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Ba…
- Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference
Jiangang Chen · 2 octobre 2026
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept an…
- LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian
Claudiu Creanga, Liviu P. Dinu · 2 octobre 2026
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation …
- When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
Esther Xin · 2 octobre 2026
Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a poli…
- From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang · 2 octobre 2026
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use …
- CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang · 2 octobre 2026
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make…
- On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller · 2 octobre 2026
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed sig…
- Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed · 2 octobre 2026
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most …
- Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi · 2 octobre 2026
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, w…
- Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
Darpan Aswal, C\'eline Hudelot · 2 octobre 2026
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-ge…
- Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
Zhiyun Shi · 2 octobre 2026
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand L…
- Distilling Directional Verification
Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim · 2 octobre 2026
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore…
- Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
Keuntae Kim, Yong Suk Choi · 2 octobre 2026
An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual ques…
- The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention
Yufa Zhou · 2 octobre 2026
Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes …
- Verbalized and Internal Probabilities Are Coupled in Large Language Models
Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina · 2 octobre 2026
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that int…
- Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Peter Nutter, Dani Roytburg, Cl\'ement Dumas, Jinghua Ou, Shi Feng · 2 octobre 2026
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideal…
- Initialization Improves LLM-Driven Discovery
Mansi Sakarvadia, Marco Ciccone, Colin Raffel · 2 octobre 2026
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual disco…
- Closing the Loop: Practical Training Recipes for Looped Language Models
Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin, Inessa Fedorova, Dmitry Bocharov, Yuliana Shakhvalieva, Maria Tikhonova, Valerii Ternovskii · 2 octobre 2026
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, …
- The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar · 2 octobre 2026
Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at…
- Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li · 2 octobre 2026
Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Adversarial Robustness in Machine Learning3 552 papiers / 12 mois+118 %
- Reinforcement Learning in Robotics2 519 papiers / 12 mois+117 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
- Natural Language Processing Techniques1 595 papiers / 12 mois+69 %
