Physical Sciences › Computer Science › Artificial Intelligence
Large Language Models
7.407 indexierte Paper
Die hier versammelten Arbeiten untersuchen die internen Mechanismen und die Leistungsfähigkeit von Large Language Models, wobei sie darauf abzielen, deren Funktionsweise zu optimieren, ohne ihre grundlegende Architektur zu verändern. Einige konzentrieren sich auf Methoden zur Verbesserung der Antwortgenauigkeit, wie die Steuerung der linearen Attention oder die adaptive Kompression von Prompts, während andere analysieren, wie diese Modelle mit Abstention, der Interpretation von Daten oder der Entscheidungsfindung unter unsicheren Bedingungen umgehen. Wieder andere schlagen Frameworks vor, um ihre Zuverlässigkeit zu bewerten, Fehler zu diagnostizieren oder ihr Reasoning zu strukturieren, insbesondere bei komplexen Aufgaben wie der Beantwortung von Multi-Step-Fragen.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten41 % · 2.033 Artikel
- China37 % · 1.824 Artikel
- Vereinigtes Königreich6,4 % · 319 Artikel
- Deutschland5,8 % · 291 Artikel
- Indien4,9 % · 247 Artikel
- Kanada4 % · 199 Artikel
- Südkorea3,4 % · 171 Artikel
- Frankreich3,4 % · 168 Artikel
Über 4.994 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 100 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
Yinheng Li, Justin Wagle · 2. Oktober 2026
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of t…
- Which LLM to pick? Online Active Model Selection for Large Language Models
Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel · 2. Oktober 2026
Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult …
- Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias · 2. Oktober 2026
Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicat…
- Role-aware Heuristic Episodic Attention for Conversational LLMs
Wanyang Hong, Zhaoning Zhang, Yi Chen, Libo Zhang, Baihui Liu, Linbo Qiao, Zhiliang Tian, Dongsheng Li · 2. Oktober 2026
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attentio…
- Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations
Jungwon Park, Changin Choi, Jimyeong Kim, Nojun Kwak, Wonjong Rhee · 2. Oktober 2026
As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-inje…
- Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow
Polina Tsvilodub, Andreas Waldis, Linlu Qiu, Tal Linzen, Michael Franke · 2. Oktober 2026
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Ba…
- Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference
Jiangang Chen · 2. Oktober 2026
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept an…
- LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian
Claudiu Creanga, Liviu P. Dinu · 2. Oktober 2026
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation …
- When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
Esther Xin · 2. Oktober 2026
Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a poli…
- From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang · 2. Oktober 2026
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use …
- CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang · 2. Oktober 2026
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make…
- On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller · 2. Oktober 2026
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed sig…
- Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed · 2. Oktober 2026
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most …
- Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi · 2. Oktober 2026
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, w…
- Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
Darpan Aswal, C\'eline Hudelot · 2. Oktober 2026
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-ge…
- Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
Zhiyun Shi · 2. Oktober 2026
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand L…
- Distilling Directional Verification
Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim · 2. Oktober 2026
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore…
- Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
Keuntae Kim, Yong Suk Choi · 2. Oktober 2026
An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual ques…
- The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention
Yufa Zhou · 2. Oktober 2026
Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes …
- Verbalized and Internal Probabilities Are Coupled in Large Language Models
Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina · 2. Oktober 2026
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that int…
- Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Peter Nutter, Dani Roytburg, Cl\'ement Dumas, Jinghua Ou, Shi Feng · 2. Oktober 2026
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideal…
- Initialization Improves LLM-Driven Discovery
Mansi Sakarvadia, Marco Ciccone, Colin Raffel · 2. Oktober 2026
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual disco…
- Closing the Loop: Practical Training Recipes for Looped Language Models
Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin, Inessa Fedorova, Dmitry Bocharov, Yuliana Shakhvalieva, Maria Tikhonova, Valerii Ternovskii · 2. Oktober 2026
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, …
- The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar · 2. Oktober 2026
Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at…
- Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li · 2. Oktober 2026
Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Reinforcement Learning in Robotics2.519 Papiere / 12 Monate+117 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
- Natural Language Processing Techniques1.595 Papiere / 12 Monate+69 %
