Physical Sciences › Computer Science › Artificial Intelligence
Large Language Models
7407 artículos indexados
Los trabajos reunidos aquí exploran los mecanismos internos y el rendimiento de los Large Language Models, buscando optimizar su funcionamiento sin modificar su arquitectura básica. Algunos se centran en métodos para mejorar la precisión de las respuestas, como el control de la atención lineal o la compresión adaptativa de los prompts, mientras que otros analizan cómo estos modelos gestionan la abstención, la interpretación de datos o la toma de decisiones en contextos inciertos. Otros más proponen marcos para evaluar su fiabilidad, diagnosticar sus errores o estructurar su razonamiento, en particular en tareas complejas como la respuesta a preguntas multi-pasos.
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual - últimos 12 meses
Países de los laboratorios
- Estados Unidos41 % · 2033 artículos
- China37 % · 1824 artículos
- Reino Unido6,4 % · 319 artículos
- Alemania5,8 % · 291 artículos
- India4,9 % · 247 artículos
- Canadá4 % · 199 artículos
- Corea del Sur3,4 % · 171 artículos
- Francia3,4 % · 168 artículos
Sobre 4994 artículos de este tema con al menos un laboratorio localizado. 100 países representados.
Se trata del país del laboratorio, nunca de la nacionalidad de las personas. Un artículo firmado desde varios países cuenta para cada uno de ellos, por lo que las partes suman más del 100 %. La cobertura es parcial y el vacío no es aleatorio: un investigador cuya institución se desconoce suele publicar poco, lo que sobrerrepresenta a los laboratorios consolidados.
Últimos artículos
- LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
Yinheng Li, Justin Wagle · 2 de octubre de 2026
Jev-style decision models return categorical probability distributions over predefined options without generating free-form text, enabling software systems to act on their outputs directly. In this work, we investigate the extent to which general-purpose LLMs already possess this capability out of t…
- Which LLM to pick? Online Active Model Selection for Large Language Models
Alessandro Turrin, Patrik Okanovic, Torsten Hoefler, Nezihe Merve G\"urel · 2 de octubre de 2026
Large Language Models (LLMs) are increasingly applied to process streaming data, with practitioners relying on benchmarks to select the best model even though these signals only approximate real performance. While oracle annotations can provide reliable feedback, they are often costly and difficult …
- Sentence Specificity Scores for Collaborative Technical Documentation: A Domain-Transfer Study
Rocker D'Antonio, Thomas Benton Townsend, Dimitrios Michael Manias · 2 de octubre de 2026
Collaboration depends on shared context, and technical documentation is one way that context persists across people and AI teammates. Specificity, the amount and exactness of detail expressed in language, shapes what information documentation captures and how precisely that information is communicat…
- Role-aware Heuristic Episodic Attention for Conversational LLMs
Wanyang Hong, Zhaoning Zhang, Yi Chen, Libo Zhang, Baihui Liu, Linbo Qiao, Zhiliang Tian, Dongsheng Li · 2 de octubre de 2026
Large language models often lose track of persistent instructions and relevant information as multi-turn conversations grow. We study this cumulative contextual decay through three related failure modes: attention pollution, dilution, and drift. We propose REA (Role-aware Heuristic Episodic Attentio…
- Efficient Task Adaptation in Large Language Models: A Survey of Weight-Based, Prompt-Based, and Embedding-Based Adaptations
Jungwon Park, Changin Choi, Jimyeong Kim, Nojun Kwak, Wonjong Rhee · 2 de octubre de 2026
As large language models are increasingly deployed across diverse downstream tasks, efficient task adaptation has emerged as a central challenge. In response, a wide range of task adaptation methods have been proposed, spanning parameter-efficient fine-tuning, in-context learning, and embedding-inje…
- Bayesian Fine-tuning Yields Language Models that are as Bayesian as their Beliefs Allow
Polina Tsvilodub, Andreas Waldis, Linlu Qiu, Tal Linzen, Michael Franke · 2 de octubre de 2026
Language models (LMs) are increasingly used for tasks that require reasoning about hidden variables from a few observations, for which Bayesian inference is the normatively correct solution. While supervised fine-tuning of an LM on the outputs of an optimal $\textit{Bayesian}$ model leads to near-Ba…
- Lingtai: What Concept Geometry Reveals--and Does Not Reveal--About LLM Inference
Jiangang Chen · 2 de octubre de 2026
Observing what a large language model computes during autoregressive inference--online and without training probes--remains difficult. We introduce Lingtai, a training-free concept telemetry layer: at each generation step, residual states are projected onto a domain-specific bank of named concept an…
- LLM-as-a-Judge for Low-Resource Languages: Adapting Ragas and Comparative Ranking for Romanian
Claudiu Creanga, Liviu P. Dinu · 2 de octubre de 2026
Evaluating Retrieval-Augmented Generation (RAG) systems remains a challenge for Low-Resource Languages (LRLs), where standard reference-based metrics fall short. This paper investigates the viability of the "LLM-as-a-Judge" paradigm for Romanian by adapting the Ragas framework using next-generation …
- When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora
Esther Xin · 2 de octubre de 2026
Several recent pipelines build RLVR training data by masking a span of real corpus text and asking a language model to invent plausible wrong answers around it. The correct option is therefore genuine human prose; every distractor is synthetic. Correctness and provenance become entangled, and a poli…
- From Knowledge Access to Source Learning: Developing Source-Specific Competence
Lucheng Fu, Kejing Xia, Yiyang Wang, Yiqiao Jin, Jinjin He, Xiyuan Yang, Haoxin Liu, Ye Yu, Haibo Jin, Yijia Xiao, Wenke Lee, B. Aditya Prakash, Haohan Wang · 2 de octubre de 2026
Large language model (LLM) agents increasingly rely on persistent external sources to solve sequences of knowledge-intensive tasks. Existing methods improve how source content is accessed and organized, while agent-memory systems preserve reusable knowledge from prior interactions, but repeated use …
- CARM: Cancellation-Aware Response Masking for LLM Reinforcement Learning
Yafei Zhang, Songshuo Lu, Sicong Liao, Zhi Chen, Yaohua Tang · 2 de octubre de 2026
Recent years have witnessed the rapid adoption of reinforcement learning (RL) in large language model (LLM) post-training, with substantial gains in mathematical reasoning and code generation. In practical systems, however, policy updates and differences between rollout and training engines can make…
- On Language Drift during RLVR Post-Training
Michael Sullivan, Alexander Koller · 2 de octubre de 2026
Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed sig…
- Mem++: Non-Destructive Memory for Long-Term Organizational LLM Agents
Ahmad Yehia, Aly O. Abdelkareem, Islam Ahmed, Hesham Omran, Khaled Alashmouny, Christian Claudel, Abduallah Mohamed · 2 de octubre de 2026
Large Language Model (LLM) agents now take part in organizational work, where many authors record decisions across documents over months. Because a revised decision arrives as a new document rather than an edit, answering a question requires knowing which version held at a given time. However, most …
- Generalization Is Stability, Not Accuracy: Multi-Axis Evaluation of LLMs
Nagham Omar, Mahmoud Jabarin, Maya Rozenshtein, Rom Himelstein, Avi Mendelson, Amit LeVi · 2 de octubre de 2026
Generalization in large language models (LLMs) is the ability to produce consistent and semantically stable outputs when the same input is expressed in different ways. Existing work typically evaluates generalization through aggregate accuracy on a single prompt format, task, or set of variations, w…
- Temporally-Resolved Token Attribution Reveals the Generation Dynamics of Diffusion Language Models
Darpan Aswal, C\'eline Hudelot · 2 de octubre de 2026
This work presents Diffusion Layer Integrated Gradients (DLIG), a token attribution method for diffusion language models (DLMs) that extends Integrated Gradients (IG~\cite{sundararajan2017axiomatic}) to arbitrary layers and denoising steps. DLIG attributes a DLM's progressive commitment to a self-ge…
- Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
Zhiyun Shi · 2 de octubre de 2026
A long-term conversational assistant must recall the right memory at the right moment, yet the memory that matters most is often not similar to what the user says now. Current systems recover such associations by letting an LLM reason at write or read time, at a cost of hundreds to over a thousand L…
- Distilling Directional Verification
Jungseob Lee, Sugyeong Eo, Seongtae Hong, Seungyoon Lee, Chanjun Park, Jaehyung Seo, Heuiseok Lim · 2 de octubre de 2026
Knowledge distillation aims to transfer the factual knowledge of large language models to smaller models for efficient deployment. Yet a teacher may recall a relation in one direction while failing to generate the answer in the reverse direction. Distillation from its generated answers can therefore…
- Two Clocks in Diffusion MLLMs: When Answers Stabilize Before Rationales Unfold
Keuntae Kim, Yong Suk Choi · 2 de octubre de 2026
An answer candidate in a masked diffusion MLLM can stabilize while its rationale is still unfolding. We distinguish retrospective stabilization of the logged candidate from token commitment, and examine these two clocks relative to rationale generation. Analyzing our results across three visual ques…
- The Geometry of Contextual Relations: Language Models Address Facts by Order of Mention
Yufa Zhou · 2 de octubre de 2026
Human reasoning depends on how objects are related within propositions. \textit{How do relations organize the language representations of contextual contents?} We give an LLM a list of facts in its context (e.g., \emph{Alice eats an apple. Bob eats a pear.}) and measure how its hidden state changes …
- Verbalized and Internal Probabilities Are Coupled in Large Language Models
Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina · 2 de octubre de 2026
Large language models carry an internal notion of uncertainty in their sampling distribution, i.e., the probabilities they place on generating one answer rather than another. They can also be asked to state a confidence, in words or as a number: a verbalized uncertainty. Prior work suggests that int…
- Pre-training interventions, ex post facto: Grafting model beliefs across checkpoints
Peter Nutter, Dani Roytburg, Cl\'ement Dumas, Jinghua Ou, Shi Feng · 2 de octubre de 2026
Pre-training interventions are critical to alignment research, since beliefs formed during pre-training shape how a model generalizes from later training. One recently popular technique for such interventions is synthetic document fine-tuning (SDF), which aims to alter what the model believes. Ideal…
- Initialization Improves LLM-Driven Discovery
Mansi Sakarvadia, Marco Ciccone, Colin Raffel · 2 de octubre de 2026
Large Language Models (LLMs) have been used for novel discovery of algorithms, theorems, drugs, and other tasks through the use of harnesses that prompt an LLM to iteratively optimize an objective. In this work, we study the relationship between the population of previous iterates and eventual disco…
- Closing the Loop: Practical Training Recipes for Looped Language Models
Andrei Marchenko, Viacheslav Bezrukov, Oleg Kashurin, Inessa Fedorova, Dmitry Bocharov, Yuliana Shakhvalieva, Maria Tikhonova, Valerii Ternovskii · 2 de octubre de 2026
Looped language models increase effective depth by repeatedly applying a shared block of layers, but existing large-scale recipes require multi-stage training over trillions of tokens, while the benefits of recurrence remain difficult to separate from differences in data and training. In this work, …
- The Weakest Link: Distilling LLM Reasoning with Worst-Case Constrained Reinforcement Learning
Matthieu Zimmer, Xiaotong Ji, Tu Nguyen, Haitham Bou-Ammar · 2 de octubre de 2026
Distilling the reasoning capabilities of large language models (LLMs) into smaller students is a central challenge for efficient deployment. Current approaches face a fundamental tension: optimizing purely for verifiable task rewards (e.g., via GRPO) leads to reward hacking, where students arrive at…
- Certainty Is Not Just Correctness: Rethinking Token-Level Certainty in LLM Reasoning
Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li · 2 de octubre de 2026
Token-level certainty is widely used as a proxy for correctness in LLM training and inference. However, the performance of certainty-based methods depends both on the information in certainty scores and on how those scores are used. We therefore directly assess certainty's predictive ability through…
Otros asuntos del tema Inteligencia artificial
Los asuntos que la clasificación OpenAlex vincula al mismo tema, los más activos primero.
- Adversarial Robustness in Machine Learning3552 artículos / 12 meses+118 %
- Reinforcement Learning in Robotics2519 artículos / 12 meses+117 %
- Explainable Artificial Intelligence (XAI)2319 artículos / 12 meses+200 %
- Domain Adaptation and Few-Shot Learning2059 artículos / 12 meses+67 %
- Advanced Graph Neural Networks1926 artículos / 12 meses+38 %
- Natural Language Processing Techniques1595 artículos / 12 meses+69 %
