Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2785 artículos indexados
Este asunto y su jerarquía proceden de la clasificación OpenAlex, el catálogo abierto de la investigación científica mundial.
Volumen mensual — últimos 12 meses
Últimos artículos
- Leash: Adaptive Length Penalty and Reward Shaping for Efficient Large Reasoning Model
Yanhao Li, Lu Ma, Jiaran Zhang, Lexiang Tang, Wentao Zhang, Guibo Luo · 29 de diciembre de 2025
Existing approaches typically rely on fixed length penalties, but such penalties are hard to tune and fail to adapt to the evolving reasoning abilities of LLMs, leading to suboptimal trade-offs between accuracy and conciseness. To address this challenge, we propose Leash (adaptive LEngth penAlty and…
- Flexible Multitask Learning with Factorized Diffusion Policy
Chaoqi Liu, Haonan Chen, Sigmund H. H{\o}eg, Shaoxiong Yao, Yunzhu Li, Kris Hauser, Yilun Du · 29 de diciembre de 2025
Multitask learning poses significant challenges due to the highly multimodal and diverse nature of robot action distributions. However, effectively fitting policies to these complex task distributions is often difficult, and existing monolithic models often underfit the action distribution and lack …
- Variance-Aware Prior-Based Tree Policies for Monte Carlo Tree Search
Maximilian Weichart · 29 de diciembre de 2025
Monte Carlo Tree Search (MCTS) has profoundly influenced reinforcement learning (RL) by integrating planning and learning in tasks requiring long-horizon reasoning, exemplified by the AlphaZero family of algorithms. Central to MCTS is the search strategy, governed by a tree policy based on an upper …
- Videos are Sample-Efficient Supervisions: Behavior Cloning from Videos via Latent Representations
Xin Liu, Haoran Li, Dongbin Zhao · 29 de diciembre de 2025
Humans can efficiently extract knowledge and learn skills from the videos within only a few trials and errors. However, it poses a big challenge to replicate this learning process for autonomous agents, due to the complexity of visual input, the absence of action or reward signals, and the limitatio…
- Generative Actor Critic
Aoyang Qin, Deqian Kong, Wei Wang, Ying Nian Wu, Song-Chun Zhu, Sirui Xie · 29 de diciembre de 2025
Conventional Reinforcement Learning (RL) algorithms, typically focused on estimating or maximizing expected returns, face challenges when refining offline pretrained models with online experiences. This paper introduces Generative Actor Critic (GAC), a novel framework that decouples sequential decis…
- Accelerating Diffusion Planners in Offline RL via Reward-Aware Consistency Trajectory Distillation
Xintong Duan, Yutong He, Fahim Tajwar, Ruslan Salakhutdinov, J. Zico Kolter, Jeff Schneider · 29 de diciembre de 2025
Although diffusion models have achieved strong results in decision-making tasks, their slow inference speed remains a key limitation. While consistency models offer a potential solution, existing applications to decision-making either struggle with suboptimal demonstrations under behavior cloning or…
- Towards Optimal Performance and Action Consistency Guarantees in Dec-POMDPs with Inconsistent Beliefs and Limited Communication
Moshe Rafaeli Shimron, Vadim Indelman · 25 de diciembre de 2025
Multi-agent decision-making under uncertainty is fundamental for effective and safe autonomous operation. In many real-world scenarios, each agent maintains its own belief over the environment and must plan actions accordingly. However, most existing approaches assume that all agents have identical …
- Generalised Linear Models in Deep Bayesian RL with Learnable Basis Functions
Jingyang You, Hanna Kurniawati · 25 de diciembre de 2025
Bayesian Reinforcement Learning (BRL) provides a framework for generalisation of Reinforcement Learning (RL) problems from its use of Bayesian task parameters in the transition and reward models. However, classical BRL methods assume known forms of transition and reward models, reducing their applic…
- DATTA: Domain Diversity Aware Test-Time Adaptation for Dynamic Domain Shift Data Streams
Chuyang Ye, Dongyan Wei, Zhendong Liu, Yuanyi Pang, Yixi Lin, Qinting Jiang, Jingyan Jiang, Dongbiao He · 25 de diciembre de 2025
Test-Time Adaptation (TTA) addresses domain shifts between training and testing. However, existing methods assume a homogeneous target domain (e.g., single domain) at any given time. They fail to handle the dynamic nature of real-world data, where single-domain and multiple-domain distributions chan…
- Policy-Conditioned Policies for Multi-Agent Task Solving
Yue Lin, Shuhui Zhu, Wenhao Li, Ang Li, Dan Qiao, Pascal Poupart, Hongyuan Zha, Baoxiang Wang · 25 de diciembre de 2025
In multi-agent tasks, the central challenge lies in the dynamic adaptation of strategies. However, directly conditioning on opponents' strategies is intractable in the prevalent deep reinforcement learning paradigm due to a fundamental ``representational bottleneck'': neural policies are opaque, hig…
- Ensuring Safety in an Uncertain Environment: Constrained MDPs via Stochastic Thresholds
Qian Zuo, Fengxiang He · 25 de diciembre de 2025
This paper studies constrained Markov decision processes (CMDPs) with constraints against stochastic thresholds, aiming at safety of reinforcement learning in unknown and uncertain environments. We leverage a Growing-Window estimator sampling from interactions with the uncertain environment to estim…
- Context-Sensitive Abstractions for Reinforcement Learning with Parameterized Actions
Rashmeet Kaur Nayyar, Naman Shah, Siddharth Srivastava · 25 de diciembre de 2025
Real-world sequential decision-making often involves parameterized action spaces that require both, decisions regarding discrete actions and decisions about continuous action parameters governing how an action is executed. Existing approaches exhibit severe limitations in this setting -- planning me…
- Recurrent Off-Policy Deep Reinforcement Learning Doesn't Have to be Slow
Tyler Clark, Christine Evers, Jonathon Hare · 24 de diciembre de 2025
Recurrent off-policy deep reinforcement learning models achieve state-of-the-art performance but are often sidelined due to their high computational demands. In response, we introduce RISE (Recurrent Integration via Simplified Encodings), a novel approach that can leverage recurrent networks in any …
- An Optimal Policy for Learning Controllable Dynamics by Exploration
Peter N. Loxley · 24 de diciembre de 2025
Controllable Markov chains describe the dynamics of sequential decision making tasks and are the central component in optimal control and reinforcement learning. In this work, we give the general form of an optimal policy for learning controllable dynamics in an unknown environment by exploring over…
- Generalisation in Multitask Fitted Q-Iteration and Offline Q-learning
Kausthubh Manda, Raghuram Bharadwaj Diddigi · 24 de diciembre de 2025
We study offline multitask reinforcement learning in settings where multiple tasks share a low-rank representation of their action-value functions. In this regime, a learner is provided with fixed datasets collected from several related tasks, without access to further online interaction, and seeks …
- Performative Policy Gradient: Optimality in Performative Reinforcement Learning
Debabrota Basu, Udvas Das, Brahim Driss, Uddalak Mukherjee · 24 de diciembre de 2025
Post-deployment machine learning algorithms often influence the environments they act in, and thus shift the underlying dynamics that the standard reinforcement learning (RL) methods ignore. While designing optimal algorithms in this performative setting has recently been studied in supervised learn…
- Emergent temporal abstractions in autoregressive models enable hierarchical reinforcement learning
Seijin Kobayashi, Yanick Schimpf, Maximilian Schlegel, Angelika Steger, Maciej Wolczyk, Johannes von Oswald, Nino Scherre, Kaitlin Maile, Guillaume Lajoie, Blake A. Richards, Rif A. Saurous, James Manyika, Blaise Ag\"uera y Arcas, Alexander Meulemans, Jo\~ao Sacramento · 24 de diciembre de 2025
Large-scale autoregressive models pretrained on next-token prediction and finetuned with reinforcement learning (RL) have achieved unprecedented success on many problem domains. During RL, these models explore by generating new outputs, one token at a time. However, sampling actions token-by-token c…
- Environment Scaling for Interactive Agentic Experience Collection: A Survey
Yuchen Huang, Sijia Li, Minghao Liu, Wei Liu, Shijue Huang, Zhiyuan Fan, Hou Pong Chan, Yi R. Fung · 24 de diciembre de 2025
LLM-based agents can autonomously accomplish complex tasks across various domains. However, to further cultivate capabilities such as adaptive behavior and long-term decision-making, training on static datasets built from human-level knowledge is insufficient. These datasets are costly to construct …
- Sample-Efficient Policy Constraint Offline Deep Reinforcement Learning based on Sample Filtering
Yuanhao Chen, Qi Liu, Pengbin Chen, Zhongjian Qiao, Yanjie Li · 24 de diciembre de 2025
Offline reinforcement learning (RL) aims to learn a policy that maximizes the expected return using a given static dataset of transitions. However, offline RL faces the distribution shift problem. The policy constraint offline RL method is proposed to solve the distribution shift problem. During the…
- Offline Safe Policy Optimization From Heterogeneous Feedback
Ze Gong, Pradeep Varakantham, Akshat Kumar · 24 de diciembre de 2025
Offline Preference-based Reinforcement Learning (PbRL) learns rewards and policies aligned with human preferences without the need for extensive reward engineering and direct interaction with human annotators. However, ensuring safety remains a critical challenge across many domains and tasks. Previ…
- Deformable Cluster Manipulation via Whole-Arm Policy Learning
Jayadeep Jacob, Wenzheng Zhang, Houston Warren, Paulo Borges, Tirthankar Bandyopadhyay, Fabio Ramos · 24 de diciembre de 2025
Manipulating clusters of deformable objects presents a substantial challenge with widespread applicability, but requires contact-rich whole-arm interactions. A potential solution must address the limited capacity for realistic model synthesis, high uncertainty in perception, and the lack of efficien…
- SD2AIL: Adversarial Imitation Learning from Synthetic Demonstrations via Diffusion Models
Pengcheng Li, Qiang Fang, Tong Zhao, Yixing Lan, Xin Xu · 23 de diciembre de 2025
Adversarial Imitation Learning (AIL) is a dominant framework in imitation learning that infers rewards from expert demonstrations to guide policy optimization. Although providing more expert demonstrations typically leads to improved performance and greater stability, collecting such demonstrations …
- Learning General Policies with Policy Gradient Methods
Simon St{\aa}hlberg, Blai Bonet, Hector Geffner · 23 de diciembre de 2025
While reinforcement learning methods have delivered remarkable results in a number of settings, generalization, i.e., the ability to produce policies that generalize in a reliable and systematic way, has remained a challenge. The problem of generalization has been addressed formally in classical pla…
- From Novelty to Imitation: Self-Distilled Rewards for Offline Reinforcement Learning
Gaurav Chaudhary, Laxmidhar Behera · 23 de diciembre de 2025
Offline Reinforcement Learning (RL) aims to learn effective policies from a static dataset without requiring further agent-environment interactions. However, its practical adoption is often hindered by the need for explicit reward annotations, which can be costly to engineer or difficult to obtain r…
- Demonstration-Guided Continual Reinforcement Learning in Dynamic Environments
Xue Yang, Michael Schukat, Junlin Lu, Patrick Mannion, Karl Mason, Enda Howley · 23 de diciembre de 2025
Reinforcement learning (RL) excels in various applications but struggles in dynamic environments where the underlying Markov decision process evolves. Continual reinforcement learning (CRL) enables RL agents to continually learn and adapt to new tasks, but balancing stability (preserving prior knowl…
