Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2.785 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Training One Model to Master Cross-Level Agentic Actions via Reinforcement Learning
Kaichen He, Zihao Wang, Muyao Li, Anji Liu, Yitao Liang · 11. Dezember 2025
The paradigm of agentic AI is shifting from engineered complex workflows to post-training native models. However, existing agents are typically confined to static, predefined action spaces--such as exclusively using APIs, GUI events, or robotic commands. This rigidity limits their adaptability in dy…
- Addressing the Plasticity-Stability Dilemma in Reinforcement Learning
Mansi Maheshwari, John C. Raisbeck, Bruno Castro da Silva · 11. Dezember 2025
Neural networks have shown remarkable success in supervised learning when trained on a single task using a fixed dataset. However, when neural networks are trained on a reinforcement learning task, their ability to continue learning from new experiences declines over time. This decline in learning a…
- Efficient $Q$-Learning and Actor-Critic Methods for Robust Average Reward Reinforcement Learning
Yang Xu, Swetha Ganesh, Vaneet Aggarwal · 11. Dezember 2025
We present a non-asymptotic convergence analysis of $Q$-learning and actor-critic algorithms for robust average-reward Markov Decision Processes (MDPs) under contamination, total-variation (TV) distance, and Wasserstein uncertainty sets. A key ingredient of our analysis is showing that the optimal r…
- Goal inference with Rao-Blackwellized Particle Filters
Yixuan Wang, Dan P. Guralnik, Warren E. Dixon · 11. Dezember 2025
Inferring the eventual goal of a mobile agent from noisy observations of its trajectory is a fundamental estimation problem. We initiate the study of such intent inference using a variant of a Rao-Blackwellized Particle Filter (RBPF), subject to the assumption that the agent's intent manifests throu…
- Knowledge Diversion for Efficient Morphology Control and Policy Transfer
Fu Feng, Ruixiao Shi, Yucheng Xie, Jianlu Shen, Jing Wang, Xin Geng · 11. Dezember 2025
Universal morphology control aims to learn a universal policy that generalizes across heterogeneous agent morphologies, with Transformer-based controllers emerging as a popular choice. However, such architectures incur substantial computational costs, resulting in high deployment overhead, and exist…
- Heuristics for Combinatorial Optimization via Value-based Reinforcement Learning: A Unified Framework and Analysis
Orit Davidovich, Shimrit Shtern, Segev Wasserkrug, Nimrod Megiddo · 10. Dezember 2025
Since the 1990s, considerable empirical work has been carried out to train statistical models, such as neural networks (NNs), as learned heuristics for combinatorial optimization (CO) problems. When successful, such an approach eliminates the need for experts to design heuristics per problem type. D…
- Scalable Offline Model-Based RL with Action Chunks
Kwanyoung Park, Seohong Park, Youngwoon Lee, Sergey Levine · 10. Dezember 2025
In this paper, we study whether model-based reinforcement learning (RL), in particular model-based value expansion, can provide a scalable recipe for tackling complex, long-horizon tasks in offline RL. Model-based value expansion fits an on-policy value function using length-n imaginary rollouts gen…
- Test-driven Reinforcement Learning in Continuous Control
Zhao Yu, Xiuping Wu, Liangjun Ke · 10. Dezember 2025
Reinforcement learning (RL) has been recognized as a powerful tool for robot control tasks. RL typically employs reward functions to define task objectives and guide agent learning. However, since the reward function serves the dual purpose of defining the optimal goal and guiding learning, it is ch…
- An Introduction to Deep Reinforcement and Imitation Learning
Pedro Santana · 10. Dezember 2025
Embodied agents, such as robots and virtual characters, must continuously select actions to execute tasks effectively, solving complex sequential decision-making problems. Given the difficulty of designing such controllers manually, learning-based approaches have emerged as promising alternatives, m…
- Benchmarking Offline Multi-Objective Reinforcement Learning in Critical Care
Aryaman Bansal, Divya Sharma · 10. Dezember 2025
In critical care settings such as the Intensive Care Unit, clinicians face the complex challenge of balancing conflicting objectives, primarily maximizing patient survival while minimizing resource utilization (e.g., length of stay). Single-objective Reinforcement Learning approaches typically addre…
- Reinforcement Learning From State and Temporal Differences
Lex Weaver, Jonathan Baxter · 10. Dezember 2025
TD($\lambda$) with function approximation has proved empirically successful for some complex reinforcement learning problems. For linear approximation, TD($\lambda$) has been shown to minimise the squared error between the approximate value of each state and the true value. However, as far as policy…
- MARL Warehouse Robots
Price Allman, Lian Thang, Dre Simmons, Salmon Riaz · 10. Dezember 2025
We present a comparative study of multi-agent reinforcement learning (MARL) algorithms for cooperative warehouse robotics. We evaluate QMIX and IPPO on the Robotic Warehouse (RWARE) environment and a custom Unity 3D simulation. Our experiments reveal that QMIX's value decomposition significantly out…
- Safety-Aware Reinforcement Learning for Control via Risk-Sensitive Action-Value Iteration and Quantile Regression
Clinton Enwerem, Aniruddh G. Puranic, John S. Baras, Calin Belta · 9. Dezember 2025
Mainstream approximate action-value iteration reinforcement learning (RL) algorithms suffer from overestimation bias, leading to suboptimal policies in high-variance stochastic environments. Quantile-based action-value iteration methods reduce this bias by learning a distribution of the expected cos…
- Learning Without Time-Based Embodiment Resets in Soft-Actor Critic
Homayoon Farrahi, A. Rupam Mahmood · 9. Dezember 2025
When creating new reinforcement learning tasks, practitioners often accelerate the learning process by incorporating into the task several accessory components, such as breaking the environment interaction into independent episodes and frequently resetting the environment. Although they can enable t…
- Why Goal-Conditioned Reinforcement Learning Works: Relation to Dual Control
Nathan P. Lawrence, Ali Mesbah · 9. Dezember 2025
Goal-conditioned reinforcement learning (RL) concerns the problem of training an agent to maximize the probability of reaching target goal states. This paper presents an analysis of the goal-conditioned setting based on optimal control. In particular, we derive an optimality gap between more classic…
- Data-regularized Reinforcement Learning for Diffusion Models at Scale
Haotian Ye, Kaiwen Zheng, Jiashu Xu, Puheng Li, Huayu Chen, Jiaqi Han, Sheng Liu, Qinsheng Zhang, Hanzi Mao, Zekun Hao, Prithvijit Chattopadhyay, Dinghao Yang, Liang Feng, Maosheng Liao, Junjie Bai, Ming-Yu Liu, James Zou, Stefano Ermon · 9. Dezember 2025
Aligning generative diffusion models with human preferences via reinforcement learning (RL) is critical yet challenging. Most existing algorithms are often vulnerable to reward hacking, such as quality degradation, over-stylization, or reduced diversity. Our analysis demonstrates that this can be at…
- Small-Gain Nash: Certified Contraction to Nash Equilibria in Differentiable Games
Vedansh Sharma · 9. Dezember 2025
Classical convergence guarantees for gradient-based learning in games require the pseudo-gradient to be (strongly) monotone in Euclidean geometry as shown by rosen(1965), a condition that often fails even in simple games with strong cross-player couplings. We introduce Small-Gain Nash (SGN), a block…
- Learning-Augmented Ski Rental with Discrete Distributions: A Bayesian Approach
Bosun Kang, Hyejun Park, Chenglin Fan · 9. Dezember 2025
We revisit the classic ski rental problem through the lens of Bayesian decision-making and machine-learned predictions. While traditional algorithms minimize worst-case cost without assumptions, and recent learning-augmented approaches leverage noisy forecasts with robustness guarantees, our work un…
- Model-Based Reinforcement Learning Under Confounding
Nishanth Venkatesh, Andreas A. Malikopoulos · 9. Dezember 2025
We investigate model-based reinforcement learning in contextual Markov decision processes (C-MDPs) in which the context is unobserved and induces confounding in the offline dataset. In such settings, conventional model-learning methods are fundamentally inconsistent, as the transition and reward mec…
- Statistical analysis of Inverse Entropy-regularized Reinforcement Learning
Denis Belomestny, Alexey Naumov, Sergey Samsonov · 9. Dezember 2025
Inverse reinforcement learning aims to infer the reward function that explains expert behavior observed through trajectories of state--action pairs. A long-standing difficulty in classical IRL is the non-uniqueness of the recovered reward: many reward functions can induce the same optimal policy, re…
- Comparative Analysis and Parametric Tuning of PPO, GRPO, and DAPO for LLM Reasoning Enhancement
Yongsheng Lian · 9. Dezember 2025
This study presents a systematic comparison of three Reinforcement Learning (RL) algorithms (PPO, GRPO, and DAPO) for improving complex reasoning in large language models (LLMs). Our main contribution is a controlled transfer-learning evaluation: models are first fine-tuned on the specialized Countd…
- A-3PO: Accelerating Asynchronous LLM Training with Staleness-aware Proximal Policy Approximation
Xiaocan Li, Shiliang Wu, Zheng Shen · 9. Dezember 2025
Decoupled loss has been a successful reinforcement learning (RL) algorithm to deal with the high data staleness under the asynchronous RL setting. Decoupled loss improves coupled-loss style of algorithms' (e.g., PPO, GRPO) learning stability by introducing a proximal policy to decouple the off-polic…
- The Agent Capability Problem: Predicting Solvability Through Information-Theoretic Bounds
Shahar Lutati · 9. Dezember 2025
When should an autonomous agent commit resources to a task? We introduce the Agent Capability Problem (ACP), a framework for predicting whether an agent can solve a problem under resource constraints. Rather than relying on empirical heuristics, ACP frames problem-solving as information acquisition:…
- Delay-Aware Diffusion Policy: Bridging the Observation-Execution Gap in Dynamic Tasks
Aileen Liao, Dong-Ki Kim, Max Olan Smith, Ali-akbar Agha-mohammadi, Shayegan Omidshafiei · 9. Dezember 2025
As a robot senses and selects actions, the world keeps changing. This inference delay creates a gap of tens to hundreds of milliseconds between the observed state and the state at execution. In this work, we take the natural generalization from zero delay to measured delay during training and infere…
- Pretraining in Actor-Critic Reinforcement Learning for Robot Locomotion
Jiale Fan, Andrei Cramariuc, Tifanny Portela, Marco Hutter · 9. Dezember 2025
The pretraining-finetuning paradigm has facilitated numerous transformative advancements in artificial intelligence research in recent years. However, in the domain of reinforcement learning (RL) for robot locomotion, individual skills are often learned from scratch despite the high likelihood that …
