Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2,776 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- Hierarchical Reinforcement Learning with Optimal Level Synchronization Based on Flow-Based Deep Generative Model
JaeYoon Kim, Junyu Xuan, Christy Liang, Farookh Hussain · 28 July 2026
High-dimensional state and action spaces combined with sparse reward structures in reinforcement learning (RL) environments typically require advanced control architectures. Hierarchical Reinforcement Learning (HRL) demonstrates superior performance compared to atomic RL approaches in these challeng…
- Beyond Squared Error: Exploring Loss Design for Enhanced Training of Generative Flow Networks
Rui Hu, Yifan Zhang, Zhuoran Li, Longbo Huang · 28 July 2026
Generative Flow Networks (GFlowNets) are a novel class of generative models designed to sample from unnormalized distributions and have found applications in various important tasks, attracting great research interest in their training algorithms. In general, GFlowNets are trained by fitting the for…
- On the Impossibility of Unbiased and Length-Invariant Policy Optimization with Outcome Rewards
Fei Ding, Yongkang Zhang, Yuhao Liao, Zijian Zeng, Huiming Yang · 28 July 2026
Group Relative Policy Optimization (GRPO) is the dominant reinforcement learning algorithm for training reasoning capabilities in large language models, notably adopted by DeepSeek-R1. The recent improvement Dr. GRPO (COLM 2025) identifies the response-level length bias caused by per-trajectory leng…
- Hierarchical Soft Actor-Critic for Sparse-Reward Long-Horizon Reinforcement Learning
Zahra Abdalla Elashaal, Afef Hfaiedh, Nahla Khraief, Issmail Ellabib, Giansalvo Cirrincione · 28 July 2026
Exploration in sparse-reward long-horizon tasks poses significant challenges for reinforcement learning. To address these challenges, we propose a two-level Hierarchical Reinforcement Learning (HRL) framework. The first level handles high-level strategic planning, while the low-level uses the contin…
- Online Policy Evaluation for MDPs with Dynamic UBSR Measures
Weikai Wang, Erick Delage · 28 July 2026
Developing efficient function-approximation methods for policy evaluation is a fundamental challenge in risk-aware reinforcement learning. Existing approaches either focus on restrictive classes of risk measures or rely on access to a simulator, limiting their applicability in fully online settings.…
- Cortex: Compact Behavior Cloning for Quake with Frozen Visual Features
Dzmitry Malyshau · 28 July 2026
We study how far a deliberately simple behavioral-cloning policy can progress in a visually rich first-person game before adding reinforcement learning or explicit memory. Cortex is a compact Quake policy with 10.98 million trainable parameters in a six-layer transformer over a frozen DINOv3 encoder…
- Optimal Reward Shaping: Autonomous Car Parking Case Study
Emre \"Ozkaya, Nicolas R. Gauger · 28 July 2026
Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping fr…
- Constrained Reinforcement Learning Using Successor Representations
Michael Girstl, Alexander Mattick, Christopher Mutschler · 28 July 2026
Real-world Reinforcement Learning depends on the ability to formulate safety constraints into a policy. A common way to model such constraints is to introduce an additional cost signal in the Markov Decision Process, which notifies the agent of unwanted behavior independently of the reward signal. U…
- Finite-Time Analysis of the Natural Policy Gradient in Finite-Horizon Markov Decision Processes
Asha Barua, Sajad Khodadadian · 28 July 2026
Natural Policy Gradient (NPG) is a well-established Reinforcement Learning algorithm that underlies widely used methods such as Trust Region Policy Optimization and Proximal Policy Optimization, both of which have demonstrated strong empirical success. In this paper, we study exact NPG in finite-hor…
- Smooth Learning with Hard Constraints via Legendre-Regularized Policies
Zikun Lin, Rui Chen, Yijie Wang · 28 July 2026
We revisit contextual optimization from the perspective of policy class design. A desirable policy class should be expressive enough to learn rich context-decision relationships, should enforce hard feasibility constraints rather than soft penalty terms, and should remain smooth enough for gradient-…
- Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation
Bingnan Li, Haozhe Wang, Haozhong Xiong, Fangtai Wu, Jinpeng Yu, Yang Shi, Jiaming Liu, Ruihua Huang · 28 July 2026
On-policy distillation (OPD) adapts diffusion models by querying a teacher along trajectories generated by the current student, but how it should behave under classifier-free guidance (CFG), a default component of modern diffusion systems, remains poorly understood. Existing OPD methods naturally ex…
- Sparse Gaussian-Mixture-Model Q-Functions via Hadamard Overparametrization for Online Reinforcement Learning
Minh Vu, Konstantinos Slavakis · 28 July 2026
This paper develops an online, off-policy policy-iteration framework for reinforcement learning (RL), based on sparse Gaussian-mixture-model Q-functions (S-GMM-QFs). The framework reconciles streaming, non-stationary data with the Riemannian structure of the parameter space while handling distributi…
- Learning as Reasoning Unfolds: Progressive Rollout Allocation for Efficient Reinforcement Learning
Heyang Jiang, Henry Liu, Baharan Mirzasoleiman · 27 July 2026
Reinforcement learning with verifiable rewards (RLVR) has emerged as a highly effective framework for improving LLM reasoning, with methods such as GRPO among its most successful instantiations. However, GRPO relies on repeated generation of long chain-of-thought rollouts. Training time scales with …
- When Is a Learned Command Adapter Worth It? Closed-Loop Identification and Counterfactual Auditing of Frozen Locomotion Policies
Zongtan Li · 27 July 2026
Adding a learned adapter to a frozen, command-conditioned locomotion policy is worthwhile only if the interface exposes improvements that are both real and recoverable from deployment-time observations. We introduce an adapter necessity audit that separates global operating-point gain,same-state cou…
- Learning on the Job: Continual Learning from Deployment Feedback for Frozen-Weights Agents
Valentin Tablan, Scott Taylor, Kristoffer Bernhem · 27 July 2026
AI agents encounter learning opportunities in every episode they run, and discard nearly all of them: the underlying models are frozen at deployment, so an agent that resolves a difficult request today starts from zero when it recurs tomorrow. Yet ordinary operation already produces feedback, in the…
- Self-Guided Process Reward Optimization with Redefined Step-wise Advantage for Process Reinforcement Learning
Wu Fei, Shuxian Liang, Yibo Yang, Yang Lin, Jing Tang, Lei Chen, Xiansheng Hua, Hao Kong · 27 July 2026
Process Reinforcement Learning~(PRL) has demonstrated considerable potential in enhancing the reasoning capabilities of Large Language Models~(LLMs). However, introducing additional process reward models incurs substantial computational overhead, and there is no unified theoretical framework for pro…
- Do Modules Stay in Their Lane? Role Drift in Compound LLM Systems
Xiaoyang Cao, Siddarth Srinivasan, Michiel A. Bakker · 27 July 2026
End-to-end reinforcement learning can improve the accuracy of compound LLM systems, but it does not constrain how modules divide labor internally. We identify Role Drift, a failure mode in which modules preserve or improve end-task performance while deviating from their assigned roles through role-v…
- Deconstructing Off-Policy Ratios: Entropy-Scaled Trust Regions for Asynchronous Reinforcement Learning
Guanqun Zhao, Zijun Xie, Binbin Zheng, Enlei Gong, Jiafeng Lu, Yehan Yang, Aoqi Hu, Zeyu Chen · 27 July 2026
Asynchronous reinforcement learning (RL) accelerates large language model (LLM) post-training by overlapping rollout generation with policy optimization, but the resulting stale, off-policy data can destabilize optimization and ultimately cause policy collapse. Existing methods typically retain or d…
- Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning
Jian Hu, Huiying Li, Hao Zhang, Binfeng Xu, Yifan Zhang, Shaokun Zhang, Hemil Desai, Michael Demoret, Pavlo Molchanov, Jan Kautz, Yi Dong · 27 July 2026
Agentic reinforcement learning research is constant algorithm modification, new estimators, new pipeline stages, new rollout schemes, and in mainstream frameworks each change threads through layers of trainer, distributed backend, and rollout glue: the cost lands on the researcher at every iteration…
- Variance-Reduced Q-Learning over Static and Time-Varying Networks
Sreejeet Maity, Feng Zhu, Aritra Mitra, Robert W. Heath Jr · 27 July 2026
We investigate a decentralized reinforcement learning problem involving multiple agents that interact with the same Markov Decision Process (MDP). The agents can exchange information over a network to collectively learn the optimal state-action value function. For this setting, we introduce a novel …
- Discrete Action Space as a Prerequisite for GRPO Convergence in Small-Model Continuous Control
Dmytro Filatov, Valentyn Fedorov, Vira Filatova · 27 July 2026
We study whether Group Relative Policy Optimization (GRPO) can fine-tune small language models for simulated quadrotor continuous-control tasks. In our benchmark, vanilla GRPO fine-tuning of Qwen-0.5B for 25 Hz quadrotor velocity control collapses to the trivial zero action: 0 percent success rate, …
- Adjustment Speed as a Safety Constraint for Nonstationary Reinforcement Learning
Timothy Tomashevskiy · 27 July 2026
Ensuring safety in reinforcement learning under nonstationarity requires determining whether a learning system can safely adapt to forecasted environmental change within the required recovery horizon. Existing safe reinforcement learning methods typically assume stationary environments and do not ex…
- Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning
Julian G. Soltes · 27 July 2026
This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and aggregate an optimal prior from a baseline set of tasks. The QMC meta-prior…
- On the Identifiability of Controlled World Models
Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li · 27 July 2026
Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control. Joint-Embedding Predictive Architectures (JEPAs) provide a compelling framework for learning such models in representation space. …
- QLPO: Quadrant-weighted Sampling for Length-aware Policy Optimization
Siwei Chen, Siqi Chen, Xupeng Miao, Bin Cui · 27 July 2026
Recent large reasoning models often develop long chain-of-thought responses during reinforcement learning (RL), resulting in high inference latency and deployment cost. Existing methods for response length control typically rely on explicit length penalties or additional control modules, which requi…
