Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2.785 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- SB-TRPO: Towards Safe Reinforcement Learning with Hard Constraints
Ankit Kanwar, Dominik Wagner, Luke Ong · 2. Februar 2026
In safety-critical domains, reinforcement learning (RL) agents must often satisfy strict, zero-cost safety constraints while accomplishing tasks. Existing model-free methods frequently either fail to achieve near-zero safety violations or become overly conservative. We introduce Safety-Biased Trust …
- Learning Policy Representations for Steerable Behavior Synthesis
Beiming Li, Sergio Rozada, Alejandro Ribeiro · 2. Februar 2026
Given a Markov decision process (MDP), we seek to learn representations for a range of policies to facilitate behavior steering at test time. As policies of an MDP are uniquely determined by their occupancy measures, we propose modeling policy representations as expectations of state-action feature …
- Rethinking the Sampling Criteria in Reinforcement Learning for LLM Reasoning: A Competence-Difficulty Alignment Perspective
Deyang Kong, Qi Guo, Xiangyu Xi, Wei Wang, Jingang Wang, Xunliang Cai, Shikun Zhang, Wei Ye · 2. Februar 2026
Reinforcement learning exhibits potential in enhancing the reasoning abilities of large language models, yet it is hard to scale for the low sample efficiency during the rollout phase. Existing methods attempt to improve efficiency by scheduling problems based on problem difficulties. However, these…
- SHAP-Guided Kernel Actor-Critic for Explainable Reinforcement Learning
Na Li, Hangguan Shan, Wei Ni, Wenjie Zhang, Xinyu Li · 2. Februar 2026
Actor-critic (AC) methods are a cornerstone of reinforcement learning (RL) but offer limited interpretability. Current explainable RL methods seldom use state attributions to assist training. Rather, they treat all state features equally, thereby neglecting the heterogeneous impacts of individual st…
- Unsupervised Hierarchical Skill Discovery
Damion Harvey (University of the Witwatersrand, Johannesburg, South Africa), Geraud Nangue Tasse (University of the Witwatersrand, Johannesburg, South Africa, Machine Intelligence and Neural Discovery), Branden Ingram (University of the Witwatersrand, Johannesburg, South Africa, Machine Intelligence and Neural Discovery), Benjamin Rosman (University of the Witwatersrand, Johannesburg, South Africa, Machine Intelligence and Neural Discovery), Steven James (University of the Witwatersrand, Johannesburg, South Africa, Machine Intelligence and Neural Discovery) · 2. Februar 2026
We consider the problem of unsupervised skill segmentation and hierarchical structure discovery in reinforcement learning. While recent approaches have sought to segment trajectories into reusable skills or options, most rely on action labels, rewards, or handcrafted annotations, limiting their appl…
- Strongly Polynomial Time Complexity of Policy Iteration for $L_\infty$ Robust MDPs
Ali Asadi, Krishnendu Chatterjee, Ehsan Goharshady, Mehrdad Karrabi, Alipasha Montaseri, Carlo Pagano · 2. Februar 2026
Markov decision processes (MDPs) are a fundamental model in sequential decision making. Robust MDPs (RMDPs) extend this framework by allowing uncertainty in transition probabilities and optimizing against the worst-case realization of that uncertainty. In particular, $(s, a)$-rectangular RMDPs with …
- Stabilizing the Q-Gradient Field for Policy Smoothness in Actor-Critic
Jeong Woon Lee, Kyoleen Kwak, Daeho Kim, Hyoseok Hwang · 2. Februar 2026
Policies learned via continuous actor-critic methods often exhibit erratic, high-frequency oscillations, making them unsuitable for physical deployment. Current approaches attempt to enforce smoothness by directly regularizing the policy's output. We argue that this approach treats the symptom rathe…
- Agile Reinforcement Learning through Separable Neural Architecture
Rajib Mostakim, Reza T. Batley, Sourav Saha · 2. Februar 2026
Deep reinforcement learning (RL) is increasingly deployed in resource-constrained environments, yet the go-to function approximators - multilayer perceptrons (MLPs) - are often parameter-inefficient due to an imperfect inductive bias for the smooth structure of many value functions. This mismatch ca…
- From Absolute to Relative: Rethinking Reward Shaping in Group-Based Reinforcement Learning
Wenzhe Niu, Wei He, Zongxia Xie, Jinpeng Ou, Huichuan Fan, Yuchen Ge, Yanru Sun, Ziyin Wang, Yizhao Sun, Chengshun Shi, Jiuchong Gao, Jinghua Hao, Renqing He · 2. Februar 2026
Reinforcement learning has become a cornerstone for enhancing the reasoning capabilities of Large Language Models, where group-based approaches such as GRPO have emerged as efficient paradigms that optimize policies by leveraging intra-group performance differences. However, these methods typically …
- Learning Reward Functions for Cooperative Resilience in Multi-Agent Systems
Manuela Chacon-Chamorro, Luis Felipe Giraldo, Nicanor Quijano · 2. Februar 2026
Multi-agent systems often operate in dynamic and uncertain environments, where agents must not only pursue individual goals but also safeguard collective functionality. This challenge is especially acute in mixed-motive multi-agent systems. This work focuses on cooperative resilience, the ability of…
- Action-Sufficient Goal Representations
Jinu Hyeon, Woobin Park, Hongjoon Ahn, Taesup Moon · 2. Februar 2026
Hierarchical policies in offline goal-conditioned reinforcement learning (GCRL) addresses long-horizon tasks by decomposing control into high-level subgoal planning and low-level action execution. A critical design choice in such architectures is the goal representation-the compressed encoding of go…
- Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
Youngjoong Kim, Duhoe Kim, Woosung Kim, Jaesik Park · 2. Februar 2026
Consistency models have been proposed for fast generative modeling, achieving results competitive with diffusion and flow models. However, these methods exhibit inherent instability and limited reproducibility when training from scratch, motivating subsequent work to explain and stabilize these issu…
- Scaling Multiagent Systems with Process Rewards
Ed Li, Junyu Ren, Cat Yan · 2. Februar 2026
While multiagent systems have shown promise for tackling complex tasks via specialization, finetuning multiple agents simultaneously faces two key challenges: (1) credit assignment across agents, and (2) sample efficiency of expensive multiagent rollouts. In this work, we propose finetuning multiage…
- PPO in the Fisher-Rao geometry
Razvan-Andrei Lascu, David \v{S}i\v{s}ka, {\L}ukasz Szpruch · 2. Februar 2026
Proximal Policy Optimization (PPO) is widely used in reinforcement learning due to its strong empirical performance, yet it lacks formal guarantees for policy improvement and convergence. PPO's clipped surrogate objective is motivated by a lower bound on linearization of the value function in flat g…
- MC-GRPO: Median-Centered Group Relative Policy Optimization for Small-Rollout Reinforcement Learning
Youngeun Kim · 2. Februar 2026
Group-relative policy optimization methods train language models by generating multiple rollouts per prompt and normalizing rewards with a shared mean reward baseline. In resource-constrained settings where the rollout budget is small, accuracy often degrades. We find that noise in the shared baseli…
- Offline Reinforcement Learning of High-Quality Behaviors Under Robust Style Alignment
Mathieu Petitbois, R\'emy Portelas, Sylvain Lamprier · 2. Februar 2026
We study offline reinforcement learning of style-conditioned policies using explicit style supervision via subtrajectory labeling functions. In this setting, aligning style with high task performance is particularly challenging due to distribution shift and inherent conflicts between style and rewar…
- Automatic Constraint Policy Optimization based on Continuous Constraint Interpolation Framework for Offline Reinforcement Learning
Xinchen Han, Qiuyang Fang, Hossam Afifi, Michel Marot · 2. Februar 2026
Offline Reinforcement Learning (RL) relies on policy constraints to mitigate extrapolation error, where both the constraint form and constraint strength critically shape performance. However, most existing methods commit to a single constraint family: weighted behavior cloning, density regularizatio…
- Offline Goal-Conditioned Reinforcement Learning with Projective Quasimetric Planning
Anthony Kobanda, Waris Radji, Mathieu Petitbois, Odalric-Ambrym Maillard, R\'emy Portelas · 2. Februar 2026
Offline Goal-Conditioned Reinforcement Learning seeks to train agents to reach specified goals from previously collected trajectories. Scaling that promises to long-horizon tasks remains challenging, notably due to compounding value-estimation errors. Principled geometric offers a potential solution…
- Latent Spherical Flow Policy for Reinforcement Learning with Combinatorial Actions
Lingkai Kong, Anagha Satish, Hezi Jiang, Akseli Kangaslahti, Andrew Ma, Wenbo Chen, Mingxiao Song, Lily Xu, Milind Tambe · 2. Februar 2026
Reinforcement learning (RL) with combinatorial action spaces remains challenging because feasible action sets are exponentially large and governed by complex feasibility constraints, making direct policy parameterization impractical. Existing approaches embed task-specific value functions into const…
- PlatoLTL: Learning to Generalize Across Symbols in LTL Instructions for Multi-Task RL
Jacques Cloete, Mathias Jackermeier, Ioannis Havoutis, Alessandro Abate · 2. Februar 2026
A central challenge in multi-task reinforcement learning (RL) is to train generalist policies capable of performing tasks not seen during training. To facilitate such generalization, linear temporal logic (LTL) has recently emerged as a powerful formalism for specifying structured, temporally extend…
- RN-D: Discretized Categorical Actors with Regularized Networks for On-Policy Reinforcement Learning
Yuexin Bian, Jie Feng, Tao Wang, Yijiang Li, Sicun Gao, Yuanyuan Shi · 2. Februar 2026
On-policy deep reinforcement learning remains a dominant paradigm for continuous control, yet standard implementations rely on Gaussian actors and relatively shallow MLP policies, often leading to brittle optimization when gradients are noisy and policy updates must be conservative. In this paper, w…
- Why GRPO Needs Normalization: A Local-Curvature Perspective on Adaptive Gradients
Cheng Ge, Caitlyn Heqi Yin, Hao Liang, Jiawei Zhang · 2. Februar 2026
Reinforcement learning (RL) has become a key driver of language model reasoning. Among RL algorithms, Group Relative Policy Optimization (GRPO) is the de facto standard, avoiding the need for a critic by using per-prompt baselines and variance normalization. Yet why and when this normalization helps…
- Learning Decentralized LLM Collaboration with Multi-Agent Actor Critic
Shuo Liu, Tianle Chen, Ryan Amiri, Christopher Amato · 30. Januar 2026
Recent work has explored optimizing LLM collaboration through Multi-Agent Reinforcement Learning (MARL). However, most MARL fine-tuning approaches rely on predefined execution protocols, which often require centralized execution. Decentralized LLM collaboration is more appealing in practice, as agen…
- Exploring Reasoning Reward Model for Agents
Kaixuan Fan, Kaituo Feng, Manyuan Zhang, Tianshuo Peng, Zhixun Li, Yilei Jiang, Shuang Chen, Peng Pei, Xunliang Cai, Xiangyu Yue · 30. Januar 2026
Agentic Reinforcement Learning (Agentic RL) has achieved notable success in enabling agents to perform complex reasoning and tool use. However, most methods still relies on sparse outcome-based reward for training. Such feedback fails to differentiate intermediate reasoning quality, leading to subop…
- From Meta-Thought to Execution: Cognitively Aligned Post-Training for Generalizable and Reliable LLM Reasoning
Shaojie Wang, Liang Zhang · 30. Januar 2026
Current LLM post-training methods optimize complete reasoning trajectories through Supervised Fine-Tuning (SFT) followed by outcome-based Reinforcement Learning (RL). While effective, a closer examination reveals a fundamental gap: this approach does not align with how humans actually solve problems…
