Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2.776 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Beyond Trajectory Imitation: Strategy-Guided Policy Optimization for LLM Reasoning
Tianyuan Shi, Canbin Huang, Bei Li, Xin Chen, Xiaojun Quan, Jingang Wang, Qifan Wang · 24. Juni 2026
Distilling reasoning capabilities from strong to weak language models typically involves imitating specific solution trajectories, effectively transferring what to answer rather than how to reason. This trajectory-level imitation encourages memorization of instance-specific steps rather than acquisi…
- ASALT: Adaptive State Alignment for Lateral Transfer in Multi-agent Reinforcement Learning
Anurag Akula, Satheesh K. Perepu, Abhishek Sarkar, Kaushik Dey · 24. Juni 2026
Multi-agent reinforcement learning (MARL) addresses the problem of training multiple agents that pursue collaborative, competitive, or mixed objectives. Prior work has investigated transfer learning between source and target domains in MARL; however, the majority of existing approaches impose the co…
- KLip-PPO: A per-sample KL perspective on PPO-Clip
Riccardo Colletti, Robin Holzinger · 24. Juni 2026
Proximal Policy Optimization (PPO) is the standard policy-gradient algorithm for on-policy reinforcement learning. The literature presents it in two forms, a clipped surrogate that bounds the importance ratio between successive policies and a Kullback-Leibler penalty between them. These forms are tr…
- Reinforcement Learning for Computer-Use Agents with Autonomous Evaluation
Marta Sumyk, Oleksandr Kosovan · 24. Juni 2026
Computer-Use Agents (CUAs) execute high-level user goals by perceiving and acting directly within graphical user interfaces. However, reinforcement learning for CUAs remains difficult because open-ended desktop environments rarely provide scalable, machine-readable reward signals: task success is of…
- Safe and Generalizable Hierarchical Multi-Agent RL via Constraint Manifold Control
Zihao Guo, Jianing Zhao, Ling Li, Hao Liang, Giuseppe Loianno, Yali Du · 24. Juni 2026
Multi-agent systems are widely used in safety-critical applications that require coordinated behavior under strict safety constraints. Existing approaches face a fundamental trade-off: learning-based methods achieve strong empirical performance but lack theoretical safety guarantees, while control-t…
- LaGO: Latent Action Guidance for Online Reinforcement Learning
Kuan-Yen Liu, Ren-Jyun Huang, Ti-Rong Wu · 24. Juni 2026
Large language models (LLMs) have shown strong potential for planning and sequential decision-making, but prior work often relies on using them as direct controllers, which requires precise action generation and can be unreliable in practice. This paper proposes Latent Action Guidance for Online Rei…
- An Introduction to Causal Reinforcement Learning
Elias Bareinboim, Junzhe Zhang, Sanghack Lee · 24. Juni 2026
Causal inference provides a set of principles and tools that allow one to combine data and knowledge about an environment to reason with questions of counterfactual nature, i.e., what would have happened had reality been different, even when no data of this unrealized reality is currently available.…
- Evolving Programmatic Skill Networks
Haochen Shi, Xingdi Yuan, Bang Liu · 24. Juni 2026
We study continual skill acquisition in open-ended embodied environments where an agent must construct, refine, and reuse an expanding library of executable skills. We introduce the Programmatic Skill Network (PSN), a framework in which skills are executable symbolic programs forming a compositional…
- Reward-free Pretraining for Reinforcement Learning via Occupancy Coverage Maximization
Marco Prattic\`o, Pietro Novelli, Massimiliano Pontil, Carlo Ciliberto · 23. Juni 2026
Sparse rewards pose a central challenge in reinforcement learning, since agents receive no informative signal until they reach their goal. Intrinsic-reward methods address this issue by optimizing non-stationary objectives such as novelty, prediction error, or skill diversity, thereby injecting a su…
- Learning Process Rewards via Success Visitation Matching for Efficient RL
Raymond Tsao, Andrew Wagenmaker, Sergey Levine · 23. Juni 2026
In many modern applications of reinforcement learning (RL), the natural reward for a task of interest is inherently sparse: a reward of 0 is given everywhere except when the task is completed, when a reward of +1 is given. Training a policy to maximize such a sparse reward requires solving a challen…
- Evolutionary Discovery of Developmental Reward Schedules in Deep Reinforcement Learning
Alan Nadelsticher Ruvalcaba · 23. Juni 2026
The temporal structure of reward composition in reinforcement learning (RL) is typically hand-designed and held fixed throughout training, leaving the progression of motivational priorities largely unexplored. In this work, we propose an evolutionary framework for discovering developmental reward sc…
- You've Got a Golden Ticket: Improving Generative Robot Policies With A Single Noise Vector
Omkar Patil, Ondrej Biza, Thomas Weng, Karl Schmeckpeper, Wil Thomason, Xiaohan Zhang, Kausik Sivakumar, Robin Walters, Nakul Gopalan, Sebastian Castro, Stephen Hart, Eric Rosen · 23. Juni 2026
What happens when a pretrained generative robot policy is provided a constant initial noise as input, rather than repeatedly sampling it from a Gaussian? We demonstrate that the performance of a pretrained, frozen diffusion or flow matching policy can be improved with respect to a downstream reward …
- Formalizing Task-Space Complexity for Zero-Shot Generalization
Jung-Hoon Cho, Heling Zhang, Siqi Du, Roy Dong, Cathy Wu · 23. Juni 2026
Policies must operate across diverse conditions, yet a single policy is often conservative while fully adaptive schemes can be complex. We study zero-shot generalization in contextual dynamical systems and introduce a performance-centric, directional task dissimilarity--the signed divergence--that u…
- Horizon Adaptive Offline Policy Learning via Value Stitching
Kexin Zheng, Xianyuan Zhan, Xintao Yan · 23. Juni 2026
Learning accurate value functions plays a decisive role for reinforcement learning (RL) agents to solve long-horizon, complex tasks. Conventional temporal-difference (TD) learning objectives suffer from value-estimation bias that accumulates over the horizon, while extended-horizon modeling methods,…
- A Reward-Petri-Net Interpretation of Temporal Behavior Trees
Till Schmeil, G\"unther Waxenegger-Wilfing, Sebastian Schirmer · 23. Juni 2026
This paper introduces an interpretation of Temporal Behavior Trees (TBTs) as Reward-Petri-Nets (RPNs) for reinforcement learning (RL). Designing reward functions for complex, long-horizon robotic tasks is notoriously difficult, especially when tasks have hierarchical structure and temporal constrain…
- Imagine to Ensure Safety in Hierarchical Reinforcement Learning
Gregory Gorbov, Artem Latyshev, Aleksandr I. Panov · 23. Juni 2026
This work investigates the safe exploration problem in reinforcement learning, where an agent must maximize cumulative performance while simultaneously satisfying safety constraints. This challenge becomes even more pronounced in long-horizon tasks, where existing safe methods face fundamental limit…
- Structural Distinguishability of Static and Adaptive Policy Regimes in Agent-Based Regulatory Simulation
Roberto Garrone · 23. Juni 2026
Agent-based models are widely used to evaluate policy interventions in complex socio-technical systems, yet many policy-oriented ABMs represent regulation as a fixed scenario parameter. This limits their ability to distinguish whether regulatory conclusions depend on agent adaptation, policy adaptat…
- Select-to-Act: Hierarchical Reinforcement Learning via Adaptive Language Guidance
Hanping Zhang, Adam Koziak, Yuhong Guo · 23. Juni 2026
Reinforcement Learning (RL) has been widely applied to sequential decision-making, yet it often suffers from poor sample efficiency due to costly interactions with the environment. A limited line of recent work has started exploring improving RL efficiency by leveraging external knowledge expressed …
- On the Position Bias of On-Policy Distillation
Yan Xie, Sijie Zhu, Tiansheng Wen, Bo Chen, Yifei Wang · 23. Juni 2026
On-Policy Distillation (OPD) improves the learning efficiency of standard reinforcement learning through dense, token-level supervision from teachers. In the standard KL objective of OPD, token-level losses are uniformly averaged, implying equal weights for all tokens. However, we discover that not …
- Beyond Penalizing Mistakes: Stabilizing Efficiency Training in Large Reasoning Models via Adaptive Correct-Only Rewards
Jungseob Lee, Seungyoon Lee, Seongtae Hong, Minhyuk Kim, Chanjun Park, Heuiseok Lim · 23. Juni 2026
Training large language models to reason efficiently is a critical challenge. While integrating length-penalizing rewards into Group Relative Policy Optimization (GRPO) aims to reduce verbosity, it frequently triggers reward collapse, severely degrading reasoning capabilities. Through a systematic e…
- ARCO: Adaptive Rubric with Co-Evolution for Multi-Step LLM-Based Agents
Zihang Tian, Jingsen Zhang, Rui Li, Xiaohe Bo, Yuanzi Li, Xu Chen · 23. Juni 2026
Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods score at the trajectory level and freeze the…
- IRumAI: Reinforcement Learning for Indian Rummy
Vignesh Mohan · 23. Juni 2026
Despite its massive player base and complex hidden-information dynamics, Indian Rummy has received no reinforcement learning attention. Existing agents rely on combinatorial search, which is tactically strong but slow at inference. We present IRumAI, the first RL agent for the domain. IRumAI integra…
- Sim2O: Efficient Offline-to-Online MARL via Joint Action Composition
Bingchang Song, Yiqin Yang · 23. Juni 2026
Offline-to-online adaptation serves as a pivotal paradigm for mitigating the prohibitive cost of online exploration by bootstrapping reinforcement learning from offline datasets. While this paradigm has been extensively studied in single-agent settings, its extension to Multi-Agent Reinforcement Lea…
- PoLAR: Factorizing Extent and Mode in Latent Actions for Robot Policy Learning
Youngjoon Jeong, Jihwan Yu, Minsoo Jo, Junha Chun, Taesup Kim · 23. Juni 2026
Latent action pretraining learns representations of visual change from pairs of observations, but existing methods typically encode each transition as a single unstructured representation that entangles transition extent and transition mode. We introduce Polar Latent Actions with Radial structure (P…
- FAST: A Framework for Aligned Sampling and Training in Parallel Reinforcement Learning for Autonomous Driving
Bonan Wang, Letian Tao, Bin Shuai, Jiaxin Gao, Wenxin Zhao, Wei Xiong, Kehua Sheng, Bo Zhang, Yang Guan, Shengbo Eben Li · 23. Juni 2026
Deep reinforcement learning is pivotal for closed-loop autonomous driving yet remains constrained by severe bottlenecks in sampling efficiency. Standard parallel sampling mitigates this but suffers from the straggler effect, where the premature termination of a single environment necessitates a sync…
