Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2.776 indexierte Paper
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen — letzte 12 Monate
Neueste Paper
- Curriculum reinforcement learning with measurable task representation learning
Yongyan Wen, Siyuan Li, Mingjian Fu, Yiqin Yang, Xun Wang, Peng Liu · 25. Mai 2026
In curriculum reinforcement learning (CRL), an agent incrementally accumulates knowledge over a sequence of tasks (i.e., a curriculum), and the learning process is aimed at using the accumulated knowledge to finally solve a challenging target task. While early CRL works focus on sequencing candidate…
- Infra-Bayesian Reinforcement Learning Agents Outperform Classical RL For Worst-Case Robustness
Manish Aryal, Faiyaz Azam, Agnivo Banerjee, Sai Sidhanth Manoharan Jayanthi, Allegra Laro, Cl\'ement Legentilhomme, Andrew Lin, Florian Lorkowski, Radman Rakhshandehroo, Patric Rommel, Emanuel Ruzak, Nathan Theng, Paul Yushin Rapoport · 25. Mai 2026
Classical reinforcement learning assumes the agent interacts with a fixed environment whose behavior does not depend on the agent's policy. This assumption breaks down in non-realizable settings where other actors might anticipate the agent's behavior, including environments crucial to AI safety, wh…
- Goal-Conditioned Agents that Learn Everything All at Once
Michael Matthews, Matthew Jackson, Michael Beukman, Thomas Foster, Alistair Letcher, Scott Fujimoto, C\'edric Colas, Jakob Foerster · 25. Mai 2026
A goal-conditioned reinforcement learning agent exploring an environment will see a wealth of information throughout a trajectory, most of which is discarded when only performing on-policy updates with respect to the commanded goal. All-goals learning, where each transition is used for learning off-…
- OPPO: Bayesian Value Recursion for Token-Level Credit Assignment in LLM Reasoning
Yu Li, Rui Miao, Tian Lan, Zhengling Qi · 25. Mai 2026
Reinforcement learning with verifiable rewards has become the standard recipe for improving LLM reasoning, but the dominant algorithm GRPO assigns a single trajectory-level advantage to every token, diluting the signal at pivotal reasoning steps and injecting noise at uninformative ones. Critic-free…
- Reflex: Reinforcement Learning with Reflection Symmetry Exploitation in State-Based Continuous Control
Shuai Zhen, Yifan Zhang, Yuling Wang, Yanhua Yu · 25. Mai 2026
Reinforcement learning has long struggled with poor sample efficiency. One promising approach to mitigate this problem is leveraging group-invariant Markov Decision Processes ($G$-invariant MDPs). Existing works in this direction have primarily focused on image-based RL and rotational symmetry such …
- Dreaming Smoothly and Sample Efficiently with Gradient Penalized Latent Dynamics
Romil V. Sonigra (Texas A&M University), P. R. Kumar (Texas A&M University) · 25. Mai 2026
Model-based reinforcement learning improves sample efficiency by learning a world model. However, existing latent world models such as DreamerV3 do not explicitly enforce local smoothness in their learned transition dynamics, leaving a useful inductive bias for transition dynamics learning unexploit…
- ARMS: Automatic Reward Shaping for Sparse-Reward Multi-Agent Reinforcement Learning
Elie Abboud, Oren Gal · 25. Mai 2026
Sparse rewards are a major bottleneck in multi-agent reinforcement learning (MARL), where simultaneous learning induces non-stationarity and makes reward design especially delicate. Reward shaping can accelerate learning, but in the multi-agent setting it must preserve the strategic structure of the…
- Learning Kernel-Based MDPs from Episodic Preferential Feedback
Nikola Pavlovic, Sattar Vakili, Qing Zhao · 25. Mai 2026
Human feedback often arrives as preferences rather than calibrated numeric rewards, motivating reinforcement learning from preferential feedback, also referred to as reinforcement learning from human feedback (RLHF). We present a rigorous theoretical study of preference-only learning in episodic ker…
- Computable Fairness: Boltzmann-Softmax Control for AI Resource Allocation
Ji-Won Park, Chae Un Kim · 25. Mai 2026
In large-scale AI systems, allocating scarce resources such as GPU compute time and bandwidth among multiple agents is a critical challenge. Conventional policies focus on efficiency metrics, potentially leading to dominance concentration that undermines system diversity and stability. We propose Co…
- Score-Based One-step MeanFlow Policy Optimization
Kyungyoon Kim, Donghyeon Ki, Hee-Jun Ahn, Byung-Jun Lee · 25. Mai 2026
Diffusion and flow matching have emerged as expressive policy classes in reinforcement learning, but their reliance on multi-step denoising imposes substantial computational overhead at inference time, which is particularly problematic in online RL. MeanFlow offers a promising alternative by learnin…
- Understanding Goal Generalisation in Sequential Reinforcement Learning
Jason Ross Brown, Edward James Young · 25. Mai 2026
Reinforcement learning agents often exhibit unintended goal-directed behaviour outside their training distribution, but we currently lack a principled understanding of how such agents will generalise to novel environments based on their training history. We address this gap for agents trained sequen…
- Principled RL for Flow Matching Emerges from the Chunk-level Policy Optimization
Yifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du, Keyu Fan, Bo Li, Sinan Du, Xu Wan, Zhiyu Chen, Bo Xia, Tiantian Zhang, Yongzhe Chang, Changqian Yu, Kun Gai, Xueqian Wang · 22. Mai 2026
Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive …
- Jointly Learning Predicates and Actions Enables Zero-Shot Skill Composition
Benedict Quartey, Sebastian Castro, Eric Rosen, Wil Thomason, George Konidaris, Stefanie Tellex · 22. Mai 2026
Learning from Demonstration (LfD) enables robots to learn complex behaviors from expert examples, yet existing approaches often fail to generalize to new compositions of known skills without retraining. Modern generative policies model distributions over action trajectories alone, thus are unable to…
- Stochastic MeanFlow Policies: One-Step Generative Control with Entropic Mirror Descent
Zeyuan Wang, Da Li, Yulin Chen, Yuehu Gong, Yanming Guo, Ye Shi, Liang Bai, Tianyuan Yu, Yanwei Fu · 22. Mai 2026
Online off-policy reinforcement learning (RL) is shaped by two coupled choices: the policy class and the update rule. Gaussian policies are fast and have tractable entropy, but struggle with multimodal action distributions. Generative policies are more expressive, but often require iterative samplin…
- Reinforcement Learning with Discrete Diffusion Policies for Combinatorial Action Spaces
Haitong Ma, Ofir Nabati, Aviv Rosenberg, Bo Dai, Oran Lang, Craig Boutilier, Na Li, Shie Mannor, Lior Shani, Guy Tenneholtz · 21. Mai 2026
Reinforcement learning (RL) struggles to scale to large, combinatorial action spaces common in many real-world problems. This paper introduces a novel framework for training discrete diffusion models as highly effective policies in these complex settings. Our key innovation is an efficient online tr…
- AGPO: Adaptive Group Policy Optimization with Dual Statistical Feedback
Miaobo Hu, Shuhao Hu, Bokun Wang, Ruohan Wang, Xin Wang, Xiaobo Guo, Daren Zha, Jun Xiao · 21. Mai 2026
Reinforcement learning improves LLM reasoning, but PPO/GRPO typically use fixed clipping and decoding temperature, which makes training brittle and tuning-heavy. We propose Adaptive Group Policy Optimization (AGPO), a critic-free refinement of GRPO that uses group-level statistics to control both up…
- Decoupling Communication from Policy: Robust MARL under Bandwidth Constraints
Alexi Canesse, Beno\^it Goupil, Jesse Read, Sonia Vanier · 21. Mai 2026
Communication enables coordination in multi-agent reinforcement learning (MARL), but many real-world applications, e.g., search-and-rescue with drone swarms, operate under severe bandwidth constraints. Many communication architectures still expose a coupled bottleneck in which a shared latent repres…
- Complementing reinforcement learning with SFT through logit averaging in the post training of LLMs
Xingwei Gan, Ying Zhu · 21. Mai 2026
We introduce a novel method that averages the logits of a frozen reference policy (e.g., SFT) and a trainable policy, and incorporate the method into Group Relative Policy Optimization (GRPO). In contrast to Reinforcement Learning with Verifiable Rewards (RLVR) methods, our proposal does not involve…
- Distributed Direct Preference Optimization
Zhanhong Jiang · 21. Mai 2026
Preference-based reinforcement learning (RL) is a key paradigm for aligning policies with human judgments, yet its theoretical behavior in distributed settings where preference data are fragmented across heterogeneous users remains poorly understood. Direct Preference Optimization (DPO) avoids expli…
- Beyond the Bellman Recursion: A Pontryagin-Guided Framework for Non-Exponential Discounting
Hojin Ko, Jeonggyu Huh · 21. Mai 2026
Most value-based and actor--critic reinforcement learning methods rely on Bellman-style recursions, yet these recursions collapse under non-exponential discounting common in human preferences and survival processes. We show the breakdown is structural: exponential discounting sits at a fragile inter…
- FBOS-RL: Feedback-Driven Bi-Objective Synergistic Reinforcement Learning
Xikai Zhang, Yongzhi Li, Likang Xiao, Yingze Zhang, Yanhua Cheng, Quan Chen, Peng Jiang, Wenjun Wu, Liu Liu · 21. Mai 2026
Reinforcement learning has become a cornerstone for aligning and unlocking the reasoning capabilities of large-scale models. At its core, the training loop of GRPO and its variants alternates between rollout sampling and policy update. Unlike supervised learning, where each gradient step is anchored…
- Smaller Abstract State Spaces Enable Cross-Scale Generalization in Reinforcement Learning
Nasehatul Mustakim, Lucas Lehnert · 21. Mai 2026
While humans readily generalize abstract concepts to more complex or larger tasks, building Reinforcement Learning (RL) systems with this ability remains elusive. Here, we present the first theoretical model of how such Out-of-Distribution (OOD) generalization can be achieved in RL agents. Our appro…
- CIG: Exploration via Conditional Information Gain
Tim Joseph, Marcus Fechner, Philipp Stegmaier, Karam Daaboul, J. Marius Z\"ollner · 21. Mai 2026
Intrinsic rewards for exploration in reinforcement learning condition on different contexts: lifelong rewards score each transition against accumulated experience but ignore within-rollout redundancy; episodic rewards penalize intra-trajectory repetition but discard lifetime progress. Hybrid methods…
- ZEBRA: Zero-shot Budgeted Resource Allocation for LLM Orchestration
May Hamri, Inbal Talgam-Cohen · 21. Mai 2026
As autonomous agents increasingly execute end-to-end tasks under fixed monetary budgets, the pressing open question shifts from whether the budget is respected, to how to spend it effectively. Existing budget-aware methods typically control reasoning step-by-step within a single agent, or learn reso…
- ReversedQ: Opportunities for Faster Q-Learning in Episodic Online Reinforcement Learning
Sofia R. Miskala-Dinc, Aviva Prins · 21. Mai 2026
We study model-free Q-learning in finite-horizon episodic Markov Decision Processes (MDPs) with stationary dynamics across episodes. We identify a central issue in nascent model-free posterior-sampling works: the reliance on delayed learning in order to prove theoretical guarantees. In particular, w…
