Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2,519 papers indexed
Reinforcement learning applied to robotics explores how autonomous agents acquire complex behaviors by interacting with their environment. Recent work focuses on refining methods such as Q-learning, diffusion models, or policy optimization, addressing challenges like offline estimation, reward distribution, or managing large action spaces. These studies also examine issues of safety, fairness, and dynamic adaptation, particularly in multi-agent settings or when transferring knowledge between tasks.
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Lab countries
- United States42% · 733 papers
- China36% · 637 papers
- United Kingdom6.9% · 121 papers
- Germany6.4% · 112 papers
- Canada6.3% · 110 papers
- South Korea4.3% · 76 papers
- France3.4% · 60 papers
- Hong Kong SAR China3.3% · 58 papers
Across 1,749 papers on this subject with at least one lab located. 73 countries represented.
This is the country of the laboratory, never the nationality of individuals. A paper signed from several countries counts for each of them, so the shares add up to more than 100%. Coverage is partial and the gap is not random: a researcher whose institution is unknown usually publishes little, which over-represents established labs.
Latest papers
- Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu · 2 October 2026
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning prog…
- Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
Yu Mao, Lei Yu, Zining Zhu, Yusheng Zheng, Haohang Li, Freda Shi, Yutong Yin, Zhaoran Wang, Jingcheng Niu · 2 October 2026
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for …
- FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv · 2 October 2026
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potent…
- Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang · 2 October 2026
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. …
- iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel · 2 October 2026
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we pro…
- Iterative Policy Refinement through Semantic Rollout Analysis
Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons · 2 October 2026
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the…
- Optimal Transport Meets Reinforcement Learning: A Survey
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana · 2 October 2026
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, com…
- Does Scaling Reinforcement Learning Really Require More Training?
Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu · 2 October 2026
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible fr…
- SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard · 2 October 2026
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitatio…
- Learning Transferable Skills using Goal-Conditioned Bisimulation
Mohammad Amin Abbasfar, Farbod Azimmohseni, Mohammad Hossein Rohban · 2 October 2026
Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously un…
- Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou · 2 October 2026
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. W…
- ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning
Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev · 2 October 2026
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize …
- T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning
Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo · 2 October 2026
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successf…
- DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye · 2 October 2026
Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task succes…
- Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Beining Wu, Zihao Ding, Jun Huang · 2 October 2026
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a c…
- Q-Learning for Reachability in MEC-Free MDPs
Lu-Chin Chang, Suguman Bansal · 2 October 2026
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decisi…
- Measuring the Stability Assumption Behind Action Chunking
Aryan Goyal · 2 October 2026
Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error o…
- Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
Jude Waide, Robert Lieck · 2 October 2026
Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has propo…
- Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao · 2 October 2026
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a leng…
- Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu · 2 October 2026
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/fail…
- Calibration-risk routing for controlled world-model adaptation
Yifan Zhang, Liang Zheng · 2 October 2026
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which…
- PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu · 2 October 2026
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. I…
- Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado · 2 October 2026
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstracti…
- Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu · 2 October 2026
This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism…
- TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models
Song Jin, Juntian Zhang, Ruyu Lyu, Yong Liu, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, Wei Lin, Rui Yan · 1 October 2026
Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to su…
Other topics in Artificial intelligence
The topics the OpenAlex classification attaches to the same theme, most active first.
- Large Language Models7,407 papers / 12 months+247%
- Adversarial Robustness in Machine Learning3,552 papers / 12 months+118%
- Explainable Artificial Intelligence (XAI)2,319 papers / 12 months+200%
- Domain Adaptation and Few-Shot Learning2,059 papers / 12 months+67%
- Advanced Graph Neural Networks1,926 papers / 12 months+38%
- Natural Language Processing Techniques1,595 papers / 12 months+69%
