Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2.519 indexierte Paper
Verstärkungslernen in der Robotik untersucht, wie autonome Agenten komplexe Verhaltensweisen durch Interaktion mit ihrer Umgebung erwerben. Aktuelle Arbeiten konzentrieren sich darauf, Methoden wie Q-Learning, Diffusionsmodelle oder Policy-Optimierung zu verfeinern, indem Herausforderungen wie Offline-Schätzung, Reward-Verteilung oder die Handhabung großer Aktionsräume angegangen werden. Diese Forschungen betrachten auch Fragen der Sicherheit, Fairness und dynamischen Anpassung, insbesondere in Multi-Agenten-Kontexten oder beim Transfer von Wissen zwischen Aufgaben.
Dieses Unterthema und seine Hierarchie stammen aus der OpenAlex-Klassifikation, dem offenen Katalog der weltweiten wissenschaftlichen Forschung.
Monatliches Volumen - letzte 12 Monate
Länder der Labore
- Vereinigte Staaten42 % · 733 Artikel
- China36 % · 637 Artikel
- Vereinigtes Königreich6,9 % · 121 Artikel
- Deutschland6,4 % · 112 Artikel
- Kanada6,3 % · 110 Artikel
- Südkorea4,3 % · 76 Artikel
- Frankreich3,4 % · 60 Artikel
- Sonderverwaltungsregion Hongkong3,3 % · 58 Artikel
Über 1.749 Artikel zu diesem Thema mit mindestens einem verorteten Labor. 73 Länder vertreten.
Es handelt sich um das Land des Labors, nie um die Staatsangehörigkeit von Personen. Ein Artikel aus mehreren Ländern zählt für jedes davon, die Anteile summieren sich daher auf über 100 %. Die Abdeckung ist unvollständig und die Lücke nicht zufällig: Forschende ohne bekannte Institution publizieren meist wenig, was etablierte Labore überrepräsentiert.
Neueste Paper
- Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu · 2. Oktober 2026
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning prog…
- Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
Yu Mao, Lei Yu, Zining Zhu, Yusheng Zheng, Haohang Li, Freda Shi, Yutong Yin, Zhaoran Wang, Jingcheng Niu · 2. Oktober 2026
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for …
- FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv · 2. Oktober 2026
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potent…
- Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang · 2. Oktober 2026
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. …
- iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel · 2. Oktober 2026
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we pro…
- Iterative Policy Refinement through Semantic Rollout Analysis
Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons · 2. Oktober 2026
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the…
- Optimal Transport Meets Reinforcement Learning: A Survey
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana · 2. Oktober 2026
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, com…
- Does Scaling Reinforcement Learning Really Require More Training?
Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu · 2. Oktober 2026
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible fr…
- SHARPO: Segment-Level Credit Assignment for Agentic Reinforcement Learning
Xinchen Du, Zhengze Zhou, Wenhui Zhu, Han Yu, Sen Na, Rohit Jain, Alborz Geramifard · 2. Oktober 2026
Agentic reinforcement learning (RL) trains a large language model (LLM) to act over long, multi-step interactions. However, a single localized error can cause task failure, while trajectory-level rewards provide limited guidance for assigning credit to individual decisions. To address this limitatio…
- Learning Transferable Skills using Goal-Conditioned Bisimulation
Mohammad Amin Abbasfar, Farbod Azimmohseni, Mohammad Hossein Rohban · 2. Oktober 2026
Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously un…
- Exploring More, Reasoning Better: Stepwise Risk-Sensitive GRPO for Diffusion Language Models
Yue YU, Bowen Zuo, David Crandall, Yinglun Zhu, Dongruo Zhou · 2. Oktober 2026
Diffusion large language models (dLLMs) generate text by denoising a sequence or successive blocks, allowing several tokens to be revealed in parallel. Reinforcement learning with verifiable rewards (RLVR) reuses terminal feedback across these decisions, even as their conditioning context changes. W…
- ALER: Adaptive Learnable Experience Rewriting for Reinforcement Learning
Oleg Shchendrigin, Egor Cherepanov, Aleksandr I. Panov, Alexey K. Kovalev · 2. Oktober 2026
In partially observable reinforcement learning (RL), a later observation can make stored information obsolete or change what it implies for the next decision. Memory architectures and benchmarks for RL mostly test retention, the ability to keep information unchanged until it is needed. We formalize …
- T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning
Bo-Wen Zhang, Junwei He, Maoqi Liu, Feiran Li, Song-Lin Lv, Wentao Ma, Rongyi Lin, Shuhan Zhong, Lan-Zhe Guo · 2. Oktober 2026
Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successf…
- DriftOPD: Sequence-Level Reverse-KL Distillation for One-Step VLA Policies
Youngjun Jun, Kyumin Choi, Youngmin Kim, Seonghyun Jin, Sunwoo Park, Jangho Park, Jong Chul Ye · 2. Oktober 2026
Vision-Language-Action (VLA) models increasingly rely on action experts that generate short action chunks under receding-horizon control. While chunk-level training is convenient across robot embodiments, it optimizes local action likelihood without explicitly accounting for long-horizon task succes…
- Not All Experience Belongs in the Weights: Component Routing for Self-Improving GUI Agents
Beining Wu, Zihao Ding, Jun Huang · 2. Oktober 2026
Self-improving GUI agents keep the trajectories they produce and return them to the agent, by fine-tuning or by retrieval into the prompt, and studies that compare the two destinations disagree. We attribute this to the unit of experience: a trajectory bundles items with different properties, so a c…
- Q-Learning for Reachability in MEC-Free MDPs
Lu-Chin Chang, Suguman Bansal · 2. Oktober 2026
Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decisi…
- Measuring the Stability Assumption Behind Action Chunking
Aryan Goyal · 2. Oktober 2026
Action chunking improves the performance of policies learned by behavioural cloning, and several mechanisms have been proposed to explain why, including temporal consistency, horizon reduction, representation learning, and reduced error compounding. We instead study what happens to an action error o…
- Decision Titan: Test-Time Training for Long-Term Memory in Offline Reinforcement Learning
Jude Waide, Robert Lieck · 2. Oktober 2026
Long-term dependencies remain a major challenge for sequential decision-making in the field of AI: RNNs suffer from vanishing gradients and the limited expressivity of vector-based hidden states, whilst Transformer-based models are limited by the quadratic scaling of attention. Recent work has propo…
- Rethinking Probability-Based Reinforcement Learning From Posterior Concentration
Shiu-Hong Kao, Yubo Zhao, Zhenyu Tian, Pengzhan Sun, Yicong Li, Angela Yao · 2. Oktober 2026
Verifier-free reinforcement learning with probability-based rewards offers a promising way to train LLMs on general reasoning tasks where external verifiers are unavailable. Yet the reliability of these rewards, especially in long-horizon reasoning, remains underexplored. This work identifies a leng…
- Dependency-Aware Reward Shaping for Agentic Reinforcement Learning
Ziyi Chen, Yan Zhang, Jianhui Wei, Daoan Zhang, Zuozhu Liu · 2. Oktober 2026
When training large language models with reinforcement learning, terminal rewards provide little guidance about which steps matter. Common methods for assigning step credit overlook that work built on uncorrected mistakes is wasted while independent work remains valid. With only a final success/fail…
- Calibration-risk routing for controlled world-model adaptation
Yifan Zhang, Liang Zheng · 2. Oktober 2026
Model-based reinforcement learning (MBRL) can exploit simulated experience, but a simulator-to-target shift creates a model-selection problem: correcting the simulator and fitting the target directly can each fail under limited target data. We introduce the Model-Corrected World Model (MC-WM), which…
- PG-SFT: Balancing Capability Acquisition and Retention in Offline Agent Fine-Tuning
Ronghua Li, Zi Liang, Zhishan Li, Shinan Liu · 2. Oktober 2026
Supervised fine-tuning (SFT) on offline agent trajectories is the standard approach for training specialized tool-using agents, but forcing models to imitate reasoning and actions token by token may harm other capabilities (e.g., general reasoning, tool calling, code generation) of the base model. I…
- Learning Multiple Timescales for Goal-Conditioned Reinforcement Learning
Pedro Robles Dutenhefner, Dikshant Shehmar, Wagner Meira Jr., Marlos C. Machado · 2. Oktober 2026
Existing approaches to offline goal-conditioned reinforcement learning (GCRL) struggle with long-horizon tasks. Discounting shrinks value differences between distant states until they fall below the function approximation error, leaving the agent with no signal for ranking states. Temporal abstracti…
- Meta-Multi-Agent Reinforcement Learning for Fast Adaptation of Interactive Policies with Applications to Autonomous Driving
Huiwen Yan, Kyriakos G. Vamvoudakis, Mushuang Liu · 2. Oktober 2026
This paper develops a meta-multi-agent reinforcement learning (meta-MARL) framework to enable fast adaptation of interactive policies in a multi-agent system (MAS). Meta-reinforcement learning (meta-RL) enables agents to rapidly adapt to new tasks/environments using a bi-level optimization mechanism…
- TagPR: Tag-Guided Process Supervision for Personalization Reasoning in Large Language Models
Song Jin, Juntian Zhang, Ruyu Lyu, Yong Liu, Xun Zhang, Yufei Zhang, Fei Jiang, Guojun Yin, Wei Lin, Rui Yan · 1. Oktober 2026
Recent advancements have endowed Large Language Models with impressive general reasoning capabilities. However, these reasoning models often perform worse than non-reasoning models on personalization tasks. While some methods use outcome-based RL to improve personalization reasoning, they fail to su…
Weitere Unterthemen aus Künstliche Intelligenz
Die Unterthemen, die die OpenAlex-Klassifikation demselben Thema zuordnet, die aktivsten zuerst.
- Large Language Models7.407 Papiere / 12 Monate+247 %
- Adversarial Robustness in Machine Learning3.552 Papiere / 12 Monate+118 %
- Explainable Artificial Intelligence (XAI)2.319 Papiere / 12 Monate+200 %
- Domain Adaptation and Few-Shot Learning2.059 Papiere / 12 Monate+67 %
- Advanced Graph Neural Networks1.926 Papiere / 12 Monate+38 %
- Natural Language Processing Techniques1.595 Papiere / 12 Monate+69 %
