Physical Sciences › Computer Science › Artificial Intelligence
Reinforcement Learning in Robotics
2 519 papiers indexés
L’apprentissage par renforcement appliqué à la robotique explore comment des agents autonomes acquièrent des comportements complexes en interagissant avec leur environnement. Les travaux récents s’attachent à affiner des méthodes comme le Q-learning, les modèles de diffusion ou l’optimisation de politiques, en abordant des défis tels que l’estimation hors ligne, la distribution des récompenses ou la gestion de grands espaces d’actions. Ces recherches examinent aussi des questions de sécurité, d’équité et d’adaptation dynamique, notamment dans des contextes multi-agents ou lors du transfert de connaissances entre tâches.
Ce sujet et sa hiérarchie proviennent de la classification OpenAlex, le catalogue ouvert de la recherche scientifique mondiale.
Volume mensuel - 12 derniers mois
Pays des laboratoires
- États-Unis42 % · 733 articles
- Chine36 % · 637 articles
- Royaume-Uni6,9 % · 121 articles
- Allemagne6,4 % · 112 articles
- Canada6,3 % · 110 articles
- Corée du Sud4,3 % · 76 articles
- France3,4 % · 60 articles
- R.A.S. chinoise de Hong Kong3,3 % · 58 articles
Sur 1 749 articles de ce sujet dont au moins un laboratoire est situé. 73 pays représentés.
Il s'agit du pays du laboratoire, jamais de la nationalité des personnes. Un article signé depuis plusieurs pays compte pour chacun d'eux, les parts dépassent donc 100 % au total. La couverture est partielle et le manque n'est pas aléatoire : un chercheur dont l'institution est inconnue publie en général peu, ce qui sur-représente les laboratoires établis.
Derniers papiers
- Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
Dongsu Lee, Haoran Xu, Amy Zhang · 5 octobre 2026
Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal int…
- Reward Inflation: A Healthy Stimulus for Reinforcement Learning
Ganghun Lee, Minji Kim, Minsu Lee, Byoung-Tak Zhang · 5 octobre 2026
Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of tra…
- Co-design Gym: A Unified Benchmark for Embodiment-Policy Co-optimization
Aviraj Newatia, Yordan Tsvetkov, Leonard Pleiss, Andrew Spielberg, Rika Antonova · 5 octobre 2026
Finding an optimal behaviour policy within a given environment is a widely studied problem in domains as diverse as games, robotics, energy infrastructure, communication networks, and multi-agent systems. Numerous benchmarks have been developed to support such research, but the vast majority assume …
- OuroReward: Sequential Reward Scheduling for Reinforcement Learning in Text-to-3D Generation
Bingyang Cui, Yujie Zhang, Yiling Xu, Yunfeng Guan · 5 octobre 2026
Reinforcement learning (RL) for Text-to-3D (T23D) generation requires optimization across multiple quality dimensions such as semantic alignment and texture clarity. Existing methods typically optimize these dimensions simultaneously through multiple reward aggregation, without explicitly modeling i…
- AdaStep: Adaptive Step Credit Weighting for Agentic Reinforcement Learning
Xin Wang, Wenhao Wu, Menghao Zhang, Zhi Wang, Kun Shao, Jian Luan · 5 octobre 2026
Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of individual decisions. Step-level credit assignment provides finer-grained supervision, but its estimates can be unreliable because observed retu…
- Single or Multiple Policies for Phase-Structured Reinforcement Learning?
Guilhem Loussouarn, Nancy Nayak, Kin K. Leung · 5 octobre 2026
Many reinforcement-learning (RL) problems are non-stationary yet structured and can be decomposed into phases, each with its own transition probabilities and reward functions. When the phase sequence is known, the common solution augments the state with information to satisfy the Markovian property …
- Follow the Winners: Conservative Policy Improvement with the Cross-Entropy Method for Critic-Free RFT
Joery Ari\"en de Vries, Neil David Lawrence, Zhenwen Dai · 5 octobre 2026
Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over repeated rollouts to reduce target variance. However, this setup is ill-suited to agents acting in stateful environments such as live services o…
- Prospective Hindsight: Self-Calibrating Reinforcement Learning via Prediction-Reality Gaps
Jiaxin Zhang, Xiangyu Peng, Qinglin Chen, Yu Li, Hiroaki Hayashi, Chien-Sheng Wu · 5 octobre 2026
Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequences, leaving the agent's belief at action time invisible to the gradient. We introduce Prospective Hindsight (PH), a self-calibrating trainin…
- CriticHack: Evaluating Visual Rewards Under Robot Policy Optimization
Jiaxuan Luo, Xingguo Xu, Shanshan Wang, Yuhan Zhou, Zhen Zhang · 5 octobre 2026
Learned visual reward models are increasingly used to optimize robot policies, yet a reward model can score an execution that acts on the wrong object as highly as one that completes the task. We show that optimizing such a reward can amplify these wrong-object failures while reward and task success…
- Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
Xinjie Liu, Ruihan Zhao, Anirban Chaudhuri, Cyrus Neary, Ufuk Topcu, David Fridovich-Keil · 5 octobre 2026
Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) …
- Lexicographic Multi-Objective On-Policy Distillation
Doseok Jang, Jon Ander Campos, Youran Qi · 5 octobre 2026
Reinforcement learning from verifiable rewards (RLVR) usually optimizes answer correctness, yet useful language-model behavior also requires high-quality reasoning and concise responses. Existing multi-reward post-training methods typically scalarize rewards or combine specialists without explicitly…
- Credit Where It Matters: Dependency-Aware Policy Optimization for Terminal Agents
Yu Li, Guangfeng Cai, Long-Fei Li, Shuo Han, Shengtian Yang, Han Luo, Kaibing Yang, Lei Feng · 5 octobre 2026
Terminal-using agents benefit from reinforcement learning (RL) in coding, debugging, and other multi-step terminal tasks. In these tasks, later commands often depend on information or intermediate results produced by earlier commands. However, existing trajectory-level and step-level credit assignme…
- EVOL: Simulator-Guided Evolutionary Expert Synthesis for Deployment-Free Learning Path Recommendation
Geonwoo Bang, Dongho Kim, Moohong Min · 5 octobre 2026
Reinforcement learning (RL) for learning path recommendation (LPR) faces two coupled obstacles. First, the policy must commit to a sequence of L concepts without intermediate feedback, producing a combinatorial search space that grows super-exponentially with L and provides reward only at the final …
- RIFAR: Reliability and Forgetting-Aware Replay for Continual Robot Learning
Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xiuying Chen · 5 octobre 2026
Genuine embodied agency requires robots to turn continuous real-world experience into lasting, transferable skills. This demands continual learning that integrates new capabilities without eroding prior knowledge as tasks and environments evolve. Experience replay mitigates forgetting, but storing c…
- FSPO: Policy-Consistent Risk and Pareto-Feasible Control for Budgeted LLM RL Post-Training
Miaobo Hu, Shuhao Hu, Xiaobo Guo, Xin Wang, Bokun Wang, Daren Zha, Jun Xiao · 5 octobre 2026
Adaptive LLM reinforcement-learning post-training changes multiple training actuators online, including rollout temperature, group size, clipping, KL regularization, verifier allocation, and update budget. Three coupled issues remain unresolved. A future-risk model trained from behavior trajectories…
- VIGOR: Zero-Shot Visual Generalization via Latent-Space Consistency in Model-Based Reinforcement Learning
Mingyu Park, Samyeul Noh, Hyun Myung, Donghwan Lee · 5 octobre 2026
Model-based reinforcement learning (MBRL) achieves strong sample efficiency by planning within learned latent dynamics, yet its performance degrades substantially under unseen visual distractions such as background variations, lighting changes, or camera shifts. Unlike model-free RL, where encoder p…
- Tropical Reinforcement Learning
Arip Asadulaev, Aladin Djuhera, Karim Salta, Holger Boche, Fakhri Karray, Martin Takac · 5 octobre 2026
Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which solution actually worked, and because probabilities …
- Make Sparse Rewards Count: Density-Aware Reward Aggregation for Multi-Reward RL
Tong Zheng, Skylar Zhai, Zhan Cheng, TianMing Sha, Youling Huang, Shuo Zhou, Shaotong Qi, Jingcheng Liang, Xuwei Ding, Pengcheng Xu · 2 octobre 2026
Multi-reward reinforcement learning trains large language models to satisfy multiple behavioral objectives simultaneously. Reward-wise normalization, as used in GDPO, preserves reward-specific relative information within rollout groups, but different objectives can still exhibit uneven learning prog…
- Probe with Participation Trophies: Random-Reward RL as a Probe of LLM Capability
Yu Mao, Lei Yu, Zining Zhu, Yusheng Zheng, Haohang Li, Freda Shi, Yutong Yin, Zhaoran Wang, Jingcheng Niu · 2 octobre 2026
We connect the spurious-reward paradox to a model's reachability and propose random-reward reinforcement learning (RL) as a useful tool for the probing enterprise, addressing a decade-long debate over what probing performance actually reveals about a model. There are two prevailing explanations for …
- FERPO: Forward Entropy-Regularized Policy Optimization
Sebastian Sanokowski, Alireza Sarmadi, Majid Khadiv · 2 octobre 2026
Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potent…
- Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
Zhuoran Li, Yunzhan Li, Xun Wang, Yihan Du, Longbo Huang · 2 octobre 2026
Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. …
- iADD: Improving Alignment and Diversity in Diffusion Policy Optimization
Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel, Ajad Chhatkuli, Danda Pani Paudel · 2 octobre 2026
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we pro…
- Iterative Policy Refinement through Semantic Rollout Analysis
Feiyu Gavin Zhu, Qi Xu, Zhifei Deng, Zhigang Hua, Luke Simon, Jean Oh, Reid Simmons · 2 octobre 2026
Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the…
- Optimal Transport Meets Reinforcement Learning: A Survey
Yujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni Montana · 2 octobre 2026
Reinforcement learning (RL) algorithms frequently compare probability distributions, such as state visitation distributions induced by policies and experts, action distributions from learned policies and offline datasets, or transition distributions from learned models and environments. However, com…
- Does Scaling Reinforcement Learning Really Require More Training?
Bangji Yang, Jiajun Fan, Hongba Ma, Ruihan Guo, Ge Liu · 2 octobre 2026
Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible fr…
Autres sujets du thème Intelligence artificielle
Les sujets rattachés au même thème par la classification OpenAlex, les plus actifs d'abord.
- Large Language Models7 407 papiers / 12 mois+247 %
- Adversarial Robustness in Machine Learning3 552 papiers / 12 mois+118 %
- Explainable Artificial Intelligence (XAI)2 319 papiers / 12 mois+200 %
- Domain Adaptation and Few-Shot Learning2 059 papiers / 12 mois+67 %
- Advanced Graph Neural Networks1 926 papiers / 12 mois+38 %
- Natural Language Processing Techniques1 595 papiers / 12 mois+69 %
