Physical Sciences › Engineering › Control and Systems Engineering
Robot Manipulation and Learning
498 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume — last 12 months
Latest papers
- GarmentPile++: Affordance-Driven Cluttered Garments Retrieval with Vision-Language Reasoning
Mingleyang Li, Yuran Wang, Yue Chen, Tianxing Chen, Jiaqi Liang, Zishun Shen, Haoran Lu, Ruihai Wu, Hao Dong · 5 March 2026
Garment manipulation has attracted increasing attention due to its critical role in home-assistant robotics. However, the majority of existing garment manipulation works assume an initial state consisting of only one garment, while piled garments are far more common in real-world settings. To bridge…
- Cognition to Control - Multi-Agent Learning for Human-Humanoid Collaborative Transport
Hao Zhang, Ding Zhao, H. Eric Tseng · 5 March 2026
Effective human-robot collaboration (HRC) requires translating high-level intent into contact-stable whole-body motion while continuously adapting to a human partner. Many vision-language-action (VLA) systems learn end-to-end mappings from observations and instructions to actions, but they often emp…
- IROSA: Interactive Robot Skill Adaptation using Natural Language
Markus Knauer, Samuel Bustamante, Thomas Eiband, Alin Albu-Sch\"affer, Freek Stulp, Jo\~ao Silv\'erio · 5 March 2026
Foundation models have demonstrated impressive capabilities across diverse domains, while imitation learning provides principled methods for robot skill adaptation from limited data. Combining these approaches holds significant promise for direct application to robotics, yet this combination has rec…
- Self-Improving Loops for Visual Robotic Planning
Calvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du, Chen Sun · 4 March 2026
Video generative models trained on expert demonstrations have been utilized as performant text-conditioned visual planners for solving robotic tasks. However, generalization to unseen tasks remains a challenge. Whereas improved generalization may be facilitated by leveraging learned prior knowledge …
- How to Peel with a Knife: Aligning Fine-Grained Manipulation with Human Preference
Toru Lin, Shuying Deng, Zhao-Heng Yin, Pieter Abbeel, Jitendra Malik · 4 March 2026
Many essential manipulation tasks - such as food preparation, surgery, and craftsmanship - remain intractable for autonomous robots. These tasks are characterized not only by contact-rich, force-sensitive dynamics, but also by their "implicit" success criteria: unlike pick-and-place, task quality in…
- Learning Object-Centric Spatial Reasoning for Sequential Manipulation in Cluttered Environments
Chrisantus Eze, Ryan C Julian, Christopher Crick · 4 March 2026
Robotic manipulation in cluttered environments presents a critical challenge for automation. Recent large-scale, end-to-end models demonstrate impressive capabilities but often lack the data efficiency and modularity required for retrieving objects in dense clutter. In this work, we argue for a para…
- Learning Contact Dynamics through Touching: Action-conditional Graph Neural Networks for Robotic Peg Insertion
Zongyao Yi, Joachim Hertzberg, Martin Atzmueller · 3 March 2026
We present a learnable physics-based predictive model that provides accurate motion and force-torque prediction of the robot end effector in contact-rich manipulation. The proposed model extends the state-of-the-art GNN-based simulator (FIGNet) with novel node and edge types, enabling action-conditi…
- Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons
Anthony Liang, Yigit Korkmaz, Jiahui Zhang, Minyoung Hwang, Abrar Anwar, Sidhant Kaushik, Aditya Shah, Alex S. Huang, Luke Zettlemoyer, Dieter Fox, Yu Xiang, Anqi Li, Andreea Bobu, Abhishek Gupta, Stephen Tu, Erdem Biyik, Jesse Zhang · 3 March 2026
General-purpose robot reward models are typically trained to predict absolute task progress from expert demonstrations, providing only local, frame-level supervision. While effective for expert demonstrations, this paradigm scales poorly to large-scale robotics datasets where failed and suboptimal t…
- Mean-Flow based One-Step Vision-Language-Action
Yang Chen, Xiaoguang Ma, Bin Zhao · 3 March 2026
Recent advances in FlowMatching-based Vision-Language-Action (VLA) frameworks have demonstrated remarkable advantages in generating high-frequency action chunks, particularly for highly dexterous robotic manipulation tasks. Despite these notable achievements, their practical applications are constra…
- Apple: Toward General Active Perception via Reinforcement Learning
Tim Schneider, Cristiana de Farias, Roberto Calandra, Liming Chen, Jan Peters · 2 March 2026
Active perception is a fundamental skill that enables us humans to deal with uncertainty in our inherently partially observable environment. For senses such as touch, where the information is sparse and local, active perception becomes crucial. In recent years, active perception has emerged as an im…
- V-MORALS: Visual Morse Graph-Aided Estimation of Regions of Attraction in a Learned Latent Space
Faiz Aladin, Ashwin Balasubramanian, Lars Lindemann, Daniel Seita · 2 March 2026
Reachability analysis has become increasingly important in robotics to distinguish safe from unsafe states. Unfortunately, existing reachability and safety analysis methods often fall short, as they typically require known system dynamics or large datasets to estimate accurate system models, are com…
- FlowCorrect: Efficient Interactive Correction of Generative Flow Policies for Robotic Manipulation
Edgar Welte, Yitian Shi, Rosa Wolf, Maximillian Gilles, Rania Rayyes · 26 February 2026
Generative manipulation policies can fail catastrophically under deployment-time distribution shift, yet many failures are near-misses: the robot reaches almost-correct poses and would succeed with a small corrective motion. We present FlowCorrect, a deployment-time correction framework that convert…
- Primary-Fine Decoupling for Action Generation in Robotic Imitation
Xiaohan Lei, Min Wang, Wengang Zhou, Xingyu Lu, Houqiang Li · 26 February 2026
Multi-modal distribution in robotic manipulation action sequences poses critical challenges for imitation learning. To this end, existing approaches often model the action space as either a discrete set of tokens or a continuous, latent-variable distribution. However, both approaches present trade-o…
- Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
Zaijing Li, Bing Hu, Rui Shao, Gongwei Chen, Dongmei Jiang, Pengwei Xie, Jianye Hao, Liqiang Nie · 25 February 2026
Hierarchical Vision-Language-Action (VLA) models have rapidly become a dominant paradigm for robotic manipulation. It typically comprising a Vision-Language backbone for perception and understanding, together with a generative policy for action generation. However, its performance is increasingly bo…
- Learning Physical Principles from Interaction: Self-Evolving Planning via Test-Time Memory
Haoyang Li, Yang You, Hao Su, Leonidas Guibas · 25 February 2026
Reliable object manipulation requires understanding physical properties that vary across objects and environments. Vision-language model (VLM) planners can reason about friction and stability in general terms; however, they often cannot predict how a specific ball will roll on a particular surface o…
- Enhancing Goal Inference via Correction Timing
Anjiabei Wang, Shuangge Wang, Tesca Fitzgerald · 24 February 2026
Corrections offer a natural modality for people to provide feedback to a robot, by (i) intervening in the robot's behavior when they believe the robot is failing (or will fail) the task objectives and (ii) modifying the robot's behavior to successfully fulfill the task. Each correction offers inform…
- DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
Li Zhang, Mingyu Mei, Ailing Wang, Xianhui Meng, Yan Zhong, Xinyuan Song, Liu Liu, Rujing Wang, Zaixing He, Cewu Lu · 24 February 2026
Articulated object pose estimation is a core task in embodied AI. Existing methods typically regress poses in a continuous space, but often struggle with 1) navigating a large, complex search space and 2) failing to incorporate intrinsic kinematic constraints. In this work, we introduce DICArt (DIsC…
- AdaWorldPolicy: World-Model-Driven Diffusion Policy with Online Adaptive Learning for Robotic Manipulation
Ge Yuan, Qiyuan Qiao, Jing Zhang, Dong Xu · 24 February 2026
Effective robotic manipulation requires policies that can anticipate physical outcomes and adapt to real-world environments. Effective robotic manipulation requires policies that can anticipate physical outcomes and adapt to real-world environments. In this work, we introduce a unified framework, Wo…
- Find the Fruit: Zero-Shot Sim2Real RL for Occlusion-Aware Plant Manipulation
Nitesh Subedi, Hsin-Jung Yang, Devesh K. Jha, Soumik Sarkar · 24 February 2026
Autonomous harvesting in the open presents a complex manipulation problem. In most scenarios, an autonomous system has to deal with significant occlusion and require interaction in the presence of large structural uncertainties (every plant is different). Perceptual and modeling uncertainty make des…
- A Primer on SO(3) Action Representations in Deep Reinforcement Learning
Martin Schuck, Sherif Samy, Angela P. Schoellig · 24 February 2026
Many robotic control tasks require policies to act on orientations, yet the geometry of SO(3) makes this nontrivial. Because SO(3) admits no global, smooth, minimal parameterization, common representations such as Euler angles, quaternions, rotation matrices, and Lie algebra coordinates introduce di…
- MoMaGen: Generating Demonstrations under Soft and Hard Constraints for Multi-Step Bimanual Mobile Manipulation
Chengshu Li, Mengdi Xu, Arpit Bahety, Hang Yin, Yunfan Jiang, Huang Huang, Josiah Wong, Sujay Garlanka, Cem Gokmen, Ruohan Zhang, Weiyu Liu, Jiajun Wu, Roberto Mart\'in-Mart\'in, Li Fei-Fei · 24 February 2026
Imitation learning from large-scale, diverse human demonstrations has been shown to be effective for training robots, but collecting such data is costly and time-consuming. This challenge intensifies for multi-step bimanual mobile manipulation, where humans must teleoperate both the mobile base and …
- RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
Seungku Kim, Suhyeok Jang, Byungjun Yoon, Dongyoung Kim, John Won, Jinwoo Shin · 24 February 2026
Synthetic data generated by video generative models has shown promise for robot learning as a scalable pipeline, but it often suffers from inconsistent action quality due to imperfectly generated videos. Recently, vision-language models (VLMs) have been leveraged to validate video quality, but they …
- Issues with Measuring Task Complexity via Random Policies in Robotic Tasks
Reabetswe M. Nkhumise, Mohamed S. Talamali, Aditya Gilra · 24 February 2026
Reinforcement learning (RL) has enabled major advances in fields such as robotics and natural language processing. A key challenge in RL is measuring task complexity, which is essential for creating meaningful benchmarks and designing effective curricula. While there are numerous well-established me…
- Latent Diffeomorphic Co-Design of End-Effectors for Deformable and Fragile Object Manipulation
Kei Ikemura, Yifei Dong, Florian T. Pokorny · 23 February 2026
Manipulating deformable and fragile objects remains a fundamental challenge in robotics due to complex contact dynamics and strict requirements on object integrity. Existing approaches typically optimize either end-effector design or control strategies in isolation, limiting achievable performance. …
- SimVLA: A Simple VLA Baseline for Robotic Manipulation
Yuankai Luo, Woping Chen, Tong Liang, Baiqiao Wang, Zhenguo Li · 23 February 2026
Vision-Language-Action (VLA) models have emerged as a promising paradigm for general-purpose robotic manipulation, leveraging large-scale pre-training to achieve strong performance. The field has rapidly evolved with additional spatial priors and diverse architectural innovations. However, these adv…
