ARMFUL: Learning Whole-Arm Grasp Configurations
for Single-Arm Pickup of Large Objects

Anonymous Authors
Under Review, 2026

Real-Robot Pickups

ARMFUL on a physical Franka arm across the 10 IKEA objects.

Abstract

Many everyday objects, such as storage boxes, baskets, and chairs, are too large or heavy for a gripper to grasp, yet a single arm can still pick them up by wrapping around and supporting them. Such grasps are hard to learn: wrapping motions are difficult to teleoperate, and existing datasets contain little whole-arm pickup data. We introduce WAG (Whole-Arm Grasp Dataset), the first large-scale dataset for single-arm pickup of large objects, with 1.7 million certified joint configurations across 978 objects. We sample configurations directly in joint space, execute each through a squeeze-and-lift primitive in simulation, and certify survivors under active shaking and external forces. To demonstrate the utility of WAG and accomplish the whole-arm pickup task, we develop ARMFUL, a whole-arm pickup pipeline that pairs a diffusion model trained on WAG, generating whole-arm grasps from the object's point cloud, with a motion planner that reaches them collision-free before the same squeeze-and-lift execution. In simulation, ARMFUL reaches 80.4% lift success on seen and 72.8% on unseen objects, and it lifts the object in 70 of 100 trials on a physical Franka arm across 10 IKEA objects.

Dataset Generation

Construction of the WAG dataset: sample grasp configurations, filter penetration, then squeeze, lift and perturb in simulation.

Construction of WAG. We sample candidate whole-arm grasp configurations, discard penetrating candidates, and test the rest directly in a GPU-parallelized simulator. The arm is teleported to the candidate grasp configuration and executes the squeeze and lift primitive. Configurations that hold the object throughout the lift and maintain the grasp with first-joint shaking and external force applied to the object are added to WAG.

Pickup Pipeline

Overview of ARMFUL: diffusion-based grasp configuration learning followed by motion planning.

Overview of ARMFUL. (a) Grasp configuration learning. A diffusion model conditioned on the object's point cloud, encoded with the DP3 encoder, and on the arm's initial configuration generates whole-arm grasp configurations (the arm's seven joint angles), which are filtered for self-collision, object and table penetration, and planner feasibility. (b) Motion planning. The arm follows an RRT-Connect path to the selected arm grasp configuration, and the same squeeze-and-lift execution completes the pickup.