Structured Exploration for Dexterous Manipulation with Human Priors
TL;DR Robot hands usually explore by perturbing each joint independently, which produces unnatural motions and makes the coordinated, human-like behaviors that manipulation needs hard to discover. EigenDEXplore explores along the main directions of human hand motion, found with PCA on human hand videos retargeted to the robot, so it learns coordinated behaviors while the policy keeps full control of every joint.
A Sharpa hand on a KUKA iiwa14 arm moves four tools along trajectories recorded from human videos, reaching one tool-pose waypoint after another. Both policies were trained with RL in simulation with SimToolReal and run zero-shot on the real robot. A trial ends at the last waypoint, or when the tool is dropped or gets stuck.
Hammer the nails
Ten trials per method, sorted. Task progress is the fraction of demonstrated waypoints the robot reaches.
Standard Exploration tends to pinch the tool between the middle finger pad and the backs of the index and ring fingers, which slips and makes in-hand reorientation hard. EigenDEXplore learns a fingertip grasp with the thumb, index and middle finger pads, which holds the tool steadily and lets the hand turn it toward each next waypoint.
Reinforcement learning finds behaviors by trying random variations of its actions, and how those variations are shaped decides which behaviors are easy to find. Everything below runs live on the Sharpa hand with the paper's PCA basis.
We compare four formulations that differ in where the human directions enter: the actions, the exploration noise, or both. EigenDEXplore keeps the robot's joint actions and uses the human directions only in the noise.
SimToolReal in simulation: same task, reward and training budget, only the exploration differs. Each frame replays a real checkpoint, from untrained to 60B frames.
An Allegro hand turns a cube to a new random orientation each time it reaches the last one. Both policies were trained with RL in simulation with DeXtreme and run zero-shot on the real hand. A trial ends when the cube drops or stays stuck for 80 s.
Two runs filmed in a separate session. The goal and the successes come from the robot's log.
Reference-guided bimanual RL with DexMachina, from the task reward only: two hands learn to move a box or a notebook along a human demonstration (ARCTIC), rewarded only on how closely the object follows it, with no curriculum or helper rewards. EigenDEXplore helps most on the high-DoF hands.
Trajectory optimization with SPIDER, from a cold start: it retargets a human demonstration (OakInk2) to a robot hand by sampling candidate trajectories and refining the best, with the same budget for both methods. EigenDEXplore only changes how the samples are drawn, and lowers the cost on every hand, most on the highest-DoF ones.