HUMAN DEMONSTRATIONS → DEXTEROUS ROBOT LEARNING

Human2Dex

Dexterous Robot Learning from
Wrist-Centric Human Demonstrations

Anonymous Authors

One shared human demonstration corpus.
Two dexterous hands. Six real-world tasks.

Human demonstrationsShared wrist-centric observations
Linker O66 active degrees of freedom
Wuji20 active degrees of freedom
Human2Dex overview: wrist-mounted human demonstrations are converted into shared observations and separate policies for Linker O6 and Wuji robot hands.
From natural human demonstrations to embodiment-specific robot policies.

THE IDEA

Demonstrate once.
Learn across embodiments.

Human2Dex turns a shared corpus of wrist-centric human demonstrations into executable training data for different dexterous hands. Observation and motion semantics are shared; retargeting, action labels, and policies are specific to each hand and task.

ABSTRACT

A reusable interface for human data.

Collecting dexterous manipulation demonstrations for each robot hand requires repeated task-specific data acquisition. We ask whether one shared corpus of natural human demonstrations can instead serve as a reusable data interface for different dexterous hands, without collecting task demonstrations on each target robot.

Human2Dex aligns observations through local hand-appearance adaptation and a shared 21-point hand representation. Fused human hand motion is retargeted to each target hand, while the Grasp-Pocket Interaction Representation (GPIR) expresses tracked objects relative to a hand-centered Functional Grasp Pocket. The same human corpus is converted into separate training sets for Linker O6 and Wuji, retaining hand-specific action labels, normalization, and diffusion policies.

1,899human demonstration episodes
2dexterous robot hands
6real-world manipulation tasks
85.8%pooled success · 206 / 240

01 / METHOD

From human observation to robot action.

Three complementary parts connect shared human data to each target embodiment.

01

Align the observation

Local hand-appearance adaptation and a shared 21-point structure establish consistent observation semantics across human and robot hands.

02

Retarget the motion

Fused PICO and wrist-RGB hand estimates generate executable, native action supervision separately for each calibrated hand.

03

Ground the interaction

GPIR describes tracked objects relative to the thumb–index grasp pocket, giving the policy an explicit object–hand relationship.

Conversion interface combining wrist RGB, PICO tracking, and a wrist-image model for observation alignment, retargeting, and GPIR.
The conversion interface: shared human measurements become hand-specific training data.

GRASP-POCKET INTERACTION REPRESENTATION

Where is the object
relative to the grasp?

GPIR uses a hand-centered Functional Grasp Pocket to express object position, together with object scale and tracking reliability. This compact cue conditions the diffusion policy alongside the structural image and hand state.

Each hand–task pair has its own policy, action labels, and normalization. No target-robot task demonstrations are collected.

Functional Grasp Pocket and grasp-relative object coordinates in Human2Dex.
Explore the policy architecture DINOv2 vision features, hand state, and GPIR condition a diffusion policy to predict native robot commands.

A DINOv2 encoder, hand state, and interaction state condition a 1-D U-Net that predicts a 16-step action sequence. The auxiliary object-regression head is used during training and omitted at deployment.

02 / REAL-WORLD DEMONSTRATIONS

One human corpus. Two embodiments.

Six tasks, spanning compliant grasping, precise placement, and coordinated contact.

Explore both robot hands

01

Sponge

Compliant grasping

Pick up a compliant sponge and place it on a target plate.

Linker O620/20 successful trials
External view · Sponge
Wuji19/20 successful trials
External view · Sponge
02

Croissant

Stable grasping

Pick up a croissant and place it on a plate while maintaining a stable grasp.

Linker O618/20 successful trials
External view · Croissant
Wuji17/20 successful trials
External view · Croissant
03

Cup

Precise placement

Place a green cup into a red cup with precise grasp alignment and insertion.

Linker O616/20 successful trials
External view · Cup
Wuji17/20 successful trials
External view · Cup
04

Bottle

Object reorientation

Lift a horizontal bottle and place it upright.

Linker O619/20 successful trials
External view · Bottle
Wuji16/20 successful trials
External view · Bottle
05

Egg Carton

Multi-finger coordination

Stabilize the carton and lift its lid with coordinated fingers.

Linker O619/20 successful trials
External view · Egg Carton
Wuji13/20 successful trials
External view · Egg Carton
06

Open Drawer

Articulated objects

Engage the drawer handle and pull the drawer open.

Linker O617/20 successful trials
External view · Open Drawer
Wuji15/20 successful trials
External view · Open Drawer

Videos show representative trials. The counts summarize 20 physical trials per hand–task pair; switching views does not imply frame synchronization.

03 / EXPERIMENTAL RESULTS

Human data, measured on real robots.

Success rates across six tasks, with independent policy training for each hand–task pair.

LINKER O6

90.8%109 / 120 successes

WUJI

80.8%97 / 120 successes

20 trials per task

A successful trial must satisfy the terminal criterion for 2 seconds. Both hands reuse the same human episode groups.

Human2Dex · successful trials / total trials
Robot handSpongeCroissantCupBottleEgg CartonOpen DrawerOverall
Linker O620/2018/2016/2019/2019/2017/20109/120
Wuji19/2017/2017/2016/2013/2015/2097/120

Building the complete interface

Pooled success across both hands. Each configuration is trained independently using the same base episodes and evaluation protocol.

Human-Raw26.3% (63/240)
Human-Visual60.0% (144/240)
Human-Structured67.9% (163/240)
Human2Dex85.8% (206/240)

These are system-level configuration comparisons and do not isolate the causal effect of individual components. Pooled counts summarize the tested benchmark, rather than unseen-task performance.

04 / QUALITATIVE COMPARISONS

Different observations. Different behavior.

Representative Linker O6 wrist-view trials from the provided recordings.

Croissant

Human-Raw
Human-Structured
Human2Dex

Cup

Human-Raw
Human-Structured
Human2Dex

Bottle

Human-Raw
Human-Structured
Human2Dex

Examples are separate trials, not synchronized executions. Quantitative conclusions are based on the full evaluation reported above.

SCOPE & FINDINGS

What the experiments establish.

Reusable demonstrations

A single wrist-centric human corpus supports independently trained policies for two calibrated dexterous hands across six tasks.

Grasp-relative grounding

Whole-object interaction cues are most helpful when grasp alignment matters. Localized contacts, such as a lid edge or drawer handle, may need finer-grained representations.

Current scope

The study uses specified task objects and reliable image-space tracking. Transfer to arbitrary hand morphologies and shared policy checkpoints remains open.

Read the full paper ↗

CITATION

BibTeX

Anonymous manuscript. Citation details will be updated with the publication information.

@misc{human2dex,
  title = {{Human2Dex}: Dexterous Robot Learning from
           Wrist-Centric Human Demonstrations},
  author = {{Anonymous Authors}},
  note = {Manuscript}
}