UVTA

Unified Visual-Tactile-Action Modeling from Human Demonstrations

Hands can change. The physics of interaction does not.

Wenqiao Li Qianyou Zhao Jiawen Hao Xuezhou Zhu Tengyu Liu Kaifeng Zhang Chuan Wen Siyuan Huang

Beijing Institute for General Artificial Intelligence Sharpa
Shanghai Jiao Tong UniversitySchool of Artificial Intelligence
Shanghai Jiao Tong University

01 / Motivation

Learning across embodiments

Hands can change.The physics of interaction does not.

Dexterous manipulation needs touch. Collecting tactile demonstrations on robot hands, however, is difficult to scale.

Human interaction offers a scalable source of contact experience. UVTA combines diverse human tactile demonstrations with a smaller set of robot demonstrations to learn a shared visual–tactile–action representation.

The key idea

Scale human contact experience to learn contact-aware robot manipulation.

Human and robot hands share interaction patterns that support learning a contact-aware policy
Different hands. Shared contact physics.

02 / System and Dataset

From human interaction to training data

Capture touch.
Collect diverse interactions.

A passive 22-DoF exoskeleton synchronizes hand motion, fingertip pressure, wrist pose, and wrist-mounted RGB vision.

Tactile exoskeleton with fingertip pressure sensors, VIVE wrist tracker, and wrist-mounted RGB camera
Human demonstrations provide more diverse scenes than robot teleoperation

Two complementary data sources

Human demonstrations broaden contact diversity. Robot demonstrations provide executable action supervision.

1,000human demonstrations / task
150robot demonstrations / task
Wrist RGB20-D touch31-D action30 Hz

03 / Model

Unified Visual–Tactile–Action Modeling

One interaction representation.
Two complementary predictions.

A 384-D visual token and a 20-D tactile token form a shared representation, trained through action diffusion and future-tactile prediction.

Unified Visual-Tactile-Action model architecture
01 / Fuse

404-D interaction token

Connect wrist vision with contact measurements in a shared representation.

02 / Act

Diffusion action head

Generate 16-step chunks of relative wrist and absolute finger-joint actions.

03 / Anticipate

Tactile world head

Predict future touch to supervise contact-aware representation learning.

04 / Real-robot demonstrations

See the action.
See the contact.

Task I / Page separation and turning

Flip Page

The robot approaches the book, slows down near the page, and coordinates its fingers to rub and separate a single page before turning it to the left.

Playback speed: 1× real time.

Task descriptions follow the paper’s experimental setup. Touch panels show recorded sensor visualizations. Playback speed appears at the top right of the external view.

05 / Generalization Test

Page turning under
changing illumination.

A separate demonstration of the page-turning task with colored lighting interference, shown with synchronized views and tactile feedback.

Lighting interference / Flip Page

Generalization Test

The robot separates and turns a page while colored lights change the scene’s appearance. By relying on tactile feedback, the robot continues to complete the task reliably despite visual interference.

Playback speed: 1× real time.

06 / Results

Human touch changes
what the policy can learn.

UVTA outperforms the strongest visual-tactile baseline by 41 points on average, while future-tactile prediction contributes a 28-point gain over its ablation.

MethodFlipBulbSwitchBallLiquidAvg.
ViTacFormer040001
RDP10202018014
T-Rex432624361529
UVTA838088564470
+41points over the strongest baseline
+28points from future-tactile prediction
84%average on the three scaling tasks at 1,000 human demos
Task score improves as the number of human tactile trajectories increases
Performance continues to increase with additional human tactile trajectories and does not saturate at 1,000 demonstrations.

07 / Limitations and Future Work

The next scale of contact learning

A step toward scalable tactile learning.
Not yet its limit.

01

Fingertip-only sensing

Our current system measures touch only at the fingertips. It does not capture contact across the palm or the rest of the hand.

Future direction

Extend tactile coverage to richer, whole-hand interactions.

02

Limited data and model scale

We evaluate UVTA at a limited range of human-data volumes and model capacity. Its behavior at substantially larger scales remains untested.

Future direction

Scale up human interaction datasets and model capacity, and systematically study the resulting performance and generalization.

Broader touch. More human experience. Larger models.