Artificial intelligence

Robot-free demonstration

Definition

A robot-free demonstration records a human example without driving the target robot during data collection. Tracking devices, instrumented grippers, or video are used to recover observations and actions that can later train a robot policy.

Also known as: Robot-free demonstrations, Robot-free imitation data

Updated

Collect an example without the target robot

Robot data is often collected by teleoperation, where a person commands the deployed robot or a matched follower during the demonstration. Robot-free collection instead records the human action directly. The robot can remain elsewhere or may not yet be available.

The Universal Manipulation Interface uses hand-held grippers and onboard sensing to collect portable manipulation demonstrations. Its authors construct relative trajectories and other policy inputs so that learned behavior can later be deployed on several robot platforms. Those design choices are part of UMI, not requirements of every robot-free interface.

Human motion must become a robot action

A recording has to represent the information the policy will receive and the action the robot can execute. Camera pose, gripper state, timing, coordinate frames, and embodiment differences all affect this conversion. Motion retargeting or a relative action representation can bridge some geometric differences, but neither makes the human and robot dynamics identical.

Robot-free demonstrations are also different from ordinary internet video. An instrumented interface can provide calibrated trajectories and gripper commands that a monocular video may not reveal. Conversely, wearing or holding an interface can change how a person performs the task.

Timing and contact do not transfer automatically

A human hand and a robot can have different compliance, acceleration, tracking error, and safe contact speed. The 2026 RoboPace paper identifies inherited human timing as a problem for policies trained on robot-free demonstrations and reports a contact-aware retiming method on tested dual-arm tasks. That is evidence for the reported method and setup, not a general solution to embodiment transfer.

Missing force, tactile, or joint-state signals can also make two visually similar demonstrations physically different. Evaluation should state what was measured during collection, how actions were mapped to the robot, and whether the learned policy was tested on real contacts rather than only replayed trajectories.

Sources