Artificial intelligence

Latent action model

Definition

A latent action model infers compact action-like variables that explain changes between observations when the underlying control commands are unavailable. Robot-learning systems can use those variables to pretrain on video, then learn a smaller mapping from latent actions to commands for a particular robot.

Also known as: LAM, Latent action modelling, Latent action modeling

Updated

Inferring actions that were not recorded

Large video collections show how scenes change but rarely include robot joint commands, gripper states, or end-effector trajectories. A latent action model learns a hidden variable that helps predict the transition from one observation to another. The variable is intended to capture controllable change without requiring the original action label.

Latent actions can be continuous vectors or discrete codes. An inverse model can infer a code from two frames, while a forward model checks whether the code helps predict the later frame. The resulting representation can become a pretraining target for a policy.

From video change to robot command

Ye and colleagues introduced Latent Action Pretraining for general Action models. Their method trains a vector-quantised model to infer discrete latent actions between image frames, pretrains a language-conditioned model to predict those codes, and then fine-tunes on robot data to map the codes to physical actions. The paper reports real manipulation results for that method; it does not establish that every latent action learned from internet video is executable by a robot.

This approach can help a vision-language-action model learn temporal and semantic structure before scarce action-labelled demonstrations are added. It also offers a possible bridge for cross-embodiment learning, because the pretraining target is not tied directly to one robot's joint layout.

A latent action model is not the same as action tokenization. Tokenisation encodes known robot actions into symbols. A latent action model attempts to infer action-like causes when the action was not observed. It is also not ordinary imitation learning, which normally assumes demonstrations contain actions or trajectories that can supervise the learner.

Latent change is not necessarily control

Video changes can come from camera motion, another person, lighting, editing, or object dynamics outside the robot's control. Without suitable data and objectives, a latent code may represent those effects instead of an actionable skill. Several different hidden actions can also explain the same pair of frames, so the representation is not uniquely identifiable.

Mapping a code to a new embodiment still requires robot data and can fail when the video motion exceeds that robot's reach, sensing, or actuator limits. Fine temporal detail may be lost by frame spacing or quantisation. Evaluation therefore needs closed-loop hardware trials, not only video prediction quality or separation of latent codes.

Sources