Artificial intelligence

Action Chunking with Transformers

Definition

Action Chunking with Transformers is an imitation-learning algorithm that predicts sequences of robot actions from observations using a transformer-based conditional variational autoencoder. It is usually abbreviated ACT.

Also known as: ACT, Action chunking transformer

Updated

ACT names a particular learning method

The authors introduced ACT alongside ALOHA, a two-arm teleoperation system. ACT learns from demonstrations and predicts action chunks containing multiple future commands.

The original model uses camera images and joint positions. During training, a conditional variational autoencoder represents variation in demonstrated action sequences through a latent variable. At evaluation, the authors set that variable to the prior mean.

Overlapping predictions support smooth execution

The paper's temporal-ensembling procedure queries the policy repeatedly and combines predictions that refer to the same timestep. The aim is to incorporate observations without abrupt switches between independently generated chunks.

The demonstrated tasks include opening a condiment cup, slotting a battery, and preparing tape. These require coordinated movement and contact handling; the paper reports results for its particular hardware, training data, and task setups.

Hardware and algorithm are separate

ALOHA is the data-collection and robot system. ACT is the learning algorithm. Using an ALOHA-style robot does not require that every policy be ACT, and the term ACT does not mean any transformer that happens to predict several actions.

ACT is an imitation-learning method, so the coverage and quality of the demonstrations remain central to what the trained policy can reproduce.

Sources