Artificial intelligence
Action tokenization
Definition
Action tokenization converts robot actions or action sequences into discrete symbols that a model can predict and decode into control commands. The tokenizer defines how those symbols represent continuous or discrete robot actions.
Also known as: Action tokenisation
Updated
Tokens need a physical interpretation
A vision-language-action model that predicts discrete tokens needs a defined mapping from those tokens to robot commands. A simple approach divides each continuous action dimension into bins. Decoding a token then selects a corresponding value or interval.
The FAST paper explains why per-dimension, per-timestep binning can be inefficient for high-frequency, dexterous action data. Consecutive commands can be strongly correlated, while the model still has to predict a long symbol sequence.
A sequence can be compressed before encoding
FAST, short for Frequency-space Action Sequence Tokenization, applies a discrete cosine transform to action sequences as part of a compression-based tokenizer. It represents structure across time instead of treating every scalar command as unrelated.
For example, a smoothly changing arm trajectory contains temporal regularity that an appropriate sequence representation can exploit. FAST is one specific tokenizer, not a synonym for action tokenization generally.
Check the decoder and action interface
A token has no universal meaning across robot systems. Its interpretation depends on the action space, normalization, discretization, and decoder used by that model.
The FAST results concern the authors' evaluated models and tasks. They do not imply that discrete action tokens are always preferable to continuous generation with flow matching or diffusion.
Sources
Related terms
Vision-language-action model
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.
Action chunking
Action chunking is the prediction or organization of several future robot actions as one sequence. A policy can execute all or part of a chunk before using new observations to produce another sequence.
Flow matching
Flow matching is a generative-model training method that learns a vector field for transforming a simple probability distribution into a data distribution. In robot learning, the generated samples can be continuous action sequences conditioned on observations and instructions.