Artificial intelligence

Language-conditioned policy

Definition

A language-conditioned policy selects actions using a language instruction together with observations. The instruction specifies or modifies the behavior requested from the policy.

Also known as: Language conditioned policy

Updated

Language specifies the requested behavior

CLIPort learns manipulation from visual input and language goals. An instruction can distinguish picking a particular object, placing it in a named region, or arranging items according to a described relation.

The policy must connect those words to both scene information and actions. It does not suffice to output a fluent description of the requested task.

Conditioning does not require one architecture

A policy can use a language embedding together with robot observations. Octo supports both language instructions and goal images, illustrating that task conditioning is an interface choice rather than a synonym for a large language model.

A vision-language-action model is a closely related model type when visual input and language lead to robot action output.

Skill selection is a distinct use of language

SayCan uses language-model knowledge and skill-feasibility estimates to choose from robot skills. Selecting a skill and predicting its detailed motor commands are different levels of decision-making, even when both use language.

Evaluate instruction following through physical outcomes and the actual range of supported commands. A policy may respond correctly to familiar wording while still failing on unfamiliar objects, spatial relations, or tasks that require abilities absent from its training data.

Sources