Artificial intelligence

Vision-language model

Definition

A vision-language model processes visual information and natural language in a shared system. Depending on its design, it may connect images with text representations or generate text from visual and textual inputs.

Also known as: VLM, Vision language model

Updated

The output depends on the architecture

CLIP learns image and text representations by identifying which captions belong with which images. It can compare a visual observation with descriptions without being a conversational text generator.

PaLI instead generates text from visual and textual inputs, supporting tasks such as image captioning and visual question answering. Both are vision-language models, so the label alone does not specify the output interface.

Visual semantics can help robot tasks

A robot may need to distinguish the object described as a red cup from neighboring objects. A vision-language representation can provide semantic information for that decision.

CLIPort combines CLIP's visual-language features with a spatial manipulation architecture to learn language-conditioned picking and placing. The robot policy adds action-relevant structure beyond the pretrained visual-language representation.

Understanding an image does not execute an action

A vision-language-action model explicitly includes robot action output. A VLM can instead serve as a perception component, task planner, or source of pretrained features in a larger robot system.

Textual accuracy and physical task success therefore require different evaluations. Naming the correct object does not by itself show that the robot can reach it, grasp it, or complete the requested operation.

Sources