Artificial intelligence
Open-vocabulary object detection
Definition
Open-vocabulary object detection locates objects in an image and associates them with text-defined categories beyond a fixed closed list used during detector training. It combines spatial detection with vision-language representations that can be queried using class names or phrases.
Also known as: OVD, Open vocabulary object detection
Updated
Text expands the detector's query set
A closed-vocabulary detector predicts boxes and labels from categories fixed by its training dataset. An open-vocabulary detector uses language-aligned visual features so that a user or robot can request categories through text, including categories not present in the detector's labelled box training set.
OWL-ViT transfers image-text representations to detection and accepts text queries for target categories. Grounding DINO combines a Transformer detector with grounded language pretraining and reports evaluation as an open-set detector. The papers use related terminology, and individual benchmarks define differently which categories count as unseen.
A robot can ask for task-relevant objects
A general-purpose robot may be instructed to find a particular tool, package, or household object. Open-vocabulary detection can propose image regions associated with that phrase without adding a dedicated output class for every possible instruction. The detections can guide active perception, tracking, or a manipulation pipeline.
The output is usually a two-dimensional region and a text match score. Grasping still needs depth, pose estimation, segmentation or geometry, reachability checks, and a controller. Language-conditioned detection is one perception component, not an action policy.
An open vocabulary is not unlimited recognition
Results depend on the image-text pretraining data, detector training, prompt wording, image quality, viewpoint, and decision threshold. Rare objects, fine-grained distinctions, occlusion, and unfamiliar environments can cause missed or incorrect detections. A model may also match visual context or text correlations instead of the physical attribute relevant to the task.
"Open" therefore describes how categories are queried and evaluated, not a guarantee that the detector understands every phrase. A robot should preserve uncertainty and verify critical detections through additional views, sensors, or task-specific checks.
Sources
Related terms
Vision-language model
A vision-language model processes visual information and natural language in a shared system. Depending on its design, it may connect images with text representations or generate text from visual and textual inputs.
Active perception
Active perception is perception in which a robot chooses motions or interactions partly to obtain more useful observations. Examples include moving a camera to reveal an occluded object or touching an object to reduce uncertainty about its state.
Pose estimation
Pose estimation determines the position and orientation of an object or robot relative to a reference frame. For a rigid body in three-dimensional space, a full pose has three translational and three rotational degrees of freedom.
Robotic manipulation
Robotic manipulation is the use of a robot to change an object's position, orientation, or state through physical interaction. It includes grasping and moving objects as well as actions such as pushing or carrying them without a grasp.