Robotics
3D dynamic scene graph
Definition
A 3D dynamic scene graph is a layered graph that represents places, objects, people, and other spatial entities as nodes connected by geometric, semantic, and time-dependent relations. It gives a robot a structured scene representation above raw geometry alone.
Also known as: 3D DSG, Dynamic scene graph
Updated
A map with entities and relations
A geometric map can record points, surfaces, or occupied cells. A 3D dynamic scene graph adds named entity types and relationships at several spatial scales. A node may represent an object, a person, a room, or a larger place; an edge can record containment, adjacency, relative position, or a relation that changes over time.
Rosinol and colleagues define a layered 3D dynamic scene graph that combines metric and semantic information with static and dynamic entities. Their Kimera system builds this representation from visual-inertial data and connects components for mapping, object localisation, human pose estimation, and scene parsing.
Structure supports task-level queries
A humanoid asked to carry an item needs more than a dense point cloud. It may need to know which detected object is on a table, which room contains that table, and where a moving person is relative to both. Graph relations provide an interface for queries and planning at that level.
The Kimera paper reports using its graph for hierarchical semantic path planning. That result concerns the system and datasets studied in the paper. A 3D dynamic scene graph is the representation concept, not a claim that every graph builder understands every object or activity.
It is not a complete world model
The graph depends on upstream pose estimation, reconstruction, detection, tracking, and semantic labels. A mistaken object identity or camera pose can create incorrect nodes and edges. Dynamic scenes also make relations stale unless the system updates them and represents uncertainty.
Different systems may choose different layers, node types, and relation meanings. The term therefore describes a family of structured spatial representations rather than one universal schema. Geometry still matters for collision checking and control even when a high-level graph says two entities are related.
Sources
Related terms
Simultaneous localization and mapping
Simultaneous localization and mapping is the joint estimation of a robot's state and a map of its environment from sensor observations. It is commonly abbreviated SLAM.
Point cloud
A point cloud is a collection of points representing sampled locations in space, usually with three-dimensional coordinates. Individual points may also carry attributes such as color or return intensity.
Occupancy grid
An occupancy grid divides space into cells and records occupancy information for each cell. A two-dimensional robot map commonly distinguishes occupied, free, and unknown regions.
World model
A world model is an internal predictive model of an environment and how it changes. In robot learning, it can predict future states or observations under possible actions to support planning or policy training.