Robot control

Differential dynamic programming

Definition

Differential dynamic programming is a local optimal-control method that repeatedly approximates dynamics and cost around a nominal trajectory, performs a backward pass to compute control updates and feedback gains, and rolls the result forward through the dynamics.

Also known as: DDP

Updated

Alternate backward and forward passes

Differential dynamic programming starts from a time-indexed state and control trajectory. Around that trajectory, it builds local approximations of the dynamics and cost. A backward recursion estimates how a state change affects future cost and produces both a feed-forward control change and local feedback gains. A forward rollout applies a scaled update through the nonlinear dynamics and evaluates the new trajectory.

This structure makes DDP a shooting method for trajectory optimization: the state sequence is generated by forward dynamics, rather than treated only as unrelated decision variables. Repeating the two passes can converge to a locally improved control sequence.

From rigid bodies to multi-contact motion

Crocoddyl applies a DDP-family solver to multi-contact optimal-control problems, including legged gaits and dynamic manoeuvres in simulation. Its feasibility-driven variant modifies how gaps in an initial state-control trajectory are handled. This is one implementation and extension, not the definition of all DDP methods.

DDP is closely related to iterative linear-quadratic regulator methods. A common distinction is that full DDP retains second-order derivatives of the dynamics, while iLQR uses a first-order dynamics approximation. Nganga and Wensing study how to compute the second-order information more efficiently for rigid-body systems and report tests on robotic models.

Local models impose local limits

The result depends on the initial trajectory, cost, horizon, dynamics model, regularization, and line search. Non-convex tasks can contain several local minima, and a backward pass can become numerically unstable when the local control Hessian is not suitable. Collision, contact, and actuator limits also require explicit treatment; an unconstrained DDP update does not satisfy them automatically.

A trajectory optimized on a nominal model is not a hardware guarantee. State-estimation error, unmodelled contact, latency, and changing payloads can invalidate the rollout. Receding-horizon use within model predictive control adds feedback through replanning, but it also imposes a strict computation budget.

Sources