Robot control
Model predictive path integral control
Definition
Model predictive path integral control is a sampling-based model predictive control method that evaluates many perturbed control sequences, weights them by trajectory cost, and uses the weighted perturbations to update a receding-horizon command.
Also known as: MPPI, Model predictive path integral
Updated
Control through sampled rollouts
Model predictive path integral control starts with a control sequence over a finite horizon. It adds sampled noise to create many candidate sequences, rolls each one through a dynamics model, and evaluates its accumulated cost. Lower-cost rollouts receive greater weight when the nominal sequence is updated. The controller applies the first command, shifts the horizon, observes the new state, and repeats.
Williams and colleagues derive this sampling-based controller from an information-theoretic treatment of stochastic optimal control and demonstrate it on aggressive autonomous driving. The path-integral name refers to that derivation over possible trajectories, not to integrating a geometric path.
A sampling method inside MPC
MPPI is one way to implement model predictive control. Conventional MPC may solve a constrained nonlinear programme using derivatives. MPPI can instead evaluate rollouts without differentiating the dynamics or cost, which is useful when those functions are discontinuous or available only through a simulator.
Parallel hardware can evaluate many candidates at once. Robotics applications use the method for vehicle control, navigation, manipulation, and legged movement when a suitable forward model and cost are available. A learned dynamics model or policy can also contribute proposals, but that does not change the need to evaluate behaviour over the horizon.
MPPI also differs from offline trajectory optimisation. Both optimise time-indexed commands, but MPPI normally runs as feedback: it repeatedly replans from the latest estimate and executes only the first part of its solution.
Samples and models limit the result
The controller can miss a narrow feasible region if its samples do not explore it. More rollouts or a longer horizon increase computation, while too little of either can make the result short-sighted or noisy. Temperature, noise covariance, cost scaling, and the initial control sequence strongly affect the weighted update.
A cheap approximate model enables fast sampling but can prefer motions that fail on hardware. Hard safety constraints are not automatically guaranteed merely because collisions have a high cost; rare unsafe rollouts and limited sampling still matter. Practical systems may add constraint handling, a safety filter, smoothing, or a lower-level tracking controller. Reported performance therefore applies to the chosen model, sampling budget, costs, hardware, and environment.
Sources
Related terms
Model predictive control
Model predictive control repeatedly optimizes future actions using a system model, applies the next part of the solution, and replans from updated state information. It can account for objectives and constraints over a finite prediction horizon.
Trajectory optimization
Trajectory optimization finds a time-varying motion, and often control inputs, that minimizes an objective while satisfying specified constraints. Robot applications can include geometric, kinematic, and dynamic constraints.
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Differentiable simulation
Differentiable simulation is physical simulation that provides derivatives of simulated outcomes or losses with respect to inputs such as controls, initial states, model parameters, or robot design variables. Those gradients can drive optimisation and learning through the simulated dynamics.