← All robotics laboratories
Steampunk robot artwork from the book
Robotics & Perception · Section 3.6

Learning to Act Optimally

Value iteration, policy iteration, and Q-learning

Compare with the notebook’s Python example

Cell numbers below are zero-based indices in the notebook. For live experiments, browser calculations use the same stated parameters; seeded random realizations and pedagogical extensions are identified in the experiment.