Fill in the update rules for model-free Q-learning, Dyna-Q, and hindsight experience replay. Then watch how the agent's values, policy, and learning curve change inside a classic four-rooms gridworld.
Fill in the update rules for model-free Q-learning, Dyna-Q, and hindsight experience replay. Then watch how the agent's values, policy, and learning curve change inside a classic four-rooms gridworld.