The course is designed to build a practical understanding of the concepts of Reinforcement Learning (RL) from the fundamental level. It provides a comprehensive exploration of RL, the branch of machine learning that enables agents to learn optimal decision-making strategies through direct interaction with or modeling of their environments.
Students will be guided through the basics of RL, defining its key components such as Agent, Environment, Actions, Rewards, Policies, and Value Functions.
Using Markov Decision Processes (MDPs), sequential decision-making problems will be modeled, and the fundamental balance between exploration and exploitation in the learning process will be addressed. Several algorithms will be introduced, from the basic but useful Q-Learning and Policy Iteration to more advanced methods such as those in the Policy Gradient
family.