
Reinforcement Learning beginner to master
https://www.udemy.com/course/beginner-master-rl-1/?referralCode=376738F1E8AF47CAA6F1
Advanced Reinforcement Learning in Python: from DQN to SAC
https://www.udemy.com/course/advanced-reinforcement/?referralCode=2C96ADF61C80DD7FD392
Explore Google Colab as an online programming environment for writing and running code in the cloud. Discover notebook workflow, GPU access, minimal setup, and easy sharing via Google Drive.
Identify your prior knowledge to begin advanced reinforcement learning with dqns. Start with basics like the Markov decision process, then choose leveling modules or jump to bite, lightning.
Refresh your reinforcement learning basics, focusing on Markov decision process, and learn to convert a controlled task into an MDP to solve it with value functions, Q functions, and policy.
Explore the five core elements of control tasks in reinforcement learning: state, actions, rewards, agent, and environment, illustrated through chess, a robotic arm, and Pac-Man.
Define the Markov decision process as a discrete time, stochastic control process that has no memory, with state space, action space, rewards, and transition probabilities, to achieve goals.
Explain how the discount factor gamma discounts future rewards to favor quicker gains and maximize the long-term discounted sum of rewards in a maze control task.
Explore how an agent's policy maps states to actions, whether stochastic or deterministic, and how the optimal policy maximizes the discounted sum of rewards.
Define state value v(s) as the return from that state under a policy, and action value q(s,a) as the return after taking action a in state s.
Explore the Bellman equations for state value and action value, revealing their recursive structure through expected returns, rewards, and discounted future values under a policy.
Explain how to solve a Markov decision process by maximizing the expected return, defining the optimal policy and q-values through Belmont optimality equations.
Temporal difference methods learn from experience to update value estimates and guide policy, blending Monte Carlo and dynamic programming, with bootstrapping and generalized policy iteration.
Explore how temporal difference methods solve control tasks by estimating state-action values (Q), applying Bellman equations, and updating Q-values using the TD error and alpha-weighted rules.
Explore off-policy q-learning with two policies: a greedy target policy and an exploratory policy that collects experience, updating q-values to derive the optimal policy.
Explore how neural networks estimate value functions in reinforcement learning, bridging the shift from table-based methods to neural network approximations, and learn their structure and optimization for state-action values.
Explore function approximation as a scalable alternative to tabular methods for continuous state spaces. Build adaptive models using linear and polynomial forms with parameter vectors to fit value functions.
Explore artificial neurons as mathematical units that aggregate weighted inputs, apply activation functions like identity, rectified linear unit, and sigmoid, and propagate signals through hidden and output layers.
Represent a neural network in code as a three-layer model with three input dimensions, a six-neuron hidden layer, and two outputs to approximate the value function for each state-action pair.
Minimize cost function using stochastic gradient descent with rewards from the environment and neural network estimates. Update network parameters by stepping opposite the gradient with learning rate alpha via backprop.
Optimize a neural network to approximate Q values by tuning the W parameters to minimize mean squared error from environment samples, using reward plus discounted next Q value as target.
Apply neural network knowledge to augment Q-learning, creating deep learning to tackle more difficult problems and leverage neural networks' generalization to solve complex environments.
Master deep Q-learning by combining temporal difference with neural networks, using off-policy epsilon-greedy exploration, replay memory, and a target network to stabilize updates.
Store state transitions (state, action, reward, next state) in a replay memory. Sample a batch to compute the cost function and update the neural network.
Discover how PyTorch Lightning simplifies deep learning by providing the LightningModule and Trainer to automate training, with callbacks like early stopping and checkpointing, plus logging to TensorBoard for visualization.
Begin implementing your first deep learning algorithm with PyTorch Lightning, installing tools for environment rendering. Set up gym and a Lightning trainer to run on CPU or GPU.
design and implement a deep q-network in PyTorch to estimate q-values for each action from a given state. build a sequential network with input size, hidden units, and final layer.
Create an epsilon-greedy policy that maps environment states to actions using a neural network to estimate Q-values, selecting random actions with probability epsilon or the best action otherwise.
Implement a replay buffer to store environment observations with a fixed capacity using a deck, enabling len, append, and sample, and wrap it in a PyTorch Lightning dataset.
Create the lunar lander version two environment via a gym make call. Explore its eight observation features and four actions, render episodes, and record videos for analysis.
Define a deep q-learning class extending the lightning module, configure optimizers, set up a replay buffer and data loader, and train with an epsilon policy.
Define play_episode to sample data from the environment, reset the environment, and store (state, action, reward, done, next_state) tuples in the buffer using an epsilon-greedy policy with the Q network.
Implement the forward method to compute q-values from the environment state and configure AdamW optimizer with learning rate. Create datasets and data loaders to feed training samples into training step.
Define the train_step() to process a batch and align shapes for states, actions, rewards, and done signals using the target network. Compute state-action values with Q network; log Q error.
Implement the train_epoch_end method to sample data and update the target network at sync rate intervals; compute epsilon from start to end over epochs, then run episodes with that epsilon.
Set up visualization tools and train a deep q-learning agent on lunar lander version two with a trainer, early stopping, and a 400-step time limit.
Watch how the reinforcement learning Q-network learns to estimate action values and refine the policy, enabling the rocket to land between the flags after 13 minutes of training.
Optimize deep learning hyperparameters with Optuna through automated search. Define studies and trials, and use samplers and pruners to find learning rate, network size, and replay buffer capacity.
Learn to tune reinforcement learning hyperparameters with the Abdullah library by selecting gamma and learning rate, using moving averages of the last 100 episode returns to compare parameter sets.
Define the objective function to optimize learning rate and gamma via a trial, using a logarithmic scale and pruning callback in a lunar lander reinforcement learning setup with 1000 epochs.
Import the successive holding pruner from Abdullah's Pruners module and create the study to manage the hyperparameter search, maximizing the running average of returns while pruning unpromising trials.
analyze twenty deep learning runs, identify the best hyperparameters from the study results, and retrain with those values via keyword arguments to reproduce the top performance.
Address maximization bias in deep Q-learning by using a double deep learning approach that splits action selection and target evaluation between the main and target networks for more robust training.
Apply double dip learning to the double deep q-learning algorithm to stabilize training by having the main network select actions and the target network estimate their values, reducing maximization bias.
Explore the dueling deep q-network architecture that separates state value and action advantages to speed up reinforcement learning by focusing on when actions differ.
Implement a dueling DQN by building a value and advantage streams network that decomposes state value and action advantage, then combines them to compute Q-values.
Apply observation normalization to stabilize learning by converting state features to zero mean and unit variance using running means and variances. Normalize rewards to emphasize above-average returns and speed training.
Solve Flappy Bird Version zero with the Pay Game Learning Environment built on gym, converting image observations to a state vector via a wrapper that enables a two-action policy.
Create a Flappy Bird environment using the make function from the Bigham Learning Environment Library, then wrap it with normalize observation and normalize reward to stabilize learning.
Build a deep Q-learning agent by creating a Deep Learning class that extends a lightning module, wiring the network, policy, replay buffer, and training steps for Flappy Bird.
After training, test the resulting agent for ten episodes by rendering environment frames and using the trained policy to act, demonstrating fluent play of Flappy Bird.
Explore prioritized experience replay to bias sampling by TD error, adjust with alpha and beta, and apply importance sampling to speed up learning while correcting distribution bias.
Implement a prioritized expedient replay DQN to solve Flappy Bird with visuals. Build a two-part CNN-based Q network that extracts image features and estimates value and advantage to compute Q-values.
Implement and understand a prioritized experience replay buffer that uses alpha and beta to sample high-learning-potential experiences, update priorities, and return training batches with importance-sampling weights.
Create the Flappy Bird environment and normalize observations and rewards with wrappers. Apply max and skip, warp frames to 42 by 42 grayscale, and reorder axes for efficient input.
Implement the deep q-learning algorithm with prioritized experience replay inside a lightning module, updating replay buffer priorities from per-sample errors and scheduling alpha and beta.
Launch training process by configuring a deep learning agent for Flappy Bird version zero, set epsilon, gamma, and replay buffer, and run trainer across GPUs for 3000 epochs with debugging.
This is the most complete Advanced Reinforcement Learning course on Udemy. In it, you will learn to implement some of the most powerful Deep Reinforcement Learning algorithms in Python using PyTorch and PyTorch lightning. You will implement from scratch adaptive algorithms that solve control tasks based on experience. You will learn to combine these techniques with Neural Networks and Deep Learning methods to create adaptive Artificial Intelligence agents capable of solving decision-making tasks.
This course will introduce you to the state of the art in Reinforcement Learning techniques. It will also prepare you for the next courses in this series, where we will explore other advanced methods that excel in other types of task.
The course is focused on developing practical skills. Therefore, after learning the most important concepts of each family of methods, we will implement one or more of their algorithms in jupyter notebooks, from scratch.
Leveling modules:
- Refresher: The Markov decision process (MDP).
- Refresher: Q-Learning.
- Refresher: Brief introduction to Neural Networks.
- Refresher: Deep Q-Learning.
Advanced Reinforcement Learning:
- PyTorch Lightning.
- Hyperparameter tuning with Optuna.
- Reinforcement Learning with image inputs
- Double Deep Q-Learning
- Dueling Deep Q-Networks
- Prioritized Experience Replay (PER)
- Distributional Deep Q-Networks
- Noisy Deep Q-Networks
- N-step Deep Q-Learning
- Rainbow Deep Q-Learning