04
Tabular methods for RL CS@THU RL2025-Fall 4<br>
05
Function approximation for RL CS@THU RL2025-Fall Figure credit: David Silver, “Value Function Approximation” 5<br>
06
Supervised machine learning Classification
Discrete response Regression
Continuous response CS@THU RL2025-Fall 6<br>
07
Supervised machine learning Fundamental assumptions
Independent and identically distributed instances
Empirical risk minimization
Typically under a stationary distribution CS@THU RL2025-Fall Both are violated in an RL setting! 7<br>
08
Value function approximation CS@THU RL2025-Fall 8<br>
09
Value function approximation CS@THU RL2025-Fall Goal of value function learning? Guide policy optimization Mean square error is NOT necessarily the best objective for value function approximation, whose aim is to help policy optimization. 9<br>
10
Incremental update Update the approximation after every action taking
Any online optimization method would work
The objective function is not necessarily convex
Stochastic gradient descent is typically the choice CS@THU RL2025-Fall 10<br>
11
Incremental update CS@THU RL2025-Fall 11<br>
12
Incremental update CS@THU RL2025-Fall 12<br>
13
Linear models One of the most well-studied approximation methods
Guaranteed convergence under gradient descent
Needs an unbiased gradient estimation, e.g., MC methods CS@THU RL2025-Fall 13<br>
14
Linear models under semi-gradients Under TD(0) semi-gradient CS@THU RL2025-Fall When the estimation converges 14<br>
15
Recap: Bellman equation for value function CS@THU RL2025-Fall 15<br>
16
Recap: value function approximation CS@THU RL2025-Fall 16<br>
17
Linear models under semi-gradients Under TD semi-gradient CS@THU RL2025-Fall Approximation error obtained by TD Approximation error obtained by MC 17<br>
18
Incremental update CS@THU RL2025-Fall 18<br>
19
Many other different types of approximations Non-linear models
Regression trees
Neural networks, a.k.a, Deep RL
Kernel methods CS@THU RL2025-Fall Almost all methods we have learned in machine learning apply here! All aforementioned methods apply to action value function approximation! 19<br>
20
Batch update CS@THU RL2025-Fall Solution: from online update to batch update 20<br>
21
Experience replay CS@THU RL2025-Fall Figure credit: MC.AI Prioritize by TD error (proportionally) 21<br>
22
Prioritized experience replay CS@THU RL2025-Fall Figure credit: Schaul et al., “Prioritized experience replay”, ICLR’2016 Q-learning with tables Q-learning with linear approximation 22<br>
23
Control with value approximation Under generalized policy iteration CS@THU RL2025-Fall 23<br>
24
On-policy control with approximation CS@THU RL2025-Fall 24<br>
25
Sarsa with action value approximation CS@THU RL2025-Fall Three possible actions:
full throttle forward (+1)
full throttle reverse (-1)
zero throttle (0)
Reward is -1 on all time steps until reaching the goal 25<br>
26
Sarsa with action value approximation CS@THU RL2025-Fall 26<br>
27
Off-policy control with approximated function value Now we have shifted target value and target distribution CS@THU RL2025-Fall 27<br>
28
Target value correction Importance sampling
Importance weight
Gradient correction CS@THU RL2025-Fall The previous learnt off-policy methods, e.g., Q-learning, Treeback also apply for this purpose, as they directly provide us the target value. 28<br>
29
Challenges in off-policy control with approximation Behavior policy is unaware of the quality of approximation
A one-dimensional case CS@THU RL2025-Fall 29<br>
30
Baird’s counterexample CS@THU RL2025-Fall Reward is zero everywhere 30<br>
31
Baird’s counterexample CS@THU RL2025-Fall Deadly triad:
Function approximation
Bootstrapping
Off-policy learning 31<br>
32
Linear value-function geometry CS@THU RL2025-Fall What is a better objective for off-policy learning? Expectation of TD error! 32<br>
33
Takeaways Function approximation generalizes value estimation across states
On-policy prediction and control are typically performed as SGD for parameter estimation
Off-policy with value approximation is much more challenging CS@THU RL2025-Fall 33<br>
34
Suggested readings Chapter 9: On-policy Prediction with Approximation
Chapter 10: On-policy Control with Approximation
Chapter 11: Off-policy Methods with Approximation CS@THU RL2025-Fall 34<br>
35
Emphatic-TD methods CS@THU RL2025-Fall 35<br>
36
Emphatic-TD methods CS@THU RL2025-Fall Baird’s counterexample 36<br>