David Silver 强化学习 (知识点分解)

1912
21
2018-12-17 11:13:46
24
15
117
8
本专辑是David Silver 强化学习课程的个人剪辑笔记版本 想看原视频(同等清晰度)的小伙伴,可请前往https://www.bilibili.com/video/av24060851,或其他B站链接 大家可以对照问答文本,来有选择的学习不同知识点 https://zhuanlan.zhihu.com/c_180551461
As Simple As Possible
视频选集
(5/58)
1.1 admin textbook lecture notes
06:15
1.2 RL lays at the core of many sciences
03:01
1.3 what special about RL
02:54
1.4 RL examples and demos
09:37
1.5 define Reward in RL problem
03:44
1.6 reward examples
02:17
1.7 maximizing reward can be tricky
01:46
1.8 define agent and environment why reward scalar
04:20
1.9 define history and state in RL problem
03:26
1.10 what is environment state
03:42
1.11 what is agent state
01:32
1.12 what is information or markov state
05:57
1.13 example of agent state for prediction
03:21
1.14 fully observable state
01:13
1.15 partial observability
04:40
1.16 major components of RL agent
01:42
1.17 what is policy
01:09
1.18 value function and example
06:39
1.19 what are model transition reward
01:36
1.20 maze example on policy value function model
02:53
1.21 value based agent categorization
02:16
1.22 policy based actor critic agent categorization
01:18
1.23 model free and model based agents
01:46
1.24 differentiat learning and planning
02:56
1.25 atari example on planning and learning
02:01
1.26 balance exploitation and exploration
03:53
1.27 prediction and control with course outline
03:54
2.1 lecture 2 outline
00:48
2.2 MDP basics
02:42
2.3 what is a markov state
01:15
2.4 what is state transition matrix
01:55
2.5 formal definition of markov decison process
01:12
2.6 student markov chain example intro
02:01
2.7 markov chain episodes and transition matrix example
03:38
2.8 Markov reward process - define reward and goal
02:32
2.9 how reward and discount construct goal
03:24
2.10 why use discount to control short far sight interest
05:03
2.11 explain value function how state sample reward discount expectation const
04:13
2.12 how value function work in short long sight mode
01:54
2.13 understand bellman equation of MRP
04:44
2.14 explain bellman equation of MRP more
05:16
2.15 explain bellman equation in matrix operation
02:15
2.16 solve bellman equation
01:53
2.17 definition of MDP with action example
03:26
2.18 policy distribution in MDP
02:54
2.19 relation between MRP and MDP
01:53
2.20 state and action value functions
02:52
2.21 Bellman expectation equation on state and action
06:32
2.22 Example on Bellman expectation equation
02:46
2.23 solve Bellman expectation equation with matrix form
01:21
2.24 QA MRP and MDP clarification
04:01
2.25 explain optimal value function with example
07:48
2.26 what is optimal policy
04:14
2.27 how to find optimal policy
02:02
2.28 explain Bellman optimality equation
06:22
2.29 Bellman optimality equation example and solution
03:57
2.30 QA to intuition of Bellman optimality equation
08:42
2.31 extention to MDP
01:49
客服
顶部
赛事库 课堂 2021拜年纪