Decision Making and Reinforcement Learning - Georgia Tech - Machine Learning Udacity https://www.youtube.com/watch?v=ZM9M2FsBA9M Transkript (automatisch erstellt) 0:00 Okay Michael so before we get started diving into a particular formalism or framework that I want to 0:06 talk about, that we are going to use for this mini course or at least the first half or 0:09 so of it. I want to remind everybody what the differences are between the three types of learning 0:14 that we said we are going to look at over this entire course. Th, you recall what they are? 0:18 >> Well reading off the slides supervised, unsupervised and reinforcement. >> That's right. And you know, again these things are all strongly 0:25 related to one other but it's very useful to think about them as being very separate. So just as a reminder, supervised learning 0:31 sort of takes the form of function approximation where you're given a bunch of x, y pairs And your goal is to find 0:40 a function f that will map some new x to a proper y, you recall that? >> Yep. 0:46 >> Unsupervised learning is very similar to supervised learning except that it turns out that you're given a bunch of x's 0:53 and your goal is to find some f. That gives you a compact description of the set of x's that 1:00 you've seen. So we call this clustering, or description as opposed to function approximation, and finally we're getting to reinforcement 1:08 learning. Now reinforcement learning is actually a, a name that means many things in different fields, and here we tend 1:14 to talk about it in a relatively specific way, and superficially it looks a lot like Supervised learning, in 1:19 that we're going to be given a string of pairs of data, and we're going to try to learn some 1:26 functions. But in the function approximation case, a supervized learning case, we were given a bunch of X and Y pairs. We were asked to learn F, but 1:35 in reinforcement learning, we were given something totally different. Were instead going to be given x's and z's, and 1:42 I'll tell you what the x's and the z's stand for in a minute, and then were going to have to learn some f that 1:47 is going to generate y's, and so even though its going to look a lot like function approximation, its going to turn out that it's going to 1:53 have a very different character. And that's really what the, the next few slides in a little bit of discussion is really about, is 2:00 understanding what that character is. You'll also notice from the title here that I have decision making reinforcement 2:04 learning, and that's because reinforcement learning is one mechanism for doing decision making. And again, I'll define 2:10 that in just a second. Okay, so you with me, Michael? >> I think so. So should y be circled? 2:15 because in some sense, you, you underlined the things we were given and you circled the things we needed to 2:19 find, so y is something that we're going to find, right? >> Yeah, I suppose that I like that. 2:22 >> All right. Interesting. >> Now it's a theta. 2:25 >> [LAUGH] It's a Y trapped inside of a theta and it's yelling, Y? >> [LAUGH] I like that. I'm a Y 2:32 trapped in a theta. Hm, I need to write a book about that. Okay, good. That's a good point Michael. So before 2:40 we were learning, effectively trying to learn one thing and here we're still learning one thing. Because it's going to produce another thing 2:46 for the deterministically, usually. But it's worth pointing out that we are going to be figuring how to produce both of these things 2:52 as opposed to be given those things. Good job. OK. Are you ready to move forward,l Michael, with an example and a quiz? 2:56 >> Awesome. >> Excellent.