Das Video kommt von YouTube: erst beim Abspielen verbindet sich die Seite mit YouTube (Google).
Decision Making and Reinforcement Learning - Georgia Tech - Machine Learning
Das Wichtigste aus dem Video
Tipp auf eine Zeit – das Video springt genau dorthin.
Transkriptautomatisch erstellt · 31 Zeilen
- Okay Michael so before we get started diving into a particular formalism or framework that I want to
- talk about, that we are going to use for this mini course or at least the first half or
- so of it. I want to remind everybody what the differences are between the three types of learning
- that we said we are going to look at over this entire course. Th, you recall what they are?
- >> Well reading off the slides supervised, unsupervised and reinforcement. >> That's right. And you know, again these things are all strongly
- related to one other but it's very useful to think about them as being very separate. So just as a reminder, supervised learning
- sort of takes the form of function approximation where you're given a bunch of x, y pairs And your goal is to find
- a function f that will map some new x to a proper y, you recall that? >> Yep.
- >> Unsupervised learning is very similar to supervised learning except that it turns out that you're given a bunch of x's
- and your goal is to find some f. That gives you a compact description of the set of x's that
- you've seen. So we call this clustering, or description as opposed to function approximation, and finally we're getting to reinforcement
- learning. Now reinforcement learning is actually a, a name that means many things in different fields, and here we tend
- to talk about it in a relatively specific way, and superficially it looks a lot like Supervised learning, in
- that we're going to be given a string of pairs of data, and we're going to try to learn some
- functions. But in the function approximation case, a supervized learning case, we were given a bunch of X and Y pairs. We were asked to learn F, but
- in reinforcement learning, we were given something totally different. Were instead going to be given x's and z's, and
- I'll tell you what the x's and the z's stand for in a minute, and then were going to have to learn some f that
- is going to generate y's, and so even though its going to look a lot like function approximation, its going to turn out that it's going to
- have a very different character. And that's really what the, the next few slides in a little bit of discussion is really about, is
- understanding what that character is. You'll also notice from the title here that I have decision making reinforcement
- learning, and that's because reinforcement learning is one mechanism for doing decision making. And again, I'll define
- that in just a second. Okay, so you with me, Michael? >> I think so. So should y be circled?
- because in some sense, you, you underlined the things we were given and you circled the things we needed to
- find, so y is something that we're going to find, right? >> Yeah, I suppose that I like that.
- >> All right. Interesting. >> Now it's a theta.
- >> [LAUGH] It's a Y trapped inside of a theta and it's yelling, Y? >> [LAUGH] I like that. I'm a Y
- trapped in a theta. Hm, I need to write a book about that. Okay, good. That's a good point Michael. So before
- we were learning, effectively trying to learn one thing and here we're still learning one thing. Because it's going to produce another thing
- for the deterministically, usually. But it's worth pointing out that we are going to be figuring how to produce both of these things
- as opposed to be given those things. Good job. OK. Are you ready to move forward,l Michael, with an example and a quiz?
- >> Awesome. >> Excellent.
Zum Nachlesen
Maschinelles LernenMaschinelles Lernen (ML) entwickelt, untersucht und verwendet statistische Algorithmen, auch Lernalgorithmen genannt. Solche Algorithmen können lernen, …
Bestärkendes LernenBestärkendes Lernen. Reihe von Methoden des maschinellen Lernens, bei denen ein Agent selbständig eine Strategie erlernt, um erhaltene Belohnungen zu maximieren.
KlassifikationsverfahrenDas Erzeugen von Strukturen aus vorhandenen Daten wird auch als Mustererkennung, Diskriminierung oder überwachtes Lernen bezeichnet. Dabei werden …