WEBVTT

0:00:00.240 --> 0:00:22.720
<v A>Foreign programming throwdown episode 180 reinforcement learning. Take it away, Patrick.

0:00:23.200 --> 0:00:37.310
<v B>Welcome to another episode. This is going to be a good one. Excited to be here, actually, because this is a topic I have been meaning to learn about. And Jason has agreed to be put on his professor hat robe. I don't know. What does a professor wear?

0:00:37.390 --> 0:00:43.790
<v A>I got, I got hooded. When I got the PhD, I got hooded, which I thought would be an actual hood, but it's really just a sash.

0:00:44.590 --> 0:00:49.230
<v B>Wait, what is getting hooded? That's like what you get when you get. I don't know about this.

0:00:49.230 --> 0:01:10.720
<v A>Okay, so when you get a PhD, you get hooded, which means you go through the same ceremony as the master's students, or I think the same ceremony as everybody, but you get a hood, which is actually a sash, and the, your, your PhD advisor actually puts the sash around you over you as part of the ceremony.

0:01:11.040 --> 0:01:23.760
<v B>Okay, I, I, I feel like maybe I've heard that term, but I always just kind of had some weird, probably bad association with hoodwinked. But anyways. Okay, where are we? Off topic.

0:01:23.760 --> 0:01:33.260
<v A>Anyway, so it's funny because I actually, I, actually the first thing I think of is, is, is actually because I grew up cities, I was like, okay, we're going back to my childhood here.

0:01:33.740 --> 0:01:59.420
<v B>Oh, oh, okay. Interesting. Okay. Wow. All right, so today we've learned there's many, many associations of the word hood. So, okay, we didn't even talk about cars yet. So that's true. You get a new upgraded carbon fiber hood for your car. You get hooded. Okay, why? People are like, what is going on? What are we listening to? Well, that's how you know we're not the AI. They would not be this off topic.

0:01:59.720 --> 0:02:00.120
<v A>That's true.

0:02:00.680 --> 0:02:02.040
<v B>Definitely stick to the script.

0:02:02.040 --> 0:02:03.960
<v A>The AI is not allowed to say this stuff.

0:02:05.400 --> 0:03:38.670
<v B>You would definitely be pushing the, the down vote button, you know? Oh, wait, people probably are now. That's why this is not live. Okay, so for, and I'll keep it brief because I actually want to get to the, the, the meat of the story today. Oh, no pun intended. But talking about cooking outside, I had a grill on my, like, back patio, which I would use to cook, cook food. Occasionally it was uses these little pellets of wood. So it's called like a pellet grill. So like pellets feed down and it burns it and it makes the heat and has an electronic controller. And I would do some like, you know, smoking on it and some, some grilling. Anyways, it broke and it's, it's old, so you Know, okay, fine. So I, I went to go like, okay, I can get a new grill. I did not know. There are so many different kinds of grills that are, you know, like, popular now. And I feel like growing up, my parents always just had. Yours were kind of one of two things. You had the charcoal Weber grill, you know, with the, like the bowl and the charcoals, or you had the propane grill that, you know, like you had the tank and you hooked up the hose and, and that was a two. But now there are like, you know, all sorts of things where it's, you know, infrared cookers where the propane goes into like a, some sort of like catalyst something and like turns hot. Like the patio heaters that, you know, are at restaurants sometimes. Yeah, it's like, I, I don't know. And then there's these like, egg shaped, I guess they call them komodo grills, like big green egg and komodo and they're like big ceramic things. And then you can get like various kinds of cabinet smoke.

0:03:38.670 --> 0:03:38.950
<v A>Like.

0:03:39.270 --> 0:03:58.010
<v B>Anyways, I, I just, maybe I'm naive in my. I just bought something straightforward and simple and then I like, oh, I'm overwhelmed by the tyranny of choice. That, that's just all I was going to say. If you've never kind of looked up grill technology, it's. It's actually kind of crazy. There's got a lot of choices here. So now I don't know what to pick.

0:03:58.330 --> 0:04:08.010
<v A>That's wild. I have a propane grill and then I have a smoker. I have a separate electric smoker that takes the pellets and smokes meat.

0:04:08.650 --> 0:04:09.170
<v B>Okay.

0:04:09.170 --> 0:04:14.110
<v A>But yeah, I think the green egg can do both. So it's like a two in one. And yeah, it's just wild.

0:04:14.430 --> 0:04:23.710
<v B>And then there's like pizza ovens now people are doing. So like, when I went to go look at the grills at the hardware store, it was like, there are also pizza ovens here. And yeah, I.

0:04:23.870 --> 0:04:41.290
<v A>Okay, yeah, my neighbor has a pizza oven and I think he's used it twice in four years. Well, I mean, you know, how many times do you eat pizza? I mean, and also it's like, if you eat pizza, you're often in a hurry, so you're either ordering it to go or you're doing the, the regular oven because you're in a hurry.

0:04:42.010 --> 0:04:47.930
<v B>Okay. All right, so. So you're down on the pizza oven. That's a short, short on the pizza oven stocks for, for Jason.

0:04:48.410 --> 0:05:09.850
<v A>Yeah, I mean, I'm not, not a big Fan of the pizza oven. I think, you know, we've only, I've only ever endorsed one stock on my entire programming throwdown career, and that was Datadog. And I think it's like up just as much as everything else. I made one like, like in hindsight, in five years of hindsight, like relatively neutral endorsement.

0:05:10.250 --> 0:05:12.970
<v B>Everyone's bracing themselves for the meme coin announcement now.

0:05:13.770 --> 0:05:15.690
<v A>Oh, we need a programming throwdown coin.

0:05:15.690 --> 0:05:16.610
<v B>No, we don't. I'm not.

0:05:16.610 --> 0:05:18.010
<v A>No, I'm not rug pulling people.

0:05:19.130 --> 0:05:21.530
<v B>Okay? I know. Okay, we gotta keep going. We gotta keep.

0:05:21.530 --> 0:05:28.010
<v A>We are not, we are not rug pulling people. We went the opposite way. We stopped doing ads. So it's like the opposite of rug pulling people.

0:05:29.130 --> 0:09:13.140
<v B>All right, time for, for news of the show. So I've got the first one. And this is an article entitled you can't call yourself a senior until you've worked on a legacy project. So talking about what is a senior engineer? This is like a age old debate. Whatever. Anyways, this person was kind of pointing out how they hadn't really worked on legacy code base. There's some specifics here of their, their thing and, and I, you know, if you want to go read it. Good article. And uh, the point though is, is pretty interesting that they kind of rightly wanted to avoid working on a legacy code base. They ended up kind of doing it. They were right, it didn't like it, but they actually learned a bunch of stuff that that didn't. And I, I think a couple interesting takeaways for me from the story and just, you know, thinking on the topic is about regardless of the label of senior, like just growing as an engineer is no matter what work you're doing, finding the takeaways that are applicable and, and lots of analogies. The one I've taken to using recently just for myself and for others that I talk to about this is like just really trying to compound the growth. So not just thinking like, hey, how do I do this thing? But like, how do I think about additively, like applying things I've learned before in a way that like my growth sort of grows on top of itself and you're sort of stacking it up and sometimes you need to widen the base right, of like expanding into new things. But other times you're trying to build up and trying to apply these different experiences. And so I think partly this plays into that. And then there's an observation here specifically about legacy code bases in your place of work and understanding why maybe something isn't done that way anymore. Or why the stuff that you see in like the, the kind of current pieces of tech are how they got there, right? People will say, oh, it's organic growth or you know, whatever. You kind of get there. But I think there is something different between saying this is the current recommended practice and I have done it the other way. And I will tell you the other way sucks. Like we're doing it this way. Those two things come from slightly different places. And understanding why you do something, not just there is value in actually just not knowing, oh, this is kind of bad. But when you get into style guidelines and stuff, right? I think just picking away and having everyone do it is useful because it really does matter. But then there are things also that even if you don't always know exactly why, eventually kind of figuring them out and digging in. So the one I always use in C is the ternary operator. So you can write this Boolean expression, put the question mark and then the thing that is true first and then a colon and then thing that is false if it's false second. And you can use this. And we have it banned in our code base. And the reason why is it. It. It literally does nothing unique. You can, okay, there's like some very rare, you know, use case someone could come up with that, that you know, in a const expression or something. But for most part you're just simplifying, writing an if else statement. But the cognitive load to read the ternary operator, make sure you understand what it does and that a new engineer showing up has the same practiced expertise at reading that. Like why? Like why? Just because features are there doesn't mean you have to use them. And I try to explain this to people, but I would argue even not knowing my explanation and still doing it leads to good practices. But knowing why and having tried using all of the whiz bang features from the latest, you know, C update constantly and refactoring code just to rewrite it into those features. Having done that once and burning your hands probably teaches some lessons. And so legacy code bases is. Can be really useful.

0:09:13.300 --> 0:10:42.270
<v A>Totally, totally agree. I mean the equivalent in Python, which is even more confusing, they have a ternary operator where you can say like x equals 3 if foo is true, else 5. So it's a ternary operator, but that you switch the first and the second like position. So it's like, so it's like even harder to read. And you know. Oh yeah. And so I remember like there was someone on my team who would do this a lot. Like all over the place. And you know, I let it go. I didn't really push back on it because to your point, like, it's. It's not. Until you ban it, it's not banned. And so you can't really say, like, don't do this because you have no, like, moral grounds other than your intuition. And then it was a disaster. And so, like, people just kept getting burned by these really long in line, you know, like, conditions. Right. And so now, like, I can ban it and I don't feel insecure about it or feel like hesitant about it. I. I don't feel like, oh, you know, it's not really against the rules because now it's like, no, like, I've done this. I've seen people like, cause all sorts of issues. And some issues went to prod. And so we're not doing it now that the. One of the tough things is, you know, if you're talking to folks who don't have that experience, you have to like, ban it in a way that is shows empathy and doesn't like, create any resentment or anything.

0:10:42.590 --> 0:11:26.700
<v B>And there is this balance that I think is useful but often gets brushed away that bringing new folks and you actually want them to feel empowered to question. So when they see that there's a ban on this and they say, I love the ternary operator because it makes me look cool, you know, they're going to say that part, but, you know, I love the ternary operator. You know, why is it banned? I don't think it should be banned. And you actually want to take time to explain them and in some cases be willing to hear them out and maybe, you know, adapt your practice or be flexible. But in other times, like you said, I think the word confidence there is like, no, we've done this. Like, I hear you, but you're just gonna have to trust me that like, we've tried it the other way and the other way, like, not banning it leads to problems.

0:11:27.750 --> 0:12:25.380
<v A>Yeah, yeah, totally. Yeah. I mean, I just to wrap this up, a question I always ask in interviews. If I'm doing a technical design interview, I'll always start with the question of, like, tell me a time you refactored something and why did you refactor it? Like, what led to the decision to do a big refactor? And that usually, like, opens up all sorts of interesting things because, you know, people, you know, the, the worst answer is the one that totally neglects the conflict that comes from scarcity. Right? It's like, you don't have enough time but the code is garbage. And it's like. So it says, like, that creates conflict and then you have to resolve that conflict one way or the other. That's interesting. If someone's like, yeah, you know, I rewrote it and it was the right thing to do and everyone agreed with me from day zero to day, the day I wrote it and everyone praised me at the end, it's like, okay, well, you know, that's a little unrealistic.

0:12:26.740 --> 0:12:32.380
<v B>I was going to ask you, does anyone ever tell you. Because I didn't write the code, so therefore it could of course be better.

0:12:32.380 --> 0:12:36.100
<v A>I've definitely. I've got that. Yeah. I've had people.

0:12:36.180 --> 0:12:39.220
<v B>That's how people really do. But I would be surprised if someone said it.

0:12:39.220 --> 0:13:08.250
<v A>Yeah. People say, like, oh, it was, you know, another team and that team, you know, we inherited their code and it was garbage. And so I rewrote it all and it's like, okay, not, not the best answer. My new story is Recraft might be the most powerful AI image platform I've ever used. Here's why. And it's a Tom's Guide article. Honestly, like, Recraft is definitely the most powerful AI image system I've ever used. I found out about it yesterday, so.

0:13:08.250 --> 0:13:09.370
<v B>I haven't ever heard about it.

0:13:09.530 --> 0:13:52.200
<v A>Yeah, this is very. Obscure isn't the right word. But like, like, I'm really into this stuff, like generative AI. I'm following it closely and I hadn't heard about it until y. It can do some amazing things. It, for one, it can produce vector art, like SVGs. Now the SVGs are like, if you or I were to create a stop sign, for example, we'd create like a white octagon and then we create a red octagon inside the white octagon. And that's how we'd get a stop sign with a border around it. Right. But if you use this program, it's going to give you like a red octagon and then a bunch of white polygons around it. You see what I'm saying? Like, it doesn't have a concept of layer.

0:13:52.520 --> 0:13:53.880
<v B>Like each edge would be.

0:13:53.880 --> 0:15:21.690
<v A>Yeah, yeah, yeah. It's all one layer. So that's not ideal, but it's a step in the right direction. It's the only thing I've ever seen that really will give you an svg. Like, even if they're doing it post hoc or something, it's extremely, like, responsive to the prompt. So, for example, one thing I've. I've. I've tried with a lot of these AI systems is I've said like a person holding nothing in their hand. Because like I'll, I'll, I'll say like a person doing X, Y, Z and I'll get a person like holding a phone in their hand and I'll like. And I'll type the same prompt like with nothing in their hand. And then they'll have two phones in their hand. And then it's like, it's like it has a hard time, especially with negatives. So then I'll try like a person unarmed, you know, so it's like there's not like a negative there and it'll still like not work. But with this system it's like very good at following instructions, even negatives. So it's phenomenal. The other part of it is it's got this cool workflow where you can take an image and then you can say, okay, now put a phone in their hand and it'll make like a new image of the same person with a phone in their hand as opposed to like, you know, getting a totally different person. So it's really cool, very, you know, easy to use, relatively cheap. So I would highly recommend folks check it out. It's pretty neat.

0:15:22.810 --> 0:15:35.490
<v B>What are they using? Kind of under the hood. Do you know? Like is. Are they using their own. So some of these craft up other stable diffusion, whatever, you know. Is this like a layer on top of stuff? It sounds really good, but yeah, this.

0:15:35.490 --> 0:16:13.450
<v A>Is a totally custom thing. Okay. It's a pretty big model. It's a 20 billion actually. The 20 billion is the V2. There's a V3 model which I think is even bigger. They, they're not open source, so we don't really know what they're doing. They might have a blog post about it. I haven't seen one yet. I'm assuming it's the same type of technology where you're doing like self. Self attention and then you're doing like, you know, masking and trying to uncover masks and whatever. Like the, the. I don't think they're really pushing the envelope on the, the base model, but then they built a bunch of really impressive things on top of it.

0:16:14.570 --> 0:16:35.260
<v B>That's awesome. I think that's one of the debates, like where's the magic? Is it in the UI and the like, you know, higher level abstractions or is it in the, the base model? The, the deep stuff, you know, I, I don't know that we have an answer. I think there's lots of opinions But I will say, I would say that I tried the Dolly. When was Dolly popular?

0:16:35.260 --> 0:16:37.060
<v A>Oh that's, that's a four year and a half ago.

0:16:37.060 --> 0:16:38.180
<v B>Two years ago? No. That long?

0:16:38.420 --> 0:16:40.180
<v A>I think so the first Dolly?

0:16:40.260 --> 0:16:44.260
<v B>No. I don't know. Like whenever it was making the big rounds, I feel like it was maybe two, three years ago.

0:16:44.420 --> 0:16:45.180
<v A>Okay, okay.

0:16:45.180 --> 0:17:40.090
<v B>Maybe four years ago. I. Anyways, I recently did download on. I have a MacBook Air and I downloaded one of the like on device Stable Diffusion to run Flux. And they just have like an app you can download so you don't have to do it the command line with Olama or something. And you just like download the app and it will download the model from I think a hugging face link or something and it downloads the flux and you'll generate it. And it takes, I don't know, it's like 15 to 20 seconds or something to generate an image. But it's crazy that like first of this images are much better than when, when Dolly was making the rounds initially you kind of wrote it off like it didn't really obey your prompts. It would make cool pictures. But like anyways to now the Flux stuff and it runs like on my computer. And it's free. Like the models are open source, the program's free. So it's, it's running locally. There's no subscription, you know. And it, you know, obviously it has to be using less power because it doesn't have an external GPU or anything.

0:17:41.050 --> 0:18:12.250
<v A>Yeah, the Flux models are super, super impressive. There's a, there's a package called M Flux which will. Which is Flux optimized for mps, optimized for the Apple processor. So you can, you can use mflux and it'll run in like half the time or a third of the time or something. It's really impressive. And this Recraft thing really takes it to another level. So to your point, it'll only be a matter of time before there's an open source version of Recraft. But at the moment they, they have a monopoly.

0:18:12.970 --> 0:18:16.010
<v B>Well, that's another debate. The open source was closed, but we'll just keep moving on.

0:18:16.170 --> 0:18:19.930
<v A>Yeah, right. Open. Open. Wait. Yeah. Oh.

0:18:19.930 --> 0:18:21.690
<v B>Oh, yeah. So. Okay. Oh no, no, no.

0:18:21.690 --> 0:18:22.010
<v A>Okay.

0:18:22.010 --> 0:20:08.230
<v B>No. My Next one is NASA has a list of 10 rules for software development and this person is taking the sort of like publicly disclosed list of software rules that I've bumped into before being an embedded engineer previously in my career that for writing C code, but they've tried to extend it to C as well. There's like a set of embedded guidelines for not doing. And this individual is taking sort of, I'll say a kind of critique of some of them and, and why maybe they don't make, you know, much sense or other stuff, but if you've never seen them before, I will say it is somewhat interesting to. And it's a lot harder if you use something like Python or Java to I guess derive value maybe from it. But if you've ever programmed in C or C before, or you use Rust probably applicable as well, or one of the other sort of systems programming languages, go kind of looking through there and seeing like, how would you approach if these were your rule sets? Leaving aside, I guess this, the blog post is kind of talking about how maybe the rule set could be improved or doesn't make the most sense. But also if someone, if you showed up on a job and this was the restrictions because it was a contract that you were trying to honor, like, how would you accomplish this? So things like never using dynamic memory allocation and then, you know, there's kind of two approaches you end up with. One is so like in C, the standard library generally just uses lots of allocation under the hood. So you end up with, you know, of course you can't use that. So some people do a lot of like static sizing of things up front and trying to, you know, have all of their things of, of a known size. Other people, isn't there like a concept.

0:20:08.230 --> 0:20:13.380
<v A>In C called an arena or like you define a thousand spots? Yeah. Okay.

0:20:13.380 --> 0:21:32.560
<v B>Yeah. So then other people use a memory pool. An arena is like a kind of memory pool where they basically write their own sort of very thin sort of memory management, but it doesn't necessarily suffer from some of the same problems. And so you can use containers that adapt to that. But if you think in your head, how would you make sure that your code never did. Some of these things never had a loop that couldn't exit. Right. So all loops need to have an upper bound and just code coverage. How would that work? These kinds of things, some of them are again, like pretty restrictive. Like every function has to be smaller than can fit on a single piece of paper with, you know, sizing given. And it's like, well, that is probably good practice, but maybe, you know, sometimes, you know, you, you want to change it to be one thing, thing or another. But it is worth reading if you've never read an embedded like rule set like this before. They're not uncommon. And you can occasionally, if you work in embedded space, bump into places where this is the stuff. There has been an explosion in sort of like processing power and real time operating systems and the. Just the complexity and abilities of the processors. But I still think there are some places where there are probably many lines of code being written having to follow these guidelines.

0:21:33.540 --> 0:21:52.300
<v A>Yeah, this is fascinating. I mean, this is a whole new universe for me. But this is really interesting. I mean, this is definitely a person who is kind of like. I think the general critique here is on C. Like this person's like, I draw that you should use another language. Yeah.

0:21:52.300 --> 0:21:57.380
<v B>Like use ada, not C. This is just a valid commentary, but yeah.

0:21:59.390 --> 0:23:03.450
<v A>All right, so my next News story is AMD Radeon RX 9070 XT performance estimates leaked. Okay, so I want to go into a little rant here. I hate kind of complaining about products. Like, I feel like that's maybe like not the best use of the show, but I bought a PC, like a mini PC with a Radeon and I used it for a little while. It was okay. The drivers were really buggy. I had to go into safe mode and some stuff to like get it working. But then I got it working. You know how like you can plug like USB to DisplayPort or USB to HDMI? Like you have these cables. I don't actually know how they work under the hood, but there's some magic that allows you to go from a USB C port to right into your display. Right. I think it's like something called like DisplayPort pass through or something. Anyways, I plug one of these cables in, pop, the GPU blows up. Like hardware blow up dead.

0:23:03.610 --> 0:23:04.970
<v B>Where did the cable come from?

0:23:06.170 --> 0:25:02.010
<v A>It's a cable that I use in my MacBook Pro. Like I've used my. I've used this cable for years. Okay, so is the cable is fine, at least for the MacBook. Yeah. Plugged in his mini PC and it just popped. And I guess like, I mean, I'm kind of dogpiling here. I almost feel bad. But like, you know, George Hots has this post on Twitter about this. Like he's trying to build these tiny boxes that use, like that run, that run Pytorch and they use Cuda or, or, or Rockham, which is AMD's equivalent. But like the AMD, like, you know, anything above the hardware just sucks. And it's kind of disappointing because everyone wants there to be an alternative, if nothing else, not only for price, but maybe they could do something interesting. Like maybe they could make a card that has like a gigabyte of RAM and is not very fast, but just has a ton of ram. Like, you know, when there's multiple people in the pond, like there's multiple ideas. Right. But I was just super disappointed. I mean, I'm one of like a long line of people now who will just not buy these AMD cards. And I guess maybe just to turn this into a question like, like, how does AMD kind of like recover from this? I, I kind of feel like if I was them, I would hire like some software person, like some person who's really high up in the stack to like lead like a whole branch of the company to just like go through and, and, and just ruggedize everything from a software perspective. So, so it's like I feel like this is more in your area because like, you know, when it popped I was just shell shocked. I just, you know. So like what causes you know, hardware to kind of fail like that, you know, and not be tested? And what do you think you would do like if you're CEO of amd, Patrick, how would you save this?

0:25:02.510 --> 0:27:39.750
<v B>Oh my gosh. I, you know, as much as has been in the news with Nvidia GPUs and stuff and the scarcity and the crypto stuff and now the AI stuff, I don't know, I'm not super up on like the, where the profit margins come from. I know AMD of course like makes processors as well and I actually have AMD processor and GPU and the PC I built and I feel like I got, you know, the Ryzen processor was like a good value for dollar over the intel one at the time. Maybe, you know, they used to be two separate companies. Now like they're combined. I, I don't know how internally their company is structured where the profit margins are on gpu. Like you said. I, you may say, oh, there's a, you know, large group of people who would really buy a, you know, large amount of memory, not that much, you know, processing power. But it's possible that like the D cost and the spin up for that, especially when, you know, Nvidia could basically pay premium for any foundry costs because they have, you know, far more supply than demand. If you're sitting in second place, I don't know, you may be forced to basically pay more for foundry costs and things in order to be able to like get your chips made right. So it may make it difficult. They may not have as much freedom as, you know, they otherwise would want. And I think the software drivers are hard as well because there's so many people. You need to appease, right? You need to appease like people like you're saying I just want to plug it in and get a monitor working. But then the video game people are like, I want variable frame rates. And then you need to, you know, appeal to the video game developers who are, you know, needing to optimize the outputs on your card and the, you know, wrappers that sit on top of it. So OpenGL or DirectX or whatever. Like there's all this different stuff swirling around and I actually just feel like this, I don't even like that space just seems so complicated that you know, I get to plug into a variety of motherboards with a variety of power situations, a variety of connectors, a variety of like the amount of compatibility you need on a gpu, almost honestly like on par, probably more with than almost any other part of the PC is, is just actually bonkers. That needs to be compatible with all manner of software, all manner of, you know, oss, you know, all manner of hardware internal to the computer, external to the computer. That, that's a lot. Maybe that stuff's all really robust and ruggedized, but I imagine just a ton of time to get, to get nailed down. Exactly right.

0:27:40.870 --> 0:28:27.300
<v A>Yeah, that's a really good call out. I wonder. And also like a lot of these things are very thankless. You know, it's like, like the guy who makes sure that you could plug the USBC to display port versus the. Because, because I, I had this machine working DisplayPort to display port so I know that the machine worked. But then as soon as I plugged in this other way, like I heard a visible, like I heard an audio kind of pop and it was done. So like you have to have a person to like test all these different ways and then like unless it breaks, that person's not adding any value. Like they're just reducing risk. And you don't know what's risky, what's not risky. So it's one of these things, it's almost like a really high level kind of performance management like, like value of the company kind of thing that you have to fix.

0:28:28.100 --> 0:29:39.040
<v B>And there's, and, and then on your specific issue, I mean it could be a faulty card that you know, hopefully they would warranty replace. And then the question if you did it again with the same cable with the new card, would it happen again? I don't know. I mean I would if I were you. I wouldn't risk it. But like, yeah, there's two failure modes there, there's a failure mode of an individual card. And then there's a design flaw that, like, every card where you did that, it would, it would sort of break. And I don't, you know, from where I am, I just don't know enough about which specific problem it is. But certainly, like you said, it makes people have very bad sentiment. But even if you look at the. And again, not knocking, but like, the Nvidia consumer cards that were having, like, the plug you plug into the GPU to give it extra power directly from your power supply had so much power going through it that they were like, melting. They need to, like, have cooling on it. Like, again, like you said, this is sort of thankless, right? The dude or dudette who is trying to, like, design the interface for some wire, copper cable to come in and deliver power. And all of a sudden they never. Anyone ever thought about them in their whole entire lives. And now they're like, you know, front page of social media because these expensive graphics cards are melting.

0:29:39.600 --> 0:30:45.810
<v A>You know, this is going to be a distraction within a distraction. But, like, one thing that's really interesting is, like, what jobs are of that sort where, like, if you do a good job, nobody cares. Or, like, if you do a good job, nobody notices. And what are the jobs? Because it feels to me like in every career there exists, like, maybe that's. Maybe I'm. I'm trying to stretch too far here. In many careers, there exists, like, jobs where you get praised for doing a good job and nothing happens. If you do nothing, it's like, like jobs where you're trying to, like, bring in more business or raise the profit of a company or something. And, and then there's. There's jobs where, like, you're trying to keep the lights on. And so people are kind of. It's really hard for people not to ignore you until something goes wrong. It feels like there's these two kinds of jobs, and it feels to me like the former is almost always, like, better than, than the latter in terms of, like, your satisfaction and, and a lot of other things.

0:30:46.690 --> 0:31:11.690
<v B>I was having this conversation and I won't, I won't give the details because it'll, it'll make it sound like a political statement one way or another. And we're not trying to make it. But anything where you're. And in this case it was, you know, government officials, whatever, but anything you're dealing with a probabilistic event. So something may happen, may not happen. Even if it happens, the certainty isn't well known, right? You think, like, Weather forecasting, you know, whatever, like.

0:31:12.250 --> 0:31:12.810
<v A>Yeah.

0:31:12.970 --> 0:32:14.130
<v B>And when you get it right, it, it's sort of like no one, you, you could just do nothing, right? And it would probably just be fine. And then that one time, you know, it's, it's out of standard deviations, like very high. And then everybody's like, why didn't you do. And it's like, well, it's true. Probably could have done better or done these other things, but you could have been doing those other things every single time and it either still wouldn't have made a difference or wouldn't have been relevant. And so to your point, anytime you're dealing with something where, you know, like, think of preparing for Holiday rush on a server that's doing E commerce, right? You could spend tons of money, like building up, you know, extra CDNs and having flexible compute. And then, you know, like, people come, they shop, and there's no outages. No one knows if you did a great job or a bad job. But it didn't crash. So, like, there's that. But if it crashes, certainly you're getting hauled in and told how much millions of dollars you lost. The, the website.

0:32:14.210 --> 0:32:22.830
<v A>Right? Yeah. I mean, it's pretty tragic that, you know, things are just set up that way, but I don't know if there's not, you know, any clear solution or anything.

0:32:24.750 --> 0:32:26.430
<v B>All right, well, book of the show.

0:32:27.390 --> 0:32:28.910
<v A>Book of the show. What's your book?

0:32:29.150 --> 0:32:56.730
<v B>All right, my book. This is going to be a, a little bit different, but I've not been reading as much as I, I, I should or I want to. But I did just start because I have never read any books by this author before and often see it recommended in science fiction. So I decided I'm going to read a book and tried to find the recommendation, and this is the recommendation I got. And that is a book by Ian M. Banks, and I chose the Player of Games. Have you read any Ian Mbanks books before?

0:32:56.730 --> 0:32:57.290
<v A>I have not.

0:32:57.530 --> 0:33:34.980
<v B>Okay. Yeah, neither have I. So I am trying to embark on reading this one. Apparently some of the books can be a little hard to read, which is okay. But this one, people are saying is a, is a good introduction. There's not from what I've seen online without trying to read any spoilers, apparently there's not like a strong order you need to read the book since it's not necessarily the first book, but it is an often recommended one. So I'm starting here. That's not really a good recommendation because I don't know whether to tell you, it's good or bad. Other than every other person I saw on the Internet, this seemed to be bubbling to the top as a good starting place. So I'm embarking on a journey here.

0:33:35.380 --> 0:36:44.640
<v A>Cool, that sounds awesome. Yeah, I might check that out. I have a few books queued up that I have to get to and then I'll check that out. My book of the show is Basic Role Playing Universal Game Engine, which is a reference book. It's written by a couple of people who have been making tabletop, tabletop games for their whole adult lives. And they've made a bunch of them, I think some of them even from like the 70s and the 80s. So it's kind of wild. I mean some of the stories of the, of the creators, but, but basically they synthesized. So, so there's this question in general, like, think about any art form. You always think about, like, what is the essence of this art? You know, like you might as an artist, like draw a bunch of things and then think to yourself, okay, what is like the essence? Like, if I had to reduce something to its like most basic form, what would that be? And so these guys got together and thought, well, if we had to reduce these tabletop games because there's a lot of like lore built into it. You know, like so many games have magic missile. Why? Because Dungeons and Dragons 1.0 had Magic Missile. But like, what really is a magic missile? I guess it's like an arrow made out of magic, right? Or something. But like, you know, pair these things down to their essence and, and just explain like it literally like the first, the first page of the book is like the point of an RPG is for your players to have fun. So it's like, you know, it's like, let's start from like the first principle. The first principle is that people should be having fun. And it kind of builds up and so it's, it's a combination of an instruction guide to making a game engine and a reference manual of like, here's a list of like hundreds of skills from like all these skill based games we've ever seen. I, I've kind of, to be honest, been skipping over a lot of the, you know, here's a list like, of of like a million different types of armor. Because it's not what I'm interested in. What I'm interested in is like, how do people create these game engines and how do they keep them balanced and how do they keep them interesting? And so when I say game engine, just to be clear, it is programming throwdown. But this has nothing to do with programming. It's literally like the math that, that like you could use this for like a card game or a tabletop game or really anything. It's like the math that keeps people kind of on the edge of their seat. Right. So it's like, how do you have all these different options for your players but still keep them kind of on the edge of their seat? And at the same time, how do you do that in a minimalist fashion where it's not like, okay, I'm now gonna have to roll like 45 dice to like build my character or whatever. So. So these people talk, tackle a lot of that. And I'm about, I want to say maybe about halfway through, as I said, skipping a lot of the pure reference stuff. And it's really interesting. I'm having a really good time reading it.

0:36:45.440 --> 0:37:51.080
<v B>I've never played an in person tabletop RPG or even. I mean, I probably played a video game that somewhere under the hood was running some sort of like roll checks and chance checks or something. I learned the other day a spoiler, something I'm gonna talk about in a few minutes. That Pokemon was actually doing that. When the Pokeball rattles, it's doing like a, you know, probability check and it can fail at each of the things. And then that's when the Pokemon came out, which I didn't know. And, and maybe I'm completely wrong, but that's sort of like what the Internet was telling me, which is in line to your point with rolling a dice and getting certain values, but never had the occasion to play one. But I'm endlessly fascinated by, like you said, the, the sort of crafting of the stories and the storytelling and the fact that it's less game than, you know, a board game with rigid rules and more about, like you said, having an adventure together, making it fun and entertaining and collaborate, like collaboratively doing something which is somewhat still gaming, but is also. You are sort of being flexible on the fly as well to, to, you know, keep it fun.

0:37:51.960 --> 0:38:11.400
<v A>Yeah, exactly. Yeah, exactly. Like, how do you let each person at the table have a unique character that brings something unique while still like being able to handle a person not being there? It's like, it's like, oh, you know, Jim, Jim's wife is having a baby, so this chest has to stay locked.

0:38:12.360 --> 0:38:16.440
<v B>You know, he was the one that had the keys. He's sleeping in the inner.

0:38:16.910 --> 0:38:31.950
<v A>Yeah. Or he's the only one with lock picks or something. Yeah, yeah. So, yeah, the, the, the Book is really interesting. I'd recommend, folks, check it out. It's. If nothing else, it's a nice book to have on your coffee table because it has kind of a provocative title, you know, Universal Game Engine.

0:38:33.870 --> 0:40:05.670
<v B>All right, well, I spoiled it. But Tool of the show for me is a video game. And I, for whatever reason, skipped every, like, modern Pokemon video game. And so I think the last one I actually like, legit played was when I got Pokemon Red in my Game Boy as a, as a child and played that to no end using. And I was trying to describe this to my, to my kids and like, I had to go to when we would go shopping at like, the Walmart or Kmart and I would like, look in the strategy guide for why I was stuck. And so I would go with my mom so I could go to the, you know, video game section and like, open the strategy guide and like, look. Because I wouldn't just buy it. I probably should have just bought it. But anyways, and then go home and, and you know, get through anyway, so. So Pokemon Red. Anyways, I, I've been aware I've, you know, dabbled various times, but I hadn't really sat down and played, but I was sitting down and playing Sword and Shield. I was playing the, the shield variant, but not super important on the switch and just hadn't done that in a really long time. And I, I, you know, I know it's a pretty big departure for this series, but it was really kind of fun. Like, I was really into it. I realized now that the game is easy. Like, it's not supposed to be challenging to actually, you know, quote unquote beat the game. So it's not that much of an accomplishment. But I had a great time and if you've been. Ever been interested and you, you know, have a switch or whatever, we definitely recommend checking out one of the newer ones, Sword and Shield. I guess the other one I'm going to try now is Scarlet. Scarlet and I think Violet.

0:40:05.670 --> 0:40:09.950
<v A>It is Pearl or something. Oh, yeah, okay.

0:40:10.590 --> 0:40:11.110
<v B>I did.

0:40:11.110 --> 0:40:12.710
<v A>Or that's a different. I think Pearl.

0:40:12.710 --> 0:40:45.460
<v B>I think that's a remake. I think that's a remake. They did like a remake. Yeah, like, so anyways, if you hadn't checked one of those out, they definitely went and some of them like a little bit more with open world sections and you can kind of control how often you get into a battle versus, you know, just wandering around in a set of grass until it happens. Uh, so. So definitely some quality of life improvements over the old ones. That make it less frustrating and you know, ability to save sort of everywhere you want. And so if you never checked one out, I, I guess this is me telling you the obvious thing of like it's a thing and it's, it's kind of fun.

0:40:45.860 --> 0:41:34.580
<v A>Yeah, I played it with my kids and it's been maybe a year or two and, and you know, they would get frustrated and so I'd help them like kind of optimize their characters a little bit, but more or less they could get through it eventually. And yet the other thing is it's all the bosses and everything. As far as I know, they have static levels. So you know, if, if you're kind of like, like my kids are just running around kind of aimlessly for a while. So their Pokemon were like super over leveled and that made the game even easier than if you're trying to like speed run it. But yeah, that game is awesome. I think, I think the open world added a lot actually. Like being able to, to, to really like see the enemies and they run into you physically and then the fight starts. Like it a lot, I think. Yeah.

0:41:34.580 --> 0:42:07.090
<v B>And I, I think that it goes crazy deep though. Once you like look on the Internet, there's all this like each Pokemon you catch has different stats and, and there's what we're talking about. I never paid attention to it other than like it has a type and certain moves or whatever, very basic level strategy. And that was fine to get through the game, but when you look online you find out, oh yeah, the competitive stuff and people playing online and whatever, that each Pokemon you catch has like different base stats that have been rolled for that character. I mean it's not actual dice, but probabilistically generated. And so some are better than others even if they're the same levels.

0:42:07.250 --> 0:42:11.250
<v A>All right, so Patrick, do you want me to waste hours and hours of your life?

0:42:12.690 --> 0:42:14.130
<v B>Is it going to be fun?

0:42:14.370 --> 0:43:31.480
<v A>It's going to be fun. Okay, so later on, go on YouTube, okay, there is this guy who really understands the Pokemon mechanics and purposely. So there's a, there's a Pokemon. I think it's a web based game and it's probably not legal, it's probably already shut down or something, but. Or maybe it's sanctioned, I don't know. But there's this web based game where you can just do Pokemon battles with other people and there's a, the elo, like for chess and everything. You rise up the ranks and so it's, it's literally just the battling part of Pokemon and So this guy who really understands the mechanics he makes builds that are very unintuitively strong and he plays people who, and he must play a lot of people, but inevitably he ends up playing someone who, who starts off like making fun of him and like, oh, like you just have one Pokemon, like why didn't you build the other four Pokemon? Haha, you're so trash or whatever. And then he wrecks them and they start raging and they start like. And then they won't make their final moves. And he's like, hey, your time's running out. And they just get so pissed and everything. It's like the people who get the scammers upset or whatever, it's that but for, for gaming trolls. And it is hilarious.

0:43:31.960 --> 0:43:36.800
<v B>Oh dear. Okay, now you have to send this to me, but I feel like I'm going to not like you for doing.

0:43:36.800 --> 0:46:48.980
<v A>It, but yeah, I mean, I don't know how many videos he has. I'm pretty sure I've watched like five or six of them. They're really, really funny. All right, so. Oh my tool of this show is features and labels or Foul AI. There's a bunch of alternatives. There's together AI, there's fireworks AI, there's a bunch of them. But basically these are people who are kind of a middleman between you and the AI models. So they'll host the open source ones, they'll often have agreements with the closed source ones. So you can run like the Google Imagen, you know, which you, you, Otherwise you'd have to use some proprietary Google API, whatever. So think of these as like a middle layer and they often charge you, you know, per thing that you do versus like you having to rent a machine for an hour. Right. The thing that, that so the reason I picked foul is I actually know the founders, so I, I, I'll just put it right out there and say, I don't know if FAL is any better than any of the other ones, but the user interface is really nice. They have like a playground mode where you can just build things on the web and then you can click on the API button and get the Python code if you wanted to make that programmatic. The other thing they did, which I thought was really clever UX, you know, as engineers, especially as people who have GPUs or maybe an M2 MacBook or something, we think to ourself, like, yeah, I mean, I should just run flux myself. Like Patrick's run flux, I've run flux. Right. But when you go to foul, they're like yeah, so you can run this model like 87 times for a dollar. Like it basically for every model it tells you how many times you can run it for a dollar. And that to me is like really powerful because like often like I have some code that I have right now and I have a local version of, of Flux and then I have the FAL version of Flux and you can just like toggle between one and the other and. And you know, like I'll want to run sometimes I'll think oh, I'll run the local version cuz I'm going to work and I'll just let it run and it'll be done when I get back. But then I'm like yeah, it'll be done when I get back. Or I could spend like $7. And this is just like done in, in like a second. So it's like. So it's like they did a good job of kind of like really laying out the economics which are themselves startling. You know, how the economics have changed for AI, but, but they just put it right out there. And they recently had something, they had something where they, they're able to do some caching of the, I think caching of the tiles. The way the, the image transformers work is, you know, breaks your image up into tiles and I think they're caching tiles that are very similar or something. I don't know. It's, it's. It's something I don't remember off the top of my head, but. But it made the price even cheaper. So I guess, long story short, check out these folks. They're all awesome. I know the fireworks people too, that all these services are great and there's an economy of scale that you can really take advantage of.

0:46:50.660 --> 0:47:52.230
<v B>Yeah, I mean I think all of it from like someone was asking me with 3D printing, how much would it cost you to print this? You know, I saw it in a shop or something and I was like, well, most obvious thing is how much plastic it takes. But then you start thinking about it. There's depreciation of your machine, like wear and tear on your machine. There's like the power to run it. There's my time to like walk out. Mine's in the garage. Like walk out to the garage and like get it off or clean the, you know, build plate. And so I think what you're saying is interesting too, that running it locally to me is. I guess I just cheap. Like I don't want to give places my credit card. I don't know. And I have stuff. And so I feel like I should use the stuff I have. But you're right, like by the time you factor in the power to run it, like it's not free and you know, your computer getting hot and the time taken and so the economies of scale, these really big, you know, server clusters dedicated to this AI stuff, it's really kind of amazing. Even with how expensive those really high end GPUs are.

0:47:53.030 --> 0:51:49.690
<v A>Yeah, yeah, totally. I think running it yourself is great. Everyone should learn how to do it. Definitely not discounting that. But check out these folks and they're similar folks too. It's a really neat service. If you have something that you then say, oh, I need to run this like 200 more times. Your time is also really important. All right, on to our topic, reinforcement learning. So a bit of background here is like the opposite of like assembly language show, where in this case, like this is my background, my area I know a lot of about. I'll kind of dive into it and then, and then Patrick is going to play the role of you folks and, and stop me anytime I say something that is a buzzword in my community or doesn't make sense or something. So I'll start really broad. There's, there's basically three types of AI. There's supervised learning, there's unsupervised learning, there's. And there's reinforcement learning. So supervised learning is where you have the right answer is right there. So, so someone gives you a picture and they draw a box around the stop sign. They're like, there's the stop sign. And your job is to learn a function that maps the picture to the bounding box of the stop sign. And you're given a lot of these as ground truth, right? And then you're also given a second set that you're purposely not meant to train on, called the holdout set. And if your training did really well, then you're able to interpolate between all the other stop signs that could exist in the universe. And so after training, if I was to give you a new image you've never seen before with a stop sign in it, you could draw the box around the stop sign. That's supervised learning. And under the hood, that works through what we call a loss function. So a loss function takes the output of your model. So in this case, maybe it's a bunch of hypothetical bounding boxes. It takes the ground truth, which is the actual bounding box, and it, it turns all of that into a number where the further the number is Away from zero. The worse you got, the worse you did. And so zero. A zero loss would be. You perfectly nailed that bounding box. Now, a zero loss might not be good, right, because you want the model to have some uncertainty. Like, for example, imagine we're playing paper, rock, scissors, right? And you play paper, and I play scissors with like 100% certainty. That's actually not good, right? Because. Because although I won in this game, you know, we know that someone who plays scissors a hundred percent of the time is not playing an optimal paper, rock, scissors, right? You could just, you could learn that and then just play. Wait, did I get it wrong? Anyways, you know the analogy. You could just play whatever counter is what I just said and then just win, right? So so often you're going to output a mixed answer, right? A distribution of answers. And so you're always going to have some amount of loss. But through, through what's called the learning rate, you don't like, totally change your line of thinking every time an example is presented, right? You're just slowly moving in different directions as examples are presented. And if the learning rate is low enough and all the, a hundred other things kind of stars align, then you'll create like a mixed, you know, a mixed response that is optimal. So that's supervised learning. Did I get that right? Any questions about that part of it, or did that make sense?

0:51:50.650 --> 0:52:19.380
<v B>So you said, the only question I had is you were saying, so this makes sense that you have the thing that you're trying to match and then you're tested on something else. But you said interpolate between the results. But it should be possible even with supervised learning, not strictly like, interpolation to me means like between the points given. But you should even like in your stop sign example, like you were mentioning, for stop signs that are new, the idea is hopefully you would also understand that th. Those should be. Have the bounding boxes put around them.

0:52:20.340 --> 0:52:51.630
<v A>Yeah, right. So. So what you're hoping is that you could imagine like a manifold, like a stop sign space. And in that space there's like a whole bunch of different kinds of stop signs. And, and so inside of that space, there's stop signs that like, look really different, but hopefully they're like, they're within the space of stop signs that you've already seen. So like an example where extrapolation doesn't happen.

0:52:53.640 --> 0:52:53.880
<v B>Is.

0:52:55.640 --> 0:56:22.490
<v A>So do we've seen this with Waymo, where people will wear a T shirt with a stop sign on it. And that's not. So that's an Example of extrapolation. And in that case, the model doesn't really know what to do. So it, it, it thinks it's a stop sign. So, so really. So there's a whole area around what's called out of bounds detection and out of bounds prediction. And long story short, that's a very, very hard topic, but, but really important. But by default, supervised learning will interpolate, you know, at a really high dimensional space, right? Interpolate between all these things that's seen. But, like, if you give it something totally new, it's going to have trouble. Got it. Okay. So unsupervised learning is where you don't have a loss, like there's not a ground truth, but what you do have is something that's kind of stateless and easy to evaluate. So, like, the most common example is clustering. So there often isn't like, known a perfect clustering. Like, you might have millions of documents and you want to break them into a thousand clusters, each one having 3,000 documents. And you want the, the entropy of each of those clusters to be really small. So you want all of them to be like really close together. Right. So you might never know the perfect clustering like you do at supervised learning, but it's like trivial to like, evaluate. So I can like, show you a set of clusters. You could put the documents in the clusters and come back with a score. This clustering has a score of seven. And I can make some changes. They say, oh, this clustering has a score of 8. It's a little better. I make some changes. Clustering has a score of 9, et cetera, et cetera. And so that's an example of unsupervised learning. And so you're not even really trying to figure out the best way to cluster, like a human is doing that, but the computer is just kind of following the instructions and then over time getting a better and better clustering. And you can measure that. And so you might never get to the optimal, but you can get closer. So unsupervised learning is a little bit trickier in a sense that you don't have a ground truth. Now, reinforcement learning is, in my opinion, like, the hardest of these areas. Not in terms of, like, you have to be the smartest to do it or anything, but the hardest in terms of, you know, like, getting, getting good results is the most difficult. And this is because you have all the challenges of unsupervised learning, where you don't have a perfect game of Go or a perfect game of chess. To reference, but you also are making decisions. In the unsupervised learning case, you're not really making any decisions. There's a human making decisions or a human written algorithm making decisions, and you're just evaluating them. But here you have to make the decision. So it's like, here's a set of clusters. How do I make them better? And then you do that, and then did they actually get better? So if you were actually on the fly designing your own clustering algorithm with AI, then that's reinforcement learning.

0:56 --> 0:56:37.080
<v B>But the stuff that we talk about when we say at a high level, like, oh, flux or this or whatever, it may be using components that were trained with a variety of these techniques or use a variety of these techniques.

0:56:37.080 --> 0:56:37.360
<v A>Right?

0:56:37.440 --> 0:56:54.250
<v B>So it's not necessarily that a whole. I know what the distinction there is. Like a whole program application is one of these. You're. You're sort of talking like a little lower level. You're saying like one part of that pipeline was done this way, kind of.

0:56:54.250 --> 0:59:42.460
<v A>So, so in the case of flux, that's all supervised learning. So in the case of flux, it's what's called self supervised learning, where you hide part of an image and you ask the AI to draw it. And then because you hit it, you know what it used to be. And so you show that to the AI and say, hey, you know, this pixel actually should be red, but you drew purple or something. And so, yeah, so that's pure supervised learning. There are like, you know, recently, some. So another way of saying it is reinforcement learning, like, does stuff like takes actions. And supervised learning and unsupervised learning kind of reveal knowledge. So in the case of the stop sign, you know, drawing the bounding box around the stop sign kind of reveals or synthesizes knowledge. Like now you went from pixels to there's a stop sign there, but it doesn't tell you how to. It doesn't drive a car or turn a camera or take any action. So as soon as you want to take an action, now either the humans have to write that code, but as soon as you want AI to take an action, now you're doing reinforcement okay, so there's a bunch of different kinds of reinforcement learning algorithms, but there's basically two axes that you need to think about. One is offline versus online. And this is just a fancy way of saying, can I make mistakes? So for example, the AI that plays Go AlphaGo in the beginning of training, let's just stick with AlphaGo Zero. It's all pure reinforcement learning. So in the beginning of training, it's just playing garbage games of Go and that's fine because it's playing against itself and it's, it's, you can't embarrass the computer. So, so it just plays garbage games of Go and it gets better and better. But like you couldn't for example, like drive a self driving car randomly until it got better. Like, you know, you, you can't do that. You crashed a car, people would die, be like total mess. Right. So, so offline reinforcement learning is where whenever you make decisions in the real world, they have to come with some kind of guarantee. In the case of online reinforcement learning, you can just make decisions in the real world whenever you want at any, at any point. And so that, you know, that's like a subtle difference, but it has like pretty big consequences and the algorithms and everything else.

0:59:44.220 --> 1:00:01.290
<v B>So online it's able to change itself and like update. And then offline it's sort of like you're wanting to make guarantees. You want to know, like I, I understand what it's going to do, I've tested it in some way and I don't want it sort of like changing what it's doing.

1:00:01.770 --> 1:00:07.370
<v A>Right, Right. So online you, you're willing to put any model in production.

1:00:12.170 --> 1:00:15.010
<v B>So yeah, I think it's.

1:00:15.010 --> 1:04:45.530
<v A>Sometimes they call it on policy versus off policy, but it gets the nomenclature there doesn't matter as much. Those are the, the two kinds of. Okay. And so then there's a second axis or second kind of switch here which is value value based or policy based. So I'll go into this. So let's say you have to make some decisions, right? And when you go to make a decision, like imagine a choose your own adventure book. And whenever you go to make a choice, I was to tell you like, if you make this choice, you have like this percent chance of reaching the best ending. And if you make this choice, you have this other percent chance. Like you would just choose the highest percent. Right. And you would just do that. It would be like solving a maze with no walls. Right? You would just, you just like pick the highest percent every time until it's a hundred percent and then you would win. Right. And so the idea with value based reinforcement learning is if I know the total value of a decision and I know that for all my choices, then I've solved the problem. I just picked the one with the highest value and that's just the optimal policy. And so value based kind of ignores the whole decision part of it somewhat. And Says the game here really is figuring out the expected value. Because once I have that, I'm. I'm set. Now, here's where it gets tough, right? Is let's say AlphaGo places a stone somewhere on the go board to start the game, and it's playing itself or some other world champion or something, right? It places that stone and its value is about 0.5. It has like a 50, 50 chance of winning the game when it just started, right? But let's say I, Jason, go and play the world champion of Go, and I put the same stone in the same position. Just coincidentally, for my first move, I have a zero percent chance of winning, right? Because I'm not even close to a world champion. I'm gonna get wrecked, right? So, so you have this paradox where, like, the value is based on the policy, but if the policy is based on the value, you can see how this is like cyclical reason reasoning, right? And so getting the value is actually really, really hard for this reason. And there's, there's several algorithms. The simplest one is called Sarsa, which basically says you make a bunch of moves, the game ends or the episode ends, you stop driving the car or whatever it is, then you just go back and you know what happened. So you assign the expected value. So, you know, I turn the steering wheel here, I turn the steering wheel there, I hit the brakes, I hit the gas, and then I made it home safe. Therefore, all those actions are plus one, right? You know, turn the steering wheel, I hit the gas, I crash into a wall. Therefore, all those actions are minus one. And then you feed that into your neural network, you do your training, and that's pretty simple. You know, that completely ignores the thing we just talked about. There's other algorithms like Q learning and stuff, that try to address some of these challenges. It's really difficult. It's not. Doesn't mean value optimization doesn't have its place. But, you know, ignoring the, the, the policy makes the. Makes it really difficult to optimize. It's good in situations where, like, there's clearly one good action at any given time, you just don't know what it is. Like, Atari is a great example where, like, the actions are binary. There's usually, like one good action. Like if you're playing Mario or something and you're about to run into a Goomba, you either jump or die. And so you jump, right? But as soon as you get into environments where you need a mixed response, like poker or driving a car or really doing anything in the real world, it becomes difficult. Any questions about value optimization or did that, did that make sense?

1:04:45.530 --> 1:05:17.390
<v B>So I guess it makes sense. So I think what you're, in these cases you're trying to, like you were saying is there's like, understand the outcome of a game. So you're playing a game, you don't know what's going to happen, so you don't know if it's a good or bad move until sort of like it's too late. So but by observing many, many, many games to their conclusion, you're hoping that when you go to do it, I guess that's, that's offline but for real, that you've built up an estimate of, given us context, what is the likelihood that each decision is good or bad?

1:05:18.460 --> 1:08:37.400
<v A>Yeah. Right, Right. And as your values improve, your policy improves, which means your values now are all inaccurate. And so you're just like kind of iterating on this over and over again. Yeah, but you're totally, you totally nailed it. Um, okay, so the other type of algorithm is policy optimization. And in this case you say, at least in the most naive example, you say, I don't even really care how good this action is. Like, I, I don't need to know what's my expected value of taking this action or anything. All I want to know is I want to take actions that are good and I don't want to take actions that are bad. It's like, I cooked some eggs, they were delicious. I want to do more of that. I touched the stove with my hand. Not delicious. Don't want to do that anymore. Right. Very simple. So, so policy gradient is basically. And I'm going to try to. I'm doing a lot of hand waving here, but basically it says get the expected value of this action. So you do have to kind of like play out a whole series of events. But then when you go back and you look at what happened, take the things that were good, that had positive, positive value and do more of them. And do them in proportional to how positive the value was. Take the things that had negative value, do less of them and do it in proportion. So things that are really negative value really do it a lot less. Right. It's a very simple concept. One challenge right off the bat that you can see is your expected value has to be centered at zero. Like, in other words, if all you can do is get points, but you can never have a negative score, then your system's gonna say, do everything infinite amount of time, and it's just not gonna be able to Learn. So, so you need what's called a baseline, such that you hope that roughly half the time you're getting a positive score and half the time you're getting a negative score. Now for something like Go, it's trivial because you play a game and you either win or lose. And so unless you're playing like someone way out of your league in either direction, you're, you're hopefully going to win and lose about half the time. You get a negative one for losing, positive one for winning, you're all set. So Go makes this very easy, but in the real world it doesn't, doesn't work that way. And so it actually, like, like figuring out how you can get half of the expected values to be positive is really hard. And so you actually use a second neural network just to figure that out. And that's called the critic. And so the, the, if you hear the term actor, critic, the actor is just the policy gradient that I talked about earlier. And the critic, all it's doing is it's trying to figure out the baseline, trying to figure out the average move, what would be the expected value of that? So you can subtract that out and hopefully get us balance between positive and negative.

1:08:41.880 --> 1:08:47.880
<v B>And so this is the neural network you would use on a Go board to tell you like, how good your situation is.

1:08 --> 1:10:23.240
<v A>Yeah, exactly. So if you were so using Go for an example, let's say you're trying to solve Go with a policy gradient. So you would say, I took a bunch of actions, I won. Now I'm going to look at this one action. I got an expected value of 1 because I won the game. Now what was the average expected value? Actually, sorry, what was the advantage is what I need to know. So I need to know was this one, like, for example, like, is this, am I playing someone who's like a total chump? And so like, even though I won, I can't really learn anything, right? Or did I play someone who's like a grand master and actually learned a ton by winning? That's your sort of advantage function. And so you're going to take the value function of the current state. So, so at this current board state, what's the probability I win? And then you're gonna take the action you took and see what's the value at that state. So if, if, if the probability of winning the game is 50%, but after I took my action it jumped up to 60%, then I know that taking that action like caused an extra 10%. And so that's My advantage. And so you know, if you're playing someone of your level, half of your actions are going to cause your win probability to go down and half of them are going to cause your win probability to go up.

1:10:24.440 --> 1:10:25.080
<v B>Got it?

1:10:25.640 --> 1:10:58.040
<v A>Yeah. So as I said, for go you don't need it as long as you're doing self play. But for, for, you know, something like Atari, where there is never a negative score, you need to know like what is was a bad action. And so a bad action is one where your expected score went down after you took the action. It's like, oh, I took the action to run into the goomba and now my expected score is a lot lower cuz I have one less life to go and collect points with.

1:10:58.280 --> 1:11:06.510
<v B>So is there a conversion between lives and points then? Or it's just simply that because you lost a life, you're maximum point that you can get is reduced.

1:11:07.070 --> 1:12:23.260
<v A>It's the latter. Yeah. So all of that has to be inferred. So now you could make it explicit. So you could say, and this is something that we should talk about. It's called reward shaping. So let's say in Mario your goal is to get the most points. But that's kind of a really weird goal, right? Because often like when we play as humans, we don't even look at the score. Right. So you might come up with proxy goals. You might say, well every time I eat a, well every time you eat a mushroom you get points actually. But, but let's pretend you did it. It's like every time I eat a mushroom I'm going to give myself extra points. Maybe you don't get enough points for eating a mushroom. And if you gave yourself more points for eating a mushroom, then the AI learns easier and, and to your point, like you don't lose points when you die, but maybe you should, like maybe if you lost 10,000 points every time you died, the AI would learn a lot easier. And so this is called reward shaping. It's where you instead of learning the task at hand, you learn a new task. Because those two tasks kind of go up and to the right at the same time. They're correlated and the new task is just easier to learn.

1:12:24.250 --> 1:13:09.490
<v B>Yeah, so I mean, I don't know the scoring of Mario either. I never paid attention. But I guess like you could get weird states otherwise where you try to get a bunch of one up mushrooms and then on a particularly high value level, like basically keep dying almost at the end repeatedly in order to like keep gaining points for that level. Assuming you don't lose points when you die. And so you could get this very, like, undesired behavior because it realizes the way to get the maximum score is to keep dying and replay the level and gaining those. Gaining those points when the actual thing you wanted was like, kind of get through the levels as fast as possible and you didn't really care about the score and you used it as a. But it was like a bad proxy for what you wanted.

1:13:10.290 --> 1:15:29.310
<v A>Yeah, exactly, exactly. And then on the flip side, like, you might say, well, my. My goal is to get as far to the right as I can. Right. But if you don't take score into account, it might just be really hard for the AI to do that. Like, the AI might just like, desperately do, like, kamikaze jumps to the right where it gets killed because it didn't. Or the AI might just not be incentivized to get mushrooms and make it more survivable, that kind of stuff. So. So, yeah, reward shaping is a really big part of the problem. And, and so when people go from manually designing systems, like, if this ad has this chance of getting liked, then put a highlight around it or something. When people, like, build these things by hand, they. They have to deal with, like, all these conflicting metrics. And how do you, like, reconcile, like, you know, oh, we showed more ads, but there was, like, more kind of, like racy, unethical kind of photos. And how do I deal with that? And so reinforcement learning moves that problem to the reward shaping phase, but it doesn't really get rid of it. You're always going to need to, like, better and better understand kind of, like your goals and the nature of the problem and how to kind of how to best, like, solve that problem. Okay, so. Okay, so let's dive into the offline part. So, you know, one thing a lot of people wonder is, yeah, like, AlphaGo plays against itself. And so at the end, after you've used, like, a zillion GPU hours, it's like a world champion. But, like, how do we do that in the real world? Like, clearly, like, babies don't just, like, run into walls or, like, fall down. Actually, they do kind of fall downstairs if you let them. But that's a bad example. But like, like, you know, us humans, like, the way we drive a car is, you know, we have a person helping us, but we're not just, like, randomly jerking the steering wheel until we figure it out. Like, we have the sort of, like, base of common sense. Like, we kind of draw understanding, like.

1:15:29.310 --> 1:15:30.790
<v B>A model of how the Car should work.

1:15:31.190 --> 1:19:04.070
<v A>Exactly, exactly. And we kind of like, kind of project into using our mind. We kind of simulate the driving experience as best we can from watching other people drive, watching our parents drive. And we've built a simulation of that on day one. And so there's a question of, like, how do we do that with reinforcement learning. And a big part of that is using what's called a trust region. And so a trust region is basically, it works like this. So let's say I play, I play a bunch of games of Go, and I'm a decent player. I play a bunch of games of Go, and now I go back and I watch all of my games, right? This is just me as a human. I watch all of my games and I look to myself, I say, oh, I would have done maybe this move differently. I would have done that move a little differently. But I'm not going to say, like, I would have done every move differently. Like, as a person, like, we can't. That would put us into a really weird state where, like, we wouldn't really know what to do, right? So we would pick, like, a few key things that we would do differently, and then we would wait until it's the next tournament, exercise those differences, and then we'd repeat this process. And so, you know, with, with, with, with reinforcement learning, if you take a bunch of data and have computers try to do policy optimization, they'll just hallucinate, just like we see with chat, GBT and these other things. Like, they'll start hallucinating like, oh, if I play this move, I'm going to get every single Go piece on the board because there's, like, some inaccuracy in the model. And the other thing is, like, it only needs one action to be inaccurate on the positive side to throw everything off, right? All your values are now thrown off. Everything, right? So, so it's inherently kind of unstable. And so what Trust region, policy optimization, and proximal policy optimization. What these things do is they basically say, we're going to keep track of the actions that were taken in the real world and what the model is doing. And if the model doesn't match the real world enough times, we're going to stop training. So, you know, in the beginning of training, the model is going to match the real world perfectly because it's the same model, right? Like you, you, you, you rolled this model out in the real world, collected a bunch of data, and. And at the very first mini batch of training, the model hasn't changed. And so it's going to output the same distribution. Right. Over time, the distributions are going to start diverging. And because, you know, because you own the model, you can actually keep track of the entire distribution. Right. So even though you actually press the gas, you know that the model was like 50, 50 about pressing the gas or not. And that's what you're going to log, right? So now you're training and you say, oh, well, the model that drove the car press the gas 50% of the time, but the new model wants to press the gas a hundred percent of the time. That's a pretty big difference. And so maybe this would be a good time to like, stop training and go drive with the new model. And so that's, you know, there's a lot more math than that. That's effectively what's going on, is these are like halting, halting criteria. So you might not be able to train that much before you have to stop and go to the real world.

1:19:05.060 --> 1:19:19.460
<v B>Got it. So it's basically like you've, you're, you're too far away from what we know. So you need to go try again. You've changed a bunch of stuff a little, but your outputs are now very different. So we need to go try again and see if it actually got better.

1:19:20.020 --> 1:20:16.730
<v A>Yeah, exactly. And then the last thing that I'll kind of COVID here, there's a couple of other things. So one is imitation learning. And so this is pretty simple. The idea is, you know, we talked about supervised learning, right? And the stop signs. Right. But I could do the same thing with decisions. I could say, hey, when you see this situation, press the break. Right? And I'm just saying, as an expert, like, it's a ground truth. Like, like unambiguous, you know, press the brake here, press the gas pedal here. That's called imitation learning. And that's just supervised learning. So, you know, you could imitate a person and, and then all the regular things apply there of interpolation and everything we talked about, but that's not really reinforcing learning.

1:20:16.970 --> 1:20:25.380
<v B>And that's the AlphaGo not zero, where it was trained on all the human games. And they basically said, we want you to imitate the person who won.

1:20:26.260 --> 1:25:19.160
<v A>Right, Exactly. And so AlphaGo did that as like a bootstrapping phase, and then did, did reinforce and learning after that. The alpha goes zeros, where they got rid of the bootstrapping phase. Um, so that's. So imitation learning is a good way to bootstrap. And that's kind of what, what we do. Um, another Thing I want to cover is model based reinforcement learning. We're basically, you know, in the case of AlphaGo, it can play itself because it's just a game, right? Like it's an artificial environment. But if you want to do, for example, a self driving car, as we talked about, you can't just drive randomly while you learn what to do. And so you have to construct a model. You have to construct like a virtual environment and then play within that virtual environment. You know, the challenge now is of course, like, what happens when the virtual environment doesn't match the real environment. And so there's different ways to deal with that. There's something called a joint embedding. But long story short, with model based reinforcement learning, you have this sim to real problem. So if you look up sim to real, you'll find like a zero, a zillion papers on it. But it's like, how do you take something that was trained on a simulator and bring it to the real world and back and forth and back and forth. Okay. Yeah. The last thing to cover is policy evaluation. So, you know, with supervised learning, you have the truth. So it's like, oh, I didn't draw the bounding box around the stop sign. That's bad. I drew the bounding box around stop sign. That's good. And you can, there's a million different ways you want to count those, those errors, but you can count them those ways and, and just output that. Right. In the case of reinforcement learning, you don't really have like a perfect game of Go or anything like that. So what you have to do is, is several different things you can do. You can either use a simulator and say, oh, in the simulator, my model got better. That's what AlphaGo does, right? But in case of, let's say, self driving, where maybe you don't want to trust the simulator, there's another thing you can do where you run two models and the first model actually controls the car, and the second model just says what we call counterfactuals, which is just a fancy word for I would have done this. Right? So it's just like, it's literally a backseat driver. Literally. So. So you then take those counterfactuals and what actually happened, and you can figure out if the new model is better than the old one. So, for example, let's say the old model doesn't hit the brakes and the new model really, really wants to hit the brakes. And then like half a second later, the old model slams on the brakes. Well, that probably could have been avoided by the new model because the new model was breaking earlier. Right? It was, it was like more predictive. Right. And so that would be a sign that the new model is a step up. Similarly, if the old model slams on the brakes to avoid a collision and a new model would have hit the gas, right. That's a bad sign. That means that new model probably would have got you in a bad position. So again, a lot of math behind that. But effectively, that's the intuition behind policy evaluation. And one last thing about that is policy evaluation is harder than solving the problem, because if you have a perfect policy evaluator, then you also have a perfect policy. You just take the action that the evaluator gives the highest score to. So, so because of that reason, like versus like in supervised learning, you could just measure accuracy. Just take the times you were right and, and divide it by the total time. Very easy. Just algebra. Right? But here it's like, not only is it hard to evaluate the policy, but it's actually harder than solving the problem. And in many cases it's impossible. So, so that's a big challenge. And that continues to be a challenge today. And so that's. That's Paul. So I've kind of covered all of the technical stuff. I'll dive into a little bit of the large language model stuff. But before I do that, any questions about the technical st.

1:25:20.420 --> 1:26:19.300
<v B>So I guess the thing I'm missing is, is a bit. Is. I guess it's a little application, which is understanding. So some problems, like you're saying, are clearly not a fit for supervised or unsupervised learning. And so you could think about like, oh, maybe this is a reinforcement learning task. When do you. But then some things, maybe there's multiple approaches to solving. And so, you know, you could use a reinforcement learning. You could try something else. Um, and then from like a toolbox standpoint is like, reinforce. So even today, like, I can just Google, like your stop sign example and there's a hundred tutorials for opening up, you know, TensorFlow, PyTorch, whatever. Give it the images, give it the labels. We talked about labeling, but, you know, give it the labels and you know, get. Monitor your loss function. Is it the same. Is there like the same set of tools for doing reinforcement learning? Or is there also some of like a canonical example that you would sort of go to, to like, kind of do like the simple case?

1:26:20.500 --> 1:29:21.980
<v A>Yeah, it's a really good question. Okay, so, okay, the first part of the question, I think that my general philosophy is to Use the simplest tool for the job. Right? So, for example, I'll. I'll give you a really concrete example. There was a place I worked at. I could probably say this. I'll just say it. I don't think it's going to be that controversial or anything, but it's not that much of an expose. But, you know, when I worked at Meta, you know, we released the Oculus Store, right? And so you can go right now to the Oculus Store and buy games for the Oculus Quest, right? And so I talked to the product managers. They asked to meet with my team, and they wanted to do reinforcement learning to figure out what items to put on what places on the storefront. So when you go to, like, oculastore.com or whatever it is, the URL, like, what should they just show right there on the banner? They call that the hero position. What should they put in the hero position, et cetera. And my response to them was like, not only should you not use reinforcement learning, you should also not use AI, Right. What you should do is, like, take the app that sold the most and put it in the hero position just manually and run that way for a month. And then if you realize that, like, if you just have this intuition like, oh, there's. There's so many people with so many different interests, and we're showing everyone beat saber, and it's not going well, and. And so we need to do some AI, then. Then let's. Let's go there, right? So, like, simplest tool for the job. Like, the simplest thing was just like a YAML file with beat Saber in it, right? And so, like, they launched that. And then. And then I would say, you know, if you can do something simple around the decision, like, say, okay, in certain countries I'll show beat saber, and other countries I'll show other stuff. And then now I'm dividing by some other demographics, and the next thing you know, you're kind of like building a decision tree by hand. Okay, let me use a decision tree, Right? Yeah. And so. And then at some point, you run into, like, competing interests where, you know, I want the store to do well, but I also want game publishers to share the benefit. I don't want to just king make big beat saber, right? So now I have this, like, completing competing economic model that's very complex. Now we're starting to talk about reinforcement learning and some of that, right. So I would say, you know, stick with the simplest tool for the job. Reinforcement learning often is much simpler than trying to, like, take Actions by hand and stuff like that. So for, for running a marketplace, for driving a car, you know, reinforcement learning is a great choice.

1:29:24.460 --> 1:29:24.820
<v B>Yeah.

1:29:24.820 --> 1:29:56.440
<v A>And as far as tooling, the tooling is way, way behind. There's a lot of reasons for this. One of the biggest reasons is reinforcement learning can't really be commoditized because it's too, it's too close to the decisions that companies make, which are sensitive. And so it's just very hard to commoditize. I mean, you know, we rolled out reagent, which was the most popular reinforcer learning platform for a while.

1:29:58.180 --> 1:29:58.380
<v B>Now.

1:29:58.380 --> 1:33:16.340
<v A>There's, you know, there's OpenAI baselines, so there's a bunch of places where you can get the algorithms right. But if you want like the real techniques, like how do I do offline evaluation? You know, a lot of these are proprietary. You know, reagent actually has policy evaluation and all of that, so folks can definitely check that out. And the code base, as far as I know, is still active. But, but I think the field is just still too new for, for there to be, you know, kind of really good practices there. But yeah, those are two awesome questions. Okay, so I'll move on to RLHF. So a lot of people found out about reinforcement learning when ChatGPT rolled out RLHF, which is what makes the chat part of ChatGPT. What took it from GPT to chat GPT. RLHF is, is, is a pretty simple idea. The idea is you. So the idea is, is thinking about it this way. GPT is imitation learning. So a person wrote, you know, the frog, the fox jumped over the dog, or whatever that is. You know, when you pick a font, it always shows you that same sentence. It's like the quick fox jumped over the lazy dog. So that's probably all over the Internet, right? Because it's in every font. So, so GPT will, will imitate a human quote unquote. If a human is, you know, all the content on the Internet averaged, right? And so if you say like the fox, the quick fox jumped, GPT will respond with like, over the lazy brown dog, right? And this was trained in a supervised way. But when you think about it as like it's a decision to put that token there, like it's a decision to put that word there, then actually GPT is making decisions. And so, and so it becomes a reinforcement learning problem when it becomes multi step. So for example, you know, if I just need to predict the next word and I know exactly what it is that's supervised learning. But if I have like 10 different answers from GPT and I want to pick the best answer, like an entire answer only gives one score. Now it's a reinforcement learning problem because I have to figure out, okay, this answer is better than that one, therefore all the tokens that generated that answer are a little bit better, but we don't know how much. And so RLHF is just a pretty simple algorithm where you say, give two answers if you, if the, if the system is more likely to pick the wrong answer, it gets a negative point, it's more likely to pick the right answer, it gets a positive point. And now you do your policy gradient. And so RLHF has been a part of these LLMs for a very long time. The thing Deepseek did that made reinforcement learning. Oh yeah.

1:33:16.340 --> 1:33:20.060
<v B>What is RLHF? What does it actually stand for? Reinforcement Learning something.

1:33:20.060 --> 1:35:26.900
<v A>Reinforcement learning from human feedback. Ah, there we go. Okay, so, yeah, the person, a human is actually saying, this answer is better than that one. Okay, so the thing that Deepseek did that was pretty amazing is it took the human part out. And so it's just rlf. And so the idea is it'll generate an answer to a question that is easily verifiable. So, for example, they give it a word problem, and they know the steps of the word problem and they know the answer. And so they output two, two hypothetical answers. And then it comes back and says, hey, this one's better than that one. But it's all algorithmic. So in math, there's systems that are, you know, they're very expensive to run and they're very specific to math. Right. They only solve math problems, but they're, they're totally autonomous. So I can give you like, not just the answer, like 20 or something, but I can give you like the whole reasoning and the answer to a word problem. And this system will actually verify the entire thing. And so they replaced the human feedback with this expensive system. And then they ran this a zillion times. And what they found was the model that came out of it not only could do math problems better than anything we've ever seen, but it became like, very thoughtful and reflective. And so the, the, the reality is, what they found is if you treat every question like a math word problem, then you become like, much more reflective and, and, and like thought provoking and, and interesting in your answers. And so that's, that's basically what the, the Deep SEQ folks have done, which is definitely like a Huge leap forward and really exciting. Does that part make sense? The. The RLF part of it?

1:35:28.980 --> 1:35:41.140
<v B>Yeah, I think, I think it makes sense. I think they're. But ultimately they're going back and tuning the output of what the LLM is doing, or they're tuning something that comes after the LLM.

1:35:41.540 --> 1:41:06.280
<v A>Yeah, no, they're. They're modifying the LLM itself in all these cases. Okay, Yep. So, so if you think about it like regurgitating the next token is, is actually a, a form of imitation learning. So you're saying, like, these humans that have written this stuff on the Internet, they're experts. And I'm trying to out do the same action they're doing where action is writing letters. And so then when I change the goal to be like, solve this reinforcement learning problem, it's still like a set of actions and so you can use the same model. Oh, another thing I should mention is, so we talked about actor critic, and we talked about how to get this policy gradient stuff to work. You need to have positive values half the time, negative values half the time. So the challenge here is these models now are huge, right? Like, these LLMs are enormous. And so if you need a second enormous LLM, then that's gonna be really problematic. And so what they did, which is really interesting, and, and I'll. It actually only works because there's no intermediate rewards. But that's kind of a detail is they said, okay, we can't afford to have a second model. So what we're gonna do is we're going to get the expected value of all these different answers to this math problem. So we're going to generate 10 answers. We're going to get the expected value of all 10 of them, and then we're going to get. And then we're, we're going to basically normalize that number. So, for example, I get the expected value of all 10 answers, and let's say the expected value is all 1 except for the 10th answer, which is 2. So I'm just gonna normalize that so that all the ones become like negative 0.8 and the two becomes positive 0.8 or something like that. So, so they replaced like an entire neural network with some simple algebra. And, and so that's, that's the GRPO or Group Relative Policy Optimization. So it's one of these things that's like a really clever trick. I have kind of mixed feelings about it. I do think that, I think that with intermediate rewards, it's going to struggle. I think that like, maybe coincidentally or Maybe on purpose, but, but, but the fact that in this particular domain you just get a reward at the very end is, is one of the important causes that for this, this approach working over like PPO or these other alternatives. Another interesting thing where they've kind of diverged is generally what people have done in situations like this where you need a large actor model and a large critic model is they've had the two share the same backbone. So for example, you have a neural net where the current state of, of your, of your universe goes into the net. And then the neural network outputs two things. It outputs the distribution of actions you should take, that's your actor output. And then it outputs the expected value of the current state. That's your critic output. And so you still just have one model. It just has one tiny extra output on it. So you might say to yourself, well, that's like pretty awesome, right? I mean, I mean that seems like a no brainer, but the problem is that even though it's just adding one node, both of those nodes are sharing that network and so they're kind of competing with each other. You know, the critic model is going to be steering the entire network towards producing better values. The actor, model, actor part of the model is going to be steering the network to producing, you know, a better policy. And they're going to be causing corruption in each other. And so although you will see like people have success with this for Atari and other domains, I think it's actually super destructive. And my guess is that the folks at Deepseek actually tried this. I mean it's the most intuitive thing. So I'm almost 100% sure. I mean a lot of circumstantial evidence that the Deep SEQ folks tried to have like just one network, one single network that does the policy and the actor, sorry the, the actor and the critic in just one network. And they realized that doesn't work, that you know, it just causes corruption and it just never converges and it just is a mess. And so they ended up falling back to this approach where they said, okay, well we can't have a separate critic model. It's too big. We can't put a critic head on the LLM because that causes too much corruption. And so we're just going to abandon the entire idea of a critic and just come up with baselines on the fly. And, and that worked for them, which is really cool.

1:41:07.880 --> 1:41:17.880
<v B>Was that something that was unexpected? Like, was that like a sort of, I want to call it like an innovation to make that Leap. Or was it just sort of like. Nah, it was pretty obvious once I got there.

1:41:19.510 --> 1:42:35.020
<v A>Yeah. So this is where it gets interesting is. I mean, so there's a lot of theories. So I'll say, you know, I, you know, I'm not at Meta anymore. I don't work at OpenAI or these places. And so I don't, you know, I don't really know, like, what's on the very cutting edge that hasn't been released to the public. There's speculation that OpenAI was already doing something like this, but they hadn't published it. And so Deep Seat kind of scooped it. There's even more, like, even more speculative is the idea that somebody stole the idea from OpenAI and gave it to Deep Seek. That is pure speculation. But. But I would say the fact that OpenAI has a reasoning model now that is comparable so quickly makes me think that, like, either they worked around the clock or, you know, they were coming to the same idea. Right. And it's probably the latter. Probably deepseek saw where the wind was blowing and. And they both kind of came to that answer around the same time. That would be my guess.

1:42:37.950 --> 1:42:41.470
<v B>It makes sense. So I. There's lots of cases like that.

1:42:41.470 --> 1:42:41.710
<v A>Right.

1:42:41.710 --> 1:42:44.990
<v B>I don't know. Online I saw someone using the term nerd sniped.

1:42:46.030 --> 1:42:47.150
<v A>Wait, what does that mean?

1:42:47.630 --> 1:43:30.840
<v B>Same kind of idea. Like, people or. I think since, like, you're a YouTuber, you're, like, working on some, like, new cool project, you know, you think is, like, crazy and innovative, and someone else just releases a video of the same thing because you didn't get out fast enough, or like you said, there are, I mean, what, half dozen, dozen. Probably like a half dozen super serious competitors, and like, a dozen, like, within striking range of doing these kind of similar. I don't, I don't want to demean them by saying those, like, chat, but, like, question and answer, AI agents, agentic stuff, the reasoning, like, all of these. And so like you said, you're. You're hard at work trying to refine a project, you're not sure if it's a big enough innovation, whatever, and then someone else just goes ahead and, you know, releases it.

1:43:31.390 --> 1:43:32.310
<v A>Um, yep. Yeah.

1:43:32.310 --> 1:43:37.390
<v B>And so you get sort of sniped. Sniped out of it. Right? Like someone got it before you, just before you were gonna do it.

1:43:38.750 --> 1:44:11.010
<v A>Yeah, yeah, totally. I mean, I think one of the trends was that, like, math, answering math questions was becoming, like, a big benchmark that was very important. And so I think that led a lot of people to the same conclusion. Like if, if, if the goal, like if the, if the metric had been like write the best play or something, then we might have ended up with a totally different system. But I think once people got excited about solving these like high school math problems, I think then, then that kind of set the course for, for all these companies.

1:44:11.730 --> 1:44:41.080
<v B>The answering math stuff was really bad for a long time. So. Yeah, yeah, it's kind of thing and I, I feel, and maybe I'm kind of wrong. I feel a story like coding maybe is, is one of those things that when you think through like leap code style problems, like there are definite like setup where you're given a very high level question and then, and there are benchmarks that already have these in there. But I still feel like performance isn't amazing once you get off the bench like that they've not seen before.

1:44:41.080 --> 1:44:41.360
<v A>Right.

1:44:41.360 --> 1:44:53.850
<v B>So when you give these sort of high level problems and you kind of have a very specific known output for the program and it should be compilable and it should be. It's harder. Yeah. Of course in the math problem, but it seems within striking distance.

1:44:54.730 --> 1:45:23.280
<v A>Yeah. I mean, you know, traditionally in machine learning we have this concept called leaking the label, which means like you know, if you, if, if you took, okay, if you took the examples you trained on for the stop sign trainer and you just fed them back in and you get them all right. That doesn't mean you have a perfect system because you're kind of cheating. Right. Like you might, you might have just memorized all those examples and you can't know anything else. It's possible, right?

1:45:23.360 --> 1:45:23.920
<v B>Yep.

1:45:25.040 --> 1:46:07.420
<v A>But the problem is how do you not leak the label when you're training on the entire Internet? And so I think what they've found is a lot of these cases where like the AI solves math problems or the AI solves leetcode problems is they've leaked the label and the AI is literally outputting an answer that somebody else, some other human wrote to that leetcode problem. And so they've done experiments where they've released like things that they know are not on the Internet and AIs have struggled like mightily with it. I feel like until we get proper calculator use and tool use more broadly, I think it's going to be very hard to, to for AI to solve these problems.

1:46:10.630 --> 1:46:29.350
<v B>Well, this is, this has been a great topic and very timely. I know you've been working on reinforcement learning for a long time, but I feel it has, like you said, it's kind of reached a certain hubbub in the everyday discussions recently. So I. I'm happy to have a sort of great overview of what it is and what it's about.

1:46:30.070 --> 1:46:44.710
<v A>Yeah, totally. If folks have any questions, they can just reach out on our discord or email or in my case, social media. We need a reinforced learning algorithm to get Patrick on X. That. That needs to be the next thing.

1:46:44.790 --> 1:47:01.990
<v B>Okay. But yeah, I look. So I did look it up. So Dolly one was four years ago. You were right. That was very good. And then Dolly 2 is what I was first trying. And that was three years ago and so nice. So you were very accurate despite it just being off the top of your head.

1:47:02.150 --> 1:48:06.010
<v A>Well, I remember. You know, it's one of these things where like you connect it to stories. Like, I remember there's just Woman. She's very influential in AI. Her name is Fei, Fei Li. And I remember being in this dinner and her and her student were there and she said something like. And this was again a long time ago, but she said something like, oh, it was. Oh. We were writing captions from images. So basically given an image, write a caption so that we could. For accessibility reasons. And Facebook I think still has that in the product today. It's like if you're blind or something, you can click on an image and it'll say what is going on in the image. I remember her saying, that's cool, but it'd be really cool if you could go from the description and create the image. And that always stuck with me. I mean, that was like a decade ago. That always stuck with me. And then I remember when Dolly came out, I was like, wow. It's like the something that I thought was like a joke, but then it really happened. Like is. Is for me, it was like an amazing experience. That's why I remember it.

1:48:06.880 --> 1:48:40.880
<v B>I think people have started to get a little fatigued on the AI thing. And it's hard to know, is it you always hit plateaus? Is it like plateauing in terms of like actual functionality? Is it on. We're on an exponential. And exponentials always look self similar no matter where you look. Right. And you just sort of like we can't feel the growth. And then you tell stories like you're saying, or even about Dolly being, you know, four years ago only and like talk about the. The repaint or the flux now versus Dolly, you know, just three or four years ago. It's not that long and there are lots better.

1:48:41.680 --> 1:49:42.720
<v A>Yeah, I mean, yeah, actually that's a good Point I'll end with. With my. Where I think this is going. I think that despite loving reinforcement learning and everything, I don't think that AI should be making a lot of decisions in isolation. I think that it should be working together with people. And so, you know, the recraft is a great example where it's not just an API you call and get an image, but it's like a experience. And, like, you iterate and you say, hey, I want this to be all different, or, hey, I want this. I want an ax in this person's hand or a phone or whatever. Right. And so I think it's going to be really about collaboration. And reinforcement learning is always going to be really important, but it's going to be important in the way that the actions are more like suggesting things to people. So in other words, reinforcement learning to, like, book a flight for you, probably not a good idea, because if one out of a hundred times, you go to Tokyo by accident. Right. You're gonna be pretty upset.

1:49:43.199 --> 1:49:44.800
<v B>No, I might be happy. That sounds great.

1:49:45.440 --> 1:49:57.100
<v A>Yeah, actually, Tokyo is amazing. Yeah, we're actually. I don't want to say anywhere we don't want to go because we have a listener almost like. Anyways, um, so Antarctica.

1:49:57.500 --> 1:49:58.540
<v B>Okay, so.

1:49:59.340 --> 1:50:22.980
<v A>But. But you still need reinforcement learning to, like, suggest, like, you know, come up with, like, three different hypotheticals, send it to the person, should you text them or email them, et cetera. Like, there's still a lot of decisions to be made, but I don't believe tons of people are going to lose their jobs entirely. I think that work is going to change just like it did with. With the invention of the motor and stuff, so.

1:50:24.740 --> 1:50:31.140
<v B>Well, you heard it here first. The Jason future AI is not too scary coin.

1:50:31.780 --> 1:50:55.840
<v A>Yeah. Yeah. Don't be worried about it. Just be adaptable. If you're adaptable, I think you'll be just fine. And, and, oh, and. And coding will probably be one of the last jobs to be eliminated, by the way. So if you're worried about. If you're worried about. Yeah, please stay in coding. Not just because we want you to keep listening to the podcast, but tell all your friends to get into coding.

1:50:55.840 --> 1:50:56.840
<v B>Stay in coding.

1:50:56.840 --> 1:51:10.200
<v A>If they're worried about their job being eliminated, they should be a coder. That's like, one of the last jobs that's going to go. I mean, trust me on this. Like, we are going to lose so many doctors. We should probably lose all the CEOs before we lose the coders.

1:51:10.360 --> 1:51:12.310
<v B>All right, all right, we gotta wrap. We gotta wrap, guys.

1:51:12.380 --> 1:51:12.940
<v A>We interact.

1:51:13.740 --> 1:51:17.340
<v B>They're phoning me from the other room and telling us we're out of time.

1:51:18.300 --> 1:51:21.820
<v A>That's right. Hey, I didn't say anything about hr.

1:51:23.100 --> 1:51:25.420
<v B>What's that? Oh. Oh, yeah. We're wrapping up.

1:51:25.900 --> 1:51:35.740
<v A>All right. This was so fun. Thanks, everyone, for tuning in. And thanks, Patrick, for bearing with me and my rants on Reno.

1:51:36.140 --> 1:51:39.260
<v B>This is great. Very illuminating. This is awesome. I learned a lot today.

1:51:40.450 --> 1:52:20.820
<v A>Cool. All right, everyone, we'll catch you later. Music by Eric Barndaler. Programming throwdown is distributed under a Creative Commons and attribution share alike 2.0 license. Feel free to share, copy, distribute, transmit the work, to remix, adapt the work. But you must provide attribution to Patrick and I and share alike in kind.

