WEBVTT

0:00:16.956 --> 0:00:22.474
<v A>Welcome to another episode. This is going to be

0:00:23.318 --> 0:00:37.307
<v B>a good one. Excited to be here actually because this is a topic I have been meaning to learn about, and Jason has agreed to put on his professor hat? Robe? I don't know what it is, a professor where?

0:00:37.307 --> 0:00:43.820
<v A>I got hooded when I got the PhD. I got hooded, which I thought would be an actual hood, but it's really just a sash.

0:00:45.964 --> 0:00:49.220
<v B>Wait, what is getting hooded? That's like what you get when you get—I don't know about this.

0:00:49.360 --> 0:01:10.754
<v A>Okay, so when you get a PhD, you get hooded, which means you go through the same ceremony as the master's students, or I think the same ceremony is for everybody, but you get a hood, which is actually a sash, and your PhD advisor actually puts the sash around you over you as part of the ceremony. Okay?

0:01:11.074 --> 0:01:19.799
<v B>I feel like maybe I've heard that term, but I always just kind of had some weird, probably bad association with 'hood.' Winked, but uh anyways. Okay, where are—

0:01:22.482 --> 0:01:34.969
<v A>We off topic? Anyway, so it's funny because I actually, I actually the first thing I think of is is actually because I grew up in the inner cities. I was like, okay, we're going back to my childhood here. Oh oh, okay.

0:01:34.969 --> 0:03:33.060
<v B>Interesting. Okay, wow. All right. So today we've learned there are many, many associations of the word 'hood.' Uh so um, okay, that we didn't even talk about cars yet, so um that's true. You get a new upgraded carbon fiber hood for your car—you get hooded. Okay? Um way of people are like, what is going on? What are we listening to? Uh well, that's how you know we're not the AI; they would not be this off topic. That's true. They would definitely stick to the script. The AI is not allowed to say this stuff. You would definitely be pushing the downvote button. You know. Oh wait, people probably are now. Um yeah, that's why this is not live. Okay? So for and I'll keep it brief because I actually want to get to the meat of the story today. Oh no, pun intended, but talking about cooking outside. I had a grill on my like back patio which I would use to cook food occasionally. It was uses these little pellets of wood, so it's called like a pellet grill, so like pellets feed down and it burns it and it makes the heat and has an electronic controller. Um and I would do some like you know smoking on it and some grilling anyways. It broke, and um it's old, so you know. Okay fine. So I went to go like okay, I can get a new grill. I did not know there are so many different kinds of grills that are you know like popular now, and I feel like growing up my parents always just had a yours are kind of one of two things: you were the charcoal Weber grill, you know, with the like the bowl and the charcoals, or you had the propane grill that you know like you had the tank and you hooked up the hose. And that was two, but now there are like you know all sorts of things where it's you know infrared cookers where the propane goes into like some sort of like catalyst something and like turns hot, like the patio heaters that you know are at restaurants sometimes. Yeah, it's like is that better? I don't know. And then there's these like egg-shaped—I guess they call them Kamado Grills, like Big Green Egg—and

0:03:33.060 --> 0:03:58.035
<v B>Kamado, and they're like big ceramic things. Um then you can get like various kinds of cabinet smoke, like anyways I I just maybe I'm naive in my I just bought something straightforward and simple, and then I like oh, I'm overwhelmed by the tyranny of choice. That that's just all I was going to say if you've never kind of looked up grill technology, it's it's actually kind of crazy. There's got a lot of choices here, so now I don't know.

0:03:58.457 --> 0:04:12.970
<v A>What to pick? That's what. I have a propane grill and then I have a smoker. I have a separate electric smoker that takes the pellets and smokes meat. Um okay, but yeah, I think the Green Egg can do both, so it's like a two-in-one. Um

0:04:12.970 --> 0:04:24.023
<v B>And yeah, it's just wild. And then there's like pizza ovens now people are doing. So like when I went to go look at the grills at the hardware store, it was like there are also pizza ovens here. And yeah. Okay. Yeah.

0:04:24.647 --> 0:04:41.337
<v A>My neighbor has a pizza oven, and I think he's used it twice in four years. Well, I mean, you know how many times do you eat pizza? I mean, and also it's like if you eat pizza, you're often in a hurry, so you're either ordering it to go or you're doing the regular oven because you're in a hurry. Okay?

0:04:42.045 --> 0:04:48.036
<v B>All right. So you're down on the Pizza Oven? That's a short on the Pizza Oven stocks for Jason.

0:04:48.374 --> 0:05:09.839
<v A>Yeah, I mean, I'm not a big fan of the Pizza Oven. I think we've only ever endorsed one stock on my entire programming throwdown career, and that was DataDog. And I think it's like up just as much as everything else. I made one, like in hindsight—in five years of hindsight—like relatively neutral endorsement everyone's

0:05:10.345 --> 0:05:17.989
<v B>Bracing themselves for the meme coin announcement now? Oh, we need a programming throwdown coin? No, we don't. I'm not rug pulling people. Okay?

0:05:19.255 --> 0:05:27.878
<v A>I know. We know. Okay, we got to keep going. We got to keep—we are not rug pulling people. We went the opposite way; we stopped doing ads, so it's like the opposite of rug pulling.

0:05:28.975 --> 0:07:27.360
<v B>People. All right. Time for news of the show. So I've got the first one, and this is an article entitled 'You Can't Call Yourself a Senior Until You've Worked on a Legacy Project.' So talking about what is a senior engineer—this is like an age-old debate, whatever. Anyways, this person was kind of pointing out how they hadn't really worked on legacy code base. There's some specifics here of their thing, and I, you know, if you want to go read it, good article. And the point though is pretty interesting that they kind of rightly wanted to avoid working on a legacy code base. They ended up kind of doing it. They were right; it didn't like it, but they actually learned a bunch of stuff that didn't. And I think a couple interesting takeaways for me from the story and just you know thinking on the topic is about regardless of the label of senior, like just growing as an engineer is no matter what work you're doing finding the takeaways that are applicable. And lots of analogies—the one I've taken to using recently just for myself and for others that I talked to about this is like just really trying to compound the growth. So not just thinking like, 'Hey, how do I do this thing?' but like, 'How do I think about additively?' Like applying things I've learned before in a way that my growth sort of grows on top of itself and you're sort of stacking it up. And sometimes you need to widen the base, right? Of like expanding into new things, but other times you're trying to build up and trying to apply these different experiences. And so I think partly this plays into that, and then there's an observation here specifically about legacy code bases in your place of work and understanding why maybe something isn't done that way anymore or why the stuff that you see in like the kind of current pieces of tech are how they got there, right? People will say, 'Oh, it's organic growth,' or you know whatever. You kind of get there. But I think there is something different between saying this is the

0:07:27.360 --> 0:09:12.906
<v B>current recommended practice and I have done it the other way, and I will tell you, the other way sucks like we're doing it this way. Those two things come from slightly different places, and understanding why you do something—not just there is value, and actually just not knowing, 'Oh, this is kind of bad.' But when you get into style guidelines and stuff, I think just picking away and having everyone do it is useful because it really does matter. But then there are things also that even if you don't always know exactly why eventually kind of figuring them out and digging in. So the one I always use in C++ is the ternary operator. So you can write this boolean expression, put the question mark, and then the thing that is true first, and then a colon, and then the thing that is false if it's false second. And you can use this, and we have it banned in our code base. And the reason why is it literally does nothing unique. You can't. Okay? There's like some very rare, you know, use case someone could come up with that you know in a constant expression or something, but for the most part, you're just simplifying writing an if-else statement, but the cognitive load to read the ternary operator—make sure you understand what it does and that a new engineer showing up has the same practiced expertise at reading that. Like why? Like why just because features are there doesn't mean you have to use them. And I try to explain this to people, but I would argue even not knowing my explanation and still doing it leads to good practices. But knowing why and having tried using all of the whiz-bang features from the latest, you know, C++ update constantly and refactoring code just to rewrite it into those features—having done that once and burning your hands—probably teaches some lessons, and so legacy code bases can be really useful.

0:09:13.362 --> 0:10:43.964
<v A>Totally, totally agree. I mean, the equivalent in Python, which is even more confusing, they have a ternary operator where you can say like x equals three if foo is true else five. So it's a ternary operator, but that you switch the first and the second like position. So it's like so it's like even harder to read. And you know, oh yeah. And so I remember like there was someone on my team who would do this a lot, like all over the place, and you know, I let it go. I didn't really push back on it because to your point, like it's not until you ban it—it's not banned—and so you can't really say like 'Don't do this' because you have no like moral grounds other than your intuition and then it was a disaster. And so like people just kept getting burned by these like really long in-line, you know, like conditions, right? And so now like I can ban it, and I don't feel insecure about it or feel like hesitant about it. I don't feel like, 'Oh, you know, it's not really against the rules,' because now it's like no. Like I've done this. I've seen people like cause all sorts of issues, and some issues went to prod, and so we're not doing it. Now that the one of the tough things is, you know, if you're talking to folks who don't have that experience, you have to like ban it in a way that shows empathy and doesn't create any resentment or anything. And there is

0:10:43.964 --> 0:11:26.843
<v B>this balance that I think is useful but often gets brushed away—that bringing new folks and you actually want them to feel empowered to question. So when they see that there's a ban on this and they say, 'I love the ternary operator because it makes me look cool.' You know, they're going to say that part. But you know, I love the ternary operator. Why is it banned? I don't think it should be banned. And you actually want to take time to explain them and in some cases be willing to hear them out and maybe you know adapt your practice or be flexible. But in other times, like you said, I think the word confidence there is like, 'No, we've done this.' Like, 'I hear you, but you're just gonna have to trust me that like we've tried it the other way and the other way like not banning it leads to problems.'

0:11:27.619 --> 0:12:25.230
<v A>Yeah, yeah, totally. Yeah, I mean, I just to wrap this up, a question I always ask in interviews: If I'm doing a technical design interview, I'll always start with the question of like, 'Tell me a time you refactored something and why did you refactor it?' Like what led to the decision to do a big refactor? And that usually opens up all sorts of interesting things because, you know, people—you know, the worst answer is the one that totally neglects the conflict that comes from scarcity. Right? It's like you don't have enough time, but the code is garbage. And it's like so it's like that creates conflict, and then you have to resolve that conflict one way or the other. That's interesting if someone's like, 'Yeah, you know, I rewrote it, and it was the right thing to do, and everyone agreed with me from day zero to the day I wrote it, and everyone praised me at the end.' It's like, okay, well, you know, that's a little unrealistic. I was

0:12:26.884 --> 0:12:32.655
<v B>gonna ask you, does anyone ever tell you because I didn't write the code so therefore it could of course be better? I've definitely

0:12:32.655 --> 0:12:33.120
<v A>I've got

0:12:35.372 --> 0:12:35.963
<v A>That. Yeah, I've

0:12:35.963 --> 0:12:39.135
<v B>had people that's how people really do, but I would be surprised if someone said it. Yeah.

0:12:39.440 --> 0:13:52.322
<v A>And people say like, oh, it was you know another team and that team, you know, we inherited their code, and it was garbage, and so I rewrote it all. And it's like, okay, not the best answer. All right, my new story is Recraft might be the most powerful AI image platform I've ever used. Here's why. And it's a Tom's Guide article. Um honestly, like Recraft is definitely the most powerful AI image system I've ever used. I found out about it yesterday, so I haven't ever heard about it. Yeah, this is very obscure—isn't the right word—but like, I'm really into this stuff, like regenerative AI. I'm following it closely, and I hadn't heard about it until yesterday. Um it can do some amazing things. For one, it can produce vector art, um like SVGs. Um now, the SVGs are like if you or I were to create a stop sign, for example, we'd create like a white octagon, and then we'd create a red octagon inside the white octagon, and that's how we'd get a stop sign with a border around it, right? But if you use this program, it's going to give you like a red octagon and then a bunch of white polygons around it. You see what I'm saying? Like it doesn't have a concept of

0:13:52.322 --> 0:13:54.449
<v B>each edge. Like each edge would be. Yeah, yeah, yeah.

0:13:54.819 --> 0:15:21.642
<v A>It's all one layer, so that's not ideal, but it's a step in the right direction. It's the only thing I've ever seen that really will give you an SVG, like even if they're doing it post-talk or something. Um it's extremely like responsive to the prompt. So for example, one thing I've tried with a lot of these AI systems is I've said like a person holding nothing in their hand because like I'll say like a person doing XYZ, and I'll get a person like holding a phone in their hand. Now, like, and I'll type the same prompt like with nothing in their hand, and then they'll have two phones in their hand. And then it's like it's like it has a hard time especially with negatives. So then I'll try like a person unarmed, you know? So it's like there's not like a negative there, and it'll still like not work. But with this system, it's like very good at following instructions, even negatives. Um so uh it's phenomenal. The other part of it is it's got this cool workflow where you can take an image and then you can say, 'Okay, now put a phone in their hand,' and it'll make like a new image of the same person with a phone in their hand as opposed to like, you know, getting a totally different person. So um it's really cool. Um very, you know, easy to use. Um relatively cheap. Uh so I would highly recommend folks check it out. It's pretty neat what.

0:15:22.941 --> 0:15:35.091
<v B>Are they using kind of under the hood? Do you know like is are they using their own? So some of these Craft up other Stable Diffusion, whatever, you know. Is this like a layer on top of stuff? It sounds really good.

0:15:35.091 --> 0:16:13.330
<v A>But yeah, this is a totally custom thing. Okay. Um it's a pretty big model. It's a 20 billion—actually, the 20 billion is the V2. There's a V3 model which I think is even bigger. Um they they're not open source, so we don't really know what they're doing. They might have a blog post about it. I haven't seen one yet. Um I'm assuming it's the same type of technology where you're doing like self-self-attention and then you're doing like, you know, masking and trying to uncover masks and whatever. Like the—I don't think they're really pushing the envelope on the base model, but then they built a bunch of really impressive things on top of it.

0:16:13.459 --> 0:17:40.034
<v B>That's awesome. I think that's one of the debates like where's the magic? Is it in the UI and the like you know higher-level abstractions, or is it in the base model, the deep stuff? You know, I don't know that. We have an answer. I think there's lots of opinions, but I will say I would say that I tried the DALL-E when was DALL-E popular? Oh, that's—that's four years ago. Two years ago? No, that long. I think so the first DALL-E. No, I don't know like whenever it was making the big rounds. I feel like it was maybe two or three years ago. But okay, okay, maybe four years ago. Anyways, I recently did download one. I have a MacBook Air and I downloaded one of the like on-device Stable Diffusion to run Flux, and they just have like an app you can download so you don't have to do it the command line with Ollama or something. You just like download the app, and it will download the model from—I think a Hugging Face link or something. Okay? And it downloads the Flux, and you'll generate it, and it takes I don't know, it's like 15 to 20 seconds or something to generate an image, but it's crazy that like first of these images are much better than when when DALL-E was making the rounds initially. You kind of wrote it off like it didn't really obey your prompts. It would make cool pictures, but like anyways, so now the Flux stuff and it runs like on my computer, and it's free. Like the models are open source, the program's free, so it's it's running locally. There's no subscription, you know? And it, you know, obviously it has to be using less power because it doesn't have an external GPU or anything.

0:17:41.164 --> 0:18:13.400
<v A>Yeah, the Flux models are super, super impressive. There's a there's a package called M-Flux which will—which is Flux optimized for MPS, optimized for the Apple processor, so you can you can use M-Flux and it'll run in like half the time or a third of the time or something. It's really impressive. Um and and this Recraft thing really takes it to another level. So to your point, it'll only be a matter of time before there's an open source version of Recraft, but at the moment they—they have a monopoly. Well, that's.

0:18:13.400 --> 0:20:07.082
<v B>Another debate: was open source or closed? But we'll just keep moving on. Yeah, right. Uh open wait. Yeah. Oh, oh yeah. So okay, no, no, no. Okay, okay. Um my next one is NASA has a list of 10 rules for software development. Uh and this person is taking the sort of like publicly disclosed list of software rules that um I've bumped into before being an embedded engineer previously in my career that for writing C code, but they've tried to extend it to C++. As well. There's like a set of embedded guidelines for not doing, and um this individual is taking a sort of—I'll say a kind of critique of some of them and why maybe they don't make you know much sense or other stuff. But if you've never seen them before, I will say it is somewhat interesting to uh and it's a lot harder if you use something like Python or Java to I guess derive value maybe from it, but if you've ever programmed in C or C++ before, or you use Rust—um probably applicable as well, or one of the other sort of systems programming languages. Go. Um kind of looking through there and seeing like how would you approach if these were your rule sets? Um leaving aside I guess this the blog post is kind of talking about how maybe the rule set uh could be improved or doesn't make the most sense, but I'll say if someone—if you showed up on a job and this was the restrictions because it was a contract that you were trying to like how would you accomplish this? So things like never using dynamic memory allocation. And then, you know, there's kind of two approaches you end up with. One is so I can see plus plus the standard library generally just uses lots of allocation under the hood, so you end up with—you know, of course you can't use that. So some people do a lot of like static sizing of things up front and trying to, you know, have all of their things of a known size. Other people isn't there.

0:20:07.082 --> 0:20:13.512
<v A>Like a concept in C++ called arena where like you define like a thousand spots. Yeah? Okay. Yeah. So then

0:20:13.512 --> 0:21:32.672
<v B>Other people use a memory pool. An arena is like a kind of memory pool where they basically write their own sort of very thin sort of memory management, but it doesn't necessarily suffer from some of the same problems, and so you can use containers that adapt to that. But if you think in your head how would you make sure that your code never did some of these things? Never had a loop that couldn't exit, right? So all loops need to have an upper bound. Um And just code coverage—how would that work? Um These kinds of things uh some of them are again like pretty restrictive. Like every function has to be smaller than can fit on a single piece of paper with uh you know sizing given. And it's like, well, that is probably good practice, but maybe, you know, sometimes you know you want to change it to be one thing or another. But it is worth reading if you've never read an embedded like rule set like this before. They're not uncommon, and you can occasionally if you work in embedded space bump into places where this is the there has been an explosion in sort of like processing power and real-time operating systems and the the just the complexity and abilities of the processors. But I still think there are some places where uh there are probably many lines of code being written having to follow these guidelines. Yeah, this.

0:21:33.550 --> 0:21:52.365
<v A>Is fascinating. I mean, this is a whole new universe for me, but this is a really interesting. I mean, this is definitely a person who is kind of like—I think the general critique here is on C. Like this person is like, 'Uh, you should use another language.' Yeah.

0:21:52.365 --> 0:21:56.837
<v B>Like use Ada, not C. This is just this is a valid commentary, but yeah.

0:21:58.575 --> 0:23:58.379
<v A>All right, so my next news story is AMD Radeon RX 9070 XT performance estimates leaked. Okay, so I want to go do a little rant here. I hate kind of complaining about products. Like I feel like it's maybe not the best use of the show, but I bought a PC, like a mini PC with a Radeon, and I used it for a little while. It was okay. The drivers were really buggy. I had to go into Safe Mode and some stuff to get it working, but then I got it working. You know how you can plug USB to DisplayPort or USB to HDMI? Like you have these cables. I don't actually know how they work under the hood, but there's some magic that allows you to go from a USB-C port to right into your display. I think it's like something called DisplayPort Pass-Through or something. Anyways, I plug one of these cables in, pop, the GPU blows up, like hardware blow up dead. Where did the cable come from? It's a cable that I use in my MacBook Pro. Like I've used this cable for years. Okay, so it's the cable is fine at least for the MacBook. Yeah, plugged into his mini PC and it just popped, and I guess like—I mean, I'm kind of dogpiling here. I almost feel bad, but like you know George Hot says this post on Twitter about this. Like he's trying to build these tiny boxes that use like that run that run PyTorch, and they use CUDA or or or Rockthem, which is AMD's equivalent. But like the AMD—like, you know anything above the hardware just sucks, and it's kind of disappointing because everyone wants there to be an alternative. If nothing else, not only for price, but maybe they could do something interesting, like maybe they could make a card that has like a gigabyte of RAM.

0:23:58.379 --> 0:25:02.024
<v A>And is not very fast, but just has a ton of RAM. Like you know when there's multiple people in the pond—like there's multiple ideas, right? But I was just super disappointed. I mean, I'm one of like a long line of people now who will just not buy these AMD cards. And I guess maybe just to turn this into a question, like how does AMD kind of recover from this? I kind of feel like if I was them, I would hire like some software person, like some person who's really high up in the stack, to lead like a whole branch of the company to just go through and—and just ruggedize everything from a software perspective. So it's like I feel like this is more in your area because like you know when it popped, I was just shell-shocked. I just, you know, so like what causes hardware to kind of fail like that? You know, and not be tested? And what do you think you would do if you're CEO of AMD, Patrick? How would you save this?

0:25:05.601 --> 0:27:04.485
<v B>Gosh. I, you know, as much as has been in the news with NVIDIA GPUs and stuff, and the scarcity and the crypto stuff and now the AI stuff, I don't know. I'm not super up on like where the profit margins come from. I know AMD of course makes processors as well, and I actually have AMD processor and GPU and the PC I built, and I feel like I got, you know, the Ryzen processor was like a good value for dollar over the Intel one at the time. Maybe you know they used to be two separate companies now like they're combined. I don't know how internally their company is structured where the profit margins are on GPU, like you said. You may say oh there's a, you know, large group of people who would really buy a, you know, large amount of memory, not that much, you know, processing power, but it's possible that like the die cost and the spin-up for that, especially when you know NVIDIA could basically pay premium for any foundry costs because they have, you know, far more supply than demand. If you're sitting in second place, I don't know, you may be forced to basically pay more for foundry costs and things in order to be able to get your chips made, right? So it may make it difficult. They may not have as much freedom as they otherwise would want. And I think the software drivers are hard as well because there's so many people you need to appease, right? You need to appease like people who are saying like, 'I just want to plug it in and get a monitor working.' But then the video game people are like, 'I want variable frame rates.' And then you need to appeal to the video game developers who are needing to optimize the outputs on your card and the wrappers that sit on top of it—so OpenGL or DirectX or whatever. Like there's all this different stuff swirling around, and I actually just feel like this—I don't even like that space just seems so complicated that.

0:27:04.620 --> 0:27:39.660
<v B>You know, I gotta plug into a variety of motherboards. There's a variety of power situations, a variety of connectors, a variety of like the amount of compatibility you need on a GPU—almost honestly, like on par probably more with than almost any other part of the PC is just actually bonkers that needs to be compatible with all manner of software, all manner of, you know, OSs, you know, all manner of hardware internal to the computer, external to the computer. That's a lot. Maybe that stuff's all really robust and ruggedized, but I imagine just a ton of time to get to get nailed down exactly right.

0:27:39.660 --> 0:28:27.443
<v A>Yeah, that's a really good call out. I wonder. And also like a lot of these things are very thankless. You know, it's like the guy who makes sure that you could plug the USB-C to DisplayPort versus the—because because I had this machine working DisplayPort to DisplayPort, so I know that the machine worked, but then as soon as I plugged in this other way, like I heard a visible, like I heard audio kind of pop, and it was done. So like you have to have a person to test all these different ways, and then unless it breaks, that person is not adding any value. Like they're just reducing risk, and you don't know what's risky, what's not risky. So this is one of these things. It's almost like a really high-level kind of performance management, like value of the company kind of thing that you have to fix.

0:28:28.236 --> 0:29:39.162
<v B>There's and then on your specific issue—I mean, it could be a faulty card that, hopefully, they would warrant to replace. And then the question: if you did it again with the same cable with the new card, would it happen again? I don't know. I mean, I wouldn't risk it. But like yeah, there's two failure modes. There's a failure mode of an individual card, and then there's a design flaw that like every card where you did that—it would sort of break. And I don't, you know, from where I am, I just don't know enough about which specific problem it is, but certainly, like you said, it makes people have very bad sentiment. But even if you look at the and again not knocking, but like the NVIDIA consumer cards that were having like the plug—you plug into the GPU to give it extra power directly from your power supply—had so much power going through it that they were like melting. They need to like have cooling on it. Like again, like you said this is sort of thankless, right? The dude or dudette who is trying to design the interface for some wire copper cable to come in and deliver power, and all of a sudden they never anyone ever thought about them in their whole entire lives, and now they're like on front page of social media because these expensive graphics cards are melting. You know.

0:29:39.735 --> 0:30:45.919
<v A>Man, this is gonna be a distraction within a distraction. But like one thing that's really interesting is like what jobs are of that sort where like if you do a good job nobody cares? Or like if you do a good job nobody notices? And what are the jobs because it feels to me like in every career there exists, like maybe—that's maybe I'm trying to stretch too far here. In many careers there exists like jobs where you get praised for doing a good job and nothing happens if you do nothing. So like jobs where you're trying to bring in more business or raise the profit of a company or something. And then there's—there's jobs where like you're trying to keep the lights on, and so people are kind of it's really hard for people not to ignore you until something goes wrong. It feels like there's these two kinds of jobs, and it feels to me like the former is almost always better than the latter in terms of your satisfaction and a lot of other things I was having.

0:30:46.847 --> 0:32:14.580
<v B>This conversation, and I won't—I won't give the details because it'll make it sound like a political statement one way or another, and we're not trying to make it. But anything where you're—and in this case it was, you know, government officials, whatever—anything you're dealing with is a probabilistic event. So something may happen, may not happen. Even if it happens, the certainty isn't well known, right? You think like weather forecasting, you know, whatever. Like yeah, and when you get it right, it's sort of like no one—you could just do nothing, right, and it would probably just be fine. And then that one time, you know, it's out of standard deviations, like very high, and then everybody's like, 'Why didn't you do?' And it's like, 'Well, it's true, probably could have done better or done these other things.' But you could have been doing those other things every single time, and it either still wouldn't have made a difference or wouldn't have been relevant.' And so to your point, anytime you're dealing with something where you know like think of preparing for holiday rush on a server that's doing e-commerce, right? You could spend tons of money like building up, you know, extra CDNs and having flexible compute. And then, you know, like people come, they shop. There's no outages. No one knows if you did a great job or a bad job, but it didn't crash. So like there's that. But if it crashes, certainly you're getting hauled in and told how much millions of dollars you lost the website, right?

0:32:15.863 --> 0:32:22.400
<v A>Yeah, I mean, it's pretty tragic that you know things are just set up that way, but I don't know if there's not any clear solution or.

0:32:22.400 --> 0:33:35.024
<v B>Anything. All right. Well, book of the show, book of the show. What's your book? All right, my book. This is going to be a little bit different, but I've not been reading as much as I should or I want to, but I did just start because I have never read any books by this author before, and I often see it recommended in science fiction. So I decided I'm going to read a book and try to find the recommendation. And this is the recommendation I got, and that is a book by Ian M. Banks, and I chose The Player of Games. Have you read any Ian M. Banks books? I have not. Okay. Um yeah, neither have I. So I am trying to embark on reading this one. Apparently some of the books can be a little hard to read, which is okay, but this one people are saying is a good introduction. There's not from what I've seen online without trying to read any spoilers, apparently there's not like a strong order you need to read the book since it's not necessarily the first book, but it is an often recommended one. So I'm starting here. That's not really a good recommendation because I don't know whether to tell you it's good or bad other than every other person I saw on the internet this seemed to be bubbling to the top as a good starting place. So I'm embarking on a journey.

0:33:35.412 --> 0:35:30.460
<v A>Here. Cool, that sounds awesome. Yeah, I might check that out. I have a few books queued up that I have to get to, and then I'll check that out. My book of the show is Basic Role-Playing Universal Game Engine, which is a reference book. It's written by a couple of people who have been making tabletop—tabletop games for their whole adult lives, and they've made a bunch of them. I think some of them even from like the '70s and the '80s, so it's kind of wild. I mean, some of the stories of the creators, but basically, they synthesized. So there's this question in general: like think about any art form. You always think about like what is the essence of this art? You know, like you might as an artist like draw a bunch of things and then think to yourself, 'Okay, what is like the essence?' Like if I had to reduce something to its most basic form, what would that be? And so these guys got together and thought, well, if we had to reduce these tabletop games—because there's a lot of lore built into it, you know, like so many games have Magic Missile. Why? Because Dungeons & Dragons 1.0 had Magic Missile. But like what really is a Magic Missile? I guess it's like an arrow made out of magic, right? Or something. But like you know pair these things down to their essence and just explain like it literally. Like the first—the first page of the book is like 'The point of an RPG is for your players to have fun.' So it's like, you know, it's like let's start from like the first principle. The first principle is that people should be having fun. And it kind of builds up, and so it's a combination of an instruction guide to making a game engine and a reference manual of like here's a list of like hundreds of skills from.

0:35:30.460 --> 0:36:46.060
<v A>Like all these skill-based games we've ever seen—I've kind of to be honest been skipping over a lot of the, you know, here's a list like of of like a million different types of armor because it's not what I'm interested in. And what I'm interested in is like how do people create these game engines and how do they keep them balanced and how do they keep them interesting? And so when I say 'game engine,' just to be clear, it is programming throwdown, but this has nothing to do with programming. It's literally like the math that—that you could use this for like a card game or a tabletop game or really anything. It's like the math that keeps people kind of on the edge of their seat, right? So it's like how do you have all these different options for your players but still keep them kind of on the edge of their seat? And at the same time, how do you do that in a minimalist fashion where it's not like, 'Okay, I'm now gonna have to roll like 45 dice till I build my character or whatever?' So these people tackle a lot of that. And I want to say maybe about halfway through, as I said, skipping a lot of the pure reference stuff, and it's really interesting. I'm having a really good time reading it. I've never

0:36:46.060 --> 0:37:51.102
<v B>played an in-person tabletop RPG or even—I mean, I probably played a video game that somewhere under the hood was running some sort of like roll checks and chance checks or something. I learned the other day, a spoiler, something I'm gonna talk about in a few minutes, that Pokémon was actually doing that when the Pokéball rattles. It's doing like a, you know, probability check, and it can fail at each of the things, and then that's when the Pokémon came out, which I didn't know. And maybe I'm completely wrong, but that's sort of like what the internet was telling me, which is in line to your point with rolling a dice and getting certain values, but never had the occasion to play one. But I'm endlessly fascinated by, like you said, the sort of crafting of the stories and the storytelling and the fact that it's less game than, you know, a board game with rigid rules and more about, like you said, having an adventure together, making it fun and entertaining and collaborative—like collaboratively doing something which is somewhat still gaming but is also you are sort of being flexible on the fly as well to keep it.

0:37:52.013 --> 0:38:11.436
<v A>fun. Yeah, exactly. Yeah, exactly. Like how do you let each person at the table have a unique character that brings something unique while still being able to handle a person not being there? It's like it's like, 'Oh, you know, Jim, Jim's wife is having a baby, so this chest has to stay locked,' you know?

0:38:12.449 --> 0:38:13.950
<v B>it was the one that had the

0:38:15.115 --> 0:38:31.700
<v A>keys. He's sleeping in the inn? Yeah, or he's the only one with lock picks or something? Yeah, yeah. So you know, the book is really interesting. I'd recommend folks check it out. It's if nothing else, it's a nice book to have on your coffee table because it has kind of a provocative title: Universal Game Engine.

0:38:31.700 --> 0:40:06.700
<v B>All right. Well, I spoiled it, but tool of the show for me is a video game. And for whatever reason, skipped every modern Pokémon video game. So I think the last one I actually legit played was when I got Pokémon Red and my Game Boy as a child and played that to no end using. And I was trying to describe this to my kids, um, and like I had to go to when we would go shopping at like The Walmart or Kmart, and I would look in the strategy guide for why I was stuck. So I would go with my mom so I could go to the video game section and open the strategy guide and look because I wouldn't just buy it—I probably should have just bought it—but anyways, and then go home and get through anyway. So Pokémon, right? Anyways, I've been aware, I've you know dabbled various times, but I hadn't really sat down and played. But I was sitting down and playing Sword and Shield. I was playing the Shield variant, but not super important on the Switch, and I just hadn't done that in a really long time. And I know it's a pretty big departure for this series, but it was really kind of fun. Like I was really into it. I realized now that the game is easy—like it's not supposed to be challenging to actually beat the game. So it's not that much of an accomplishment, but you know I had a great time. And if you've been ever interested and you have a Switch or whatever, we definitely recommend checking out one of the newer ones: Sword and Shield, I guess. The other one I'm going to try now is Scarlet—Scarlet—and I think Violet it is, um, but

0:40:06.962 --> 0:40:07.587
<v A>Pearl or something.

0:40:07.924 --> 0:40:45.404
<v B>Oh. Yeah, okay. I did, or that's a different—I think Pearl. I think that's a remake. Boy, I think that's a oh they did like a remake. Yeah. Like so anyways, if you hadn't checked one of those out, they definitely went and some of them like a little bit more with open-world sections, and you can kind of control how often you get into a battle versus, you know, just wandering around in a set of grass until it happens. So definitely some quality-of-life improvements over the old ones that make it less frustrating and ability to save sort of everywhere you want. And so if you never checked one out, I guess this is me telling you the obvious thing of like it's a thing, and it's kind of

0:40:45.944 --> 0:41:33.143
<v A>fun. Yeah, I played it with my kids, and it's been maybe a year or two, and they would get frustrated, and so I'd help them like kind of optimize their characters a little bit, but more or less they could get through it eventually. And yet the other thing is it's all the bosses and everything as far as I know, they have static levels. So if you're kind of like my kids are just running around kind of aimlessly for a while so their Pokémon were like super over-leveled, and that made the game even easier than if you're trying to speed run it. But yeah, that game is awesome. I think. Um, I think the open world added a lot. Actually, like being able to really see the enemies and they run into you physically and then the fight starts. Like that really added a lot, I think.

0:41:34.290 --> 0:42:09.036
<v B>Yeah, and I think that it goes crazy deep though once you look on the internet. There's all this like each Pokémon you catch has different stats, and those are talking about—I never paid attention to it other than like it has a type and certain moves or whatever very basic level strategy, and that was fine to get through the game. But when you look online, you find out, 'Oh yeah, the competitive stuff,' and people playing online and whatever that each Pokémon you catch has like different base stats that have been rolled for that character. I mean, it's not actual dice but probabilistically generated, and so some are better than others even if they're the same level.' All right. So Patrick, do you want me

0:42:09.036 --> 0:42:11.297
<v A>to waste hours and hours of your life is?

0:42:12.833 --> 0:43:30.964
<v A>It going to be fun. It's gonna be fun. Okay. So later on go on YouTube. Okay, there's this guy who really understands the Pokémon mechanics, and oh dear purposely. So there's a—there's a Pokémon I think it's a web-based game. It's probably not legal. It's probably already shut down or something, but there or maybe it's sanctioned, I don't know. But there's this web-based game where you can just do Pokémon battles with other people, and there's the Elo like for Chess and everything—rise up the ranks. And so it's literally just the battling part of Pokémon. And this guy who really understands the mechanics, he makes builds that are very unintuitively strong, and he plays people who and he must play a lot of people but inevitably he ends up playing someone who starts off like making fun of him and like, 'Oh, you just have one Pokémon? Like why didn't you build the other four Pokémon? Haha, you're so trash whatever.' And then he wrecks them, and they start raging, and they start like and then they won't make their final moves. And he's like, 'Hey, your time's running out!' And they just get so pissed and everything. It's it's like the people who get the scammers upset or whatever. It's that but for gaming trolls, and it is hilarious. Oh.

0:43:30.964 --> 0:43:37.326
<v B>Dear, okay. Now you have to send this to me, but I feel like I'm gonna not like you for doing it, but

0:43:37.326 --> 0:45:31.259
<v A>Yeah, I mean, I don't know how many videos he has. I'm pretty sure I've watched like five or six of them. They're really funny. Okay, all right. So, oh my tool of this show is Features and Labels, or Fal.ai. There's a bunch of alternatives: there's Together.ai, there's Fireworks.ai, there's a bunch of them. But basically these are people who are kind of a middle man between you and the AI models. So they'll host the open source ones; they'll often have agreements with the closed source ones so you can run like Google Image, and you know it's you otherwise you'd have to use some proprietary Google API or whatever. So you think of these as like a middle layer, and they often charge you per thing that you do versus like having to rent a machine for an hour, right? The reason I picked Fal is I actually know the founders, so I'll just put it right out there and say I don't know if Fal is any better than any of the other ones, but the user interface is really nice. They have like a playground mode where you can just build things on the web, and then you can click on the API button and get the Python code if you wanted to make that programmatic. The other thing they did which I thought was really clever UX, you know, as engineers—especially as people who have GPUs or maybe an M2 MacBook or something—we think to ourselves like, yeah, I mean, I should just run Flux myself, like Patrick's run Flux. I've run Flux, right? But when you go to Fal, they're like, 'Yeah, so um you can run this model like 87 times for a dollar.' Like it basically for every model, it tells you how many times you can run it for a dollar, and that to me is like really powerful because like often I have some code that I have right.

0:45:31.259 --> 0:46:48.891
<v A>Now, and I have a local version of Flux and then I have the Fal version of Flux, and you can just like toggle between one and the other. And you know, like I'll want to run sometimes I'll think, 'Oh, I'll run the local version because I'm going to work, and I'll just let it run, and it'll be done when I get back.' But then I'm like, 'Yeah, it'll be done when I get back,' or I could spend like seven dollars, and this is just like done in a second. So it's like so it's like they did a good job of kind of really laying out the economics, which are themselves startling. You know how the economics have changed for AI? But they just put it right out there. And they recently had something—they had something where they're able to do some caching of the I think caching of the tiles. The way the image transformers work is you know breaks your image up into tiles, and I think they're caching tiles that are very similar or something. I don't know. It's it's something I don't remember off the top of my head, but it made the price even cheaper. So I guess long story short, check out these folks. They're all awesome. I know the Fireworks people too. That all these services are great, and there's an economy of scale that you can really take advantage of.

0:46:50.680 --> 0:47:52.223
<v B>Yeah, I mean, I think all of it from like someone was asking me with 3D printing how much would it cost you to print this? You know, I saw it in a shop or something, and I was like, 'Well, the most obvious thing is how much plastic it takes.' But then you start thinking about it. There's depreciation of your machine, like wear and tear on your machine. There's like the power to run it. There's my time to like walk out mine's in the garage, like walk out to the garage and like get it off or clean the, you know, build plate.' And so I think what you're saying is interesting too that running it locally to me is—I guess I just cheap. I don't want to give places my credit card. I don't know. And I have stuff, and so I feel like I should use the stuff I have. But you're right. Like by the time you factor in the power to run it, like it's not free. And you know your computer getting hot and the time taken. And so the economies of scale these really big server clusters dedicated to this AI stuff—it's really kind of amazing even with how expensive those really high-end GPUs are.' Yeah.

0:47:53.100 --> 0:48:43.620
<v A>Yeah, totally. I think running it yourself is great. Everyone should learn how to do it, definitely not discounting that. But check out these folks, and they're similar folks too. It's a really neat service if you have something that you then say, 'Oh, I need to run this like 200 more times.' You know, your time is also really important. So all right, on to our topic: Reinforcement Learning. So a bit of background here is like the opposite of like assembly language show where in this case, like this is my background, my area. I know a lot about—I'll kind of dive into it, and then Patrick is going to play the role of you folks, and stop me anytime I say something that is a buzzword in my community or doesn't make sense, or

0:48:45.160 --> 0:50:42.540
<v A>something. So I'll start really broad. There's basically three types of AI: there's Supervised Learning, there's Unsupervised Learning, and there's Reinforcement Learning. So supervised learning is where you have the right answer; it's right there. So someone gives you a picture and they draw a box around the stop sign. They're like, 'There's the stop sign,' and your job is to learn a function that maps the picture to the bounding box of the stop sign, and you're given a lot of these as ground truth, right? And then you're also given a second set that you're purposely not meant to train on, called the holdout set. And if your training did really well, then you're able to interpolate between all the other stop signs that could exist in the universe. And so after training, if I was to give you a new image you've never seen before with a stop sign in it, you could draw the box around the stop sign. That's supervised learning. And under the hood, that works through what we call a loss function. So a loss function takes the output of your model—so in this case maybe it's a bunch of hypothetical bounding boxes—it takes the ground truth, which is the actual bounding box, and it turns all of that into a number where the further the number is away from zero, the worse you got, the worse you did. And so zero, a zero loss would be perfectly nailed that bounding box. Now, a zero loss might not be good, right? Because you want the model to have some uncertainty. Like for example, imagine we're playing Paper Rock Scissors, right? And you play paper and I play scissors with like 100 certainty. That's actually

0:50:42.540 --> 0:51:49.654
<v A>not good, right? Because although I won in this game, you know we know that someone who plays scissors 100 of the time is not playing an optimal Paper Rock Scissors, right? You could just—you can learn that and then just play. Wait, did I get it wrong anyway? You know, the analogy, you could just play whatever counter is what I just said and then just win, right? So so often you're going to output a mixed answer, right, a distribution of answers. And so you're always going to have some amount of loss. But through what's called the learning rate, you don't like totally change your line of thinking every time an example is presented, right? You're just slowly moving in different directions as examples are presented. And if the learning rate is low enough and all the 100 other things kind of stars align, then you'll create like a mixed—you know, a mixed response that is optimal. So that's supervised learning. Did I get that right? Any questions about that part of it or did that make sense? So

0:51:50.684 --> 0:52:19.388
<v B>You said don't the only question I had is you were saying so this makes sense that you have the thing that you're trying to match and then you're tested on something else. But you said interpolate between the results, but it should be possible even with supervised learning not strictly like interpolation to me means like between the points given, but you should even like in your stop sign example, like you were mentioning for stop signs that are new, the idea is hopefully you would also understand that those should be heavy bounding boxes put around them. Yeah.

0:52:20.384 --> 0:52:46.557
<v A>Right? So so what you're hoping is that you can imagine like a manifold, like a stop sign space, and in that space there's like a whole bunch of different kinds of stop signs. And so inside of that space there's stop signs that like look really different but hopefully they're like they're within the space of stop signs that you've already seen.

0:52:48.700 --> 0:53:36.660
<v A>So like an example where extrapolation doesn't happen is—so do we see this with Waymo where people will wear a T-shirt with a stop sign on it? And that's not so that's an example of extrapolation. And in that case the model doesn't really know what to do, so it thinks it's a stop sign. So really, so there's a whole area around what's called out-of-bounds detection and out-of-bounds prediction, and long story short, that's a very very hard topic, but really important. But by default supervised learning will interpolate, you know, at a really high dimensional space, right? Interpolate between all these things that's seen. But like if you give it something totally new, it's

0:53:39.072 --> 0:54:19.605
<v A>going to have trouble. Got it. Okay, so unsupervised learning is where you don't have a loss, like there's not a ground truth, but what you do have is something that's kind of stateless and easy to evaluate. So like the most common example is clustering. So there often isn't like known a perfect clustering. Like you might have millions of documents and you want to break them into a thousand clusters, each one having three thousand documents, and you want the entropy of each of those clusters to be really small. So you want all of them to be like really close.

0:54:21.529 --> 0:56:18.740
<v A>together, right? Um so you might never know the perfect clustering like you do at supervised learning, but it's like trivial to evaluate. So I can like show you a set of clusters. You could put the documents in the clusters and come back with a score. This clustering has a score of seven, and I can make some changes. They say, 'Oh, this clustering has a score of eight.' It's a little better. I'll make some changes. Clustering has a score of nine, etc., etc. And so that's an example of unsupervised learning. And so you're not even really trying to figure out the best way to cluster like a human is doing that, but the computer is just kind of following the instructions and then over time getting a better and better clustering, and you can measure that. And so you might never get to the optimal, but you can get closer. Um so unsupervised learning is a little bit trickier in a sense—you don't have a ground truth. Um now reinforcement learning is, in my opinion, like the hardest of these areas, not in terms of like you have to be the smartest to do it or anything, but the hardest in terms of getting good results. It is the most difficult. And this is because you have all the challenges of unsupervised learning where you don't have a perfect game of Go or a perfect game of Chess to reference, but you also are making decisions. In the unsupervised learning case, you're not really making any decisions; there's a human making decisions or a human-written algorithm making decisions, and you're just evaluating them. But here you have to make the decision. So it's like, 'Here's a set of clusters. How do I make them better?' And then you do that, and then did they actually get better? So if you were actually on the fly designing your own clustering.

0:56:18.740 --> 0:56:22.440
<v A>algorithm with AI, then that's reinforcement learning.

0:56:22.440 --> 0:56:52.122
<v B>But the stuff that we talk about when we say at a high level, like oh Flux or this or whatever it may be, using components that were trained with a variety of these techniques, or use a variety of these techniques, right? So it's not necessarily that a whole—I know what the distinction there is—like a whole program application is one of these. You're sort of talking about like a little lower level. You're saying like one part of that pipeline was done this way kind.

0:56:53.674 --> 0:58:08.785
<v A>Of so, in the case of Flux, that's all supervised learning. So in the case of Flux, it's what's called self-supervised learning where you hide part of an image and you ask the AI to draw it, and then because you hit it, you know what it used to be. And so you show that to the AI and say, 'Hey, you know this pixel actually should be red, but you drew purple.' And so yeah, so that's pure supervised learning. Um there are like, you know, recently some—so another way of saying it is reinforcement learning, like does stuff like takes actions, and supervised learning and unsupervised learning kind of reveal knowledge. So in the case of the stop sign, you know drawing the bounding box around the stop sign kind of reveals or synthesizes knowledge, like now you went from pixels to 'There's a stop sign there.' But it doesn't tell you how to—it doesn't drive a car or turn a camera or take any action. So as soon as you want to take an action, now either the humans have to write that code, but as soon as you want AI to take an action, now you're doing.

0:58:11.586 --> 0:59:42.260
<v A>Reinforcement. Okay. Um so there's a bunch of different kinds of reinforcement learning algorithms, but there's basically two axes that you need to think about. One is offline versus online, and this is just a fancy way of saying can I make mistakes? So for example, the AI that plays Go—AlphaGo, in the beginning of training—let's just stick with AlphaGo Zero. It's all pure reinforcement learning. So in the beginning of training, it's just playing garbage games of Go, and that's fine because it's playing against itself, and it's you can't embarrass the computer. So so it just plays garbage games of Go and it gets better and better, but like you couldn't for example like drive a self-driving car randomly until it got better. Like, you know, you can't do that; you'd crash the car, people would die. Be like total mess, right? So so offline reinforcement learning is where whenever you make decisions in the real world, they have to come with some kind of guarantee. In the case of online reinforcement learning, you can just make decisions in the real world whenever you want at any—at any point. And so that, you know, that's like a subtle difference, but it has like pretty big consequences and algorithms and everything else.

0:59:44.044 --> 1:00:01.527
<v B>So online, it's able to change itself and like update, and then offline it's sort of like you're wanting to make guarantees. You want to know like I—I understand what it's going to do. I've tested it in some way, and I don't want it sort of like changing what it's.

1:00:03.029 --> 1:00:07.298
<v A>Doing? Right, right. So online, you're willing to put any model in production. Um so

1:00:14.250 --> 1:00:34.585
<v A>Yeah, I think it's sometimes they call it on-policy versus off-policy, but it gets the nomenclature there doesn't matter as much. Those are the two kinds. Um okay, and so then there's a second axis or a second kind of switch here, which is value and value-based or policy-based. So I'll go into this. So

1:00:38.432 --> 1:01:47.670
<v A>Um let's say you have to make some decisions, right? And when you go to make a decision, like imagine a choose your own adventure book, and whenever you go to make a choice, I was to tell you like if you make this choice, you have like this percent chance of reaching the best ending, and if you make this choice, you have this other percent chance. Like you would just choose the highest percent, right? And you would just do that. It would be like solving a maze with no walls, right? You would just—you just like pick the highest percent every time until it's a hundred percent and then you would win, right? And so the idea with value-based reinforcement learning is if I know the total value of a decision and I know that for all my choices, then I've solved the problem. I just pick the one with the highest value, and that's just the optimal policy. And so value-based kind of ignores the whole decision part of it somewhat and says the game here really is figuring out the expected value because once I have that, I'm—I'm

1:01:49.864 --> 1:02:10.350
<v A>Set? Um now here's where it gets tough, right? Is let's say AlphaGo places a stone somewhere on the Go board to start the game, and it's playing itself or some other world champion or something, right? It places that stone, and its value is about 0.5. It has like a 50/50 chance of winning the game when it just started, right?

1:02:11.886 --> 1:04:09.560
<v A>But let's say I Jason Go and play the world champion of Go, and I put the same stone in the same position just coincidentally for my first move. I have a zero percent chance of winning, right? Because I'm not even close to a world champion. I'm gonna get wrecked, right? So so you have this paradox where like the value is based on the policy, but if the policy is based on the value, you can see how this is like cyclical reasoning, right? Um and so getting the value is actually really, really hard for this reason. Um and there's—there's several algorithms. The simplest one is called SARSA, which basically says um you make a bunch of moves, the game ends or the episode ends. You stop driving the car or whatever it is, then you just go back and you know what happened. So you assign the expected value. So, 'I turn the steering wheel here, I turn the steering wheel there, I hit the brakes, I hit the gas,' and then I made it home safe. Therefore all those actions are plus one, right? You know, turn the steering wheel, I hit the gas, I crash into a wall. Therefore all those actions are minus one.' And then you feed that into your neural network. You do your training, and um um that's pretty simple. Um you know, that completely ignores the thing we just talked about. There's other algorithms like Q-learning and stuff that try to address some of these challenges. Um it's really difficult. Um it doesn't mean value optimization doesn't have its place, but you know ignoring the—the policy makes the makes it really difficult to optimize. It's good in situations where like there's clearly one good action at any given time. You just don't know what

1:04:09.560 --> 1:04:37.602
<v A>It is, like Atari is a great example where like the actions are binary. There's usually like one good action, like if you're playing Mario or something and you're about to run into a Goomba, you either jump or die. And so you jump, right? But as soon as you get into environments where you need a mixed response—um like Poker or driving a car or really doing anything in the real world—it becomes difficult. Um

1:04:42.141 --> 1:04:45.550
<v A>Any questions about value optimization, or did that make sense? So I guess it

1:04:45.550 --> 1:05:17.400
<v B>Makes sense. So I think what you're in these cases, you're trying to, like, you were saying is there's like understand the outcome of a game. So you're playing a game; you don't know what's going to happen. So you don't know if it's a good or bad move until sort of like it's too late. But by observing many, many, many games to their conclusion, you're hoping that when you go to do it—I guess that's that's offline, but for real—that you've built up an estimate of given a context what is the likelihood that each decision is good or bad.

1:05:18.338 --> 1:05:31.467
<v A>Yeah, right. And as your values improve, your policy improves, which means your values now are all inaccurate, and so you're just like kind of iterating on this over and over again. Yeah, but you're totally—you totally.

1:05:33.694 --> 1:05:55.952
<v A>Nailed it. Um okay, so the other type of algorithm is policy optimization. And in this case, you say at least in the most naive example, you say I don't even really care how good this action is. Like I don't need to know what's my expected value of taking this action or anything. All I want to know is um I

1:05:59.209 --> 1:06:23.121
<v A>want to take actions that are good and I don't want to take actions that are bad. It's like I cooked some eggs; they were delicious. I want to do more of that. I touch the stove with my hand—not delicious. Don't want to do that anymore, right? Very simple. So so policy gradient is basically, and I'm gonna try to do a lot of hand-waving here, but basically it says get the

1:06:24.825 --> 1:06:58.694
<v A>expected value of this action. So you do have to kind of like play out a whole series of events, but then when you go back and you look at what happened, take the things that were good that had positive um um positive value and do more of them and do them in proportion to how positive the value was. Take the things that had negative value, do less of them and do it in proportion. So things that are really negative value, really do it a lot less, right? Um

1:07:01.073 --> 1:08:37.260
<v A>it's a very simple concept. Um one challenge right off the bat that you can see is your expected value has to be centered at zero. Like in other words, if all you can do is get points but you can never have a negative score, then your system is going to say do everything infinite amount of time and it's just not going to be able to learn. So so you need what's called a baseline such that you hope that roughly half the time you're getting a positive score and half the time you're getting a negative score. Now for something like Go, it's trivial because you play a game and you either win or lose, and so unless you're playing like someone way out of your league in either direction, you're hopefully going to win and lose about half the time. You get a negative one for losing, positive one for winning. You're all set. So Go makes this very easy, but in the real world it doesn't—doesn't work that way. And so it actually like like figuring out how you can get half of the expected values to be positive is really hard. And so you actually use a second neural network just to figure that out, and that's called the Critic. And so the the if you hear the term Actor-Critic, the Actor is just the policy gradient that I talked about earlier, and the Critic all it's doing is it's trying to figure out the baseline. It's trying to figure out the average move what would be the expected value of that so you can subtract that out and hopefully get us balanced between positive and negative.

1:08:37.260 --> 1:08:48.128
<v B>And so this is the neural network you would use on a Go board to tell you like how good your situation is. Yeah, exactly.

1:08:49.090 --> 1:10:23.151
<v A>So if you were—so using Go for an example, let's say you're trying to solve Go with a policy gradient, so you would say um I took a bunch of actions. I won. Um now I'm going to look at this one action. I got an expected value of one because I won the game. Um now what was the average expected value? Um actually—uh sorry, what was the advantage is what I need to know. So I need to know was this one like for example, like is this am I playing someone who's like a total chump and so like even though I won, I can't really learn anything, right? Or did I play someone who's like a grandmaster and actually learned a ton by winning? Um That's your sort of advantage function. And so you're going to take the value function of the current state. So at this current board state what's the probability I win? And then you're going to take the action you took and see what's the value at that state. So if if the probability of winning the game is 50, but after I took my action it jumped up to 60, then I know that taking that action like caused an extra 10, and so that's my advantage. And so you know if you're playing someone of your level, half of your actions are going to cause your win probability to go down and half of them are going to cause your win probability.

1:10:25.682 --> 1:10:58.099
<v A>To go up. Got it. Yeah, so um as I said for Go you don't need it as long as you're doing self-play, but for for you know something like Atari where there is never a negative score, you need to know like what is was a bad action? And so a bad action is one where your expected score went down after you took the action. It's like oh, I took the action to run into the Goomba and now my expected score is a lot lower because I have one less life to go and collect points with.

1:10:58.369 --> 1:11:06.030
<v B>Is there a conversion between lives and points then, or it's just simply that because you lost a life your maximum point that you can get

1:11:06.030 --> 1:11:52.200
<v A>is reduced. It's the ladder. Yeah. So all of that has to be inferred. So um now you could make it explicit. So you could say and this is something that we should talk about—it's called reward shaping. So let's say in Mario your goal is to get the most points, but that's kind of a really weird goal, right? Because often like when we play as humans we don't even look at the score, right? So you might come up with proxy goals. You might say well every time I eat a—well, every time you eat a mushroom, you get points. Actually, but but let's pretend you did it. It's like every time I eat a mushroom, I'm going to give myself extra points. Maybe you don't get enough points for eating a mushroom, and if you gave yourself more points for eating a mushroom, then the AI learns easier. Um and

1:11:53.939 --> 1:12:23.166
<v A>and you're to your point like you don't lose points when you die, but maybe you should like maybe if you lost 10,000 points every time you died, the AI would learn a lot easier. And so this is called reward shaping. It's where you instead of learning the task at hand, you learn a new task because uh those two tasks kind of go up and to the right at the same time; they're correlated, and and the new task is just easier to learn. Yeah. So

1:12:24.179 --> 1:13:09.400
<v B>I mean, I don't know the scoring of Mario either. I never paid attention, but I guess like you could get weird states otherwise where you try to get a bunch of one-up mushrooms and then on a particularly high value level, like basically keep dying almost at the end repeatedly in order to like keep gaining points for that level, assuming you don't lose points when you die. And so you could get this very like undesired behavior because it realizes the way to get the maximum score is to keep dying and replay the level and gaining those getting those points when the actual thing you wanted was like kind of get through the levels as fast as possible and you didn't really care about the score, and used us it, but it was like a bad proxy for what you wanted.

1:13:09.400 --> 1:14:38.020
<v A>Yeah, exactly. Exactly. And then on the flip side, like you might say, well my goal is to get as far to the right as I can, right? Um, but if you don't take score into account, it might just be really hard for the AI to do that. Like the AI might just like desperately do like kamikaze jumps to the right where it gets killed because uh it didn't or the AI might just not be incentivized to get mushrooms and make it more survivable—that kind of stuff. So so yeah, reward shaping is a really big part of the problem. And um um and so when people go from manually designing systems, like if this ad has this chance of getting liked, then put a highlight around it or something, when people like build these things by hand, they they have to deal with like all these conflicting metrics. And how do you like reconcile? Like, you know, oh, we showed more ads, but there was like more kind of like racy unethical kind of photos, and how do I deal with that? And so reinforcement learning moves that problem to the reward shaping phase, but it doesn't really get rid of it. You're always going to need to like better and better understand kind of like your goals and the nature of the problem and how to uh uh kind of um how to best like solve that problem.

1:14:40.562 --> 1:14:44.967
<v A>Um okay, so um okay, so let's dive into the offline part.

1:14:46.975 --> 1:16:42.000
<v A>So you know one thing a lot of people wonder is, yeah, like AlphaGo plays against itself. And so at the end after you've used like a zillion GPU hours, it's like a world champion. But like how do we do that in the real world? Like clearly, like babies don't just like run into walls or like fall down. Actually, they do kind of fall downstairs if you let them, but well, that's a bad example. But like, you know, us humans, like the way we drive a car is uh, you know, we have a person helping us, but we're not just like randomly jerking the steering wheel until we figure it out. Like we have the sort of like base of common sense. Like we kind of draw understanding, like a model of how the car should work exactly. Exactly. And we kind of like kind of project into using our mind. We kind of simulate the driving experience as best we can from watching other people drive, watching our parents drive, and we've built a simulation of that on day one. Um and so there's a question of like how do we do that with reinforcement learning? And a big part of that is using what's called a trust region. And so a trust region is basically it works like this. So let's say I play, I play a bunch of games of Go, and um I'm a decent player. I play a bunch of games of Go, and now I go back and I watch all of my games, right? Um this is just me as a human. I watch all of my games and I look to myself. I say, oh, I would have done maybe this move differently. I would have done that move a little differently. But I'm not going to say like I would have done every move differently, like as a person. Like we can't—that would put us into a really weird state where like we wouldn't really know what to do, right? So we would pick like a few key things that we would do differently.

1:16:42.000 --> 1:18:40.380
<v A>And then we would wait until it's the next tournament. Exercise those differences and then we'd repeat this process. And so you know with with with reinforcement learning, if you take a bunch of data and have computers try to do policy optimization, they'll just hallucinate, just like we see with ChatGPT and these other things. Like they'll start hallucinating like, oh, if I play this move, uh I'm gonna get every single Go piece on the board because there's like some inaccuracy in the model. And the other thing is like it only needs one action to be inaccurate on the positive side to throw everything off, right? All your values are now thrown off everything, right? So so it's inherently kind of unstable. And so what trust region policy optimization and proximal policy optimization—what these things do is they basically say we're going to keep track of the actions that were taken in the real world and what the model is doing. And if the model doesn't match the real world enough times, we're going to stop training. So you know, in the beginning of training, the model is going to match the real world perfectly because it's the same model, right? Like you you you rolled this model out in the real world, collected a bunch of data, and at the very first mini batch of training, the model hasn't changed. And so it's going to output the same distribution. Right? Over time, the distributions are going to start diverging. And because you know because you own the model, you can actually keep track of the entire distribution, right? So even though you actually press the gas—you know that the model was like 50/50 about pressing the gas or not—and that's what you're going to log, right? So now you're training and you say, oh, well, the model that drove the car pressed the gas 50 of the time, but the new

1:18:40.380 --> 1:19:04.065
<v A>model wants to press the gas 100 of the time. That's a pretty big difference. And so maybe this would be a good time to like stop training and go drive with the new model. And so that's you know there's a lot more math than that. That's effectively what's going on is these are like halting halting criteria. So you might not be able to train that much before you have to stop and go to the real

1:19:05.010 --> 1:19:19.574
<v B>world. Got it? So it's basically like you've you're too far away from what we know. So you need to go try again. You've changed a bunch of stuff a little, but your outputs are now very different. So we need to go try again and see if it actually got

1:19:20.029 --> 1:20:16.662
<v A>better? Yeah, exactly. Um and then the last thing that I'll kind of cover here um, but there's a couple of other things. So one is imitation learning. And so this is pretty simple. The idea is um you know we talked about supervised learning, right? And the stop signs, right? But I can do the same thing with decisions. I could say, hey, when you see this situation, press the brake, right? And I'm just saying as an expert like it's a ground truth, like like unambiguous, you know, press the brake here, press the gas pedal here. That's called imitation learning, and that's just supervised learning. So you know, you could imitate a person, and and then all the regular things apply there of interpolation and everything we talked about. Um but that's not really reinforcing learning.

1:20:17.084 --> 1:20:25.454
<v B>And that's AlphaGo Zero where it was trained on all the human games and they basically said, 'We want you to imitate the person.'

1:20:26.280 --> 1:21:58.440
<v A>Who won, right? Exactly. And so AlphaGo did that as like a bootstrapping phase and then did reinforce learning after that. Um the AlphaGo Zeroes where they got rid of the bootstrapping phase. Um so that's uh so imitation learning is a good way to bootstrap, and that's kind of what we do. Um another thing I want to cover is model-based reinforcement learning. We're basically um you know in the case of AlphaGo, it can play itself because it's just a game, right? Like it's an artificial environment. Um but if you want to do for example a self-driving car as we talked about, you can't just drive randomly while you learn what to do. And so you have to construct a model. You have to construct like a virtual environment and then play within that virtual environment. Um you know the challenge now is of course like what happens when the virtual environment doesn't match the real environment? And so um there's different ways to to deal with that. There's something called a joint embedding, but long story short, uh with with model-based reinforcement learning, you have this sim-to-real problem. So if you look up sim-to-real, you'll find like a zillion papers on it, but it's like how do you take something that was trained on a simulator and bring it to the real world and back and forth and back and forth?

1:22:02.215 --> 1:22:38.378
<v A>Um okay. Yeah, the last thing to cover is policy evaluation. So you know with supervised learning, you have the truth. So it's like, oh, I didn't draw the bounding box around the stop sign—that's bad. I drew the bounding box around stop sign—that's good. And you can there's a million different ways you want to count those those errors, but you can count them those ways and just output that. Right? In the case of reinforcement learning, you don't really have like a perfect game of Go or anything like that. So

1:22:39.930 --> 1:22:42.700
<v A>what you have to do is is um

1:22:44.419 --> 1:24:39.625
<v A>Several different things you can do. You can either use a simulator and say, 'Oh, in the simulator my model got better.' That's what AlphaGo does, right? But in case of, let's say self-driving where maybe you don't want to trust the simulator, there's another thing you can do where you run two models. The first model actually controls the car, and the second model just says what we call counterfactuals, which is just a fancy word for 'I would have done this.' Right? So it's just like it's literally a backseat driver, literally. So you then take those counterfactuals and what actually happened, and you can figure out if the new model is better than the old one. For example, let's say the old model doesn't hit the brakes, and the new model really wants to hit the brakes. And then like half a second later, the old model slams on the brakes. Well, that probably could have been avoided by the new model because the new model was braking earlier, right? It was it was like more predictive, right? And so that would be a sign that the new model is a step up. Similarly, if the old model slams on the brakes to avoid a collision and a new model would have hit the gas, right? That's a bad sign. That means that new model probably would have got you in a bad position. So, so you know again a lot of math behind that, but effectively that's the intuition behind policy evaluation. And one last thing about that is policy evaluation is harder than solving the problem because if you have a perfect policy evaluator, then you also have a perfect policy. You just take the action that the evaluator gives the highest score to.

1:24:39.659 --> 1:25:06.439
<v A>Because of that reason, like versus like in supervised learning, you could just measure accuracy, just take the times you were right and divide it by the total time. Very easy, just algebra, right? But here it's like not only is it hard to evaluate the policy, but it's actually harder than solving the problem, and in many cases it's impossible. So that's a big challenge, and that continues to be a challenge today.

1:25:09.325 --> 1:25:18.740
<v A>So that's—so I've kind of covered all of the technical stuff. I'll dive into a little bit of the Large Language Model stuff, but before I do that, any questions about the technical?

1:25:20.344 --> 1:25:21.542
<v B>Stuff? So I guess the

1:25:23.280 --> 1:26:19.424
<v B>thing I'm missing is a bit—is I guess it's a little application which is understanding. Some problems like you're saying are clearly not a fit for supervised or unsupervised learning, and so you could think about like, 'Oh, maybe this is a Reinforcement Learning task.' When do you? But then some things, maybe there's multiple approaches to solving. And so you know, you could use a Reinforcement Learning; you could try something else. And then from like a toolbox standpoint, it's like, reinforce—so even today, like I can just Google 'your stop sign example,' and there's 100 tutorials for opening up, you know TensorFlow, PyTorch, whatever. Give it the images, give it the labels we talked about labeling, but you know, give it the labels and you know get monitor your loss function. Is it the same? Is there like the same set of tools for doing Reinforcement Learning, or is there also similar like a canonical example that you would sort of go to to like kind of do like the simple case?

1:26:20.605 --> 1:27:27.092
<v A>It's a really good question. Okay. So okay, the first part of the question—I think that my general philosophy is to use the simplest tool for the job, right? So for example, I'll give you a really concrete example. There was a place I worked at. I could probably say this; I'll just say it. I don't think it's going to be that controversial or anything, but it's not that much of an exposé. But you know when I worked at Meta, you know we released the Oculus Store, right? And so you can go right now to the Oculus Store and buy games for the Oculus Quest, right? And so I talked to the product managers. They asked to meet with my team, and they wanted to do Reinforcement Learning to figure out what items to put on what places on the storefront. So when you go to like oculusstore.com or whatever it is—the URL—like what should they just show right there on the banner? They call that the hero position. What should they put in the hero position, etc?

1:27:29.624 --> 1:29:27.160
<v A>And my response to them was like, 'Not only should you not use Reinforcement Learning; you should also not use AI.' Right? What you should do is like take the app that sold the most and put it in the hero position just manually and run that way for a month. And then if you realize that—like if you just have this intuition, like 'Oh, there's so many people with so many different interests, and we're showing everyone Beat Saber, and it's not going well,' and so we need to do some AI.' Then let's let's go there. Right? So like simplest tool for the job. Like the simplest thing was just like a YAML file with Beat Saber in it, right? And so like they launched that, and then—and then I would say, you know, if you can do something simple, around the decision, like 'Say okay, in certain countries I'll show Beat Saber, and other countries I'll show other stuff,' and then now I'm dividing by some other demographics. And the next thing, you know, you're kind of like building a decision tree by hand. Okay, let me use a decision tree, right? Yeah. And so—and then at some point you run into like competing interests where you know I want the store to do well, but I also want game publishers to share the benefit. I don't want to just kingmake Beat Saber, right? So now I have this like competing economic model that's very complex. Now we're starting to talk about Reinforcement Learning some of that, right? So I would say, you know, stick with the simplest tool for the job. Reinforcement Learning often is much simpler than trying to like take actions by hand and stuff like that. So for running a marketplace for driving a car, you know Reinforcement Learning is a great choice. Yeah. And as far as tooling, the tooling is

1:29:27.160 --> 1:30:33.800
<v A>way way behind. There's a lot of reasons for this. One of the biggest reasons is Reinforcement—Reinforcement Learning can't really be commoditized because it's too—it's too close to the decisions that companies make, which are sensitive, and so it's just very hard to commoditize. I mean, you know we rolled out Reagent, which was the most popular Reinforcement Learning platform for a while. Now there's, you know, there's OpenAI Baselines, so there's a bunch of places where you can get the algorithms, right? But if you want like the real techniques, like how do I do offline evaluation, you know a lot of these are proprietary. You know Reagent actually has policy evaluation and all of that, so folks can definitely check that out, and the codebase as far as I know is still active. But but I think the field is just still too new for there to be, you know kind of really good practices there.

1:30:35.502 --> 1:30:58.960
<v A>But yeah, those are two awesome questions. Okay. So I'll move on to RLHF. So a lot of people found out about Reinforcement Learning when ChatGPT rolled out. RLHF, which is what makes the 'Chat' part of ChatGPT—what took it from GPT to ChatGPT. RLHF is is a pretty simple idea. The idea is

1:31:01.000 --> 1:31:04.746
<v A>You um um so the

1:31:07.109 --> 1:32:03.657
<v A>Idea is thinking about it this way: GPT is imitation learning. So a person wrote, you know, 'The fox jumped over the dog,' or whatever that is, you know when you pick a font, it always shows you that same sentence. It's like 'The quick fox jumped over the lazy dog.' So that's probably all over the internet, right? Because it's in every font. So GPT will imitate a human quote-unquote if a human is, you know, all the content on the internet averaged, right? And so if you say like 'the fox,' 'The quick fox jumped,' GPT will respond with like 'over the lazy brown dog,' right? And this was trained in a supervised way. But when you think about it as like it's a decision to put that token there, like it's a decision to put that word there, then actually GPT is making decisions. And so

1:32:07.386 --> 1:33:05.048
<v A>and so it becomes a Reinforcement Learning problem when it becomes multi-step. So for example, you know if I just need to predict the next word and I know exactly what it is, that's supervised learning. But if I have like 10 different answers from GPT and I want to pick the best answer—like an entire answer only gets one score. Now it's a Reinforcement Learning problem because I have to figure out, 'Okay, this answer is better than that one.' Therefore all the tokens that generated that answer are a little bit better, but we don't know how much. And so RLHF is just a pretty simple algorithm where you say give two answers. If the if the system is more likely to pick the wrong answer, it gets a negative point. It's more likely to pick the right answer, it gets a positive point. And now you do your policy gradient.

1:33:07.259 --> 1:33:15.730
<v A>So RLHF has been a part of these LLMs for a very long time. The thing DeepSeek did that made reinforcement learning?

1:33:15.730 --> 1:33:20.016
<v B>yeah what is rlhf what does it actually stand for reinforcement learning reinforcement

1:33:20.016 --> 1:33:22.767
<v A>Learning from human feedback. Ah, there we go.

1:33:22.767 --> 1:33:23.695
<v B>Okay, so...

1:33:23.695 --> 1:35:21.519
<v A>Yeah, the person—a human—is actually saying this answer is better than that one. Okay? So the thing that DeepSeek did that was pretty amazing is it took the human part out and so it's just RLF. And so the idea is it'll generate an answer to a question that is easily verifiable. So, for example, they give it a word problem, and they know the steps of the word problem and they know the answer, and so they output two hypothetical answers, and then it comes back and says, 'Hey, this one's better than that one.' But it's all algorithmic. In math, there are systems that are—you know—they're very expensive to run and they're very specific to math, right? They only solve math problems, but they're totally autonomous. So I can give you like not just the answer, like 20 or something, but I can give you like the whole reasoning and the answer to a word problem, and this system will actually verify the entire thing. And so they replaced the human feedback with this expensive system, and then they ran this a zillion times, and what they found was the model that came out of it not only could do math problems better than anything we've ever seen, but it became like very thoughtful and reflective. The reality is what they found is if you treat every question like a math word problem, then you become like a much more reflective and thought-provoking and interesting in your answers. And so that's—that's basically what the DeepSeek folks have done, which is definitely like a huge leap forward.

1:35:21.519 --> 1:35:27.270
<v A>Really exciting. Does that part make sense? The RLF part of it?

1:35:29.127 --> 1:35:41.108
<v B>I think it makes sense. I think they're but ultimately they're going back and tuning the output of what the LLM is doing or they're tuning something that comes after the LLM.

1:35:42.036 --> 1:36:30.771
<v A>Yeah, no. They're modifying the LLM itself in all these cases. Okay? Yep. So if you think about it like regurgitating the next token is actually a form of imitation learning. So you're saying like these humans that have written this stuff on the internet, they're experts, and I'm trying to outdo the same action they're doing where action is writing letters. And so then when I change the goal to be like solve this reinforcement learning problem, it's still like a set of actions. So you can use the same model. Oh, another thing I should mention is—so we talked about Actor-Critic and we talked about how to get this policy gradient stuff to work. You need to have positive values half the time, negative values half the time. So...

1:36:32.864 --> 1:36:35.277
<v A>the the challenge here is um these

1:36:38.280 --> 1:37:47.232
<v A>And so what they did, which is really interesting, and all it actually only works because there's no intermediate rewards. But that's kind of a detail. They said, 'Okay, we can't afford to have a second model.' So what we're going to do is we're going to get the expected value of all these different answers to this math problem. So we're going to generate 10 answers. We're going to get the expected value of all 10 of them, and then we're going to get and then we're going to basically normalize that number. So for example, I get the expected value of all 10 answers, and let's say the expected value is all one except for the 10th answer which is two. So I'm just going to normalize that so that all the ones become like negative 0.8 and the two becomes positive 0.8 or something like that.'

1:37:48.885 --> 1:37:52.935
<v A>So they replaced like an entire neural network with some simple algebra.

1:37:55.619 --> 1:38:10.719
<v A>And so that's the PPO or Proximal Policy Optimization. So it's one of these things that's like a really clever trick. I have kind of mixed feelings about it. I do think.

1:38:12.274 --> 1:38:37.553
<v A>That I think that with intermediate rewards it's going to struggle. I think that like maybe coincidentally or maybe on purpose—the fact that in this particular domain you just get a reward at the very end is one of the important causes that for this approach working over like PPO or these other alternatives.

1:38:39.561 --> 1:39:21.647
<v A>Another interesting thing where they've kind of diverged is generally what people have done in situations like this where you need a large actor model and a large critic model is they've had the two share the same backbone. So for example, you have a neural net where the current state of your universe goes into the net, and then the neural network outputs two things: it outputs the distribution of actions you should take—that's your actor output—and then it outputs the expected value of the current state—that's your critic output. And so you still just have one model; it just has one tiny extra output on it. So you might say to yourself, well, that's like

1:39:23.520 --> 1:39:27.959
<v A>pretty awesome, right? I mean, I mean that seems like a no-brainer, but the problem is

1:39:30.810 --> 1:41:06.222
<v A>that even though it's just adding one node, both of those nodes are sharing that network. And so they're kind of competing with each other, you know? The critic model is going to be steering the entire network towards producing better values; the actor model—actor part of the model—is going to be steering the network to producing a better policy, and they're going to be causing corruption in each other. And so although you will see like people have success with this for Atari and other domains, I think it's actually super destructive. My guess is that the folks at DeepSeek actually tried this. I mean, it's the most intuitive thing, so I'm almost 100% sure—I mean a lot of circumstantial evidence—that the DeepSeek folks tried to have like just one network, one single network that does the policy and the actor—that sorry, the actor and the critic in just one network—and they realized that doesn't work. That you know it just causes corruption and it just never converges and it just is a mess. And so they ended up falling back to this approach where they said, okay, well we can't have a separate critic model; it's too big. We can't put a critic head on the LLM because that causes too much corruption, and so we're just going to abandon the entire idea of a critic and just come up with baselines on the fly, and that worked for them, which is really cool.

1:41:08.027 --> 1:41:17.984
<v B>Was that something that was unexpected? Like was that like a sort of—I want to call like an innovation to make that leap, or was it just sort of like—no, it was pretty obvious once I got

1:41:20.025 --> 1:41:28.649
<v A>There. Yeah, so this is where it gets interesting. Is—I mean, so there's a lot of theory, so I'll say you know I—um

1:41:30.707 --> 1:42:05.352
<v A>you know, I'm not at Meta anymore. I don't work at OpenAI or these places, and so I don't you know I don't really know like what's on the very cutting edge that hasn't been released to the public. There's speculation that OpenAI was already doing something like this but they hadn't published it, and so DeepSeek kind of scooped it. There's even more—like even more speculative is the idea that somebody stole the idea from OpenAI and gave it to DeepSeek. Um that is pure speculation.

1:42:06.854 --> 1:42:35.000
<v A>But I would say the fact that OpenAI has a reasoning model now that is comparable so quickly makes me think that like either they worked around the clock or you know they were coming to the same idea, right? And it's probably the latter. Probably DeepSeek saw where the wind was blowing, and and they both kind of came to that answer around the same time, that would be my guess.

1:42:35.000 --> 1:43:37.405
<v B>It makes sense. So there's lots of cases like that, right? I don't know online I saw someone using the term 'nerd sniped.' Wait, what does that mean? It's the same kind of idea like people—or I think like you're a YouTuber, you're like working on some like new cool project, you know, you think is like crazy and innovative, and someone else just releases a video of the same thing because you didn't get out fast enough. Or like you said there are—I mean what half dozen, dozen? I probably like a half dozen super serious competitors and like a dozen within striking range of doing these kind of similar—I don't—I don't want to demean them by saying those like Chat but like question and answer AI agents, agentic stuff, the reasoning, like all of these. And so like you said you're you're hard at work trying to refine a project, you're not sure if it's a big enough innovation whatever, and then someone else just goes ahead and you know releases it. Um yeah, yeah, and so you get sort of sniped—sniped out of it, right? Like someone got it before you, just before you were going to.

1:43:38.738 --> 1:44:11.205
<v A>Do it. Yeah, yeah, totally. I mean, I think one of the trends was that like math answering math questions was becoming like a big benchmark that was very important, and so I think that led a lot of people to the same conclusion. Like if—if the goal, like if the metric had been like write the best play or something, then we might have ended up with a totally different system. But I think once people got excited about solving these like high school math problems, I think that then that kind of set the course for all these companies.

1:44:11.897 --> 1:44:53.820
<v B>Answering math stuff was really bad for a long time. So yeah, yeah, it's kind of thing and I feel—and maybe I'm kind of wrong—I feel the story like coding maybe is one of those things that when you think through like LeetCode style problems, like there are definite like setups where you're given a very high-level question, and then and there are benchmarks already have these in there. But I still feel like performance isn't amazing once you get off the bed like that they've not seen before, right? So when you give these sort of high-level problems and you kind of have a very specific known output for the program and it should be compilable and it should be—it's harder. Yeah, of course in the math problem, but it seems within striking distance.

1:44:53.820 --> 1:46:07.280
<v A>Yeah, I mean you know traditionally machine learning we have this concept called leaking the label, which means like—you know if you if you took okay if you took the examples you trained on for the stop sign trainer and you just fed them back in and you get them all right. That doesn't mean you have a perfect system because you're kind of cheating, right? Like you might you might have just memorized all those examples and you can't know anything else. It's possible, right? Yep. But the problem is how do you not leak the label when you're training on the entire internet? And so I think what they've found is a lot of these cases were like the AI solves math problems or the AI solves LeetCode problems is they've leaked the label and the AI is literally outputting an answer that somebody else, some other human wrote to that LeetCode problem. Um and so they've done experiments where they've released like things that they know are not on the internet, and the AIs have struggled like mightily with it. I feel like until we get proper calculator use and tool use more broadly, I think it's going to be very hard for AI to solve these problems.

1:46:07.280 --> 1:46:29.378
<v B>Well, this has been a great topic and very timely. I know you've been working on Reinforcement Learning for a long time, but I feel it has like, you said it's kind of reached a certain hubbub in the everyday discussions recently. So, I'm happy to have a sort of great overview of what it is and what it's

1:46:30.137 --> 1:46:43.485
<v A>About. Yeah, totally. If folks have any questions, they can just reach out on our Discord or email, or in my case social media. We need a Reinforcement Learning algorithm to get Patrick on X that needs to be.

1:46:43.485 --> 1:47:01.896
<v B>the next thing okay um but yeah i look so i did look it up so dolly one was four years ago you were right that was very good and then dolly two is what i was first trying and that was three years ago and nice so you actually you were very accurate despite it just being off the top of

1:47:02.149 --> 1:48:05.800
<v A>your head. Well, I remember—you know, it's one of these things where like you connect it to stories. Like I remember there's this woman who's very influential on AI. Her name is Feifei Lee, and I remember being in this dinner, and her and her student were there, and she said something like, and this was again a long time ago, but she said something like, oh, it was—um, oh, we were writing captions from images. So basically, given an image, write a caption so that we could for accessibility reasons. And Facebook I think still has that in the product today. It's like if you're blind or something, you can click on an image and it'll say what is going on in the image. I remember her saying that's cool, but it'd be really cool if you could go from the description and create the image. And that always stuck with me. I mean, that was like a decade ago. That always stuck with me. And then I remember when DALL-E came out, I was like, wow, it's like the something that I thought was like a joke, but then it really happened. Like for me, it was like an amazing experience. That's why I remember.

1:48:06.966 --> 1:48:40.969
<v B>It. I think people have started to get a little fatigued on the AI thing, and it's hard to know is it—you always hit plateaus? Is it like plateauing in terms of actual functionality? Is it on we're on an exponential? And exponentials always look self-similar no matter where you look, right? And you just sort of like we can't feel the growth. And then you tell stories like you're saying or even about DALL-E being, you know, four years ago only, and like talk about the Midjourney or the Flux now versus DALL-E, you know, just three or four years ago. It's not that long, and there are lots better. Yeah.

1:48:41.745 --> 1:49:00.426
<v A>I mean, yeah, actually that's a good point. I'll end with where I think this is going. I think that despite loving Reinforcement Learning and everything, I don't think that AI should be making a lot of decisions in isolation. I think that it should be working together with people. Um, and

1:49:01.945 --> 1:50:22.860
<v A>so you know, the Recraft is a great example where it's not just an API you call and get an image, but it's like an experience, and like you iterate, and you say, 'Hey, I want this to be all different,' or 'Hey, I want this—I want an axe in this person's hand or a phone or whatever.' Right? Um, and so I think it's going to be really about collaboration. And Reinforcement Learning is always going to be really important, but it's going to be important in the way that the actions are more like suggesting things to people. So, in other words, Reinforcement Learning to like book a flight for you probably not a good idea because if one out of a hundred times you go to Tokyo by accident, right? You're gonna be pretty upset. I might be happy. That sounds great. Yeah, actually, Tokyo is amazing. Yeah, we're—I actually don't want to say anywhere. We don't want to go because we have a listener almost like anyway. Um, so Antarctica. Okay. So but you still need Reinforcement Learning to like suggest, like, you know, come up with like three different hypotheticals, send it to the person—should you text them or email them, etc? Like there's still a lot of decisions to be made. But I don't believe tons of people are going to lose their jobs entirely. I think that work is going to change just like it did with the invention of the motor and stuff. So.

1:50:24.919 --> 1:50:26.269
<v B>Well, you heard it here.

1:50:28.277 --> 1:50:30.920
<v B>First, the future AI is not too scary, Coin.

1:50:30.920 --> 1:50:42.440
<v A>Yeah, yeah. Don't be worried about it. Just be adaptable. If you're adaptable, I think you'll be just fine. And oh, and coding will probably be one of the last jobs to be eliminated, by the way.

1:50:45.287 --> 1:51:09.435
<v A>So if you're worried about—if you're worried about, yeah, please stay in coding. Not just because we want you to keep listening to the podcast, but tell all your friends to get into coding, stay in coding if they're worried about their job being eliminated. They should be a coder. That's like one of the last jobs that's gonna go. I mean, trust me on this. Like we are going to lose so many doctors, we should probably lose all the CEOs before we lose all.

1:51:09.435 --> 1:51:17.333
<v B>Right. All right, we gotta wrap. We gotta wrap guys. We gotta wrap. They're phoning me from the other room and telling us we're out of time.

1:51:20.084 --> 1:51:21.822
<v A>That's right. Hey, I didn't say anything.

1:51:22.935 --> 1:51:25.315
<v B>About HR? What's that? Oh, oh yeah, we're wrapping up.

1:51:27.458 --> 1:51:36.385
<v A>All right. Um, this was so fun. Uh, thanks everyone for tuning in and uh, thanks Patrick for bearing with me and my rants on Reinforcement Learning. This is

1:51:36.385 --> 1:51:39.574
<v B>Great! Very illuminating. This is awesome. I learned a lot today. So,

1:51:40.350 --> 1:51:42.940
<v A>Cool. All right, everyone. We'll catch you later.

1:51:59.031 --> 1:52:01.022
<v B>Music by Eric Barmdoller.

1:52:02.642 --> 1:52:20.760
<v A>programming throwdown is distributed under a creative commons attribution share alike 2.0 license you're free to share copy distribute transmit the work to remix adapt the work but you must provide attribution uh to uh patrick and i and uh share alike in kind

