WEBVTT

0:00:15.285 --> 0:00:22.845
<v A>Hey everybody, I had a really

0:00:23.892 --> 0:00:59.139
<v B>Interesting discussion on LinkedIn. This is like a meta-meta thing, but I post everything on LinkedIn and Twitter. And my Twitter, nobody follows my Twitter. And the weird thing is, I am my ex, yeah. Maybe it's going to be my ex social network. I just can't get anybody to follow me on there. Everyone follows me on LinkedIn, which is fine. But I even tried putting my Twitter link on presentations I give and stuff like that. And people would rather just find me on LinkedIn.

0:01:00.949 --> 0:01:12.560
<v A>I find this amazing that you give presentations where people actually could join and follow you. I always see people do that. I've never done it myself, just not on social media. But I also find it interesting that you give presentations where that would even be an opportunity.

0:01:12.660 --> 0:02:39.160
<v B>I gave an amazing presentation to SUNY Purchase, which is a university in New York. And it was on the AI singularity. And it was a lot of fun. Had a great time there. Or I wasn't in person, but had a great time speaking. Oh, nice. Now I kind of want you to give it here. I actually would love, I asked if they would make the video public. They said no. It was actually technically one of the lessons. Like it was part of a course. And so for that reason, they can't just post it on the internet. But I was really happy with it. The questions were extremely interesting. The students were very engaging. And yeah, it's a shame that we can't just share it with everybody. But the thing that really kind of took off on LinkedIn sometime over the weekend was I posed this question, you know, is work from home really work from city? And what I meant by that is how many people, you know, during the pandemic, you know, really moved to a different city. And, you know, they're not really interested in working from home. That's not the spirit of what transpired. It's actually that they want to work from another place. And so maybe I'll pose a question to you, Patrick. So if your office was there in your city, would you go to it? Why and why not?

0:02:41.180 --> 0:02:42.580
<v A>Oh, this is interesting. Interesting.

0:02:44.300 --> 0:04:38.974
<v A>I think I see your point, and I have seen the statistics, although I'm not sure how trustworthy they are. I don't know the questioning, to be honest, but that not everyone, to your point, is actually happy about work from home. Personally, you know, relocating to a different city for a plethora of reasons, to your point, it's not necessarily that it was only the office. It was also the geographic location. And as much as we are moving to online people, you know, having family and, you know, just a laundry list of different personal reasons for myself, but for others as well, there can be specific cities you want to live in or don't. That being said, I, the question, so it's hard because the place I chose to live is different than where I would potentially have lived if I had tried to locate next to an office. That is where an office most likely would be near me would be far, and therefore I would not want to go to it. Oh, I see. But if there had been an office and I was living close to it, it's not that I mind going into an office occasionally. I will turn the flip around, which I'll say by not commuting, personally have found ways of—I don't know if I say exploiting, that sounds negative—I have found ways of parlaying that time that I would normally have spent in the car into other industrious activities. So things like, you know, being more active and exercising and going out or meeting with people sort of early in the morning because I'm on the East Coast and work with people on the West Coast. And so going and, you know, meeting with people I otherwise wouldn't have because I'm, you know, working late, basically meeting with them in the morning. So replacing commute time by having a time-shifted schedule and, you know, finding these other things for me means work from home is really about work from anywhere. But I do see your point. There is a gradient between office for your full work week, hybrid work from city, like down the, down the list. I think there's sort of like a spectrum, and I think it's—it's not a, it's not a binary choice. Yeah, that

0:04:38.974 --> 0:05:57.640
<v B>Makes sense. You know, one of the reasons, so I was working in Austin. So like when I, when I moved to Austin, I did work in an office here. And then recently I switched to working from home, which wasn't a choice. It was part of just things that transpired at my company. And so now all the awesome folks are remote. There's definitely pros and cons. I'm a bit ambivalent to it, but what I really want to do is sort of blunt this argument that—I feel like when they create this dialectic of you have people in the office and people at home, then it's easy for like, I think Elon Musk is one of these people who says, 'Oh, people working from home are lazy. You know, they just want to sit, eat Doritos all day, whatever.' And so I feel like, you know, the entire premise of that argument is false in the sense like—I would, I wonder if the majority of people actually, you know, they're not working from home to work from home. They're working from home to live somewhere else. And so then, you know, the whole this whole stereotype of like, 'Oh, this person doesn't want to get out of bed,' is really not, not even true on the premise.

0:05:59.687 --> 0:06:00.362
<v A>I think, yeah,

0:06:01.898 --> 0:07:01.560
<v A>I feel like it's one of those cases where there's no universal truth. I mean, I could imagine a theoretical person that this is the class that goes and spends all their time at the water cooler, right? And actually causes a distraction of other employees trying to work. And basically finds ways of not being at their desk and not working because they're not busy and they're trying to cover for it throughout my career. I've always known people that have been more or less like that. And that for that person going home, if they're doing that specifically to avoid the appearance of you not being busy, then going from his is a revelation to them because they can just do whatever they want and no one knows if they're working or not. And then you get all the like mouse tracking software that you see companies roll out or whatever to combat this. So I do think there are people like that, but like you said, I don't know that that's universally true that just because you work from home means you're like eating Cheetos and don't have any pants on.

0:07:04.859 --> 0:07:04.892
<v A>Yeah.

0:07:05.200 --> 0:08:04.560
<v B>All right. Well, we'll see where that goes, but there's a lively discussion. If you want to follow me on LinkedIn where people actually respond to my inane questions, you can follow me on LinkedIn and see a really interesting discussion unfolding. One thing is, you know, and I try—I knew this would happen, so I tried to avoid it, but you can't completely avoid it. You know, some people took it as like an attack on the Bay Area. And that wasn't really the point. You know, and I tried to make a point of saying, look, there's just as many people who would be running the other direction. If the tech companies weren't already there, you know, like if you were to flip the script, there'd be just as many people in the Bay Area working from home while their HQ is in Florida saying like, 'Oh, you know, it'd be the same thing.' Right? So it's not, it's not about which place is better or any of that. It's just about why are people leaving?

0:08:05.180 --> 0:08:23.969
<v A>And for a subtly different question. I mean, do you have a like if a company wanted to offer work from city as a flexibility, like is that something you imagined? I mean, we work is now, I think going towards bankruptcy, I think is last I saw, but I mean, what what is your like do you have a thought towards like what that ends up looking like? Yeah. I

0:08:23.969 --> 0:09:51.620
<v B>I loved it. So we had a—we worked, basically the short story here is I was part of a startup. As you know, a startup has relatively few rules. I mean, obviously it's still a corporation and everything, but it's a small company. And so we had a—we worked, the startup was acquired by a much bigger company, and the bigger company has a bunch of rules around what constitutes an office, and we worked wasn't able to, you know, fit those rules. And so all the we works got shut down. Basically it's what happened. But I loved it. You know, I would go in, there would be a bunch of people from all sorts of different industries and different micro offices. That was interesting. I mean, you know, there's nice common areas where you can meet people. So I was a big fan. I actually, I would ride my bicycle to downtown, which is something like 12 miles each way. And that would be my whole exercise for the day. So I would bike in downtown. I would work, bike back. And yeah, it was a lifestyle that I was enjoying. But working from home, I think has been pretty much fine. Yeah. Maybe I'm pretty ambivalent to this; you know, it doesn't seem all that different. I think a lot of the people in my team were not in my city anyways. So I was spending a lot of time on the internet with them anyways.

0:09:54.672 --> 0:09:55.026
<v B>So, well,

0:09:55.026 --> 0:11:52.180
<v A>Check it out on LinkedIn. The discussion continues there, but you have to make an account, Patrick. I think I have one somewhere; I'll have to dig it out. Time for a news of the show. So, fitting with not being a news story, we should just really rename these sections. This one was an article that was titled 'Falsehoods Junior Developers Believe About Becoming Senior,' and the person's name—at least I assume his name is Vadim Vadim. I'm not sure how to say it. Sounds right. Okay. I apologize if you're listening. And they had some great points here; this is just a thought-provoking article, which is, you know, I have been in my career a little bit longer. I can sometimes forget that what it was like when you're sort of new and you don't understand a lot of stuff. First of all, I mean, I personally think the junior developer/senior developer thing is overdone at some companies. It's a very strong delineation at other companies. It's not. I think there are people who have characteristics of being 'junior' and 'senior.' And so maybe my mindset already sort of like answers some of these, but these are great to point out. I don't know. I hit on all of them. But an example is that a senior developer just knows all the answers. And I think that's obviously not true. Like senior developers, or at least—I've not met a senior developer who actually knows all the answers. I've met a few who thought they knew all the answers, but they don't actually know all the answers. Similarly, there's a belief, and we bump into this from time to time, that oh, you're a senior developer. Like you must work like the cutting-edge stuff and insert whatever language is hip at the moment or constructs in the language you're working in. So I work in C++. Like, 'Oh, you must use all the exotic C++ stuff.' And it's actually like, no, I use probably one of the more basic subsets of the language.

0:11:52.180 --> 0:12:45.640
<v A>because it's important for me to get collaboration and get help from teammates. And so the person just pointing out that there's not this belief that you somehow fundamentally shift that one day you're doing sort of junior tasks and fetching coffee—that's not equivalent to writing documentation in the code. And then one day you're not, you're doing only cool stuff and you're just pawning off all the mundane things and whatever. There may be sadistic or problematic senior developers who do that, but that's in practice not really true. And it really is a gradient. As you move up, you may be solving bigger problems, but bigger problems are really just large collections of smaller problems in my experience for the most part with very rare, occasional, singular tough nuts to crack. So interesting article. I didn't read them all. I think there's like 10 here. But check it out. We'll have a link in the show notes. What are your thoughts? This is great. Yeah.

0:12:45.970 --> 0:13:54.160
<v B>I think I actually have seen things degenerate the other direction where, you know, as someone becomes a Tech Lead, they say, 'Well, you know, my team's welfare is now really important. Therefore I'm going to do the really—you know, crappy for lack of a better word. I'm going to do the really crappy work that no one else wants to do. That way my team is happy.' And then, you know, a year later they're completely burnt out and they hate their life and everything. And you kind of like start diving into it and you find out, 'Oh yeah, they've been doing the worst work for a year,' and that's just not sustainable. So I think it actually goes in the other direction more often than not. And yeah, I guess maybe the thing takeaway is like, as you go up in level, you become like more and more of a servant to more and more people. And so, yeah, it doesn't go that way. But yeah, I also felt the same way as a junior developer. So this is interesting that how do you correct the record there? It's not; it's not there.

0:13:56.920 --> 0:14:15.500
<v B>I guess with—I guess if, as you said, if you're sadistic, you have now the power and the influence to be a sadist, but if you're sadistic, you probably have a really hard time getting promoted. Not impossible, but it's harder. So it kind of works itself out.

0:14:17.280 --> 0:14:44.819
<v A>Yeah. I mean, culture and management is the fallback here, which is depending on how your team is run determines how much if you were one of a very small number of software developers at a non-engineering, non-software engineering company, I think it's a little different than if you're at a sort of big tech company. I think the how those situations develop and evolve can be very different. Yep. Yep.

0:14:45.699 --> 0:16:45.680
<v B>All right. My news story is Pure Pursuit. This is really cool. It's a relatively simple way of designing a robot vehicle trajectory optimizer. So for example, imagine like this is a good example; take a racing video game, right? For most people, a racing video game is kind of a non-starter because they have no idea how to build the opponent AI. And it seems like when you watch—you know, even Mario Kart or something, when you watch them, when you watch these racing games, it can seem really hard. Like, 'How do I code something like that up?' And what if they get totally bumped by the player way off track? How do they get back on track? And yeah, I guess the thing is like they, as opposed to like Mario or the enemies, basically have their own physics and their own universe in this case. You know, everyone's a car with the same—you know, if you can't really bend the rules; if the AI just teleports in the middle of the track or something, it completely destroys the immersion. You can make the AI go faster and slower, and there's rubber banding and some of that, but still like, it has to kind of follow the same kind of physics constraints as the player. Otherwise it ruins the game. And so, it turns out there's a whole bunch of relatively simple methods for doing kind of basic robot navigation and trajectory planning. And this is one of them called Pure Pursuit. So I included two links. One is a pretty lengthy tutorial that has a bunch of Python code where you can follow along and see what they're doing. The other one is a YouTube video where it's, as you'd imagine, much more visual. And yeah, it's a lot of fun. If you're looking for kind of a neat thing to do, I think coding something like this

0:16:45.680 --> 0:16:55.180
<v B>up and having a bunch of little robots race each other, you know, it could be kind of a fun thing to add to your GitHub account.

0:16:56.952 --> 0:18:21.300
<v A>I think I saw there was—I guess that's reinforcement learning. Someone wrote a tutorial or like a—not a tutorial, one of those funny videos, I guess you watch mostly for entertainment, about TrackMania, which is some race car game that is not super hyper accurate, but it's a little bit more. And so they were kind of showing how over generations trying to, I guess you'd call it evolve or train an ML agent to get the best times. And they were talking about some of the subtleties about what moves or inputs you are aren't allowed to do and whether it was allowed to drift or whether it tried to avoid drifting. As an example, it's a very complex topic. And I agree with Jason. I mean, I think it's pretty interesting and it has a lot of—I don't want to say like complete real-world application, but there's a lot of thinking through these kinds of things that help you in other domains that would be sort of adjacent. So anything from like building a little flying airplane or a quadcopter is going to have some of these same control loops and dampening and making sure you don't get into oscillations and these kinds of things as well. So yeah, I don't—is there a—I don't want to call it like a playground or like a something where it's like pretty well set up where all you have to do is really code in the kind of driving. There must be no one off the top of my head.

0:18:21.580 --> 0:18:49.373
<v B>Yeah, there was a—let me see if I can find it. Oh, I forgot what it was called. Let me see if I can. Yeah, there's Carla. Oh, Torx. Torx is what I was thinking of. It stands for the Open Racing Car Simulator, and Torx was designed for AI. Like you could actually play it, but it's really meant for AI to play it. Nice.

0:18:49.980 --> 0:18:56.840
<v A>So I was going to say that if you want to get down to the specific cut of the problem may get you there faster than writing Mario Kart from scratch.

0:18:57.760 --> 0:19:15.159
<v B>Oh yeah, totally. I mean, I think, you know, the Mario Kart from scratch, you would make just like a really fake physics engine, and it's just really silly. But yeah, if you wanted to ultimately control a real car, if you wanted to ladder up to that, then yeah, Torcx would be a good starting point or Carla would be another one.

0:19:16.540 --> 0:21:13.340
<v A>So my next news article unintentionally, I guess, is a piece of what you might use to kind of get there. And that is something called a control loop or a controller and sort of figuring out, given some observations of the real world, affect some output, and you can get as fancy or non-fancy of these words as you want. I'm pretty non-fancy. And that is, I call it PID Controller. So the article I have here is PID without a PhD. Jason was telling us at the pre-show; he says PID, which I've heard before as well. So, we'll have to ask the AI which one is correct. But PID without a PhD is a little PDF with—I will say not the most elementary explanation, but without getting into my background and makes it such a, if you just look at the Wikipedia page for PID Controller, you're quickly going to get masked out. Well, that's not good grammar. Oh, well, you're going to get into math over your head, or at least math I'm not comfortable with for like a casual reading. And so this PDF is a great introduction and—I did want to give a shout-out to it, but also just to talk about, similar to what Jason is saying, a useful tool in the toolbox that kind of pops up in a surprising number of places. And so just in brief, a PID Controller is an acronym stands for Proportional Integral Derivative. And that is that you have some value you want to achieve in your output. Say you're heating a pot of water and you have a thermometer in it and you're controlling the heater underneath. If you crank the coil underneath all the way to full power, the water is not going to instantaneously boil. If you've ever watched a pot of water, it doesn't—I'd be nice if it did that, but it doesn't. And so you have some temperature you're trying to achieve and you have the thermometer that's giving you feedback, but there's this latency. And as you get closer, you want to start modulating the power down. So you don't necessarily like overshoot your temperature.

0:21:13.840 --> 0:22:41.978
<v A>You know, cause if you were cooking—well, no, I said water—but if you were cooking something in the water that you didn't want to be overheated, like burn your food or whatever the equivalent of over-overheating, it would be—you don't, you can set this without going into it, the proportional integral and derivative terms that will help you to kind of control the behavior of getting to as quickly as possible. And then sort of having good behavior as you sort of have this latent response to your inputs. And this is a form of a control loop because you're kind of sitting there looping over and over again, take the measurement, control the measurement to what you want the output to be, decide what you want to do to your settings. And so they grow from here. It's kind of the most trivial one. There are all sorts of more advanced optimizations that you can do with PID Controllers. And then ultimately even moving on to other controllers, once you get past sort of like the trivial things, and even the online discussion around this article had a lot of debate about whether PID is sort of 99% of the time okay, or whether it's the worst thing ever because it just not the most optimal answer. My take is, you know, it's a good tool in the toolbox on learning. And there are often things where you want to apply some sort of move from A to B, but in a controlled response. And it sort of ends up being in the same bucket of problems, I will say. And so being aware of these things is good to know your way around an unfamiliar territory.

0:22:41.978 --> 0:23:25.020
<v B>And I haven't read this article, but I think the P is basically how far away from the target you are. The I is what has happened recently. And the D is your kind of more immediate gradient, like what direction you're headed. And so they're all useful in different ways. Like, if you're screaming towards the finish line, then that proportion better be really big. Otherwise you have to start slowing down. So that's where the derivative comes in. And if you're waffling back and forth, it's clear that like, you need to slow down because you've overshot six, seven times in a row. And that's where the integral comes in.

0:23:27.020 --> 0:24:21.159
<v A>Even setting the values has a whole bucket of theories around. And it is really interesting. Just use a PID Controller to set the values. Gee, I can't see where this might become recursive. Stack Overflow. I think that if you once you learn about these, you will notice like cruise control in your car and how does it behave? You can sort of think about this, not necessarily that they do it, but you'll start to see also some, if you have like a 3D printer or as an example, it will not know how much kind of mass the heating, the heated bed has at the bottom or the part that squirts out the plastic. So often you'll see during when you're turning it on, it'll do some calibration. And part of its calibration is heating, cooling, heating, cooling to understand the response of those devices in their current configuration and their current environment and setting the parameters of these PID Controllers.

0:24:22.159 --> 0:24:22.679
<v B>Yep.

0:24:23.280 --> 0:24:24.899
<v A>Yep. That makes sense. Very cool.

0:24:25.360 --> 0:24:38.567
<v B>Yeah, folks definitely read the—I mean, PID Controller is, as Patrick says, like the simplest thing it's in every house or your thermostat. You know, it's almost certainly using a PID, maybe not nowadays.

0:24:38.567 --> 0:24:40.103
<v A>I think, but yeah, yeah.

0:24:40.680 --> 0:26:19.220
<v B>I mean, but if you have like one of these old Honeywell ones, you know, whatever. But definitely worth learning about. I think it's a great foundation. All right. My new story is Google releasing Gemma. So, most people have heard about Llama. Llama is this open-source LLM from Facebook. Someone made a project called llama.cpp, which is kind of a weird name of a project, but basically it's a way to run these Llama models really fast. There's no PyTorch. They're loading the PyTorch model, but it's all done in C++ and everything is super customized and optimized for all these different architectures. And it gets to the point where like on your MacBook, you can run these large language models almost in real time. That's been really exciting. There's been a whole ton of research that's come out of that. So Google released Gemma. This is also kind of an area of debate, but okay, well just basically the Gemma models are also these models that are small enough that you can run them locally. In fact, they even have one even smaller than the smallest Llama model. So it's so as a result, it can run on even less hardware, like maybe on your phone in real time. They claim the performance is much better. There's a big debate because the way they're kind of.

0:26:21.015 --> 0:28:19.120
<v B>It kind of gets in the weeds here, but you know, the seven—if you look at like the Llama 7B or Gemma 7B, that seven B number means 7 billion. So there's 7 billion parameters in the model, 7 billion weights that have to be tuned. And whenever you execute the model to generate a token, you have to at least use all of those 7 billion weights. So, the smaller that number gets, the less weights you have to use, the faster everything gets. Now the problem is like, that's not the only; it's not so one-dimensional, right? So the Gemma models have a much larger embedding and I'm not going to get too much in details here because it's not relevant, but basically because they have such a large embedding size, they could be a lot slower than the Llama model of the same size. But they can also perform better. So there's a little bit of advertising talk around, 'Oh, we perform better at 7 billion.' But the cool thing is, it's a whole new set of open-source models that we have at our disposal. One of the most interesting large language models that have come out that you could run on commodity hardware is this one called mixed drill, where I think they mixed several different open-source models to make one kind of supervisor model. And I think that's fascinating. And so now you have another model you can add, add to the mix literally. So, I think there's a lot of potential here and folks should check it out. The llama.cpp, despite the project name, actually runs a lot of these open-source models.

0:28:19.120 --> 0:28:31.020
<v B>And I'm sure they're feverishly working by the time this podcast is out, they'll have the Gemma models in llama.cpp. So you could run them on your laptop. So, definitely something to check out.

0:28:32.185 --> 0:28:45.720
<v A>Is the number of parameters also like, sort of related to how much memory it takes to use? Because that's one of the things that they always make a big deal about is not just needing a GPU, but a GPU with very large amounts of memory.

0:28:46.500 --> 0:29:13.320
<v B>Yep. Yep. And so there's a bunch of tricks you can do. They got to the point where they're doing four-bit quantization. So they're only allowing each weight to be one of 16 different values. But that's what Llama.cpp project is. That's one of the things they do. But you're right. You know, the more parameters, the more either CPU RAM or video RAM you need to run the model.

0:29:13.320 --> 0:29:51.020
<v A>Very cool. And then does it do, or maybe we can move on, but fine-tuning, fine-tuning of these, like, is it some models are easier to fine-tune? Like, so the training, so running them and executing them, presumably like Gemma coming from Google is related to, you know, the ones that you can just go online and use. So if you don't have an internet connection or something, it feels useful and being open source. But to me, the power, and I think you've talked about trying that before is like adding your own inputs and doing some additional training or switching or customization to, are these ones equivalent when it comes to that aspect or is like some are better for that and some are worse?

0:29:51.600 --> 0:30:58.780
<v B>Yeah. So the big difference between Gemma and Llama is the embedding layer of Gemma is enormous. So, and we'll talk about this actually later in this episode, but basically the embedding layer is the layer where you switch from understanding what you just said to deciding what to say next. And so that layer is really important. And so that layer is enormous in Gemma. So I would expect it to be harder to fine-tune just because of that, needing more data, more iterations. But in general, yes, I mean, it's much easier to fine-tune the 7 billion model and the 70 billion. Yet the other challenge about fine-tuning is you can't do quantization while you're training. And so now you need to store the 32 bit or maybe even 64 bit potentially float for each weight. And so that gets really expensive. So that's where you need like the 64 gig GPU to do the training.

0:30:58.780 --> 0:31:06.962
<v A>This is a fascinating time. Fascinating time. I'm excited. Need to learn here. Good thing we have a

0:31:06.962 --> 0:31:19.028
<v B>topic queued up for this. Yeah. You could start a large language model startup, but I think they took all the names. There's no names left in the entire language. There's so many of them. All right. Time for book

0:31:19.028 --> 0:33:01.880
<v A>of the show. All right. What's your book of the show? Mine is The Eye of the World by Robert Jordan, which is the first book in the Wheel of Time series. I guess I'm late to this, you know, classic fantasy. Most people probably already heard of this. Also, if you have not seen any of the large amounts of advertisement Amazon Prime has done for Amazon Prime Video, they have a Wheel of Time series, a TV series, I guess it's called. Oh, cool. And so that's actually, I had long known about this. We talked a lot about Brandon Sanderson books and Brandon Sanderson ended up writing the ending of the Wheel of Time book series because I didn't know that Jordan passed away before the series could be completed. So he left his notes and Brandon Sanderson sort of released it. So very adjacent to this series, which is never, if you've ever seen the books, they're intimidatingly large. There being, you know, I don't, I should have looked at, I think there's 12 of them, but being there are so many, you know, I was always a little intimidated to pick it up, but I actually really enjoyed the TV show. I've watched there's two seasons now and decided it's finally time to pick the book up being aware that, you know, books and TV shows are not necessarily the same thing. But interesting enough in the world and, you know, seeing some conversations, I guess Brandon Sanderson is actually one of the consultants for the Wheel of Time show. And, you know, you see arguments, well, I guess I'm nerdy, but I, and there's parts of the internet that we all spend time on. You'll see people kind of debating about the Wheel of Time TV show. And so it piqued my interest. And so now I have to like know for myself and go look at the book so I can judge if the TV series is a good reflection of the book or not.

0:33:01.880 --> 0:33:47.940
<v B>Very cool. I read this book probably when I was like 14 or something, but you've not seen the TV series. I've not, I didn't even know there was one until you just mentioned it. I'm going to have to catch up. I'm one of those people that I don't watch TV, but I do watch YouTube. And I watched the premium, so I only get the ads that the creators actually put in themselves. But yeah, actually you and I probably have very similar YouTube interests lately. I've been binging on this guy, Blacktail Studio who makes coffee tables. Yes. Have you seen this guy? Yeah. But yeah, I don't have, I have Amazon Prime, but I've never watched videos on it, but I will check that out. That sounds cool.

0:33:48.642 --> 0:34:49.940
<v A>So my new, so sorry on a side topic here, watching YouTube a lot. I have apparently like too many interests slash hobbies that YouTube doesn't like, it's not able to hold them all in its recommended video list. So I have subscriptions of course, but like, if I go to my front page, like whatever certain topics, if I ever watch a video. So we were talking about power world. I watched a video about power world instantly. Like my whole feed became like two thirds dominated by power world. And I had to like go remove watching power world videos from my watch history to get back to any semblance. But in just in general, as I like rotate through my interests that like, I noticed that it feels like it can't hold them all. And it's like, you know, consideration matrix or whatever, however it's doing, you know, not to anthropomorphize it, but it just feels like, yeah, if I once you move to a new thing, you get lots of videos about that and you stop getting videos about the other, even if you were watching them when offered. And so I don't, I don't know, maybe that's a personal problem just because I should focus, but.

0:34:50.580 --> 0:35:02.420
<v B>Well, I think it just, yeah, there's a movement of the masses there. Like, like probably people get on a topic and binge it. And then that changes the behavior of the system for everybody else.

0:35:04.461 --> 0:37:01.740
<v B>My book of the show is How to Make a Video Game All by Yourself. This is a very short book, extremely useful book I've been going to over the past, you know, since COVID let up a bit, I've been going to a lot of video game developer meetups. As people know, I've made this AI hero game. But I've also been talking to a lot of other game developers in the area. And one thing I've noticed is a lot of them don't see themselves as producers. A lot of them are software engineers and they end up with like a really cool tech demo or engine. One guy built this thing where it was like kind of like a Minecraft world, but you had rocket launchers for arms and every rocket actually caused a crater in the world. And so you flew around like blowing craters in the world and they even had the water working. So like if you, if you're underground and you blew a hole up and you accidentally blew a hole into the bottom of the ocean, like water would just start filling in. It was cool. I mean, I was super impressed. And as the water filled in like voxels of air, like floated up. So he had naturally had bubbles, like you didn't have to design the bubble separately. It was fascinating, but what it wasn't was it was not a game, like, like I had a lot of fun, but I wouldn't say I was playing a game. I would say maybe I was, it was more of like an art kind of thing. And I feel like this book would have been perfect for this person because it really starts off with the first principle of like, you are a video game producer. And so it lists all these things that ultimately don't really matter.

0:37:02.320 --> 0:37:23.240
<v B>Like which engine you pick or these other things, right? Or they're secondary to your goal. I thought it was really well written. I didn't bother to research the author, but I'm assuming it's somebody who's done a bunch of indie games. And yeah, it was a short read, but I had a lot of fun.

0:37:24.799 --> 0:37:34.500
<v A>That sounds cool. I'm fascinated by this as well. One day I should go to a video game meetup. That actually sounds like it'd be really cool. It's a blast. And I would start writing video games.

0:37:35.120 --> 0:37:38.020
<v B>It's a total blast. I've met some really nice people there.

0:37:39.240 --> 0:39:37.660
<v A>Very cool. Time for tool of the show. All right. So I'm up first again. My tool is not so much about the tool, but just a shout out that like, I wish more companies kind of did this thing. So, we talked about it a while ago and I recently used this tool. Google released an unlock tool for their online streaming gaming service Stadia, which shut down. They had run sales where you could get a Chromecast and a Stadia controller, which is what you needed to play Stadia games. The Stadia controller interestingly connects straight to Wi-Fi. So you are Wi-Fi controlling your server box in the cloud that is running your game. And then the results of the video would stream down to your Chromecast and onto your TV. So you were not interacting, like you may be sitting in front of your TV, but you could have been from another city, you know, playing your video game on the TV in your house or whatever. It wouldn't care because your controller just connects up to the internet via Wi-Fi. And then, you know, happens to control those devices. There were some UI loops, if you had multiple of them, make sure that you were controlling the right TV and that kind of thing. But they kept running specials trying to get people to sign up for their service. I just wanted the Chromecast. So I ended up having a number of the Stadia controllers laying around, but there's nothing really do with them because they connect to Bluetooth and the service shut down. So they released on the radio chip that they had in there; it also had Bluetooth. They released a website that you can plug your controller in, follow the onscreen instructions, and we'll basically install a new firmware that removes the Wi-Fi and adds Bluetooth connectivity. You just use it as a Bluetooth gamepad, and you can play on your Steam Deck. I paired it to my Steam Deck and played it on my Steam Deck or my computer as well. Lots of things have Bluetooth these days. So it's really easy to use it as a game controller. And I just thought they didn't need to do that.

0:39:37.660 --> 0:40:08.540
<v A>Like it was custom hardware for their thing, but they found a way that was like easy enough for them and straightforward to kind of do this. If you are like me, you probably have a drawer full of things that correspond to services which are dead and not able to be used anymore. So it's not always possible. I understand the complexities of it, but it would be really cool to see this be like a thing that people try to do. Like at least give some functionality to devices when they reach their sort of end of service life. Yeah, totally.

0:40:09.279 --> 0:40:13.965
<v B>Yeah, that's amazing. I kind of wish I had bought some of these, but

0:40:16.092 --> 0:40:17.037
<v B>maybe you can get them

0:40:17.037 --> 0:40:25.964
<v A>on sale on eBay or something. I don't know if they went up in price or down in price after this. I haven't followed the eBay price trajectory. All right. My

0:40:25.964 --> 0:42:24.980
<v B>tool of the show is Fuse and SSHFS. This is one of these kind of table stakes things where it's really good to have this in our toolbox. I used this recently. Just a recap. So Fuse is a way for it stands for file system in user space, something like that. I don't know if I'm getting that totally right, but basically it's a way for you to mount a file system without being the root user. If you are logging into a system at work, you probably don't have the root account. So you need to use something like Fuse. Even if you are doing something in your house where you do have root access, it could be cumbersome to have to sudo all the time and put in the root password and all that. You might just want to mount a directory right there, as a subdirectory of your home directory. Like imagine mounting your Google Drive or something like that. You want that to just happen. You don't want to have to put in your root password. So Fuse lets you do that. SSHFS is an interesting thing where you can SSH into a machine and you have now remote access. You can run an editor, do all that stuff. There's also something called SCP where it uses SSH, but instead of giving you a pseudo terminal, it gives you a file. So you can take a file from that remote computer and put on your computer. So you can like SCP food.txt to my home directory. And it'll actually or sorry, SCP user@server:food.txt. And so I'll actually go to that server, find the food.txt file, bring it to your computer. That's okay, but it could be kind of cumbersome. I was having a situation where I was creating files on a remote computer and I was wanting

0:42:24.980 --> 0:43:28.880
<v B>to look at them on my laptop and work pretty quickly. I didn't want to have to keep copying them to the laptop and all of that. Also the files were kind of big and all I needed to look at the first part of them, et cetera. So, so I just set up this SSHFS system. And so I mounted a directory on the target computer as a directory on my laptop. And I could just look at all of those files. I could read, you know, the first hundred kilobytes or 10 megabytes of the file without having to copy the whole thing. It worked well. The challenge is, you know, if your laptop goes to sleep, just like any SSH connection, it when it comes back, the file system's like broken and you have to unmount it or remount it. So, you know, it's not perfect. But but it's extremely useful, especially at a pinch. And it's one of these things that's almost ubiquitous because almost any machine you can SSH into. So you're only required to install things on your local machine.

0:43:28.880 --> 0:44:40.820
<v A>I didn't know you were going to talk about this. I didn't know this is. I actually use this for the first time a couple of days ago. No way. Interestingly, there is a Windows version as well. That will allow you to mount another Linux computer over SSH to a drive letter in Windows. And I was on a Windows computer and my I wanted to copy some files from my Steam Deck memory card, which is ext3 formatted. So my Windows computer couldn't see it. Well, don't ask anyways, long story, but the Steam Deck allows you to just enable SSH pretty easily. So I just did that. And this was the this was the path. The path was to basically turn on SSH file system, have it on windows show up as a as mount. And then I could have used, I actually did later end up using the SCP method you talked about. It was a little cleaner, but mounting as a file system was also something that was imminently doable. And then other applications could have pointed at it. So it has this advantage. But yeah, so I did end up using, as you said, so many things have SSH on them, or you can SSH into, and if you can, and it has files, this is a good way of getting access to those files.

0:44:41.980 --> 0:44:50.600
<v B>Yeah, totally. All right, let's jump into Transformers and Large Language Models.

0:44:52.600 --> 0:46:06.440
<v B>Yeah, I mean, there's it's one of these things. It's actually agoraphobic, like there's just so much content and so topical, but we'll start with the basics, which is you know how neural networks store information. If you don't know what a neural network is, we actually had some AI, an AI two-part series. It's a bit dated, but I don't it covers that in pretty good detail. But you know, basically you have these layers and each layer you do a bunch of dot products. This almost becomes like a tensor product to produce the next layer. And so the weights, the things that you're multiplying by, you can change those as part of training this model. By default, you know, it goes from one layer to the next, to the next is like this, think of it as like a DAG, right? And it can split and it can rejoin and there's convolutions, but effectively it's because it's acyclic. It just goes in one direction and then it ends with some target. You compare that target to your expected value for that target and you use that difference to adjust all the weights.

0:46:09.977 --> 0:47:04.062
<v B>Now, you know, people very quickly wanted a way to store information. They wanted these neural networks to be stateful. You know, imagine if you are training a neural net to solve a maze. You could, if it's the same maze every time, you could just train the neural net to solve that maze, and the weights of the neural net will just hold really specific information about that maze. Like, oh, when I see this intersection turn left. But if you wanted to train like a generic neural network to solve any maze and to solve it over time, it has to keep a representation of the maze in the neural network. Right? And so it's a memory that is sort of online, that's independent of training. So this is the goal. And there's a

0:47:05.783 --> 0:47:52.307
<v B>lot of different ways to do it. The kind of obvious thing would be, well, make it cyclical, like take some of the weights and instead of making it an acyclic graph of operations that just ends in this point, make some of those weights just point backwards. And so you execute the network. And when you execute it, some of the data is kind of leftover. Right? And so they call us a Recurrent Neural Network. The problem with this is they're incredibly hard to train and to learn anything meaningful. And there's a ton of reasons for this. But the biggest one is this problem called the vanishing gradient problem. And so the idea is,

0:47:56.796 --> 0:49:56.575
<v B>basically if you multiply a lot of numbers less than one, you very quickly get zero, right? That's basically the gist of the vanishing gradient problem. There's more complexity than that. But if you multiply a bunch of numbers together that are bigger than one, then you quickly go to some huge number, right? That approaches infinity. And so that's not going to happen because you have something called regularization. So that's not an issue. But the other thing where you multiply numbers smaller than one over and over again, and you get zero, that happens. And so you're kind of having this sort of dilemma. It's like either everything goes to infinity or everything goes to zero. Either way, it's not really very usable. So someone came up with LSTMs, Long Short-Term Memory. And they basically said, let's have one process that's going to infinity and let's have another process that's going to zero and then add them together and hope that the two problems cancel each other out. And again, it's one of these things that yes, in theory, you have these long-term gradients, these short-term gradients and the problems of both of them cancel each other out. But in practice, it just is really hard to get it to do things. Throughout my career, tons of people have tried to do LSTMs for all sorts of practical things that all these companies have worked at and it's never worked. There's a time at Facebook, someone came up to me and said, 'Hey, I have this idea. I think we'll do an LSTM to predict the effects over time of people watching things on Facebook.' I was like, 'Forget it. I'm not interested. It's not going to work. Like I've just seen it fail too many times.' And it's like, you know, they call this like a tar pit idea

0:49:56.946 --> 0:49:58.110
<v B>because, you

0:49:59.880 --> 0:50:31.607
<v B>know, you don't really realize you're stuck and then, and it seems, it doesn't seem like it's a problem, maybe a quicksand idea. It's like, it doesn't seem like there's anything bad about that idea. Once you get into it, it just sucks up all of your time. You don't get anything. So LSTMs, you know, not much success. But then something interesting kind of happened: Differentiable Algebra kind of took off. So what that means is, you used

0:50:35.674 --> 0:51:31.024
<v B>to, and Patrick, you might've done this like electrical engineering where they have you kind of derive all of the updates for a neural network. So they show you like, you know, here's how you calculate like the derivative of this type of activation layer and you get these like exact numbers and exactly like, you know, if my answer was four and the answer should have been five, then this weight and this neural network needs to go up by exactly like 1.2 times the learning rate or something. And so you have these like very specific formulas. And as long as you follow the formulas, you'll eventually get to the right place. The problem is the formulas only work in certain circumstances. So you couldn't, for example, say,

0:51:34.028 --> 0:52:30.981
<v B>you couldn't, for example, say like take the maximum of these three values because now like you can't differentiate the max function. Like there's no derivative of the hard max function. Right? And so what came out, what got popular in around 2015 was this idea of numerical differentiation. It's like, instead of trying to come up with the derivative of all of these functions, let's numerically differentiate all of them. And so now you don't, you can actually have a gradient of the hard max function or any function. It's just a numerical gradient, a numerical approximation of a gradient. And so what that means is now you can write pretty much any code, virtually any code and,

0:52:33.660 --> 0:53:09.640
<v B>and backpropagate through it. So, you know, you're not going to change that hard max function. It is what it is, but you'll be able to differentiate through it, and things that happen before and after it that you can change will start changing. So for example, let's say I have three neural networks and then I take the max of the outputs of those three neural networks. And then I have a fourth neural network. All four of those networks can be updating and learning, even though they have this function in the middle that's not differentiable.

0:53:11.279 --> 0:54:27.920
<v B>So all of that leads into this concept called attention layers. So an attention layer is a set of algebra that you apply on three things. One is your query, which is what are you interested in right now? Your keys, which is how does the thing you're interested in now relate to the other things in your list, and your values, which are the other things in the list. So for example, you might say 'The cat jumped over the moon.' That's my those are my values. My query is going to be, let's say dog. And then my keys are going to be what the relationship that I think the other words have to dog. So like cat and dog probably have a pretty close relationship, even though we joke about cats and dogs, but they have a close relationship because they're both animals. But you know, dog and the probably don't have a good relationship because the is just related to everything and it just washes out. Right.

0:54:29.730 --> 0:54:56.940
<v B>So given your query, your relationship of that query with each of these items, and something that describes each of these items, you can then create a total amount of attention. So you get, so you can say like, you know, dog has this much, is capturing this much energy from that sentence.

0:54:59.869 --> 0:55:28.460
<v B>So that gets into self-attention, which is just a fancy way of saying given, for example, a sentence, take every single word in the sentence and find out how much attention the sentence offers each of those words. So if you say, 'The cat jumped over the moon,' for each of those words, for the cat jumped, how much attention am I getting from the other words in that sentence?

0:55:29.147 --> 0:55:45.260
<v A>And for that Jason? So if you're saying like cat related to dog, I imagine you can look across a kind of training corpus and kind of say how often do they appear together? Or, you know, appear next to other words. But for self-attention, how do you figure out that? Like what's holding the weight in a sentence?

0:55:46.460 --> 0:55:46.600
<v B>Yeah.

0:55:51.102 --> 0:56:06.160
<v B>So, the way this works is, the keys, like the weight between two words that you're going to learn over time. So in the beginning is going to be just random.

0:56:08.449 --> 0:58:02.980
<v B>But then when you calculate the attention, then you take all of, and these are actually stacked on top of each other. So you take these stacked attention layers. So, you know, you're learning keys, you're doing the attention algebra, and then you're learning a new set of keys and then you're doing another attention algebra step. And then at the end, you know, all of it is hopefully in service to some task. By default, most of these models are what's called self-supervised models or forward models. So what that means is they're trying to predict the next word in the sentence. So, we'll just walk through like the very first, you know, training step. Everything is totally random. Oh, one thing I didn't mention is usually you have a token embedding that you train somewhere else. There's some other process that takes a word and turns it into a vector of numbers. That could be even all part of the same thing, but usually it's broken up in two systems. So now what I have is for each word, I have a vector of numbers. And so I pass in, let's say 'the cat jumped over.' I pass in that matrix, right? So it's a set of vectors of numbers. And, in the beginning, it's going to say, well, they're all just randomly interacting with each other. So I'm going to get a bunch of random numbers out of that, calculate the attention and do this a bunch of times. And then, that's the encoding step of the Transformer. So what comes out of that whole process we just talked about is a single vector. And this is true.

0:58:03.040 --> 0:58:23.279
<v B>If you're doing Llama Chat, GPT, all of these things, after all these attention layers, what comes out is just a single vector of numbers that describes the entire context of what you've seen so far.

0:58:24.799 --> 0:59:43.400
<v B>Now you have to do something with that embedding. And so what you usually do is you say, 'Okay, I want to predict the next word.' So I'm going to take this embedding. I'm going to create a decoder. I'm going to create a function that takes the embedding as an input and outputs which word I think should go next. And so in the beginning, it's going to be totally random. Even the decoder is going to be random and it's going to output whatever, 'foobar.' It's literally going to pick a random word to output next. Then it's going to look at the actual word. So I think I did 'the cat jumped over.' The next word is going to be... as it's going to say, 'Oh, you got it wrong. It wasn't foobar. It was the,' and so what needs to change? So that next time I ask you that question, you say 'the,' and so that what needs to change, or that's the loss. And that loss is going to be propagated all the way back. And it's going to change all of the attentions among all the pairs of objects and every attention layer. It's going to change the entire encoder. Like everything is going to change a little bit. And this process then repeats a zillion times.

0:59:44.540 --> 0:59:45.300
<v A>A zillion.

0:59:46.640 --> 0:59:47.120
<v B>Yeah.

0:59:49.900 --> 1:00:04.280
<v A>And so then we talked about the trying to get it is saying the comes from some training. So you're just sticking in texts, books, whatever, and learning, 'Hey, try to guess this next sentence.'

1:00:05.560 --> 1:01:01.889
<v B>Yeah, that's right. There's been a bunch of work on this. So sometimes they try to predict sentence fragments. Sometimes they try to predict single words and they run it each time for a different word. But you're right. I mean, people are going through all of Wikipedia. You can download all of Wikipedia. There's something called Common Crawl, which has like gigabytes and gigabytes of text from the internet. And these systems are going through all of this, taking the first words and trying to predict the next one. And this is happening at massive scale. It's just ingesting huge volumes of text. This also works on images and video and all that as well, but it's ingesting huge volumes of text and trying to predict the next thing.

1:01:03.610 --> 1:01:18.038
<v B>And so, yeah. And so, it's kind of remarkable that it works at all, but there's a lot more complexity around if you dive layers deeper. For example,

1:01:19.675 --> 1:01:23.168
<v B>the sentence needs to make sense. And so if

1:01:24.990 --> 1:02:10.180
<v B>you're just predicting one word at a time, you might end up with things that the system might paint itself into a corner. And it might realize, 'Oh, actually, I outputted this word, but now that I've kind of gone three, four words in, I realized I made a mistake three words ago.' You know, we even do this as humans when we're typing, right? So there's, it used to be beam search. Now they're doing something else, but basically there's a way where you kind of look ahead and based on that, you can kind of go back and make some changes. And so it almost becomes like a search system in the decoder that's doing like a best-first search.

1:02:15.900 --> 1:02:55.560
<v A>Yeah. So the, if you ever saw one of those, I guess that's the Hidden Markov Model train something would, you know, you'd feed all the Harry Potter books to a Hidden Markov Model and ask it, you know, or on your keyboard, on your phone, if you, it has like word suggestion. And if you just keep tapping the word suggestions, you get what I would call sentences, but yeah, they don't go anywhere. You're just inserting like plausible next words right after each other. And so you end up with something akin to a sentence structure, but it there's no like story or progression or statement. It just, you know, words that come close to each other, you know, just happen to appear. Yep. Yep.

1:02:55.643 --> 1:03:19.335
<v B>And so, you know, by default, if you put like, uh, is it the cat or the cow jumped over the moon? Right. Isn't that the thing? Yeah. I've been saying 'cat' the whole time, but if you put, you know, 'the cow jumped over the,' it's going to say 'moon' with super high accuracy because it's just seen that a bunch of times on the internet. Right.

1:03:21.124 --> 1:04:34.395
<v B>But so that works pretty well. And so a lot of these things, when you, for example, ask a question to ChatGPT, there's actually, it's embedding your question in a prompt, right? Because otherwise, if you just took one of these naive forward models and you asked and you typed a question and said, 'Generate some texts,' it'll probably generate more questions, right? Like, you know, it's whatever people who wrote that kind of question would write next. And so sometimes it'll generate an answer. Sometimes it might generate more questions. And so typically, you know, it might not be obvious from the user interface, but what usually happens is you will put 'Question: [your question]' and then new paragraph 'Answer:'. And then the model will start generating tokens. So you didn't see the question, the answer, but they're there. And that tells the model like, hey, this is the broader context of what's going on. So this works pretty well. The thing about it is, you know,

1:04:35.982 --> 1:05:43.740
<v B>It's we want to be able to make improvements. And so it's such a huge corpus of data that we can't, it's not like normal machine learning stuff where you say, 'Oh, I'm going to just change my data set to improve the model' because the data sets are like the entire internet. So like, you can't really do a whole lot with that. So there needs to be some way to fine-tune things after the fact. And, once you have this forward model. The first attempt at this, which is what OpenAI did with their initial models was called RLHF—Reinforcement Learning from Human Feedback. And basically the way this works is they would have ChatGPT, and this is before it was released or anything, generate four or five different answers. They would give those answers to people. The people would score the answers, and then the ChatGPT model would optimize for that score and try and get the highest score possible.

1:05:46.266 --> 1:05:48.342
<v B>The problem with that

1:05:49.894 --> 1:07:16.060
<v B>is numeric scores only make sense when there's a real unit of measure, you know, like is a score of four really twice as good as a score of two? It's very hard to get people to think in such linear ways, right? Such proportional ways. And so people are just basically the human raters are just scoring everything either, you know, a five or a one, very few two, three, and fours. And even if they were, it wasn't very proportional, right? A five wasn't five times better than a one, et cetera. So what's become more popular now is Direct—I think it's called Direct Policy Optimization, but basically it's pairwise ranking. So you have the system generate two answers. You give it to a person and the person decides which answer is better. And then the system is encouraged to give the better answer and discouraged from giving the weaker answer directly. And so there's a paper that came out, I think only six months ago, maybe a year ago on this. But as much as I love Reinforcement Learning—you know, that was my PhD—I also was pretty skeptical of this. I kind of felt like this was a poor use of Reinforcement Learning. And so I was not very surprised when Direct Policy Optimization came out.

1:07:18.488 --> 1:07:37.920
<v B>And so, you can imagine this is how this is how a lot of these things are ironed out. So a lot of ambiguities and sentences and things like that, you know, things that require a lot of common sense reasoning, those are ironed out through DPO.

1:07:37.920 --> 1:08:08.819
<v A>So the, but this works. So they do the initial training you described, and then they sort of go into this fine-tuning step. But the answers—I mean, I guess they could become part of like a future training corpus, but they aren't trained on the same way because if you have it generate an answer from its current weights, and then you try to tweak them, that's different than showing answers from another version of ChatGPT one or whatever and saying, 'Hey, here are the responses they got.' Or like, how does that work? You just sort of accumulate all of them over time.

1:08:11.000 --> 1:09:02.920
<v B>Yeah, it's a good question. Let me see if I understand. So, in the beginning, you're just trying to predict the next token and there's no preference, right? Then later on, you're given a preference, but for an entire answer. And so there is somewhat of a credit assignment problem there. It's like, which token caused that answer to be good or bad? But again, with enough data, it works itself out. But you're right. The loss is really different. And so the way that the network is changing in that second phase is very different. So, yeah, you can't really go back once you've started this preference approach. I mean, you could—there's nothing to stop you—but it's like the two are kind of disrupting each other. And so you just disrupt the first with the second and then you get what you want.

1:09:03.720 --> 1:09:20.967
<v A>So, yeah. So if you did like 10,000 rankings, I don't know how many they do at the end. They don't, it is not necessarily reusable. Like if they had a new update to the corpus or whatever, they would have the weights, they'd fine-tune rolling forward, but they would go do another 10,000. They wouldn't just like play back the previous 10,000. Oh, now

1:09:20.967 --> 1:09:46.559
<v B>I see what you're saying. This is really interesting. So, okay. So you're saying, I have ChatGPT-6 and it generates a totally different answer than the ones than either of the ones I sent to the human last time. Yeah. So yeah, this gets complicated, but there is something called Importance Weighting. And basically the gist of it is like this.

1:09:48.135 --> 1:09:49.823
<v B>When a token is generated,

1:09:51.544 --> 1:11:17.780
<v B>It's generated from a distribution. So it's not like the neural network runs and then at the end it says 'cow'; it actually runs and it outputs a probability over the entire space of all words that it could output. And it just happens that 'cow' had the highest value, right? But you can normalize that output vector. And now what you get is a probability mass function over all the words. So there is—for example, there is a chance that ChatGPT-6 and ChatGPT-5, or ChatGPT-5 with two random seeds, two different random seeds, there is a chance that they will both generate exactly the same answer. It might be a low chance, but it's there. And so there's actually, there is a chance that it will generate anything, right? Any any sequence of words. And so because that chance is non-zero, you can do something called importance weighting where you say, 'Okay, the new ChatGPT, although it generated nothing like the two answers I sent to a human last time, it still has a chance of generating those two answers.' And so if it's—it should be—it should be more likely to generate the better answer, even though it really didn't want to generate either of them.

1:11:18.660 --> 1:11:32.940
<v A>Oh, so I think it doesn't give the—yeah, like it doesn't pick the final answer. It goes and looks at the distribution and says, was your distribution closer to the better one or the worst one? And then you can somewhat reuse. I see. Oh, that's interesting. Yeah, exactly.

1:11:33.120 --> 1:12:02.393
<v B>Now the thing about importance weighting is as two problems. One is—you know, the easier problem is—those numbers are going to be small. You know, the chance that you do anything like with that complexity is going to be small, and you're dividing small numbers against each other. And so things get a little crazy there. And the other thing is, oh yeah, you might have to move a lot or actually rather—it might

1:12:04.080 --> 1:12:07.962
<v B>be that the answer is really orthogonal to either.

1:12:08.040 --> 1:12:11.260
<v A>It could be like almost like equidistance from both of them.

1:12:11.580 --> 1:12:23.019
<v B>Yeah, exactly. And so in that case you end up—doesn't matter what answer you pick then. So it's not perfect, but it is; you still get a lot of value out of it. It's not like that old work was wasted.

1:12:25.039 --> 1:12:25.107
<v B>Oh,

1:12:25.107 --> 1:12:45.780
<v A>That's cool. Yeah. I mean, cause I guess like when you're moving up versions, you could ask it a question. And in one case it generates two different diatribes about cows and moons. But then in the next one, it generates a picture of a cow jumping over a moon, and actually that is better. But then now you're stuck with this problem of like, oh, I have a picture versus two sentences. And so, yeah, I hear you. Yeah. Makes sense.

1:12:46.140 --> 1:13:03.000
<v B>But yeah, the data moat is real. You know what I mean? All this data and all that human effort that's gone into labeling and ranking all of that data is just permanently valuable. So, so yeah, it's a huge boon for a lot of these companies that have started early. Yeah.

1:13:04.300 --> 1:13:27.920
<v A>Oh, that was—so the large language model, and then we talked about all this. So the large there just comes from the fact that the parameter count has gotten so numerous, and it happens to be that these large language models today have a very similar architecture that you're describing, and not so much that there's like, you couldn't do it a different way. Right? We just don't know of a different, better way.

1:13:28.560 --> 1:13:58.122
<v B>Yeah, exactly. So, the recurrent neural network, you couldn't make it large because the gradients would vanish. The LSTM, you couldn't make it large because it was an unstable equilibrium. And so the current—the gradients would either vanish again or explode. And so, you couldn't do it. But with the attention layers in the beginning, a lot of people, myself included, were a little skeptical of the attention approach because,

1:14:01.142 --> 1:15:09.480
<v B>it felt like a strange compromise between convolution where you have like a relatively small mask, but because the mask is small, it can kind of roam roving eye, you know, through an image—kind of a compromise between that and an LSTM, which in theory has like an infinite horizon. Like you could, in theory, you could feed—you know, a near-infinite amount of—not literally an infinite, but you can feed an incredibly large amount of data to LSTM. It could remember all of it. There's no limit. So attention was strange in that, in that, you know, it had a lot of the features of long term, short term, but it didn't have the benefits. But you know, the fact that it was a stable training, that was really incredibly useful. It's one of these things that's hard to know because in theory everything is stable, right? It's hard to know in practice, like what will this do? Would you have 7 billion worth of this? Right? And it held up really well at that scale.

1:15:11.440 --> 1:15:45.340
<v A>Yeah, that's really interesting. So it's like when you grow really big, it was hard to predict what would happen, but that, uh, it has the attributes that sort of end up working at that scale and the compromise. But I guess like to the same thing, it doesn't mean it's none of this means that it is the sort of like optimal or correct answer. It's just the best we know today. So if we found a new LSTM, I mean, maybe not, but like LSTM, like you found a new way of, you know, solving the problems or adding another thing. It could do better, but we just don't know that today. So someone would have to figure it out.

1:15:46.280 --> 1:15:46.800
<v B>Yep. Yep.

1:15:49.615 --> 1:15:53.125
<v B>Now here's where it gets really tricky is,

1:15:54.644 --> 1:16:20.901
<v B>You know, all of this is very direct. Like, you know, you have this direct policy optimization. You know, the RLHF was technically Reinforcement Learning, but it was a very direct approach. And so a lot of these systems, you know, they can't actually do things like they can't get rewarded or punished really. And if

1:16:22.825 --> 1:17:28.900
<v B>they do, it's this really sort of brute way where, you know, you could tell ChatGPT, like go invest in stocks, like, oh, you're bankrupt. That was bad. You know, it's like, but it can't, it can't, it can't really reason in a way, like it can't sort of build a world model that it can then use to reason and to make decisions. So, so I think the future of this is this really interesting work coming out of Facebook called JEPA, which is I think Joint Embedding Policy Architecture. But it's basically a way to combine a lot of these embedding approaches, whether it's a large language model or a large image model, or maybe it's a large, you know, actuator motor response module or something. It doesn't necessarily have to be language, but a way to sort of combine those with decision making. And I think that's really going to be super exciting, but it's going to take a while for that research to mature.

1:17:30.560 --> 1:17:57.220
<v A>Yeah. So I guess like to your point, we were talking about earlier picking a racing line and driving your car towards it. You could tell ChatGPT or one of these large language models how you want, and it maybe it could generate your code or something, but it can't actually go play the game itself. Like it doesn't, it doesn't have the hooks. It doesn't have the inputs to go do it. There's no module. And this, the architecture isn't really set up to be a Mario Kart AI agent.

1:17:57.960 --> 1:18:41.280
<v B>Right. And then, you know, people joke about telling ChatGPT like your answer is terrible. And then it generates another one. And yeah, that's funny, but really it's not, it's not really trying to optimize for some objective. Like it will pivot. But even that pivot is kind of artificial. It's not really directionally heading towards some place of greater value. And so, you know, yeah, you can't plug ChatGPT straight into Mario Kart and like over time, get something that get it to produce Python code that gives you a higher and higher score over time. Like there's not an easy way to do that yet.

1:18:41.800 --> 1:19:36.972
<v A>Yeah. Yeah. So the I always find that funny, I guess like the, I don't know if it's anthropomorphizing, it's not the right way, but you're right. People like fuss at ChatGPT. Like I'm going to, I'm going to kill you if you, you know, I'm going to unplug you if you do one more thing, but it's all just adding to the context. It's not like you said, it's not actually evolving forward. So you could just written all of that out and claimed what it told you like is what it told you, even though it didn't and just put it all into prompt and it would do the same thing. Right. In other words, like if someone comes back to you and says, Jason, you're bad because you did this thing and you didn't do that thing. You're going to be like, what? No, like you're just going to ignore it. But with ChatGPT, if it gives you an output, if you were to start a new session and copy that, what it told you and what your response was into the very beginning, it's the same. Like it's not actually like remembering and evolving in that way. It's just building this ongoing context that kind of feels similar. Right. And,

1:19:36.972 --> 1:20:40.240
<v B>All of this, you know, even what's coming in the future is not going to kill coding. Like we should spend probably the last five minutes of the show talking about this, but if you are a follower of the show, you know, maybe you are working professional or not yet, you're in college and you're thinking, oh, if I major in Computer Science, there won't be a job for me because ChatGPT will have my job. That is not going to happen. You know, I think coding is just a way of solving problems at the end of the day. And so, you know, the media might change, but the need to solve hard problems is not going away anytime soon. And if anything, ChatGPT will actually automate everyone else's job. Like people who are solving easier problems. Those are the people who really should be worried. If you're in here, you know, listening to the show, going to college or even on your own learning to be a programmer, you're doing it because you want to solve hard problems. And that's doing hard things, really mentally difficult things is going to be one of the last things to get automated.

1:20:40.240 --> 1:21:56.241
<v A>Yeah, I mean, I think they've always talked about it for a long time and accuracy hasn't been there. But the one I always think about is like, I don't know, just using some like X-ray tech, like you go in and, you know, you have a broken arm and you get your X-ray and the job of the person who I don't know what the right word reads the X-ray, like looks at the X-ray is supposed to look at the notes from the doctor about what they think was wrong, look at the X-ray and check for problems and then like write a response. But if you gave it to 100 doctors, they won't all say the same thing. But there is a correct reading of the chart. And hopefully most of them would give the same answer. This one feels like very difficult before you would just given it to like, you know, something that would use like Jason was mentioning a Convolutional Neural Network or something, try to highlight where the fracture is. But now you're getting to the point where you could give the context of the doctor's notes and hey, I was in an accident and my arm is hurting. And you know, here's my, yeah. And so it would, you know, kind of understand what it's attempting to do. And you're right, those things are problem solving, but not really in the same way that like you said, hey, I want to build a game where you're a go-kart racing turtle and throw shells at each other. And like, you know, that kind of problem solving is a fundamentally different approach. Right, right. And also, you know,

1:21:57.895 --> 1:23:26.000
<v B>When you're building anything, this is true of anything you're building, software or anything. There's really two things that you're constantly adapting to. One is like product market fit, you know, like are people enjoying my game? Who is enjoying the game? And then the second one is, you know, quality of life. So maybe my game is actually fun, but there's too many buttons in the menu and people like are like not even getting past the main menu. Right. And so you have to constantly adapt to all of these changes and you have to decide like what parts to be flexible, what parts of the code should be inflexible and written quickly. And these are all things that, you know, trade-offs that ChatGPT is not going to make very effectively or any, any AI is not going to make very effectively. So, you know, and who's to say what's going to happen way out in the distant future. But I would say, you know, if you're listening to this show, if you're interested in this topic, your job is extremely safe. I think that you can, if, if, if you said, oh, I'm going to pivot to accounting. Well, that's, that's probably not on why. No, I mean, I brought their laws and accountants, so it's not an insult to accountants, but when you say, you know, you are in a very safe profession, this idea that coding is dead or will be automated is absurd. And, don't worry about it.

1:23:26.519 --> 1:23:31.787
<v A>And yeah. And like normal term, I think there's a caveat there. Like who knows what happens in a thousand years,

1:23:32.969 --> 1:23:49.099
<v B>But yeah, exactly. I—but I—again, I think you know your job will be one of the last ones to go. So by then I saw this crazy stat that like 99% of job titles didn't exist a hundred years ago, something like that. Oh, interesting.

1:23:49.719 --> 1:24:38.660
<v A>Yeah. Yeah. I mean, I think like taking as an example, like someone who is an actor or something, which I know that there's a bunch of politics around that and fighting whatever, but taking an actor and their voice and their body image and like the way they—this is very easy, easy to drain on. And then get them to do new things and become an AI agent of some sort. And you know, probably like they call those people quote unquote talent, but like their talent is something very, very specific, and mostly a function of how they look and how they sound. So even the things that they do on camera, all scripted, right? The writers and the producers and all of that stuff. And so, yeah, I think you're right without saying when or if—I mean, saying one of the last to go is a reassuring fact.

1:24:39.340 --> 1:24:48.080
<v B>Yeah. I mean, you would be late enough that you would see the writing on the wall and you would pivot to one of the 99% of jobs that are coming out in the next hundred years.

1:24:48.740 --> 1:24:55.660
<v A>I've seen the Matrix. Your job is to become a heater, eat food, walk around in the Matrix and provide a warmth for the robots.

1:24:55.880 --> 1:25:30.280
<v B>The robot army. So good. All right, folks. I think we'll put a wrap on that. If you have any questions about LLMs, join our Discord. I do have Discord is one of the few apps actually have notifications turned on. So when people post in Discord, I do see it right away. Join our Discord, you know, support us on Patreon. We really love and thank all of our supporters. You know, we're putting all that money back in this show, trying to get more people—kids, adults—into programming. And we will catch everybody next show. Thanks everyone.

1:25:30.280 --> 1:25:44.840
<v A>Music by Eric Barndoller.

1:25:46.599 --> 1:26:07.060
<v B>Programming Throwdown is distributed under a Creative Commons Attribution Sharealike 2.0 license. You're free to share, copy, distribute, transmit the work, to remix, adapt the work, but you must provide an attribution, uh, to, uh, Patrick and I, and, uh, share alike in kind. Thanks Ann

