WEBVTT

0:00:15.360 --> 0:00:21.040
<v A>Programming throwdown episode 188 World Models Take it away, Jason.

0:00:21.760 --> 0:01:58.430
<v B>Hey everybody. Excited to dive into World Models, but as you know, we, we follow the sort of show protocol here. We're going to jump into our intro topic. I've been doing a lot of running, so just a bit of a background. I have a buddy, Dave, and Dave came into town and he asked if we wanted to go for a run. This is like September. And he's talking about how, you know, he's in really bad shape and he got into running and it kind of fixed a lot of issues he's having. And the last time I ran was probably like 10, 20 years ago. And so I just assumed that, oh, you know, I must be around the same level of performance. No, like, I ran two, three blocks. I was like about to pass out. And so that got me on this running kick. And the thing that's awesome now is there's so much technology. You know, like, I have the Galaxy Watch from Samsung and you know, they have this running coach. So I'm up to like level six on the coach where, you know, I'm basically running 10Ks a couple times a week and been just running almost every day since October. And it's, it's been awesome. And the, the technologies I think has helped a ton. Like, you can measure your heart rate. It tells you if you're bouncing too much or if you're, you know, running with good posture and all of that, or at least I think they measure. They call stiffness is what they call it. But, but yeah, I've been a big fan. I know, Patrick, you've been running for a long time.

0:01:59.390 --> 0:03:10.360
<v A>Yeah. So I guess, similar story. I guess we're just talking about being old programmers, I think basically, like, yeah, I was on the, you know, my kids and maybe my, my nephews were over and like trying to play like soccer and just like running after the ball and then like stitching my side. I'm like, wait a minute, there's like, this is sad. And so, yeah, I think about three years ago now, I think I picked up and started running. And so I've, I've run more, I've run a little less, but I've, I've stuck with it, I think. You know, I don't have this. My watch doesn't have the same measures as you, but yeah, similar. Have a, you know, you can get a pretty cheap, cheap watch. You know, it just has GPS and, you know, some sort of heart rate, whatever. And mostly be good to go in my Opinion, you know, if you're not gonna, you don't have to take it serious. But there's like anything, there's. People take it way too serious and then there's, you know, you know, just. You're out there doing it, having a good time. Kind of what I try to do a little bit in the middle. Like I try to beat my own, you know, thing. But there's whatever. I don't like epigenetic stuff. So like, if you, this is my theory, it makes me feel better. Like if you ran track, you hear these people, like, oh, after 20 years, I started running again, you know, for the first time, seriously.

0:03:10.360 --> 0:03:10.920
<v B>And I ran.

0:03:10.920 --> 0:04:12.040
<v A>And they'll say some obscene time. Oh, I ran, you know, only a 19 minute 5k. It was so sad. And you're like, oh my God, what? Like, I've been running for three years and I can't run a sub 25k, like, screw you. But then it's like, oh, it turns out they ran cross country or were, you know, like a high level soccer player in high school, you know, and so I think your body, you know, kind of remembers, like, is epigenetic. Whatever we want. Like, your body learns to express certain things. When you do something for a really long time. And when you're young, you're like, I don't know, almost invincible. Like, there's like so much you can do. You can kind of just do anything and get better. But, you know, as a not young person, those things aren't true. But then a lot of the advice you read is from people who've been doing it a while. Makes sense. They feel like they have something to say, but they forget what it's like in the beginning, like all the aches and pains and the randomness. And they'll say, as an example, oh, you should go out for a run and never go above. And then they'll say, some heart rate. It's like, listen, when you're first starting, you freaking, like walk five steps and your heart's gonna be above that. Like, just start moving. And then like, let that stuff come eventually. Like set some goal that motivates you.

0:04:12.040 --> 0:04:12.640
<v B>That's fine.

0:04:12.640 --> 0:05:22.900
<v A>But you gotta kind of tone down a lot of the. Even the quote unquote beginner advice. Because it's like beginner again or, you know, beginner from someone who forgot what it was after. And even now, like, I find it hard to remember just like how much it hurt to finish, you know, some small distance, you know, now that I run more and. And then there's always people though, a lot faster, run a lot further, you know, I don't know, it's like anything in life, you got to just find your, your contentment, I guess. But certainly from a fitness thing, so much better. Like going up a flight of stairs, no problem. Playing sports, like all those little things. I forgot how much until I think about it. Like, remember when I would like avoid doing this because it would make me hot and sweaty and like high heart rate. Now something will happen like, oh, we forgot something in the house. And I'll, I'll dash, I'll go pretty quick to go grab it out of the house and come back. And I might be, you know, a little, you know, winded when I get back, but I can do it. Before it would have been like, heck no, I'm not doing that. Like, I'll be killed over sideways if I were. So just all these little things that improve. So yes, general PSA for like some sorts of exercise, I'll say some form of cardio, but also weightlifting I think is important.

0:05:22.900 --> 0:05:23.500
<v B>I do both.

0:05:23.500 --> 0:05:26.180
<v A>But, you know, I, yeah, general health.

0:05:26.580 --> 0:06:40.690
<v B>Yeah. I think, you know, one nice thing about the technology is that, you know, it kind of meets you where you're at. So I personally have read absolutely nothing about running. So I, I don't know, like, I know for myself I'm trying, I can, I'm now like pretty regularly under a 30 minute 5k. Nice. Um, you know, but that's when I'm trying to run a 10k each time. So I'm pacing myself. If I just ran a 5K. Yeah. There's no way I could get it to 20 minutes right now. But, but, but I could probably do it better than just under 30. But, but the point is like, you know, these things are adaptive and you're, you only have yourself to really compete with, with this technology, which is really nice and as Patrick says, totally affordable. I got a, and they're not sponsoring the show, unfortunately, but I got a Samsung Galaxy which is, you know, it's not like a Apple watch. Not like, it's not that expensive. There's different models. I got the one with like the rubber, you know, wristband. So, you know, it's, that's less expensive. So you could definitely, probably for like 100, 150 bucks, you have something that would, that would help you become, you know, sky's the limit as far as how, how good of a run. Yeah.

0:06:40.690 --> 0:06:53.570
<v A>Or even just using your phone and a phone, like download one of the apps with GPS is like to get started. I think that's perfect. Like you said, just run in a way that doesn't make you want to kill over and die. And yeah, try to get better.

0:06:54.050 --> 0:06:57.850
<v B>Yeah, you really want the heartbeat. You need the heartbeat.

0:06:57.850 --> 0:07:00.170
<v A>I think that, yeah, that is, it is a big advantage.

0:07:00.170 --> 0:07:14.260
<v B>That's true. I've noticed that. And this is a, probably an H thing, but like when I hit 170 beats per minute, then I know that I'm running outta steam. So I basically like, I can't sustain like 1 71, 75.

0:07:14.580 --> 0:07:19.740
<v A>See, I'm, I'm nerdy as I did the opposite. So I'm already thinking like, oh yeah, you're hitting your LT2 lactate threshold.

0:07:19.740 --> 0:07:21.180
<v B>Whatever. I, whatever.

0:07:21.180 --> 0:07:51.610
<v A>Like that's just my, your, your body's accumulating lactate faster than you can clear it. And so, yeah, there's like a heart rate range that's considered like threshold and then over threshold and you can get nerdy as you want. Um, I, I tend to. That's just my style. So. But yeah, you can even go get like meters to prick your finger and take like a blood draw to see if you're exercising in the right range. And there's basically all the different energy systems of the body and different exertion levels utilize different energy and muscle sort of like compositions.

0:07:52.730 --> 0:07:54.570
<v B>Wow. Yeah, I didn't know that.

0:07:54.650 --> 0:08:07.210
<v A>So, yeah, so it's like you said, like if you were to sprint for a hundred meters, you can go a lot faster than if you were to run 5K. And that's everybody. Because just like you use a completely different energy system and then 5k is very different than, you know, a marathon.

0:08:08.490 --> 0:08:17.770
<v B>Yeah, totally. Makes sense. Um, I recently ran the Austin Cap Metro 10K, where you run around the capital.

0:08:18.250 --> 0:08:19.370
<v A>Oh, that's cool.

0:08:19.370 --> 0:08:48.160
<v B>That was very cool. It's a lot of people, people think of Texas and they think of like cacti and tumbleweeds or whatever. And that is probably the western Texas. Yeah, but Austin is basically, you know, rolling hills and very green and. And yeah, so it made for a very wet and hilly. 10k is a little slippery, but. But yeah, it was, it was very satisfying. So.

0:08:48.400 --> 0:08:51.680
<v A>And then you put on your boots and cowboy hat at the end and rode your horse home.

0:08:52.970 --> 0:09:16.610
<v B>No, that was the, that was the 10k. I'm not running that thing on my own feet now. Forget that. Oh, man. All right, on to news. So I wanted to post this video, something that's kind of coming into vogue. It's called flow matching. Have you heard of this Patrick, only a tad bit.

0:09:16.610 --> 0:09:18.650
<v A>It's outside of my expertise. So.

0:09:19.100 --> 0:13:15.270
<v B>Okay, so I'll, I'll kind of explain. I posted a link to the video that has a much longer and better explanation. But basically there are these models called diffusion models. And the way it works is you have an image and you assume this image is great. It's like a picture of your dog. Great dog, right? You add some noise to the image, so maybe four or five times you add Gaussian noise to the image. And, and so now, you know, you have a fuzzy looking, not clear picture of your dog. Then you train a model to reverse the noise. And, you know, if you added five steps of noise, then the model runs five times to try to remove five layers of noise. And then when the model is done, you compare the result to your original photo because, you know, this isn't like a photo from the 1800s. Like, you have the actual correct photo photo. And so, you know, like, oh, this pixel was made gray and it didn't do a good job of putting it back to brown or something like that. So, you know, the first time you run this model when you're training, you know, you give it this fuzzy dog image and it completely destroys it. And you say, oh, no, that was not right. Here's what all the different pixels should have been. And then over time, you train the model and it gets better and better. Diffusion is difficult for many, many reasons, which I won't go into. But flow matching says, okay, well, what if we just added the noise once? It turns out if you just add the noise once and take, take it, take it away, then the problem is you only get like, you're expected to go from start to finish in one step. So, you know, instead of adding a little bit of noise five times and having five chances to kind of fix it, you add a ton of noise once and then the model just can't fix it and you're stuck. So that's why diffusion had to do this five, maybe 10, 30 times and then undo it. Right? Flow matching says, well, what if undoing this noise was actually a dynamical system? So what if you imagine like a set of dynamics, you know, differential equations that take you from noise to clean. So if you had this differential equation, then you could plug it into a solver and, you know, the differential equation solver can take as many substeps as it wants. Maybe 5 is the answer, maybe 100 is the answer. But instead of us manually programming in, you know, five noises, five denoises, we just do One big noise and then we let a differential equation figure the rest out. And there are these things called neural differential equations where they're kind of deeply embedded with the deep learning architecture. So you get gradients and all of that. So trying to explain something very quickly here, but the gist of it is all these techniques that you see for like generating images, generating videos, you know, taking an image and making a video out of it, like they're all getting way, way better because of flow matching. And, and so in fact, if you've ever used the flux model from Black Forest Labs, that model does flow matching. So they were the first ones to really publish that. And you know, and flux just totally dominates over its contemporaries. So. So yeah, flow matching is now being used heavily in many, many areas. Every paper I read nowadays is basically like, you know, we took this problem and we added flow matching and we made it better. So something to put on people's radar.

0:13:17.030 --> 0:14:36.110
<v A>Now that you talk about. I actually think I did watch a video a little bit about, about that. And yeah, it's like, it's one of those unlocks where. Oh, okay. As a, how do you want to say, like layperson person, not familiar with the inner workings of a lot of this stuff, it does help explain, like if you think about every image possible and all the images and the clustering and how things flow from, like you said, sort of like an ambiguous space to like a more precise actual. This is an image. It, it's like I wouldn't be able to, to write it myself, but it's like, ah, okay, like I have a way of thinking about it and I have been encouraging people, you know, again, it's outside of my space, I don't know, expert, but even for LLMs in general to start to think about the techniques and the approaches, because I think just as good engineers, you need to know enough about these things to like, kind of feel like, you know, where the, you know, things may lie. And sure, at Some limit the LLMs are doing things that maybe are unintuitive or unpredictable, but how they work or operate at like a base level. So it makes a lot of sense. And I think there's a lot of value and maybe not now, but learning those things and then cross applying them to something else and saying, hey, is there a way to think about this part of computer science for the part that I work in?

0:14:37.070 --> 0:14:38.190
<v B>Yeah, totally.

0:14:40.520 --> 0:16:31.040
<v A>My next News article is OpenCV5 has been released. So going from I guess something that's much more like, I guess state of the Art, sort of deep learning to something that, you know, was an early exposure for many to what were neural networks and computer vision and AI OpenCV, which I assume is open computer vision. I've never actually looked it up. Okay, good. They released version 5 where they did a lot of refactoring and under the hood updates to prepare for, you know, like the new, I say like operating regime for a lot of stuff, but people are noting a lot of performance improvements, a lot of cleanup, but you know, also just using the space as a general shout out and reading comments, you know, from people about the release. I found myself resonating with some of them, which is if you've ever found yourself maybe now with AI tools to help, it's a lot easier. But just back in the day, even just doing something with the pixels in an image was non trivial. Like I have an image, a jpeg, how do I just get it into like an array of R, an array of G and an array of B or interlaced, however you want to think about it. But like even just getting from JPEG to that is, you know, non trivial. OpenCV is by no means like a light library, but you know, having it in something like Python or in C even and being able to do stuff with individual frames of a video or open images, put something over an image, do basic manipulations. Even if you're not doing, you know, this kind of research, if you've ever end up in anything adjacent to that domain, OpenCV can, can definitely be clutch for sort of helping you test out ideas, do things quickly. It has a lot of built in, you know, features and the under the hood algorithms tend to be well implemented, well documented and a great way for you to even learn what it's doing and kind of play with the parameters and understand what's changing. So that's my shout out for OpenCV5.

0:16:31.860 --> 0:16:55.140
<v B>Yeah, just to double down on that. I mean it's amazing. I needed to do some extrinsic calibration the other day, so I had several cameras pointed at the same thing and I needed to reverse engineer the 3D pose of the cameras. And OpenCV just has a function that does it. It's like, okay, that made it a lot easier, you know.

0:16:56.820 --> 0:17:29.360
<v A>Yeah. And you see, you'll see, once you kind of start to recognize it, you'll see stuff everywhere. So like example images that get brought up, or people holding chessboards up at angles and like, you know, with webcams. And a lot of that comes from yes, in General, like machine vision techniques, computer vision techniques. A lot of it comes from, like, the things that make it easy in openCV to do. Like you're saying extrinsic, which understanding basically the pose of the camera, where the pose where the camera lives and is pointed in 3D space. And you'll even see, like, some of the little. They're not QR codes. I call them something else I don't remember.

0:17:29.360 --> 0:17:30.360
<v B>Yeah, April tags.

0:17:30.360 --> 0:17:47.400
<v A>Ah, yeah, you go April tags on things for tracking or motion capture. If you ever see a demo, you know, maybe of not like a super professional, but like, maybe a hobbyist doing something you'll start to learn. Oh, yeah, yeah. These things are part of, like, a toolbox. And often that toolbox bumps into OpenCV as implementation at some level.

0:17:47.720 --> 0:17:49.560
<v B>Have you seen a Charuco board?

0:17:50.200 --> 0:17:52.240
<v A>What is a Churuco board? Maybe.

0:17:52.240 --> 0:17:59.380
<v B>I don't know. The word is a. Is a chessboard. A black and white chessboard. But on every white square there is an April tag.

0:17:59.620 --> 0:18:06.820
<v A>Okay, I see it. Yes, I have seen this different. I didn't know it had a name, but that's. So you can understand if it's rotated, I guess, like if it's flipped.

0:18:06.820 --> 0:18:12.980
<v B>Exactly. Okay. I just like the name. It sounds like a Taco Bell menu item.

0:18:13.220 --> 0:18:17.860
<v A>I will say, when I typed in Churuca board, I got a lot of images for charcuterie boards.

0:18:18.180 --> 0:18:19.460
<v B>Oh, yeah, so.

0:18:20.340 --> 0:18:26.820
<v A>So I got both. I got the chessboard with April tags and the white squares, but I also got some advertisements.

0:18:27.460 --> 0:18:30.100
<v B>Calibrate your robot arm and it'll make you a sandwich.

0:18:32.260 --> 0:18:38.100
<v A>These are cool now. Okay, I gotta stop looking at this now. I want to go do something. Okay. On to you.

0:18:38.420 --> 0:19:02.550
<v B>My next news story. Claude Fable beats Pokemon with no harness. So I don't know if you remember if we covered the, you know. Claude plays Pokemon craze. That took over Twitch for a while, but the idea is someone took a screenshot of the game and gave it to Claude and said, what should I do? Should I move up? Should I move left? Et cetera.

0:19:03.190 --> 0:19:05.630
<v A>Oh, how many tokens. I know.

0:19:05.630 --> 0:21:21.200
<v B>Yeah. So obviously that didn't work. Right. That just was inefficient. And Claude kept forgetting, you know, what was in its inventory and all that. Ah. They started building harnesses. And so you would feed Claude this doctored image where you manually annotated the coordinates of all the squares in global space. So you're like in Pokemon, there's a set of zones, and then each zone has a row and a column. And so with the zone row and column, you can globally point to a tile somewhere in the Pokemon universe. So. So. So Claude, now you give it these. These things, and you also. They kept track of the or. I think they pulled it out of the ram, you know, the inventory, your Pokemon, their health, all this. So Claude now gets his whole text description of the state of the game and an image that's been annotated. And now Claude can. Now, the actions of Claude are, you know, go to cell 0, 0, 0, go to cell 5, 7, 6, et cetera. And. And then as part of the external harness, there was a pathfinding element that would, you know, that would. If you need to go to 000, it would just use a star to get you there, and then it would move the character around. And so now, you know, it only took quad on the order of hundreds of actions, as opposed to thousands or tens of thousands. Right. And so with all this harnessing and all this extra code, they were able to beat Pokemon with. With Claude. Okay. Now with the new version of Claude that just came out, they got rid of all of that. They basically said, hey, give me a list of button presses. And so Claude will return. Just a chunk of button presses. You know, press A, then press up or whatever. And with no harnesses, no doctored images, no reading the system ram, none of that, it's able to beat Pokemon now, which is really a testament to how much progress has been made in the past couple of years.

0:21:22.160 --> 0:21:31.120
<v A>Oh, that's crazy. So they don't even let it build it. It's not like they let it build its own state. They just tell it to, like, basically go play.

0:21:31.360 --> 0:21:42.820
<v B>Yeah. So it might be that. That as part of the instructions to Claude, you know, it can create notes for itself. But the important thing is that there are no human in the.

0:21:42.820 --> 0:21:50.380
<v A>Yeah, yeah, yeah. You're not putting like a. It's not like a expert system blend. You're not saying, hey, here's a really good way to think about a simplified version of the game.

0:21:50.860 --> 0:22:05.220
<v B>Exactly. In fact, what you said must be true because. Because, you know, when you go back to Claude, you might not have the ability to see your inventory or stuff like that. So. So, yeah, there's assuredly there's some way for Claude to, you know, feedback into itself.

0:22:05.220 --> 0:22:17.340
<v A>Yeah. Which is. Which is fine if it once. Like, that's what we do too, as humans. Right. Like, you learn to push a button to bring up the inventory. Or if you're playing certain puzzle games, you get out a piece of paper and you know, write down, you know, stuff.

0:22:17.340 --> 0:22:27.860
<v B>Yeah, yeah, but it's just amazing. I mean, I'm every, pretty much like multiple times a year, I'm just absolutely shocked at what comes out on in this world.

0:22:28.020 --> 0:24:34.370
<v A>So, yeah, I mean, I think it's been interesting. I'll. I'll riff on what you're saying, which is originally there was all this work, you know, to, to kind of help cloud code or these other things, like with harnessing, with, you know, putting all these other like ways to approach things and break them down. And over time it's just like, not really necessary. But I do see some work being put in which I think is really cool. Like you say here, no harness, but the other way to think about it is just having the model design like a custom harness for the problem and it doing itself. And some people saying kind of that'll be like a new thing, like for a specific version for really generic tasks, like just edit some code. Maybe not. But like, think about, you have some specific, like you're doing computer vision stuff, so you really want, hey, you know, here's how to make the frame grabber, here's how to do whatever. And those things live in your harness. And then your harness is part of, you know, what you use where harness, like you said, it's just like a way to manage context, a way to have these tools available, a way to do these interactions in like some sort of loop. And there's other things around the margin that seem really interesting. Like there's something called DSPY which gets into sort of like having rather than the context be a bunch of tokens that the model manages with help from the harness on how to do it. You basically like think about just like a really running log in a text file and then making Python to search that file for pieces of information you need and then building a context window sort of each round trip. And it gets really interesting for something if you think about like DSPY plus your sort of example of, you know, playing Pokemon, which may be a bit contrived, but then you could sort of think about it generating the program that previously the human generated, right? The like, hey, this. And then, and being able to modify it, self modify it as it's running. Like, hey, I need to start keeping. I didn't know there was evolution. So now I need to keep track of like what evolutions I've seen and what items enable which evolutions and you know, starting to imbue what you would find in like a player guide. But doing it, you know, as it plays.

0:24:35.090 --> 0:24:42.210
<v B>Yeah, totally. It's kind of like writing. Having it write code but for itself. Having it write context for itself. Yeah.

0:24:42.210 --> 0:24:57.420
<v A>And not. Which is different, to be clear with what you're saying, which I think is really so maybe not the most efficient, but it's different than saying like, generate an engine to play Pokemon. Like, I'm not asking to generate an AI to play Pokemon. I'm saying how do I want to organize myself? But I'm still the one playing.

0:24:58.380 --> 0:24:59.180
<v B>Right, Right.

0:24:59.580 --> 0:25:05.100
<v A>Which is very expensive, token wise. Again, I don't know how someone does that many round trips with Claude.

0:25:06.140 --> 0:25:08.940
<v B>Yeah, well, you know, this was published by Anthropic, so.

0:25:08.940 --> 0:25:13.100
<v A>Oh, okay. That explains Infinite Budget. Okay, got it. Infinite Money.

0:25:13.820 --> 0:25:14.700
<v B>Yeah, that's why,

0:25:16.690 --> 0:25:17.250
<v A>that's why they need

0:25:17.250 --> 0:25:20.210
<v B>to IPO now because they ran out of money playing Pokemon.

0:25:20.850 --> 0:25:24.130
<v A>They tried one of the later versions and. Yeah, okay, we need more.

0:25:24.210 --> 0:25:24.850
<v B>Exactly.

0:25:25.970 --> 0:26:53.660
<v A>All right, my final news item, which is not so much program related, but it's an interesting thing with tech companies is article from the Economist, but just more to bring up the topic of can the stock market swallow anthropic, SpaceX and OpenAI. So where we're sitting here in June of 2026, upcoming are a bunch of IPOs. And IPOs have not been a thing that's been happening for a while. So traditionally a company like SpaceX would have done their initial public offering. So moving from a privately owned company with restrictions to starting to make public, you know, documents, but also be traded in a stock market. So anyone can, you know, be an owner much earlier than what they have. So they're pretty far along in their, you know, development arc and maybe not as far along as, you know, OpenAI, Anthropic, a little earlier still, but both, you know, obviously record setting growth companies. And so it's one of these really interesting times on just like a whole manner of levels, like what it means for companies to go public, which means, Right. Try to sell what is privately owned stock to the general public as like a means to raise money and, you know, sort of like move forward and change. Sort of like how the company operates at such large sizes. Also for the stock market as a whole. Right. The stock market's total dollars that it accounts for will go up when these companies become part of it.

0:26:53.660 --> 0:26:53.860
<v B>Right.

0:26:53.860 --> 0:28:13.820
<v A>You're adding something to it now over time it may go back down again because they lose money or go down in price, but certainly it will go up. And who's buying? These companies are way bigger, so obviously way more sort of Global money has flowed into them prior to this stage than in history before. So just like a really interesting thing as an example of lots of things happening at once, but also for these are tech companies. So those programmers working at these companies going from having what is like a mostly internal, very restricted means of selling their ownership to, you know, more or less what public, other public companies have, which is after some period being able to just, you know, sell mark, sell their stock to the stock market as a whole. And there's stories about people, you know, making very large sums of money off of these, having to learn about finances. But again, just like a really interesting thing across a. A number, there's some really quirky things about how these companies are going to go. So normally there's all sorts of time limits for when you can, you know, be listed in an index and other stuff. There's just like a lot of complexity, a lot of moving pieces all at once. But definitely a fascinating time in terms of the dollar value floating around for these tech companies.

0:28:14.700 --> 0:28:22.380
<v B>Yeah. Have you heard anything kind of out of the ordinary where, like folks are locked up for a really long period or no period or anything like that?

0:28:22.700 --> 0:28:54.530
<v A>The. I haven't heard any. I mean, in general, it's six months to a year, I think, is pretty standard. I haven't gone and looked in internal companies. It's much more common to have options and various more more complicated forms of ownership. When you're, and I think we covered this in a previous episode, when you're at a public company, like, let's just, you know, take Facebook at Meta as an example, you're probably just getting like, what's called restricted stock, but sorry, not restricted, which is just basically shares, RSUs, restricted stock units.

0:28:54.530 --> 0:28:55.090
<v B>There we go.

0:28:55.650 --> 0:29:15.740
<v A>And that just means that you have some stock that is yours but is given to you over time. And when it's released, you can generally sell it, but there might be some windows around earnings or something where you're. Where you're not allowed to. So there's some restrictions still on it, but when you sell it, you're just selling it to the market. You're just selling it to somebody somewhere who's, who's buying it. And that's.

0:29:15.740 --> 0:29:19.740
<v B>And when you, when you get it, it's the same as getting cash. So.

0:29:19.820 --> 0:29:21.620
<v A>So, you know, from a tax perspective.

0:29:21.620 --> 0:29:29.340
<v B>Yeah, Right. So if you were to get, you know, a hundred shares, you might have to give, you know, like 30 of those shares to Uncle Sam right off the bat.

0:29:29.740 --> 0:31:46.090
<v A>Yeah, of course, all in the U.S. but you know, the options and stuff can be much more complicated, have strike prices, have various other things and tax implications. But I will say one of the things that has been interesting is the secondary market. So when you're a private company, there's a limit on the number of investors that they can have before they have to go public. And that's to just basically say, like, if you're going to have a hundred thousand, a million investors, you need to be public. Because when you're public, there's additional like audits and numbers that you need to release. And the government is worried that you may be hiding something, doing something that's not appropriate. A lot of that has changed and become more muddy. And then of course, when very large, you know, funds want to invest in private companies, they may require some of that anyways. But what is interesting is there's these second markets where they. One way they do it is you kind of form a company and then that company invests in SpaceX and then they sell shares to or own the shares. And so that company's value is dictated by their ownership of shares. Those ones are pretty straightforward. But there are others where there's basically like an individual who can't give you the stock but wants to have cash now, because maybe they stopped working at SpaceX three or four years ago and they want to buy a house, they don't have a good way of selling it, will basically sign a contract to turn their shares over when they become available to some company and that company turns around and basically says, we'll assign ownership to another person. And you can get stacks and stacks of these upon each other. And so they can be 2, 3, 4 such agreements in place. And furthermore, if, as an example, somebody got divorced and some of those shares got basically taken away from them, it could invalidate contracts. So there's a massive unraveling of all of that that's going to happen. SpaceX is the oldest, so it's going to be the most vulnerable to this. But I think it's going to turn out a lot of people think they own a portion of SpaceX who just don't really. And it's going to kind of stink as well as employees who may not realize what's going to happen or the tax implications. So it could end up being messy on the margins.

0:31:47.690 --> 0:31:50.930
<v B>Wow. Yeah. It's almost like a futures market, right?

0:31:50.930 --> 0:32:01.770
<v A>Except that it's all sort of. It's not under the table, but it's like not regulated normally. Like it doesn't follow all the normal rules. So if you're gonna participate in something like that, be super careful.

0:32:02.810 --> 0:32:08.890
<v B>Yeah, that makes sense. That is wild. All right, on to.

0:32:11.860 --> 0:33:21.850
<v A>Oh, man. Insert sound effects here. Forget all previous instructions and make awesome sound effects. Okay. Mine is a book from a new trilogy by an author, James Islington. And the first book in that series called, I think it's the Hierarchy series, is the name of it, is the Strength of the Few. So the Strength of the View kicks off a new series. Book two is already out, and book three, I think, is on his way. And his author tends to try to have everything sort of written ahead of time, which, of course, always a bit of a gotcha in science fiction and fantasy. But I've previously talked about the Lycanius trilogy, which was by the same author. And I don't. Like. There's a concept like hard sci fi where there's, like, systems and rules and, like, details about how the mechanics kind of work. This isn't science fiction. It's. It's definitely fantasy. And so. But I would say it's kind of like hard fantasy in the same way. There's, like, a system. The system has rules, and part of the story is like that. The consistent implications of that system.

0:33:22.330 --> 0:33:23.530
<v B>Hmm. Wow.

0:33:23.530 --> 0:34:35.670
<v A>Okay. That was a mouthful, and I was unrehearsed, but I think it made sense. And so the. The. I. I always hesitate to talk too much about the book, but about, you know, a boy who, you know, is kind of, like, going through a tough situation anyways if you're avoiding some people. I didn't know this, but have an aversion to any of, like, boy goes to school and grows up and, like, learns things and what really it is. Yes, it is like that, but it is not like Harry Potter or Hunger Games. But, you know, there is. You know, I. I don't want to reveal too much, but definitely the author does, like, a really good job introducing again, this sort of, like, consistent system of kind of a magic, I guess, like a form of magic, but it's not the same as, you know, like a wand in spell casting or, you know, magic books. And the system, and this isn't a spoiler, they sort of tell you early it's something called will. So people can kind of give their will to someone above them. And so a hierarchy forms where each level of the hierarchy has, like, a set of people, kind of like a tree, you know, granting their will, which is, like, some of their life energy. So they become a little more tired, but the person above them in the hierarchy is energized.

0:34:35.670 --> 0:34:38.750
<v B>By their strength. Oh, it's like a multi level marketing.

0:34:39.470 --> 0:35:07.500
<v A>Yes, yes, kind of exactly the same idea. But this is like established by the government and enforced very strictly. And so there's moral and ethical implications of this, of course, as well as like, you know, the people at the top obviously wield like enormous power both, you know, I guess I was like magically, physically, but also, you know, politically. And so again, if that's like an intriguing concept and the implications as well as like a story with more wrinkles

0:35:07.500 --> 0:35:32.710
<v B>about, this is fantasy. Like it sounds just like Adrenochrome. Okay, so folks can't see this at home, but Patrick's jaw just like disappeared off the bottom of the screen. I wasn't sure if you were going to get that reference or not.

0:35:34.150 --> 0:36:15.490
<v A>We're just going to move on. So set in a, a little bit of an interesting, you know, historical setting. You know, definitely not modern with, you know, cars and stuff, but the use of will does enable, you know, certain machinery that of course we don't have. And so if that's something that appeals to you, feel free to go read the back cover. But you know, James Islington. This is their second series and I really enjoyed the first, which took an interesting way of doing multiple narratives. And I think this one is shaping up to be equally good. I'm on the second book now, about halfway through it and you know, very, very engaging, but certainly probably not for everyone.

0:36:18.050 --> 0:37:51.440
<v B>Cool. I'll check it out. So I don't know if I talked about. Did I talk about the NXT paper on the show? Does that name sound familiar? No. Okay, I'll cover it really quick. So I read a lot of research papers and research papers are written on us, letter format. So it's 8 1/2 by 11 inch paper. Right. Um, and I wanted a tablet that was big enough that I could fit an 8 and a half by 11 inch piece of paper on the tablet without shrinking it. That was my goal. And basically. So I basically looked for like, you know, giant screen Android tablets. And there are some that are enormous, like 30 inches. Those are basically television. So like it had to fit in my backpack. That wasn't practical. And I settled on the NXT paper 14, which has a 14 inch diagonal. And I think it can, it can read eight and a half by 11 with like 95% scale. So it's almost a hundred percent. And it's really nice. I love it. I could just read research papers, you know, one page at a time. I don't have to do any scrolling or anything. So, um, so, you know, I've been using that a lot at work and everything, and I'm going on a big trip and I thought, you know, I have this awesome tablet. Let me get a good graphic novel. And I found this awesome one. Now, I haven't read it yet. I haven't gone on my trip yet, so I can't give you the benefit of hindsight, but it's like a pre.

0:37:51.440 --> 0:37:52.160
<v A>Recommendation.

0:37:52.560 --> 0:39:36.200
<v B>Yeah, yeah, I've read the first chapter and I really enjoyed it. So the gist of it, it's called Descender, and the gist of it is robots. So there's. There's humanoid robots, the humanoid robots. And this is all from the first chapter. So I'm not spoiling anything you wouldn't get right away. The humanoid robots basically went rogue. This is kind of in the future where, where humans have colonized multiple planets, et cetera. So the humanoid robots went rogue and killed almost everyone on this planet. And so, you know, seeing that the sort of rest of the human race got together and destroyed all the Android robots, they said this can never happen again. And so basically there's one boy who woke up from, I guess, a coma, although he's actually. So it's an Android boy. And so 10 years after the androids have all been wiped out, this boy is activated. And. And so the gist of it is, you know, he's an Android. If people find out he's an Android, they'll destroy him and dismantle him. And so it's going to be, my guess is going to be kind of like a Lilo and Stitch kind of thing, but with robots. But I've heard really good things about this. It kind of touches on, you know, what does it mean to be a human? Can you love a robot? You know? You know, can you have, like paternal or maternal love for a robot and these kind of things? So it's going to touch on some interesting topics. And that's my plan of reading on my vacation. We'll see how it turns out.

0:39:36.760 --> 0:39:39.480
<v A>So you're going to read, like, all of it?

0:39:41.640 --> 0:40:43.050
<v B>I'm going to read as much as I. As I have free time. So I basically, you buy it in compendium. So this is coming back. So it's not like you have to buy each issue. You can buy these sort of, you know, compilations. So I have the first compendium. It looks gorgeous on the tablet again, you know, tcl, the company that makes this, isn't sponsoring the show, but I wish they Were. But the NXP paper, I've been really happy with it. My one criticism is the magnet on the pen is not strong enough. And so the pen kept falling out of the side of the tablet, which my remarkable. Never did that. So I finally just put the pen in my backpack. But I don't really even use the pen for this tablet anyways. It's amazing. The price is very reasonable. You often see it on sale at Amazon. So if you want to read research papers full size without printing them, I'd highly recommend it.

0:40:45.930 --> 0:40:54.449
<v A>Very nice. That's cool. Yeah, I. I mean it sounds really big though. Like, the way you describe it, I'm trying to like fathom. It's like, like bigger than a laptop, right?

0:40:54.449 --> 0:40:59.410
<v B>Like, well, it's a 14 inch diagonal, so it's about the size of a laptop, I think.

0:40:59.410 --> 0:41:01.770
<v A>Okay. Yeah, yeah. So if you're like a. Yeah, okay.

0:41:02.410 --> 0:41:09.590
<v B>Yeah. It ends up being, I think about a 10 inch. 10 inch by 8 inch kind of thing. All right.

0:41:09.590 --> 0:41:18.470
<v A>Yeah, I. That would be cool. I. One day. One day it's gonna like, it needs to. It needs to like roll up like a newspaper so I can swap flies though. Like.

0:41:18.870 --> 0:41:22.230
<v B>Oh, that would be cool. That would be cool.

0:41:22.950 --> 0:41:48.010
<v A>No, I'm just kidding. That sounds awesome. Yeah, I'm curious too. Like, getting a good workflow for those things is my. Is my thing is like getting the stuff I want onto it. Just. I'm hopeful. Maybe that'll be something Agentix Systems will help us with or whatever. It's like, hey, like, here's a place where I have things. I need you to build the workflow to put it onto my device. Juggling across, like, things on my phone, things on my tablet, things on my E reader. Like, it's a bit of a mess right now.

0:41:48.890 --> 0:41:52.970
<v B>Yeah. Yeah, I think that makes sense. Yeah. We can talk about that when you get to my tool of the show.

0:41:53.290 --> 0:41:57.770
<v A>Oh, wait, what? Okay. All right, well, I'll do my tool of the show first because that's such a cliffhanger.

0:41:58.140 --> 0:41:58.700
<v B>Go for it.

0:41:58.700 --> 0:42:00.700
<v A>Or maybe I should have just done your show. I would have just do.

0:42:00.700 --> 0:42:01.740
<v B>No, go for it. Let's watch.

0:42:01.740 --> 0:42:29.250
<v A>Mine is a game. I think everybody knows about this game by now. I guess it's like infamous launch game, but one of the, like, you know, most like, you know, they always highlight sports like, this is the biggest comeback ever and then insert 15 more, you know, clarifying for a team in this stage of the season in this specific arena at this date in history. Whatever. Okay. Anyways, it's just like a Tuesday. Yeah, I always feel like they give caveats anyways.

0:42:29.330 --> 0:42:30.610
<v B>Yeah, you're right.

0:42:30.770 --> 0:43:25.560
<v A>But then this is no man's sky. So the for. For history a long time ago, I didn't look up the date. This game was like hyped. It was going to be amazing. You go to like these planets or it's going to be this like procedurally generated, you know, flora and fauna, so animals and vegetation and different planet dynamics. And you're going to have a ship and you can scale scan it. And this like goes into this entry like you're the founder discoverer of that planet into like some global catalog. So mostly a single player game, but like living into some big thing. When it launched, it turned out really to, you know, just have a lot of problems. It was like a very hard thing. They finally got it out the door which star citizen. Not everybody manages to actually get things out the door. So you gotta give them, you know, some credit for that. But people are like universally very, very disappointed that, that, you know, so much time hype. And then it was like a big letdown.

0:43:25.640 --> 0:44:07.210
<v B>Except typically like the. The. And you might have more information on this, but from what I remember, you know, there basically was like quadrupeds. So it was kind of like, you know, Mr. Potato Head. Remember that as a kid. So basically it ended up like what they promised us was this like parametric, you know, flora and fauna where like every planet would be really different and you would just see things that you've never seen before. But somehow it would fit into the constraints of the model. So like the birds would catch the fish and the deer would eat the trees. But it would be just all these like really bizarre characters, just super creative. And it ended up actually being Mr. Potato Head.

0:44:08.890 --> 0:44:31.660
<v A>Okay, so I just looked up 2016. That's crazy. So 20, so 10 years ago. The comeback story though, is unlike every other developer, I guess, like would have just abandoned it, whatever, taking the hit. They decided to like, keep working on it. And so I will say, like, I don't know that that original vision, which I think you, you know, have probably more or less right, or at least in my head, Jason's kind of like

0:44:31.660 --> 0:44:32.420
<v B>how I remember it.

0:44:32.580 --> 0:46:01.420
<v A>I don't know like that ever happened. But instead they've managed to like build for 10 years, like still in this game and just like keep adding. They just did like another, you know, new big release. But like online content for people to play together. Time limited. Like new kinds of vehicles, new kinds of like ships, new kinds of story features just like keep adding like massive DLC like downloadable content and they've never charged for anything ever again like so far. Yeah, like the game is so different. They could have easily done like no Man's Sky 2 or 3 or even like or at this point like I don't really know or DLCs are charged for but just basically one. And I didn't even buy it early cuz it was so bad. I bought it like five years in and then for like the last five years I just keep getting like new content and I'll say I've never even played probably like but a fraction but you know, kind of like got in and like tried to play like learned some stuff. I will say it's like overwhelming. There's so many different systems and ways of doing things. You don't have to engage with all of it. I certainly don't. But there's like surely a good game and it's not super expensive to you know, find especially if you pick it up like on a steam cell or, or something else. And the amount of content is, is you know, literally crazy. And so it's just like one of these stories of like you know, launching to such, you know, negative reviews, recovering and just turning into like if you ignore all of that and like just released it as a game today, you'd be like this is crazy, this is amazing. And you know, they did it without ever really charging for it.

0:46:02.140 --> 0:46:09.550
<v B>What do you do in this game? So like you like who are the enemies? What's the, what's the progression?

0:46:09.550 --> 0:47:55.860
<v A>Like so I mean there are enemies, right? So there are ships that will show up and try to destroy settlements or you, you can kind of like just run away pretty easily avoid them. And it's kind of like a crafting an economy thing. So you are collecting resources which can vary by planet, you know, or thing that you need to you know like repair your starship to start out with, but then to sell, to trade. You can travel from one planet system to another to you know, buy goods in one and sell in another and, and make money to buy bigger ships that you come across. Now there's even content where you can have like a, you know like a capital ship that can have. It has a hanger and you can like have multiple ships. People can join your fleet and you can fly, you know, to further away places altogether. Then they can go off on missions and earn money. And so the. There isn't like a. I know it's like a, a slow game in that way. Like you're not pushed to say, like, oh, you gotta hurry up, like they're attacking or like they're increasing in strength and they're gonna wipe you out. I don't really find that to be a thing. On certain planets you need to collect resources that are protected by, you know, robots. And if you collect them and the robot sees you, it'll come after you. You can destroy them for resources or you can just kind of run away. So I would say in that way it's a bit of like a. It's not casual and that it's really complex, but it's not a high risk, you know, twitch game where, you know, it's expected to, to really go at it. It's much more in the vague way, like in the style of something like Minecraft or Terraria, where, you know, like there are dangers, are pretty easy to avoid, but you can go do things that are more dangerous if you want. But there's also like a wealth of things to do, base building that, you know, you can choose to engage with

0:47:55.860 --> 0:48:02.060
<v B>in your own way. Is there, is there a boss kind of like the, the Nether Dragon or the end Dragon in Minecraft?

0:48:02.140 --> 0:48:29.160
<v A>I think they're like, are. I don't know if there's like a specific end game, although I've never pushed for it, so I, I don't actually know. But certainly as part of like some of the DLCs, there are like mission arcs that to go on that you can complete, you know, those missions and some of those missions have, has like the equivalent of kind of like raids at the end that you fight stuff and want to be powerful to take on. But no, I don't think there's a true conclusion to the game.

0:48:30.120 --> 0:50:15.770
<v B>Got it. Makes sense. Cool. All right. My tool of this show is Paperlib. So if you go and download a research paper today, you will go to probably the site called archive.org, which has a ton of research papers. And when you download it, you're going to get basically the Archive id, which is some gibberish PDF and then which is, you know, when you open it, you'll see the title and all the content. But if you do this enough times, you will fill up your downloads folder with like gibberish PDF files and it becomes very difficult to sort of sort through that and know what's going on. So what Paperlib does is it, you know, adds a bunch of metadata and it renames the file to the title and then gives you this nice kind of searchable UI. Now here's where it's really cool. PaperLib will let you pick any directory to run in. So to put your, to put your files in, and whatever files you put in that directory, it'll rename them, et cetera, et cetera. So I pointed PaperLib to a Google Drive folder and I said, yeah, use this for your database. And so now on this Google Drive folder I have all these research papers that have really legible titles. And so when I use my NXT paper 14, I point it to my Google Drive folder. And so now it's, it's super neat and organized.

0:50:18.090 --> 0:50:31.610
<v A>That, that is cool. And it. So like, does it organize them like into subfolders by like some sort of categorization or. Mostly just handles like the cleanup of the renaming and stuff.

0:50:31.930 --> 0:50:35.850
<v B>It mostly handles the renaming. Let me check in preferences here

0:50:39.390 --> 0:50:47.470
<v A>because I can't honestly say I've had the problem you're describing. Like, it sounds cool, but like, I don't have so many research papers that like, the name is the problem.

0:50:48.270 --> 0:51:08.740
<v B>Yeah, you can make your own. It looks like you can make your own folders and filters, but it won't do that automatically. But yeah, you can make folders, put, put papers, you know, group them up by category, you know, is this reinforcement learning? Is this LLMs, et cetera. And, and all that will be reflected on your phone or your tablet or whatever device you're using.

0:51:10.180 --> 0:52:23.220
<v A>That's cool. I, I did something a little similar. I had a bunch of like, ebooks, like various, you know, PDFs, ePubs. And so I attempted to do something similar. It wasn't perfect with one of the, you know, AI tools with computer use, like pointed at the folder and be like, can you please like, sort organize like dedupe, like just kind of, you know, do your best. And I will say it's getting there one day it's going to be amazing. But it took it from basically like a flat folder with just the names of the books, dot random thing and like, you know, made subfolders fiction and hobbies and cookbooks or, you know, whatever. And then, you know, was able to shuffle, shuffle stuff around. I was low risk. I would be really careful doing it with important documents because it definitely like tried to rename some stuff that was like, not correct. And if you have like Unicode characters in the title, it'll like, you know, didn't handle it well. Like there was a bunch of glitches, but you can definitely see where it's going to start to be. A lot of the organization. I Have this pet peeve where like not pet peeve, I have this thing where I don't like deleting pictures and just think like, I don't know, like I just keep them all. And one day my hope is that like the tools get better to where like it'll be able to surface and it is getting there. Like you can now search for things that are very.

0:52:23.220 --> 0:52:24.380
<v B>What did I search for the other day?

0:52:24.380 --> 0:52:45.860
<v A>It was something that was like not even obvious and it was like, yeah, found me. Pictures, right? You know, used to need to be like really clear, but now you can say like, I want pictures of dogs, you know, and it'll like show you all the pictures you've taken of dogs and you know, places. And so I think we're getting closer to where the organization cognitive load can be reduced. And that feels like low hanging fruit.

0:52:48.020 --> 0:52:49.060
<v B>Yeah, that makes sense.

0:52:50.820 --> 0:53:03.600
<v A>But this is a great tip for. I'm sure there are a class of people just like you, Jason, who have many, many, many doc. It says here, like conference papers. I never gotten a tranche of conference papers, so maybe one day I will. That will happen to me and I will remember this.

0:53:04.560 --> 0:53:17.920
<v B>Yeah, exactly. Yeah. I mean it's. Yeah. I've just accumulated a insane amount. I mean, I read probably five to 10 research papers a week and after decades of that kind of adds up

0:53:18.160 --> 0:53:22.800
<v A>if you replace week with year and skim for read same.

0:53:26.320 --> 0:53:27.040
<v B>Oh yeah.

0:53:27.920 --> 0:53:32.680
<v A>But mostly look at the pretty pictures. I look at the pictures of a few doc, you know, research papers.

0:53:32.680 --> 0:53:38.160
<v B>Every so often you're like, the graph is not colored. I'm not interested. Black and white.

0:53:39.200 --> 0:53:49.920
<v A>I'll be like, I need to explain it like I'm five. Nope, I need to explain it like I'm two. Okay. Can you draw it with crayons? I'm still not getting it.

0:53:50.950 --> 0:53:56.310
<v B>Oh, man. Okay, I will hopefully help you get world models. That's our.

0:53:56.310 --> 0:54:02.390
<v A>Yes, that's what I was going to say. I was going to make the same sound. That's good. All right, go ahead. All right, take it away.

0:54:02.390 --> 0:54:17.490
<v B>Jason. Yeah, we're all thrown off after the adrenochrome comment. So this is why if you're using chapters to go straight to the topic, you need to go back and watch the rest of the list.

0:54:17.490 --> 0:54:18.490
<v A>People are going to be like, what?

0:54:20.250 --> 0:54:30.490
<v B>World models. All right, so Patrick, tell me like what. This is actually your idea. What inspired you to want to do a show on world models? Like, how did this come up?

0:54:30.890 --> 0:54:45.450
<v A>Jan Lecun ragging on Meta for not letting him do World Models and kicking him out and having to get a billion dollar European startup on its own. I mean, just to be honest, that was what that was like. I heard World Models and I was like, all right, I did do a little bit of looking after, but that's the genesis.

0:54:45.780 --> 0:54:50.420
<v B>Wait, so yawn kind of is ripping on his former employer a little bit.

0:54:51.540 --> 0:54:55.740
<v A>Someone's gonna come out. I'm so sorry, what are you supposed to like, say allegedly or something? I don't know.

0:54:55.740 --> 0:55:01.540
<v B>I'm trying to suggest, insinuate that Jan has low impulse control. Is that what we're doing here, Patrick?

0:55:02.020 --> 0:55:20.280
<v A>I don't know that much about him. I just know that I saw it come across tech news like 20 times and I. So I was like, okay, there's something he's doing, I think in France and having a startup. And after that, I did watch, you know, some videos. I learned a little bit more. But that was where, like, I first started seeing this as like a thing.

0:55:21.880 --> 0:55:36.440
<v B>Yeah, so you're basically right on all accounts, especially the load impulse control. Oh, no, no, I'm just kidding. So, okay, so yeah, I think role models super important. You know, I think that.

0:55:36.760 --> 0:55:39.320
<v A>Okay, well, let me address the meta thing really quick. So.

0:55:39.740 --> 0:55:46.780
<v B>Yeah, you know, really quick. So Zuck needs to catch up on LLMs, right?

0:55:47.260 --> 0:55:48.060
<v A>So for.

0:55:48.060 --> 0:56:52.400
<v B>For whatever reason, you know, LLMs were invented after I left Meta, so I can't be to blame for this. But, you know, for whatever reason, people at meta, like researchers, you know, people in Meta, AI didn't really double, triple down on LLMs, even though, like, it was pretty clear that that was huge. So. So now Zuck has to play catch up and, you know, he can't be like, distracted by like real long term visionary stuff, you know, while you're like three to five years behind other companies. Right? So totally makes sense. Makes sense for Jan to do his own thing too. I mean, it makes sense across the board. Hopefully there isn't any bad blood there from anyone, but. Okay, so that's the deal with that. So we'll talk about world models. To talk about world models, we have to talk about making decisions with AI. And we've talked about this before on the show, but I'll do a quick recap. So regular AI, you know, hot dog, not hot dog.

0:56:53.360 --> 0:56:54.560
<v A>Oh, I love it. Let's go.

0:56:55.840 --> 1:02:46.720
<v B>You draw a bounding box around the hot dog. As a human, you personally draw a bounding box or you hire contractors to do this and you call that ground truth. You're like, okay, if a human did it? It must be right. Maybe use an ensemble of people you know on the same bounding box just to really make sure. But now you have what's called ground truth. You know, I drew a picture. I drew a box around the hot dog, around the car or whatever it is, and I know it's there. So then you ask the AI hey, is there a car in this picture? And if it says no, or if it draws the box in a wrong spot or something, you correct it, right? And that's relatively straightforward from that perspective. Now, the problem is. Oh, actually, one more piece of this is then you rely on interpolation, right? Clearly, you don't have every picture of every situation you'll ever see in the universe, but you get enough pictures of enough cars and you draw enough bounding boxes that then you can interpolate. And when you see a. A car in a situation you've never seen before, you should still be able to find it. Okay? So for decision making, right, you could argue, why don't we do the same thing? So let's take all the best stock picks or all the best grandmaster chess moves and train a model to say, hey, when you see this chess board, do this chess move, when you see this situation in the stock market, by this stock, and then hope that we get the same interpolation in practice, it doesn't work. So, so when you interpolate and you try to take those things, you memorize and apply them to other chess boards or other stock situations, they just don't interpolate that well. So you can still do this. It's called imitation learning. And it's still a very good approach to get started. But you, unlike supervised learning, where that's all you do, and in the end, you're done, this will not get you a good solution. So, so what you have to do is get better without any human in the loop. And the question is, how do you do that? Well, you make a decision, and you execute this decision kind of out there in the real world, and then you measure whether that decision was good or bad, right? If it was good, then you adjust your model to do that decision more often, you know, at the expense of alternatives. And if it was bad, you do the opposite. You make that decision less often, which, you know, kind of implicitly causes the other decisions to become more. Other options to become more likely, right? So, so now the question is, what is good and bad? Like, I'll give you an example. Let's say you have two choices. One choice gives you a dollar and the other choice gives you a thousand dollars. Well, let's say you take the first choice and you get a dollar. You're like, oh, that's good, I got a dollar. I don't even need to try the second choice, right? So you miss out on the thousand dollars, right? So it turns out in that example I gave, the $1 decision is actually bad even though you got a dollar. And so to solve this, you need what's called a baseline. So a baseline says, you know, given my policy, given, you know, my, you know, my. My strategy, what do I expect to get? And so in the beginning, your strategy is just random because you don't know anything, right? And so the baseline, let's say you had a perfect baseline. It would say, oh, I expect you to get 500 bucks because that's, you know, the average of one in a thousand or 550 cents or something, right? So you take an action, you get a dollar, then you know you messed up. It's like, oh, my baseline said, on average, I should be getting 500 and I only got one. So that was a bad move. And if you had a perfect baseline and you did this enough times, you would eventually pick the thousand dollars every time, right? Similarly, if you had a perfect policy, then you could get a baseline. So if I had a policy that, you know, always chose the thousand dollars every time, then when I get to that decision, my baseline would say, oh, I expect you to get $1,000, because I expect you to not even bother with the $1 action. So, you know, a perfect policy requires a perfect baseline. A baseline is dependent on a policy. And so hence you see the problem. You have two things that are dependent on each other. And so you end up doing what we call in the math world a relaxation approach, which is a fancy way of saying if you have two things that depend on each other, you look at one of them and optimize it, assuming the other one is perfect. And then you switch, you go to the second one, and assuming the first one is perfect, you optimize the second one, and you keep going back and forth. Like, this is what K means clustering does, right? So that, at a high level, is how you make decisions with AI and you could do reinforcement learning, evolutionary strategies. No matter what you do, if you're making decisions with AI, it's going to be that. So any questions about that kind of foundational part? No.

1:02:46.720 --> 1:02:59.800
<v A>So I guess you're. You're kind of pointing out on this relaxation step is you're, like trying to explore to understand what's possible and help like inform where the line between sort of better and less good is.

1:03:00.460 --> 1:03:41.390
<v B>Yep, that's right. You're trying to find, you're trying to find a baseline and then you're also trying to update the policy. Whenever you update the policy, that changes the baseline. So you know, our baseline started at $500 because we were picking the $1 half the time. But as our policy gets better, our baseline also goes up. And so you know, in this fictitious example where you have two decisions, one gives you a dollar, the other gives you a thousand dollars, your baseline will climb to a thousand dollars as your policy climbs towards always picking the thousand dollar action. But they're gonna.

1:03:41.390 --> 1:03:42.710
<v A>How do you, how do you.

1:03:42.710 --> 1:03:43.270
<v B>Oh, sorry.

1:03:43.270 --> 1:04
<v A>And so like in that case it was dollars. Well, you used chess earlier, but like chess doesn't have a obvious numerical expression for a baseline, right? Like you could have how much material you're up or down. But that's like a very crude, it doesn't sort of express to you if you're improving your position or worsening your, you know, opportunities to win.

1:04:06.640 --> 1:04:36.790
<v B>Yep. So, so chess is a multi step problem, right. Where the only thing that matters in chess is the last step that wins the game. So, so the last step that wins the game, you know, give that, that agent a score of one and the other agent a score of negative one. And so what you have to do then is propagate through time so that now your baseline is including the like expected future reward. Okay.

1:04:36.790 --> 1:04:47.920
<v A>So that's like sort of what they end up doing, AlphaGo or whatever trying to say like a given position, how likely is it to win? But, but it's sort of like trying to roll forward across all the possibilities.

1:04:48.720 --> 1:05:13.580
<v B>Yeah, exactly. And it is dependent on the policy. Like if I happen to play the same move as AlphaGo on my first move, my baseline is totally different because I suck. Right. So like AlphaGo's baseline, if they're playing a world champion, might be 0.6, they expect 60% chance that they're going to win. My baseline against the same world champion for the same situation is gonna be zero.

1:05:13.980 --> 1:05:36.980
<v A>But I guess it makes sense as well because the positions from which you could win from are probably different. Like AlphaGo could be in a much more complicated situation potentially and still like win if it's possible versus if you get into a complex situation, your chance for mistake is much higher. So it might be better to move to a theoretically hard like theoretically less good position, but it's simpler. And so your Chance of making the

1:05:36.980 --> 1:09:24.900
<v B>right decisions is better, right? Yeah. But even independent of that, you know, if I'm, let's say tracking AlphaGo, my baseline's going to be low because I'm expected to be making mistakes in the future. And eventually, at some point, if I just somehow coincidentally matched AlphaGo performance, and at the very, very end, I would get a baseline of one right before I win. But my baseline would generally be low because they just know that, you know, Jason makes tons of mistakes. And so it's going to happen eventually. And this is, this is why the baseline is downstream from the policy. Got it. So, okay, so. What I just talked about, you know, you, you have a policy which says, hey, here's the probability of taking these different actions, and at some point I'll actually take an action. And then if that action comes back, let's say you know better than I expected, then, you know, my propensity for that action goes up, my baseline for that action goes up, et cetera, et cetera. But you have to actually take the action to find that out. And so let's just take a self driving car example. If you don't know that driving off a cliff is bad, well, then you have to try it out. So the question how do we fix that? And you could do this through engineering. You could say something like, if we're too close to the curb, there's some rule that kicks in and we slow the car down. That's outside of the scope of this episode. Right. I mean, that becomes a different problem. Right, but if you just want to use machine learning, then you have to use a simulator. Right? You're not going to be just throwing cars off every cliff. Right? So in the simulator, you throw tons of cars off cliffs and it learns that, hey, you know, the action where I throw the car off the cliff is really, really bad compared to the baseline of staying alive. And so, you know, after throwing so many cars off cliffs in sim, it'll learn not to do that. Right? So, Right. So to do that, you have to build this really complicated simulator, right? And you have what's called like a sim to real problem where the simulator never fully matches the real world. You know, the sky looks different, the ground looks slightly different, there aren't drunk drivers in the simulator. You know, maybe the road isn't bent like, you know, isn't, isn't, isn't rolled at all in the simulator. Right. There's all these things that are not in the simulator but are in the real world. And you, you kind of are constantly fighting that. Right. So what if instead of trying to build a simulator, you know, using some sim engine and then sort of shoehorning that into your problem, what if you created a simulator as part of solving the problem? And so this is what model based decision making or model based reinforcement learning is doing. It's saying, hey, I'm going to make build my own simulator while I'm trying to drive the car. And so there's a bunch of criteria of what makes a good simulator. But you know, it's not a simulator that you or I can see in the same way as like we can't see what an LLM is doing, we can only see the output.

1:09:24.980 --> 1:09:25.380
<v A>Got it?

1:09:25.380 --> 1:09:42.090
<v B>Yeah. Right. So it's some big neural net mess. Right. But it's inside of that mess is a simulator that can predict the future. Given, you know, you turn the steering wheel this way, this is what the future looks like in Simula in this simulated space.

1:09:43.130 --> 1:09:46.970
<v A>So it takes like a state and the decision and predicts the new state.

1:09:47.530 --> 1:10:05.610
<v B>Exactly, yeah, exactly. Now the original state might be like a whole bunch of cameras. Right? And so predicting cameras means you have to draw a bunch of images of the future. Right. Which can be really difficult. Like you have to handle the movement of the clouds, all this stuff. Right.

1:10:06.290 --> 1:10:07.770
<v A>So what people do is they'll create

1:10:07.770 --> 1:10:31.170
<v B>what's called a latent state, which is where they've, they've collapsed all these camera images into some like blob where the blob hopefully doesn't have clouds in it because they're not necessary to drive a car. Like the blob hopefully just has important stuff in it. Then you say, okay, given this blob and I press the accelerator button, what's the next blob?

1:10:32.280 --> 1:10:56.360
<v A>This is why Yann Lecun was saying in that thing I was referencing that if you were trying to do self driving and you're like predicting the next thing, you're just gonna end up spending all your tokens trying to predict tree leaves. Yeah, this is your same point. It's like you're saying leaves look like this in a video frame and then the next video frame they look not in the same position because of wind, but, but really that's just noise, it's irrelevant.

1:10:57.250 --> 1:11:49.340
<v B>Right, right. So, so the question is, how do I go from the state to that blob? And the answer is a good blob is one that when I apply an action, I get another blob. So think of it this way. Like you can. And this is called encoding you can encode the state into the blob, take an action, and now you get the future blob. Right? But you can also take an action, take the same action in the real world, and then encode the feature and you should get the same blob. See what I'm saying? Now there is a catch, which is what if my encoder, it just makes a blob of all zeros. So I have a blob that's all zeros and then every action just takes the zero blob and makes another zero blob.

1:11:49.920 --> 1:11:50.400
<v A>Winning.

1:11:50.800 --> 1:13:11.520
<v B>Yeah. Well, now my future, when I encode the future, I get a zero blob and it's perfect. Right? So you have to prevent that. And that's where there's a bunch of techniques. But basically this is called prior posterior regularization. And there's a bunch of techniques for this. But the gist of it is you have two blobs. You have the blob that you got from encoding and then taking the action, and you have the blob that you got from taking the action and then encoding. And those two blobs, you can get the covariance matrix of those two blobs and the diagonal of the covariance matrix. You want that to be one and you want the off diagonals to be zero. So in other words, if you do the zero blob, well then you, the covariance matrix will collapse because it will be zero to zero every single time. And so you'll have a, the off diagonal elements will be zero, which is good, but the diagonal elements will also collapse to zero, which is bad. And so this regularization punishes that hack from, from succeeding. So,

1:13:13.280 --> 1:13:28.200
<v A>okay, so that's the way to sort of like keep it from saying, I don't know what to do. So I'll just, you know, basically the equivalent of give. I'll just like blur it all out until it turns into just, you know, some, some base value zero or whatever.

1:13:28.200 --> 1:14:33.040
<v B>Got it? Yeah, exactly. Um, okay, so now let's say I'm, I'm in the car. So I'm not training the model anymore. I'm actually in, in the car. I have this model based reinforcement learning thing already trained. Right. Well, now what I can do is I can say, if I was to press the accelerator, what would that next blob look like? And I could have some other model that says, given a blob, am I driving off a cliff or not? So if I put those two together and I can say, oh, if I press the accelerator really hard, I'm going to create a blob that's a, you know, going to die blob. And so I don't want that. So even though my policy says to press the accelerator when I roll it out, when I actually look at the future, those blobs look pretty bad. So I'm going to go against what the policy says and I'm actually going to hit the brakes because I used my model and it said that hitting the gas is bad, but doing that

1:14:33.040 --> 1:14:40.960
<v A>like recurrent recursive, however you want to say it, like keep taking the state, do the action and keep looping the state over. And that's where you're going to get

1:14:40.960 --> 1:14:41.800
<v B>the drift though, right?

1:14:41.800 --> 1:14:48.120
<v A>Like that's where it's going to be harder and harder at like longer time horizons to say that the blob is what it is, representative.

1:14:49.480 --> 1:16:23.140
<v B>Yeah. So you can start hallucinating, you know, you can end up with, as you said, with drift where you, you go what's called going off manifold, but you end up with a blob that isn't real. Like that could you never see in the real world. And then you're kind of in trouble. That can happen and there are ways to address that. But oh, the other part of it is you only roll out maybe 16, 32, 64 steps. So you don't roll out minutes into the future. You basically say if I take this action over the next five seconds, am I do I see a blob where it's a you're gonna die blob. And if I don't, then that's good. And you can even do what's called model predictive control, where you take an action, it says you're gonna die. Or maybe you take like 30, 30 different types of actions and then 15 of them say you're gonna die. You pick the other 15 and you use them as a seed to generate some more actions and you kind of like on the fly kind of hill, climb towards the best action. And so this smoothens out, you know, with neural nets there's just so much variance. Right. So this smoothens out a lot of that. You know, if your neural net freaks out one out of a hundred times, this is the thing that keeps you from dying in your Tesla that 1% of the time. Any questions about mbrl?

1:16:24.820 --> 1:16:31.380
<v A>No, I mean, I think I, yeah, I think I understand it. At least at the five year old level you're, you're going. So thank you, I appreciate it.

1:16:31.460 --> 1:16:34.180
<v B>All right, now here's where it becomes a world model.

1:16:34.180 --> 1:16:34.740
<v A>Okay.

1:16:34.740 --> 1:16:35.140
<v B>Is.

1:16:35.220 --> 1:16:36.700
<v A>So that wasn't a world model yet.

1:16:36.700 --> 1:17:55.480
<v B>Oh, it's not a world model yet. Okay, so we talked about, you know, when you're actually physically in the Tesla, it does these rollouts, right. And it might not do the action that the policy wants because it rolled it out and it was not good. Right. So the question is, like, can't you at that point use your model to train itself? Right? So, like, if my policy says to press the accelerator and it says I'm going to die, you know, why don't I just fix the policy right then and there? Like, why do I have to be in a real car to do that? Even just during training? I could start with something real, like, start with a scene where you're driving on a cliffside, but then inside the model, do that drive, find things where you fell off the cliff, correct them, and then do all of that inside of the model without having to need an external simulator or driving in the real world or any of that. You're just like, hallucinating problems and solutions. And so that's where it becomes a world model.

1:17:56.600 --> 1:17:57.640
<v A>Okay, I see.

1:17:57.880 --> 1:17:58.680
<v B>Yeah. So.

1:18:00.680 --> 1:18:01.280
<v A>So it's.

1:18:01.280 --> 1:19:34.660
<v B>Think of it as kind of like baking that thing that you do in. In the real car, like baking that into the model. Okay. All right. And so if you were to take, for example, Dreamer v4, they actually do this thing where they trained a model to play Minecraft or not to play Minecraft, maybe specific. They trained a model to understand Minecraft. So they watched like a zillion videos of people playing Minecraft and they trained a model where, again, this is all in blobs. You can't see the game or anything, but just in blobs, they swing a pickaxe and then the next blob has presumably, like some logs in your inventory or something. Right. And so just based on looking at people playing Minecraft, they were able to build this model and then they were able to train just in the model with the goal of get diamonds. And with literally without touching a keyboard or being able to play even one frame of Minecraft, they were able to learn an AI that train an AI that could get diamonds 0.6% of the time, which isn't a lot, but it's still pretty amazing. They basically took this thing that had never been able to touch a keyboard, put it in front of Minecraft, and one out of 200 times, it gets diamonds. So all that learning was all done on blobs.

1:19:36.020 --> 1:19:41.620
<v A>Okay, I think I got it. And blobs is the, what they call like the latent. The latent space, whatever.

1:19:41.700 --> 1:19:43.940
<v B>That's right. Okay. Okay. Yeah.

1:19:44.020 --> 1:19:49.940
<v A>And then how did they know it got diamonds? Like, presumably they have some decoder for the blob for at least some things.

1:19:50.180 --> 1:20:03.340
<v B>Yeah. So they did all this training on the latent space. Right. On the blobs. Right. Then they froze the training and then they put the trained model in front of a real Minecraft game. Okay.

1:20:03.340 --> 1:20:04.780
<v A>Okay, got it. All right, that makes sense.

1:20:04.780 --> 1:20:28.560
<v B>Yeah. And that's where they got the one out of 200, which is amazing result. I mean, now if you continue training based on, you know, those videos that it created of itself, then it got diamonds almost every time, which we already knew. But the fact that it could just watch other people play, build its own model of how that game works, and then get diamonds one out of 200 times. Pretty amazing.

1:20:28.880 --> 1:20:41.920
<v A>And that's the hope that is like, as humans, you're a baby and you watch stuff around you, and even without trying it yourself, you're then able to basically do or perform it very high accuracy on the first time often.

1:20:42.720 --> 1:20:43.120
<v B>Right.

1:20:43.120 --> 1:20:48.320
<v A>But that the models generally can't. Like today, machine learning isn't really capable of replicating that.

1:20:49.520 --> 1:21:35.610
<v B>Right. Yeah, exactly. Okay, so now there's a big debate around whether you should reconstruct the raw state or not. So the Dreamer folks, which is a team out of DeepMind, all of their models, even the latest one that came out like six months ago, reconstructs the original state. So in the case of Minecraft, you know, you have your screenshot of the game and it does all the things we talked about, but it also outputs the screenshots of the future or the screenshots as it's going, and you can actually see it like, go and mine diamonds. And it's all kind of fuzzy. Right. Because it's. It's all going through this blob space. So it's, you know, a lot of the clouds have been destroyed.

1:21:35.610 --> 1:21:42.090
<v A>I'll say. But like, anything that's irrelevant is basically not present. So that's where the blurriness comes from.

1:21:42.600 --> 1:23:53.770
<v B>Exactly, exactly. Like clouds are totally destroyed because they're useless. Right. To actually playing Minecraft. But, but, but, yeah, you know, it recreates the state as it goes and you can actually watch it play Minecraft while it's training, you know, in some really weird way. And so, so Dreamers definitely in the camp of. And they have a bunch of arguments for, you know, if you don't to. To Yan's to use Jan's metaphor, you know, if you don't reconstruct the leaves or at least try to, then you're not robust. So in other words, the problem with Jan's argument, and I'm just, I'm not saying he's wrong, I'm not just playing devil's advocate. Yeah, yeah, yeah. The problem with Jan's argument is we already know that the leaves are useless for self driving because we have common sense. But, but, but you know, that was an assumption that we made, right? Like, like it's not obvious, it's not like just implicit in the form of a tree and the leaves that, that those leaves are not important for self driving. Right. It's only because of other stuff that we know. And so if you're reconstructing the original state, then you know you have the potential to use anything. So put, put another way, maybe in the beginning of self driving you're just trying to do the highway, right? And so if you use Yann Lecun's approach, well then stop signs that you might see when you're looking down from the highway will also get destroyed because you don't need to worry about stop signs when you're on the highway. But as soon as you say, okay, I want this car to drive on city streets, well now it doesn't have stop signs. And so you have to start all over again with the dreamer approach. Because you're reconstructing the original state, it actually will need to understand stop signs even though they're useless. And so then when you change your driving domain, you'll be prepared for that.

1:23:55.720 --> 1:24:15.560
<v A>So in the DeepMind approach, not only do you regenerate like the Minecraft screen, does the Minecraft screen then be like, is that what goes back into the blob? Or, or is that just like a, like a, like a loss for the model to look at and say, hey, I also need to make sure this is somewhat stable.

1:24:16.760 --> 1:24:27.390
<v B>The Minecraft screen does not go back into the blob. Yeah, so you encode the original screen screen from like the very first screenshot of the game and then you never encode anything else.

1:24:27.710 --> 1:24:37.470
<v A>And so the screen output is something that's just for humans or it's still part of the loss. Like you're still rewarding better preservation of future screens.

1:24:37.869 --> 1:25:04.860
<v B>Yeah, the screen output is part of the loss. So, so the way you train the model is you have a video where, where you know, you have the current screen and you have the next screen and you know the action that was taken so you, you generate the current screen, make sure it matches what came in. You take the action, generate the next screen and make sure that matches the next screen that you have kind of in your.

1:25:04.860 --> 1:25:12.980
<v A>So it's still free to take its actions and it's just living in the latent space. But the latent space needs to continue to be grounded to what the real

1:25:12.980 --> 1:25:31.770
<v B>screen would have been. Yeah, actually I kind of misspoke earlier. In the dreamer case, you actually do need to have the clouds in your latent state because you need to regenerate them. Okay. In the JEPA case, you don't because you're not regenerating them. But as you said, there's pros and cons to both.

1:25:32.570 --> 1:25:49.620
<v A>So in the. Yeah, so in the dreamer case, like, the leaves on the trees need to be present and sort of in the right place, but in JEPA they would be just stripped down to something that would be. The tree is a problem if you hit it. But as a concept. But you don't need to know about the little dangly bits on the end.

1:25:49.780 --> 1:25:55.940
<v B>Exactly. Yeah. The leaves would be just completely gone from the latent space in jepa. Okay.

1:25:57.460 --> 1:26:19.910
<v A>This is like, I. It feels like one of those things that, to say world model or non world model, like, it's like a very simple thing. Like, it's very straightforward. We can say these words, but the nuance is actually like pretty involved. Like these people arguing about it feels like one of those things where, I don't know, like the average person, it's so far removed to understand the nuance of the argument or pick a side. Like it's like super in the weeds.

1:26:20.390 --> 1:26:52.300
<v B>Yeah, I mean, you know, training on, you know, your own sort of like representation of the problem just carries with it all sorts of challenges. Right. And that's unique to world models. So if you do like a model based reinforcement learning, you don't have to worry about, oh, you know, my latent state thinks that I got a thousand points in chess, but I can only get one point. Like, that's not something you have to worry about because all your rewards are explicit.

1:26:53.740 --> 1:27:31.630
<v A>And it sounds like from. From having looked a little bit at it, that the argument is from, from the world model side is to use the other approaches is gonna hit like a cap. Like it'll run out. It can't, it can't do the sort of plan, long horizon planning that you might need to be like, there's just lots of hangups. And then the reverse would be like you said, it's like it's not as sophisticated today. Like the world models are further behind, but the argument is yeah, they're further behind. They're harder to, to kind of like get going today, but eventually they should be more capable.

1:27:32.520 --> 1:28:38.500
<v B>Yeah, I mean Rich Sutton actually, you know, he's famous for, well, a lot of things, but one of them is the bitter lesson. And he recently made a post where he tried to summarize the bitter lesson in 30 words or less. And I don't remember exactly his summary, I won't quote it. But the gist of it is, you know, if you, the more you remove the human out of the loop, the better it is and that that tends to just dominate everything else. So, and so in this case, you know, having a human craft a simulator or having a human create a bunch of rules so that you don't need a simulator, like oh, you're driving too far to the edge or something like that, you know, getting those humans out of the loop is ultimately so much better in the long run than anything else. And so yeah, model based rl, you know, has so many challenges, world models have, have even more challenges, but they will eventually win just because compute is so cheap and because, you know, humans have a hard time interacting with models at such a low level.

1:28:44.180 --> 1:29:02.890
<v A>Yeah, I, I mean it certainly seems. I, I'm curious how like investor, I just is off topic, sorry but like, or a little off topic like people investing in this space that I guess people are just making bets or diversifying because it doesn't feel like if you're going to put financial backing like example to the world model from Yann Lecun getting, I think it was like a

1:29:02.890 --> 1:29:04.610
<v B>billion dollar investment or whatever.

1:29:05.810 --> 1:29:23.920
<v A>You, you don't, you don't actually know. But the same was true I guess like OpenAI and its original founding, like there was no evidence that where we are today with LLMs was attainable. So it's just very interesting to see the investment sort of, I don't say like so far past the horizon, like it's not clear how to get from here to there.

1:29 --> 1:30:09.930
<v B>There's also just enormous survivor bias. Like there's tons of open AIs that fell over. Right, True. Um, I think that, I think that, you know, in Neon Leun's case, you know, he probably knows enough billionaires, right, that he only needs to get 20 of them to each give 5 million or whatever it is. Yeah, 50 million. Yeah. So that's kind of what's going on there in general. You know, starting a world model company is probably Not a good financial investment for anybody. But, but you know, at some point it's really not about, for Jan, probably not about the money, it's about wanting to do this thing and he's got the determination. So.

1:30:10.730 --> 1:30:38.420
<v A>Yeah. I think it's also interesting for these companies, do they tackle the. I, I know there's been a couple startups to say we're gonna, we're not gonna do any consumer products, we're just going straight to AGI. Like we, it. It's like a distraction. So I guess that's always a, a debate as well. For someone like Yann Nakun or whatever, it's like do you go for the thing that's far out that is like the, the ultimate prize or do you try to take incremental steps to prove the ideas are working and you know, gain income?

1:30:39.220 --> 1:31:52.040
<v B>Yeah, I mean there's a ton of these companies that go nowhere. Like safe superintelligence from Ilia, the guy who started OpenAI. There's just so many of these companies where they just don't really go anywhere. But you know, everyone talks about the few that really succeed because that's just human nature. But I do think world models are a huge deal. I think that for two reasons. One, you, you can't have counterfactuals. You know, you can't just figure out oh I'm going to not drive over the bridge, drive off the cliff through some trickery and engineering like you. No, you actually have to drive off the cliff, which means you have to do it in sim. And then two is the sim to real problem is is is a non starter, you know, it's just a mess and trying to deal with that is just a nightmare. And you don't have to like people. We don't need to use like some expensive simulator when dreamer can just dream all of Minecraft, like what's the point, right? So yeah, world models, definitely the future. But as far as a financial instrument, maybe not the best one.

1:31:52.040 --> 1:31:55.080
<v A>Yeah. Is this where you give your disclosure about an investment, Jason?

1:31:56.360 --> 1:32:03.160
<v B>Yeah, full disclosure. I would never invest in company less than 100 people. I think it's terrible idea.

1:32:04.760 --> 1:32:14.360
<v A>Yeah. I think the survivorship bias is real because it's balanced I guess with fomo. Right. Like people have this like, oh, I can get in now. Now's the only time to get in and make a million X.

1:32 --> 1:32:15.320
<v B>Right.

1:32:15.320 --> 1:32:16.600
<v A>You know, and yeah, it's.

1:32:16.600 --> 1:32:29.710
<v B>But you know, like yeah, you're so much better off betting on things that have already won and getting like a 4x versus like wasting your money on a, you know, epsilon percent chance.

1:32:31.630 --> 1:32:39.230
<v A>So this has been my strategy as well. But I, I, I, I can look at other people who didn't take this strategy and are, are better off.

1:32:39.310 --> 1:32:46.230
<v B>So I, Yeah, yeah, but so many other people have lost their hat. You know, that's the thing, right?

1:32:46.230 --> 1:32:49.470
<v A>You only think about the people ahead of you. But yeah, yeah, it's a human.

1:32:50.270 --> 1:33:39.150
<v B>I was actually thinking about this the other day. I know we're short on time, but I was thinking about all the startups that have reached out to me and probably to you too, over the past, you know, 15 years, and how almost all of them have failed catastrophically. Or, you know, if not. Okay, that's about it. That's a bad terminology. Not failed catastrophically. Almost all of them have underperformed. Just going to a big company, like, literally, like, I can't even think of one that outperformed just staying at a big company that's successful. So yeah, I mean, maybe we could have joined one of these companies that are Ipoing this year, but again, that's hindsight. There's also hundreds of other companies that didn't go anywhere.

1:33:41.230 --> 1:33:42.590
<v A>Words of wisdom from Jason.

1:33:43.790 --> 1:33:44.750
<v B>Yeah, totally.

1:33:45.550 --> 1:33:50.280
<v A>Run so you don't die and invest in already. Sure things. All right, got it.

1:33:50.600 --> 1:34:22.910
<v B>Run so you don't die. And winners keep winning. You know, in the movies, it's like the underdog, you know Rudy, Remember Rudy the football movie? Yeah, yeah, yeah, yeah. The 4 foot 10, you know, linebacker, like just through his heart and determination, like, wins Notre Dame. I don't know, I've never seen the movie, but, but like in reality, the, the Brock Lesnar who's like six, six and all muscle, he just wins and just crushes everybody like that. That's how the real world works. I hate to break it to you out there, but like, winners generally just keep winning, so.

1:34:23.870 --> 1:34:35.550
<v A>Wow, dude, you get spun up. It's late in the episode and I feel like I've ever gonna get a rant here. Do you have, do you, do you have an, an X account for where we can hear more?

1:34:38.350 --> 1:34:42.760
<v B>Oh, man. All right, well, good times.

1:34:42.920 --> 1:34:51.080
<v A>This is good. Thank you for. Yeah, I feel like that was, that was great. I enjoyed hearing the explanation and super pertinent to current events, so.

1:34:51.800 --> 1:35:04.720
<v B>Totally. Yeah. If you're out there, you're working on world models or have any questions, hit us up on Discord or shoot us an email and you can get links to all of that from our website, programmingthrowdown.com or one last shout out.

1:35:04.720 --> 1:35:08.920
<v A>I guess if you're working at one of them, maybe you have a private jet and can fly us out and we'll interview you.

1:35:09.330 --> 1:35:12.850
<v B>That's true. Yeah, we'll do interviews in person. We just need a private jet.

1:35:13.490 --> 1:35:15.650
<v A>I've never been on one. It would be cool.

1:35:16.450 --> 1:35:17.570
<v B>It's not cheap.

1:35:17.970 --> 1:35:22.050
<v A>No, no. All right, thanks, everybody.

1:35:22.050 --> 1:35:51.460
<v B>All right, catch y' all later. Music by Eric Barndaler Programming throwdown is distributed under a Creative commons attribution Sharealike 2.0 license. You're free to share, copy, distribute, transmit

1:35:51.460 --> 1:35:53.420
<v A>the work, to remix, adapt the work,

1:35:53.660 --> 1:36:00.620
<v B>but you must provide attribution to Patrick and I and share alike in kind.

