WEBVTT

0:00:14.999 --> 0:00:20.618
<v A>Programming Throwdown Episode 187: Agentic Coding. Take it away, Jason. Hey.

0:00:21.664 --> 0:01:32.387
<v B>Everybody okay? So, quick intro topic here. I was talking to a lady—I won't use her name; let's call her Jane. I was talking to Jane, friend of mine, and she has been at her company, which is kind of like a hardware company, so it's not really Jane's specialty. It's a hardware company that needs some software, right? And you know, Jane's unhappy there, and there are all sorts of reasons why Jane's unhappy. It's understandable. And she's at nine months, and so they have a one-year cliff, and so she has to kind of tough it out for three more months. And it just made me think about vesting schedules because I think that companies, you know, haven't really thought this through. There's like a cargo cult mentality where, you know, oh, this company did [have] a this type of vesting schedule, so we're just going to copy them. And I think that some bad behavior has been copied, and when you start breaking it down, it never really makes sense. And so I'll start.

0:01:33.029 --> 0:01:37.770
<v A>By Chris criticizing the one-year cliff? You want to explain what it is?

0:01:38.024 --> 0:02:10.947
<v B>Though not everybody. Oh yeah, okay. So, if you're in—let's let's like wind it back. If you're in, let's say sales, you get a commission, right? So, you build a portfolio of companies that want, you know, some product, and you get a commission. And if you're doing some service-based sales, you could potentially get a commission every single year. You could have a customer that is going to use your service for the next 20 years, but because you were the salesperson that got the deal, you will just keep getting commission for 20 years potentially.

0:02:12.465 --> 0:03:43.540
<v B>That's how sales works, right? And so for engineering and other fields, you know, they want to have an incentive structure, but you know it doesn't make sense for us to do commission because we're not that close to the customer. So what they'll do in our case is they will give equity. So they'll say if it's a private company or if it's—you know, yeah, not a public company—they might give you a percentage of the company. You might get options; you might get RSUs or shares if it's a public company. You're going to get shares. So and the way this becomes an incentive is they price the shares based on your start date. So for example, if a company's stock is trading for $80 a share, they're going to tell you, 'Okay, for the next four years, you're going to get this much number of shares,' you know, every month or every three months, or there's some schedule, right? But let's say the price jumps up to $160 a share. Let's say the price doubles. You're still getting that number of shares, and so you can end up in a position where if the company does well, even if you—you know, don't get more equity later on, just that one grant could double in how much you're getting every month just because the company did well. And so you're kind of it aligns your incentives with the company's incentives.

0:03:46.071 --> 0:04:35.430
<v B>Okay, so that's—that's equity. Now, you know, they're not going to give you four years of equity right on day one, right? Because you would just leave, right, like any rational person. If they said, 'Here's four years of, you know, upfront money and there's no strings attached,' well, then yeah, you would just go to every single company collecting four years and work for a month or something, right? So obviously, so they're going to space it out. Makes a ton of sense. Now where I think a lot of companies go wrong is they have this one-year cliff, which means that for the first 12 months you won't get any equity, and then after 12 months they will give you the past 12 months worth of equity, right? Then there, and then you'll start accruing, you know, monthly or quarterly or something like that.

0:04:36.494 --> 0:04:38.974
<v A>I think this came about if I mean.

0:04:38.974 --> 0:05:14.024
<v B>I haven't done the sort of historical work on this, but I think this came about back when, you know, a lot of these companies were very small. And so each person, you know, had to be, you know, a line item on your capital table. So when you're—when you're a small company, you have what's called a cap table, and it lists out, you know, who are your investors? Basically all the people who own parts of your company. And so, you know, I guess companies wanted to sort of try people out without polluting their cap table with all these names of all these people who, you know.

0:05:14.024 --> 0:05:16.774
<v A>After a couple of months it wasn't a good fit.

0:05:17.415 --> 0:05:41.800
<v B>That might be the inspiration. Maybe you could argue that it, you know, saves a bit of money because people who leave before a year, you know, you don't have to pay them as much, right? So that's the idea behind it. In practice, it's kind of like, you know, that joke was like communism—like works in theory but it's never do you have the saying? Have you heard the...

0:05:42.492 --> 0:05:44.213
<v A>Same thing. It's never been applied, right? It's like something.

0:05:45.495 --> 0:06:43.800
<v B>Yeah, yeah, yeah. There's some saying—yeah, we're not we're not good at politics here, but there's some kind of saying where it's like, 'Oh, communism works,' that just wasn't true communism or something. So it's all these things were like the cliff never actually is implemented this way. Every single company I've worked at, you know, there have been people who for one reason or another had to go within a year. Every single time we prorated their equity, I have never in my entire career seen a time where somebody—let's say worked six months. After six months, we decided they weren't a good fit, or there's a reduction in force, or what have you, and we didn't give that person the six months of equity. So it's kind of like it's kind of like a stick, but it's never really used. And then on the flip side, you have cases like Jane here where, you know, she has to kind of stick around for three more months. So that's why I kind of rant on this. What's your take on one?

0:06:46.330 --> 0:07:29.024
<v A>Of your cliffs? I mean, I guess I've worked a lot and I've never seen the fact that like oh, in the first year you're gonna like somehow learn something about the person or like I mean, I've never seen people just be like, 'Oh, this isn't working. Let's let them go before the first year,' as like uh that way we don't have to pay them, which I guess would be like the fear as an employee, but I've just in general never seen someone let go in under a year like because you kind of give them a period of ramp up, whatever. And by the time you kind of get through that and you're like, 'Oh, they're really not sticking around,' the process to sort of let them go at most big companies is long enough that they'll get past a year easily. So if they decide to stay, they'll make it. Um, and so that's a good.

0:07:29.024 --> 0:07:29.648
<v B>Point, hearing that.

0:07:30.002 --> 0:07:32.044
<v A>Between your pro rating story? It's just I think.

0:07:34.204 --> 0:09:33.880
<v A>Yeah, like you said. I feel like it's just a stamped-out template and it's probably been, you know, legally tested enough that they're hesitant to, you know, do something different. It also by having the longer, you know, like the four-year vesting schedule rather than just giving something for one year and then just re-giving it each year, you know, as like over the next 12 months, I think is for hope that the stock will grow and sort of give you—it's not like a golden parachute, but sometimes people refer to it as like handcuffs or golden handcuffs. Is that the hope is the price goes up and therefore you you don't want to leave because you wouldn't get an equivalent pay package somewhere else? Um, but I will say just very recently a lot of Software as a Service companies had a really huge stock decline as like AI tools, like we're going to talk about today, are rolling out. Um, because people are nervous about future earnings and so there's a lot of employees who are looking to basically bail or leave because huge portion of their compensation—I mean, a lot of places it can exceed 50 percent, it's easily 30 percent—like it's a very big portion of your salary. And people at these public companies just treat it as cash. So right? The fact that it's not in fact like there's a very happy to get into the, you know, finance game theory of it, but maybe for another time. Um, even the fact that they're issued as RSUs is really an accounting trick to basically say that uh basically Wall Street agrees to treat that cost different than a salary cost and so therefore it looks a little bit better to have RSUs as a separate line item. Um, and so if all those things you got rid of it's just another sort of quirky way of paying people. Most companies already have a stock purchase program right where you can either you're encouraged or get a discount to buy stock um with a portion of your salary. So there's various ways they already get sort of joint ownership uh with the company. So yeah, I really think it's sort of time for a revamp.

0:09:33.880 --> 0:09:58.772
<v A>But it's possible to happen with these disruptions where if some companies are really like if people are really worried about what's going to happen then shifting some compensation from, you know, stock-based to cash-based for public companies. That's the point of being public, right? Uh is that these things are like mark-to-market? That's much harder for a startup that either doesn't have the cash or doesn't have a reliable market value. Yeah. Yeah.

0:10:00.375 --> 0:10:04.360
<v B>Totally right. And then the other uh thing which like kind of

0:10:06.040 --> 0:10:41.640
<v B>Got popular like a certain niche of companies um is this like 10-20-30-40 investing schedule? Um so this uh I've never worked at a place that has this. Um Amazon is definitely the most famous, I think. Snapchat had this for a while, although I think they abandoned it. But the idea was that you know you would get you would get 10—so if your if your grant is let's say 100 shares to make it simple, you get 10 of those shares your first year, uh you get 20 the second year, 30 the next year and 40 the final year.

0:10:43.407 --> 0:11:15.267
<v B>Um, and so it would kind of backload the the equity. Um I never understood that. I always felt like shouldn't it be the opposite because the first year you know you don't have any refreshers or anything? And so it never really made sense to me that that's one that just boggles me. I really don't know why uh they do that. I mean, people have said cynically, 'Oh, it's because I think the average tenure at Amazon is very low.' So the

0:11:15.267 --> 0:11:16.026
<v A>Average tenure? Okay.

0:11:17.039 --> 0:12:17.480
<v B>Kind of side right here. But like people say the average tenure at Bank is is four years, but it's actually much longer than that. That's only if you measure all the people who have left, right? You can't really do it that way. Um you have to assign a tenure to the people who haven't left yet. You know, if you give them a tenure of infinity, well now the overall the average tenure is also infinity, right? If you give it zero, well then if you don't count it, well then you get four years, but neither of those are really appropriate. Um so the tenure is pretty long, you know. Um at Amazon people said it was I think 18 months or something. Again, it's probably longer than that, but it's definitely a lot less than the other companies. But I feel like that's probably not the reason they did the 10-20-30-40. It would be because they plan on getting rid of people after two years, but I think it's kind of what you said Patrick that you know they just came up with something and then it's just too hard to change.

0:12:19.797 --> 0:14:18.614
<v A>I do have a separate thing we should move on, but the it's like reverse survivorship bias. So there's this thing where people first start thinking about you know writing code to backtest stock trading. So they take the S&P 500, you know top 500 companies, and you look back in time. But what you are missing out by doing that is like the stock index changes over time. So if you take companies that exist today, you're not looking at all the companies that have failed. So if you want to go 10 years ago and play forward, you need to start with the set of companies that were in existence 10 years ago, not the set of companies in existence today because when you go backwards in time, you're going to lose companies that haven't yet started. But that's fine, but you're also going to lose companies that failed, which is hugely impactful. And so they call that survivorship bias. The same thing as planes which returned from World War II and analyzing where they had bullet holes misses the fact of like that's not where you should, you know, patch up the bullet holes. There's like a famous meme about this, right? Because you have the ones that went down, right? Like those are the ones. Um But these come up like you were you were sort of mentioning in in tenure computation which I didn't actually realize. So unless your company is like super old to where people retired and everything, you don't like sort of reach steady state. Um And then you then you can't account for for sort of flux. The other one I saw it is someone was talking about um you know we've talked previously about starting to to run and learning to run or whatever, uh you know maybe not when we were younger, and people are talking about uh it's just standard of qualifying for the Boston Marathon where you have to run pretty fast in a marathon, and they were talking about the number of years it takes to call like train before you qualify for the Boston Marathon, but they could only take people who who had qualified for the Boston Marathon. Um And saying how long they trained for? But like me, I've run for a few years. One, I'm not anywhere close to qualifying, and even if it was my goal to qualify, I don't know that I could. Um It just takes a lot of work to get there, and I'm not particularly naturally talented for it, so I would never show up in the numbers even if I

0:14:18.951 --> 0:14:38.222
<v A>Had it as a goal and trained for 10 years. So what does it mean to say on average people took two years? It's on average people who succeeded succeeded in two years, right? Like what about all the people who didn't? You have to put that number in or else unless you know that's going to happen to you. You don't know if that's a valid statistic. Yeah.

0:14:38.948 --> 0:14:55.992
<v B>We call this a Type II error or a 'near good.' Yes, it's all the things that went wrong that you never saw. So all the people who didn't qualify for the Boston Marathon because they sprained their ankle and you just never saw it because they never applied. Yep. All right.

0:14:59.789 --> 0:15:00.750
<v A>All right. So actually,

0:15:00.750 --> 0:15:38.213
<v B>Well, we should at least just put a bow on this. I mean, oh sorry, you go for it. Amazon. Well, no. I mean, so okay, the question is: Should you blacklist or should you just not apply at any company? Well, okay, almost every company has a cliff. You're not getting around that. Maybe they'll all listen to us and they'll get rid of their cliffs. I know OpenAI got rid of their cliff—their one-year cliff—about a year ago. But okay, so we're not gonna get away from that. But like the 10, 20, 30, 40—should people still go to those companies? I guess it's... I guess maybe it's just not the most important thing. You know, it's not.

0:15:38.213 --> 0:16:05.010
<v A>a that's what i was gonna say if you're comparing two companies i think the incentives are poor for fit and change ability for the backloaded things like 10 20 30 40 so i would be it would be a notch against but i don't think it's a if that's the way to break into the industry to get into a tech company if that's the people that are gonna you know pay you what you feel you're worth then like i don't think it's a reason to not go there yeah

0:16:06.597 --> 0:16:07.960
<v B>I think that's fair. So unless you think,

0:16:07.960 --> 0:16:12.760
<v A>there's foul play unless you have evidence that they're purposely churning people out within a year or two.

0:16:13.720 --> 0:16:32.360
<v B>Yeah, I think we're in the same boat. I mean, I would say maybe a little stronger. You know, for me it's definitely a yellow flag—companies that do 10, 20, 30, 40. It is definitely a concern, but again, it's if you have a good boss, a good team, it seems like the work you're interested in, then it could easily overcome it.

0:16:34.137 --> 0:16:35.605
<v A>All right. On to,

0:16:35.605 --> 0:16:37.225
<v B>News. Patrick, what's your news?

0:16:37.309 --> 0:18:34.120
<v A>All right. So my news probably... most people—well, actually, I don't know. I'm a little curious. I don't feel like it got as much coverage that just as it should, but it is now passed, which is the Artemis II mission, which was the United States NASA trying to send people not in orbit, but in a loop around the Moon. So they were influenced by the Moon's gravity to sort of slingshot around the backside of the Moon and back to Earth. And this happened as of this recording like last week, and it was a very exciting thing, but it was a little... I don't know, underappreciated initially. I think once it was happening, the news coverage kind of picked up, but a lot of people just weren't talking about it. It didn't have a hype. If you ever talk to someone who was around for the original Moon landings—which are now, you know, whatever, 70 years—it was everything. It was a major deal. It was like absolutely wall-to-wall coverage. And I feel this was just... you know, that's cool and people just kind of moved on from it. And I've been trying to come up with a theory as to why. And there are ones, you know, just the world is a little bit of a different place politically. Our country here in the United States is... you know, in a turmoil, a bit of turmoil, I guess internal stuff. So it's hard to get everybody to rally on any cause. But then I think that as well. I feel with the amount of launches for things like Starlink and the International Space Station, I think it's a bit just... it feels like yeah, okay, we do this all the time. And I think people don't understand the sort of energy difference needed to go to the Moon versus go to Low Earth Orbit. So anyways, I just want to give a shout out for the work of getting back to the Moon. And you know, it's a little controversial whether or not a Moon base will be set up, but definitely as far as,

0:18:34.120 --> 0:18:58.772
<v A>like technology demonstration and the ability to push the envelope, as well as potentially unlock lots of new resources for manufacturing. It's a definitely exciting time, and you know it's going to be a little bit of a gap before the next Artemis—Artemis III. But you know Artemis II was an exciting watch if you got the opportunity to tune in during it. If not, plenty of YouTube retrospectives go back and check out all.

0:18:59.633 --> 0:19:21.199
<v B>That happened. So... they so this already—I don't follow space, and for some reason space doesn't follow me either. Like I don't recommend it anything about space. But so that the whole project started and ended, so the astronauts just slingshotted the Moon came back, and everyone's safe. And yes. Yeah. How did? Yeah. I mean,

0:19:21.199 --> 0:19:21.975
<v A>feel like that.

0:19:21.975 --> 0:19:43.480
<v B>Should even, like, us—you know, non-Space enthusiasts like that should have showed up. On you. I have the um oh man, I'm not going to go on a huge tangent on how people get their news. I get my news from swiping left on my Android phone, and I just get the Google News that's kind of everybody gets. So, like really, uh that could have been a good place for it.

0:19:43.480 --> 0:21:01.184
<v A>Yeah, um, I maybe that is a—maybe it's more of a commentary on people's individual funnels or, however you want to call, like individual filters that we all have into our news systems now where they're all biased. But I guess if you do not that, people below a certain age turn tune into the news. But here since they launched—I mean, I live in Florida, so they launched from Florida—so we got some local news coverage of it. So we do sometimes watch the local news just to hear, you know, local events. Um, but yeah, in general national news, it seemed a little undercovered in my opinion. Yeah, yeah, totally fully agree. They're back successful. Yeah, so four astronauts from the United States and one from Canada—so three from the United States and one from Canada—went into orbit, a sort of half-orbit around the Moon. They went again like kind of slingshotted around the backside. Um, so they took a bunch of pictures. They didn't get super close; they got close enough that it was very big in their field of view, but um they could kind of see the whole thing from side to side, um not super low. And so by virtue of being so far away from the Moon when they went around the backside, they actually had a record for farthest travel for a human away from Earth. Oh wow! And so um they went further than any human has gone before. If we want to start trying to change Star Trek, no, that's.

0:21:02.044 --> 0:21:19.959
<v B>Cool, you know, but you know, to your point when they—when the SpaceX people caught the rocket, you know, when it landed in the chopsticks or whatever you all call that thing. Okay? Yeah, that I saw like that was all over my Google News. So I do think that this wasn't able to get the kind of attention it deserved for whatever.

0:21:19.959 --> 0:21:57.209
<v A>Reason. Agree. But a shout out either way. You know, I think it's something exciting, and you know, I'm here for the—I there's a lot of people, young people who are have their imagination captured by these sort of like big technological feats more so than increasingly like, yeah, another smaller, you know, phone or, you know, a better chatbot. Right? I mean, people are gonna start coming of age; you just grew up with that stuff. And so I—I'm hopeful that some of the at least even growing up for me, you know, going deep under the ocean, going into outer space, there's always things that just feel a little special.

0:21:57.884 --> 0:23:23.080
<v B>Yeah, yeah, that makes sense. Um, cool. All right, my news story. First news story is the Gemma 4 release. Um so Gemma is a family of vision language models or multimodal models that um that are totally open source, open weight, um and are small enough that you can even run them on your phone. Uh you can definitely run them on your desktop or laptop. Um and uh I actually did something kind of cool with uh Gemma 4 already. I—uh um I got it to—I fine-tuned a very small Gemma 4 model to try to correct grammar, okay? And uh it actually works really, really well. I tried this in the past with really small models, and I never got good behavior. So my vision was to build something into my phone where no matter what app I'm on, it would just—maybe it would be a custom keyboard—I don't know, but it would just correct my grammar. So it'd look at like the entire, you know, content of what I was saying and it would go through and fix grammatical things. Um and uh the models were never very good, and they were hard to fine-tune, etc., etc. But Gemma 4 actually kind of crossed that barrier where I ran through a fine-tuning on my desktop, and—and

0:23:25.195 --> 0:25:18.797
<v B>Okay, real quick thing on fine-tuning. So you know fine-tuning basically just means continuing training. And so even though it's a small model, um it's small because of all these tricks, but the tricks don't work at training time. So think of it as like—um like you bake a cookie and then you put icing on it. But if you put the icing first and then bake it, or if you try to re-bake a cookie with icing on it, you know the icing gets all hard and it gets kind of, you know, not very good. So it's kind of two different processes. And so what you need is to—it might be that to bake the cookie or in this case, fine-tune the model, you need like a pretty beefy setup. Um and then you can sort of, you know, shrink it later. So um so I got my desktop together. I—I you know fine-tuned this model. Um oh yeah, so there's there's a whole bunch of tricks now: low-rank adaptation, all these other tricks where you can fine-tune it even if you don't have like a really beefy desktop. Mine's okay; it's got 16 gigs of VRAM, which is pretty good but not enough to just do a pure fine-tuning. Um long story short, the tooling is amazing. It's come so far. You don't have to be like a super read in expert to do these things now. I literally just handed it a text file that I got off the web, a CSV that had, you know, a bunch of grammatical mistakes and then the corrected sentences, and then it went off and it just—the validation loss went from like 50 to 80 percent or validation, you know, accuracy. So the fine-tuning definitely did something, and it might actually work. So I haven't finished it yet, but hopefully I'll have something that just corrects grammar for every app that would be pretty neat.

0:25:19.692 --> 0:25:26.847
<v A>Was going to ask how you got good training did it because I would give it poor training if I just gave it my own conversations. Oh yeah, it doesn't.

0:25:26.847 --> 0:25:55.500
<v B>Seem correct. I just went on online and said, 'Hey, I need a giant data set full of grammar errors.' I found one. It's called the—I think the C200, C200M data set. But if you just go online look up a grammar error correction data set, you'll find it's the first result, and it's got way more grammar corrections than I'll ever use in my entire life. It's like 200 million of them or something. Oh wow.

0:25:57.762 --> 0:26:32.490
<v A>Yeah, I mean, I would love to try the fine-tuning. Unfortunately, I've always—we've talked about this before. I've always lagged in playing video games. I always play older video games. I never had a good GPU. Now I kind of want a good better GPU than when I have for both playing video games and for playing with like model training, and it's just extortion. I like—I know like I could buy one, but then I see articles saying, 'Oh, this is a you know 500 GPU,' and I go on eBay and it's like 800. I'm like, no, I'm not. I'm not doing that. Like, I'm sorry, it's like it's a principled thing like I just—I can't. Yeah.

0:26:33.115 --> 0:26:42.840
<v B>I know. If I had to start over, I would—well, I use my GPU for gaming too, but but if I didn't have the GPU, I would probably just use an EC2 instance for me. Yeah.

0:26:42.840 --> 0:27:36.885
<v A>I that's what I need to do. It's just one more step, one more you know unbounded like it's—it's a rather than I spend this money and I do whatever I want. You know, it's like okay each time I do this if I make a mistake is you pay a little more psychologically, but you're absolutely right. I mean, way cheaper to do that stuff often with rented and probably better hardware um yeah than you would have at home. So yeah, that's something I definitely would love to get into. And I think there's something to be said to the approach you're taking and like small models in general run a lot faster, run in more places. So even if you have the ability to run a big model or subscription having like a small model that you can use for specific tasks, uh and we're going to talk about, you know, potentially future more organic sort of work having things offloaded to scripts or to small models to do, and just the speed—the tokens per second that you get is just so much faster. I think you're going to see more of that blended stuff. Yeah.

0:27:38.084 --> 0:27:39.518
<v B>Totally agree. Turning the

0:27:40.598 --> 0:28:03.480
<v A>Corner to yet one, which is—oh, the news is all over the place this time. It's okay. Uh, it's a small you got to click the link in the show notes or just search 'Hatreds.' I'm i hate this game so bad. I guess that's why it's named Hatreds. First of all, I watch the—I don't know if it's summoning salt whoever does the, you know, Tetris like infinity scores.

0:28:03.480 --> 0:28:05.759
<v B>And the broken levels? Oh yeah, runners. Well, I

0:28:05.927 --> 0:29:16.313
<v A>love these. I watch them. I'm like, 'I'm gonna play some Tetris.' I suck. Like, oh, this is the worst. It's so bad. Like, I you know that I'm like, you know, Tetris strategies anyways. So, I've been playing—I think it's a Game Boy Advance homebrew game, Apotris, on a little retro handheld that I have. Um, so if you're interested, there's like sort of a modern interpretation. So, it has like some niceties and and a lot of customization. So, I've been playing that a little bit but still suck at it. Um, so Hatreds, though, instead of—so okay, people don't want Tetris. There's all the different shape pieces. There's various ways of selecting them. So sometimes it can be literally just whatever the next piece is is random. Sometimes it's selected from what's called like a draw bag. So they'll put all the pieces in a bag metaphorically and then draw one out at a time. So you kind of the longest, you have to wait for a certain piece is bounded. Um that's what more sort of modern interpretations do. Hatreds instead says, 'I'm gonna run a little like, you know, program heuristic machine learning thing whatever the computer and it's gonna pick the worst next piece.' So you

0:29:19.924 --> 0:30:05.369
<v A>start off getting one of the—I don't know what they're called, like the Z pieces, and you kind of just keep getting Z pieces which are kind of hard to put together. Yeah, but then unless you force it into an option where there's two ways for you to score a line, it's just going to give you whatever piece won't arrange to fit in the notch that you have left. So you have to like carefully plan and to make it worse, it lets you go as slow as you want, right? So you feel like this has got to be easy, but I literally it's like getting one point and then I'm like flipping the table excited. You know, the whole thing going off. It's amazing. But yeah, zero is where I was started, and maybe I'm just bad. So you can be like, 'Oh man, that Patrick guy like it clearly just sucks. Like, I'm way good at this. Go play it.'

0:30:05.706 --> 0:30:07.022
<v B>Maybe you'll score a lot of

0:30:07.022 --> 0:30:20.840
<v A>points, but it is enormously frustrating. I can only play for like two or three times and then I'm like, 'No, I'm done. I can't. I can't.' I've got to try. All right. Hey, true hilarious if you get more than zero. Don't email me because I don't want to know.

0:30:24.437 --> 0:30:25.973
<v A>Oh man, this is so good. Yeah.

0:30:25.973 --> 0:30:36.874
<v B>I'll have to try a report back. All right. All right. He's gonna be doing it well. I ought to go on a little monologue and give him time to play just at the end of our episode. It's like, 'Oh, I got 10.' Oh dude. Yeah. Yeah.

0:30:37.364 --> 0:30:38.280
<v A>I'll hang up.

0:30:40.469 --> 0:30:46.814
<v B>Oh man. All right. My my second news is um drip warts school of drip. So this is uh

0:30:48.535 --> 0:30:52.028
<v A>I don't have to click the YouTube link. I already—I

0:30:52.028 --> 0:31:12.784
<v B>Already know what it is? Do you really? Yes. Okay. So I play board games every Monday with a group of guys—a very fun group—and one of them showed me this. This has got to be the most viral AI video that's been created. I mean, I don't think anything has even come close to this in terms.

0:31:13.257 --> 0:31:38.350
<v A>Of virality? I disagree. I think there's been viral videos—aren't clearly AI—that are AI that just... This is the most overtly AI thing to go viral. Like it's clear that, oh, not bad AI. Like it's but it's clear that it's AI, like it's not real, right? I think there are ones that have been purported to be real that were just viral and like, 'Oh okay, yeah, that's true.' Yeah, I don't have evidence. Okay, so.

0:31:38.350 --> 0:31:59.960
<v B>So yeah, what I mean specifically is like this is one where the whole point is it's AI. No producer would ever make this, and we all know that. And so, and it's just hilarious amazing premise: you know Harry Potter but if it's kind of like gangster style.

0:32:00.000 --> 0:32:01.030
<v A>But also.

0:32:01.182 --> 0:32:41.074
<v B>Like, you know, high fashion, you know, all mixed together. Very very funny. I just burst out laughing almost immediately. This is great. I mean, I feel like there—I'm actually surprised that OpenAI killed Sora because I do think that there is a play there. You know, there's, you know, really like opening it up and letting the whole world think about what are funny things that we could mash together and sharing that. That seems like super powerful. I mean, maybe that's what TikTok is going to become. Is—is that okay? But

0:32:44.010 --> 0:33:20.781
<v A>You—I don't know. I don't know how much conspiracy theory tin foil hat we want to go about. Why? Yeah, go all okay. So no, no, I also thought the same thing because Sora was like trying to do a thing. Like my kids kind of knew what it was, which is always like a pretty big thing for tech. Yeah. Um so we don't have it in the news of the episode now because like I think it's still in flight, but Anthropic has been teasing this Mythos, right? Their next sort of LLM version after Opus 4.6, I guess. And it's supposedly, you know, earth-shattering, which to be fair, they all claim to be before they come out. I'm making no

0:33:21.119 --> 0:33:24.038
<v B>Statement there, but the rumor is that it was sort of

0:33:24.038 --> 0:34:03.846
<v A>Like 10x more training than the last one, but that there's some sort of—do you know like the thing that happens with Grokking? So when a machine learning model trains and it hits some sort of like asymptote and it seems like it's just sort of stabilized in performance, but actually under the hood it's sort of like self-organizing and then it's able to like sort of reach a new level after it sort of gets through this barrier of not actually getting better loss, right? So your loss, your validations aren't getting better, but the model is sort of actually organizing itself to the point where then it sort of like unlocks additional capacity and then trains further. That's Patrick's non-mathematical. Yeah, it's

0:34:03.846 --> 0:34:08.520
<v B>Like defragmenting your hard drive kind of, and then you can it can run fast enough for you.

0:34:09.027 --> 0:34:48.360
<v A>To do something else, right? And so supposedly the rumors go that Anthropic's Mythos took like whatever 10 or 100x what the last one did, which was already insanity, but that it was a huge unlock. Right? So there's all these scaling rules and so it would break the sort of projected—sort of—regressions fit to the sort of like how much training versus performance you get. And supposedly if that's true, right, then OpenAI needed to free up the Sora services so that they could use the additional hardware to get the training budget they needed to basically not get leapfrogged.

0:34:50.084 --> 0:34:54.117
<v B>Wow. Yeah. I mean, that would be remarkable if true. I mean, to

0:34:54.117 --> 0:35:11.970
<v A>Be fair, it's like a rumor about a rumor that is like it could have just been—it's not making money, but like I feel there had to have been some reason to sort of like deprecate it and not just wait for the new version to be better. Yeah. I mean, I well, okay.

0:35:11.970 --> 0:35:59.879
<v B>I'll tell you my—I don't know if this is a conspiracy theory or just a theory. My theory was that, you know, OpenAI's brand recognition is suffering a lot. You know, like Sam Altman's name recognition is suffering a lot. Did you see The Onion interview of Sam Altman? No, oh my god, it's hilarious. I mean, obviously completely fake. It's just—you only see a transcript; there's no video or anything, but very, very funny. Um, but I felt like Sora is kind of a big brand risk, like even this Harry Potter thing absolutely hilarious, but I don't think the brands, you know, Balenciaga, I think Louis Vuitton, or literally by name—I don't think they're happy with that. So I felt like they were trying to just limit their exposure, but but they

0:35:59.879 --> 0:36:06.038
<v A>had a big partnership with Disney, right? Like a billion-dollar potential value. Yeah, yeah, and it's

0:36:06.949 --> 0:36:22.279
<v B>all just gone. It's kind of wild. I mean, I—yeah, I would chalk it up to the same thing, like a brand thing even for Disney, but maybe maybe it really is the compute. We'll have to wait and see when this new mythos comes out. What's going on there?

0:36:24.786 --> 0:36:46.454
<v A>All right, talking about science fiction, it's time for Book of the Show. We're going to take turns this time, so I'm going to do a Book of the Show and Jason's going to do Tool the Show. So my Book of the Show—Late Better Late Than Never, they say. Project Hail Mary. So if you've well timed with Artemis, which I think was coincidental. Okay, real quick, not to

0:36:46.757 --> 0:36:50.048
<v B>interject too much, but but I am just starting this book, so you

0:36:50.048 --> 0:38:01.615
<v A>can't try not to spoil it. Yeah, yeah, no, I'll be good. I'll be good. I always always undershoot, I think, but it's also there's movies, there's trailers, very hard to avoid. That's true. There's like levels of spoilers, but but I'll—I'll be even better than I think the trailers. I think trailers are really good; they reveal too much. Um, but Project Hail Mary is a book about, you know, science fiction, about a person out trying to, you know, save a dying Earth, right? And there's lots that goes into that. It's a bit of—if you've, you know, kind of gone through Andy Weir's other book—oh, why is the name escaping me now? The one Martian, right? Oh, *The Martian*. Thank you. I was going to say a different sci-fi Mars book. I was like no, that's wrong. Um, thank you. *The Martian* it's a little bit the same, right? It's like sort of tech science grounded. He does a pretty good job about that, but then also like this sort of hopeful, like, you know, you want to root for the good guy, you know, and not have like sort of the same chaotic, you know, black mysterious dark, like, you know, is this good or bad? It's just like you want to root for someone. And so it's kind of in the same theming. Um, and I will say I was encouraged the book for a very long time did like *The Martian*.

0:38:02.307 --> 0:38:05.024
<v B>Man, I don't know. Like the book was good. It's a

0:38:05.024 --> 0:39:36.014
<v A>short read compared to most of my recommendations. Most of the books that I read very easy read. I went on a vacation recently and actually like first two days of the vacation, I basically read the whole book. Oh wow, we did a very long flight, you know, leaving the country going to a different continent. So oh wow, it was a very long flight. So to be fair, okay, it was still a lot of reading. Um, book was good. Book was like, you know, pretty good for what it is, like, you know, not super deep but, you know, well-grounded. Had some some issues, but, you know, just around the edges. I'm not super critical of science stuff generally other than just being like, yeah, right, that's come on, that's—that's nobody knows such different disciplines to this level. Um, you know, that's just unrealistic, but other than that, you know, plausible, I guess. But I, you know, I don't the movie people love the movie, but I feel like I wasn't as keen on the movie as I was on the book. The book was definitely better than the movie, but the book was not as good as it was hyped up to me. But it was so short, it's got to be worth it for just like sometimes you need the casual stuff, right? Like you can love the deep grindy this is really making me think, but then, you know, sometimes just taking that—that sort of like I say I'm that way about Marvel movies. A lot of people rag on Marvel movies, superhero movies. Yeah, I just I want to go in and watch something I don't have to think so hard about, like, you know, I don't want it to be, you know, some deep brooding, you know, at the end was he in the dream or was he awake? Or, you know, okay. I don't—I don't whatever. I don't want to know. Like just tell me what happened. Yeah, I'm

0:39:37.971 --> 0:39:53.496
<v B>right there with you. I mean, you know, I go into a movie, you know, I saw the Super Mario Movie with The Boys. Oh, I—I want to see this. Yeah, yeah, they they loved it. Um, and you just you have to go in with the right expectations. You're not going in to expect something really high concept.

0:39:54.120 --> 0:40:40.915
<v A>Yeah, so Project Hail Mary is like a near-near future very grounded, not super, you know, far-flung, um but definitely definitely worth a read. It's pretty easy read. I—you know, on a scale of such books, I guess, and reasonably short. So I've heard the audiobook is also really good. I normally listen to audiobooks; I actually read this one on my Kindle. So oh, um I don't know if the audiobook is as good as it's crap. I'm digging the audiobook. I mean, I've only just started, but the voice is fine. So yeah, I think it's all right. Shout out for the audio, but yeah, I most would probably know this now because the movie is out. My daughter's reading the book after the movie and having a good time as well. So I don't think, you know, it's one of those they're just different; they're the same story. It's pretty true to each other, but they're still pretty different in the, you know, depth of content.

0:40:41.370 --> 0:40:55.005
<v B>They have so you shared your Project Hail Mary literature with your daughter and I share my Drip Warts School of Drip with my son. Which one of us is the better father? I showed my kids that.

0:40:55.005 --> 0:41:05.480
<v A>video too just to be clear. Just typically there is some vulgarity, so if you like just be mindful if you—if that's something you're very concerned about, there is a there are a few.

0:41:05.586 --> 0:41:15.400
<v B>Minor profanities. There's definitely yeah, there's probably some F-bombs and stuff like that. To be fair, I only showed it to my to my 12 year old, so I don't think I would show it to a six year old.

0:41:16.319 --> 0:41:18.090
<v A>But just watch it first. That's what.

0:41:18.090 --> 0:42:16.731
<v B>We're saying yeah, totally definitely a viewer. What is it? Discretion advised or something? Cool. All right. Yeah, and if you do read the book and get it from the library instead of paying for expensive movie tickets, you could turn around and give that extra money to us by following us on Patreon. We do really appreciate all of our patrons. All the money just sits in an account where we use it to help out the show. We try and get more folks interested in the show, especially folks who are starting their career. And so all of us collectively really appreciate your donations. Okay? So tool of the show. Patrick's gonna skip this time. Okay, Patrick, this is gonna blow, either gonna blow your mind or you already know about it. I wanted to play Final Fantasy VI—that's the one with Edgar and Sabin and Tifa. Basically, it's actually Final Fantasy III in the US, but they call it.

0:42:16.731 --> 0:42:18.604
<v A>Okay, I know which one this is, so.

0:42:18.722 --> 0:43:06.200
<v B>It's with you know Tifa. She's like got a connection with the Espers, and it turns out—I'm not gonna spoil it—but basically love the game, love the story. Haven't played it probably 20 years, so I thought more than 20 years. So I thought I want to play it, but like I know I'm just gonna beat it. I mean, if I could beat it at 12, I'm sure I could beat it at 40, whatever. So, so I was like how can I play this game and get the story and experience but like still be challenged, right? Okay, oh man. So I found out that people make you—I've always done like ROM hacks for translation so I could play like English versions of Japanese games, but there's a whole ROM hacking community just for making games either harder or more interesting or both. Or what?

0:43:09.162 --> 0:44:05.288
<v B>Have you? And so I played Ogre Battle. It's not Ogre Battle Hard Type. There's an Ogre Battle mod which I can look up, but basically made the game harder but also balanced all the units and everything. There's separately an Ogre Battle Hard Type, but that was just very frustrating like at some point what I don't want to do is to play the same game but have to grind for like a hundred hours, you know? Like that's not fun. What I want is just like a harder experience but taking roughly the same amount of time. You just have to be more strategic, right? So I finished that Ogre Battle mod, and then I found this amazing Final Fantasy VI mod. It's called T edition, and it adds an insane amount of content. For example, there's achievements, there's like all these extra side quests. There's just a ton of content, there's new bosses.

0:44:06.925 --> 0:45:19.065
<v B>There's actually like a—there's like different mechanics. So for example, you know generally in Final Fantasy, you almost never—you almost never like cast certain spells, like you know they're kind of in the game for continuity, but like when do you really ever like poison your enemy? It's pretty rare, right? But now it's like there's different bosses that have certain weaknesses, and that kind of encourages you to use the whole gambit of spells. Phenomenal. I mean, I'm about halfway through it. I was worried that, you know, when you play a ROM hack, you know there's always the risk that the the rom hack—the hack designer just isn't as thorough as the original game designer, and like you'll get maybe 40 hours into the game and now you're just stuck like you you just you can't make progress. And so you you're you wasted your time. But this T edition is super popular. It's got a thriving community. It's been around forever. Tons of people have beaten it, and so you know that it's like you're gonna get to the end. But it's super super fun. I've been playing it on my phone, and it's it's a blast.

0:45:20.415 --> 0:45:26.946
<v A>Was awesome. I have never played Final Fan—I would always call it Final Fantasy III, but VI. Yeah, yeah, wait.

0:45:29.022 --> 0:45:30.359
<v B>So you've never played III?

0:45:31.907 --> 0:45:33.139
<v B>No. Oh man.

0:45:33.139 --> 0:45:39.974
<v A>oh seven i did seven but never three so i actually stopped at three because i never had a

0:45:40.359 --> 0:45:47.483
<v B>PlayStation. I missed out. I actually played VII a year ago, but you know my childhood, I missed out on VII is.

0:45:47.854 --> 0:46:09.420
<v A>It's worth, like, I guess. So when sometimes when you play the old games, like if you don't have nostalgia, they're really tough to play for, like, quality of life reasons. Just like lots of random stupid stuff you got to do or just, like, you know, they tried to make the game longer. You know, I don't know. Stuff like that. Like is this game still playable? Or do you gotta have nostalgia for it? Or is this going to be like no there's

0:46:10.315 --> 0:46:13.521
<v B>Better games to play. It's really hard to say. Okay.

0:46:15.192 --> 0:46:24.360
<v B>I would say the story is phenomenal and probably worth playing even if it's the first time. Yeah, the story is very, very

0:46:26.751 --> 0:46:35.813
<v B>good. I feel like the pace is good. I'm debating whether if you go straight to T edition that might be difficult because you

0:46:38.614 --> 0:47:22.489
<v B>do you are kind of expected to? Yeah, I would play the regular edition. Okay, T edition would really be for people who want to play it a second time, but I would definitely go back and play if you're going to play the regular version. There's some Android ports or iOS ports that are probably better than playing it on emulator. But I think it's a great game, very solid. It shocked me the first time I played it where I thought I was, you know, at the end of the game and it turns out you're only at the halfway point. Kind of like if you ever played A Link to the Past Super Nintendo. Yeah, remember when like you fight Ganond and then like he throws you into the Dark World or whatever they call that, the other world? Do we

0:47:22.489 --> 0:47:24.936
<v A>have to give spoiler alerts for like 30-year-old guys?

0:47:26.489 --> 0:47:43.077
<v B>It's kind of like that where although in that game is more obvious because, like, okay, clearly there aren't just three things to do in the whole world, right? Um here it's less obvious, right? I literally thought, 'Well, the game's over,' and I was only halfway done. Nice. Yeah, I'm

0:47:43.920 --> 0:47:54.180
<v A>ashamed to me. I also have only ever made it like through the very beginning of Chrono Trigger, and everyone's like it's the most amazing RPG ever, but yeah, I'm bad with RPGs.

0:47:55.345 --> 0:48:16.236
<v B>Yeah, Chrono Trigger was very good. I think the um I could see people getting stuck in Chrono Trigger because there's some parts where you just have to kind of persevere to make the other parts kind of worth it. Versus this, I would say Final Fantasy III. There's constantly something interesting going on. All right. All right.

0:48:16.557 --> 0:48:17.350
<v A>Well, I am.

0:48:19.000 --> 0:48:26.969
<v A>Trying to go back and play more of these, so I will. I'll add it to the list. We'll see if I get to it. You should totally play the ROM hack. It's weird. It says

0:48:26.969 --> 0:49:16.244
<v B>that it doesn't there's glitches on emulators, but I think that was back in the past. Just quick history of emulators. So you know the machine has its own instruction set, and Patrick probably knows this way better than I do. So I'm gonna try my best. You know, it has like, you know, move things over here or this other type of instruction or whatever, like at really low level. And your computer, you know, doesn't have all the same instructions, right? Your phone probably has different instructions in your desktop. And so it like it has sort of translate the instructions from, you know, what it would give a real Nintendo to what it has to give your computer. And back in the day, like they couldn't just translate it perfectly because it became really expensive. Like there might be an instruction that was super fast on Nintendo but it'd be really slow on your computer.

0:49:17.813 --> 0:49:35.937
<v B>And so when they made these hacks, I guess the hacks only worked on the Nintendo without glitches. But now the emulators are like absolutely perfect because the machines are just so fast. So there were these warnings about, 'Oh, if you're playing on emulator, the sound won't work,' but everything worked perfectly.

0:49:36.578 --> 0:49:47.969
<v A>Almost all old games with the sound off because I find the repetitive chip tune thing like only listenable in very small doses. So I basically play with everything on mute. So yeah.

0:49:47.969 --> 0:49:48.930
<v B>Me too, actually. Yeah.

0:49:49.555 --> 0:49:51.720
<v A>But even on my Switch, that's like my case is.

0:49:51.720 --> 0:49:56.279
<v B>Because I'm in the car or you know, I'm in the passenger seat or something where I don't thank you.

0:50:00.017 --> 0:50:02.869
<v B>For clarifying. Yeah, I just don't want to bother people.

0:50:04.810 --> 0:50:11.965
<v A>Usually. Yeah, all right. It is time for the agents to code to the rest of our podcast. Yeah, I mean, why are we even here? Like can't?

0:50:11.965 --> 0:50:12.920
<v B>I just press a.

0:50:17.736 --> 0:50:51.250
<v B>Button on Generate Script. Oh man, I mean what a crazy time we're living in. So quick history lesson. You know, IntelliSense has been around forever. I don't know if folks remember the word IntelliSense, but and I don't know how it worked. I mean, I can take a guess. I mean, I guess that you know it ingests all of your code on your local computer and then it builds up some kind of statistical models. And then when you start typing, you know,

0:50:52.853 --> 0:52:10.360
<v B>or maybe it doesn't even need statistics. Maybe it's totally deterministic. But you start typing, you know my object dot c a c and it just fills in, you know, cache cache triangle or something like that because it knows that that function exists and there aren't any other functions that start with c a c. And so, it'll just pop that up next to your cursor, and you could hit Tab and it'll just auto-complete that function name for you. So that's been around forever, and that's been great. I had like an Emacs extension at some point a long time ago that did this. So even in the terminal, you could do it, and everything that's been around forever. And then you know the Large Language Models started to get really good, and um and so yeah, there was they started doing auto-complete on steroids where you could use GitHub Copilot and you could start typing, you know for, and it would just auto-complete, you know, the entire for loop and the contents of the for loop or something if it was pretty easy to infer from other parts of the code base. So that was really cool. And then a Cursor came out, and so Cursor.

0:52:12.570 --> 0:54:08.772
<v B>had this neat feature where it would auto-complete at your starting at your cursor, but then it would jump around. So for example, you know in most languages you have some type of import. So in TypeScript it's called 'import,' and Python it's called 'import.' C++ it's '#include', right? And so let's say as part of the auto-complete it ended up needing to use a library that you weren't yet importing. You could tab to complete that code, and then it would give you a little notification type thing that there's more work to do somewhere else in the file. So you could type Tab again, and it would jump to the top of the file and show you what it wants to auto-complete over there. And then you could type Tab a third time, and you would do the imports as well. And so, you know this kind continued to the point where I think at some point Cursor had multi-file like you could just kind of keep hitting Tab and it would jump around your code base doing various things. And so this is all great. Um and so then so people were you know using that. I was using it. It's very exciting. And then Claude Code came around, and that was a huge game changer. So this is where, you know instead of just extending an idea that you had partially written, you could kind of ask it in English to do something with your code, and it would go off and do it. And that was pretty wild. I remember when that first came out. I posted actually, I put my most popular post on LinkedIn was basically talking about how awesome Claude Code is, and a whole bunch of people like ripping on me. This one guy I'll never forget it.

0:54:09.278 --> 0:55:46.280
<v B>If you're out there listening to us, you're a jerk. Stop listening to us. But this one guy was like, 'You should be better than this.' Like, do better. You know, like I can, you know, like Claude Code and these things are—he called it a stochastic parrot, which I found out later is like a common pejorative for LLMs. He's like, 'You should do better than this.' You have a programming podcast, and you're trusting these stochastic parrots,' and just insulting my character and everything. That guy was a jerk. But but it was a very controversial post. There's one of these things that like I didn't do it on purpose, but it got—it got I think like one and a half million views. And there's definitely like pro your proponents and antagonists. Um but I was right in hindsight. I told everyone early on that this Claude Code thing is amazing. Actually what I'd said that I think was so provocative was, 'I have not written a single line of machine code in my whole life. I've run 100 percent of my code through the compiler,' and so I'm already not writing pure code. So if I if I start running a hundred percent of my code through Claude, and I just tell it in English what to do, it's really not changing anything.' And that's kind of where I still stand. I mean, I use it constantly now, you know. Even today there are times I have to go into the code, and we'll talk about all of that, but in general kind of where I'm coming from is I'm a big fan, and I think it's pretty amazing, pretty amazing times we're living in.

0:55:46.280 --> 0:57:44.760
<v A>So maybe to expand a little on the switch from those early sort of like you kind of explained how Cursor was in the early days to what Cloud Code or Codex or the Gemini solutions are today. I think we've also seen a lot of iterations around the edges, so things like what the MCPs—you know, around retrieval augmented graphs (RAGs)—around, like, basically, in my opinion these are kind of how to let the LLM, which is just next token prediction, right? How to allow it to do tool calling, which is something we talked about in the podcast a long, long time ago where we said, 'Hey, these chat things would be really cool if they could reach out and do web searches or connect to Wolfram Alpha.' Or okay, well, anyways, turns out people are already working on that; we just didn't know. So everybody had the same good idea—good, but then you know, so basically how to interact with tools, and there's been a lot of iterations, solutions through, you know, how that works to what it is today. I don't say it's settled today, but where it is today. And then the other one is around sort of managing the context window, right? The hey, how much of your code base, of your problem space, of your conversation—how much of that can be held for the context? I think it's been interesting that the models themselves have gotten bigger, but one of the things that isn't super obvious is that even if you see something like Gemini can handle a million tokens as context, that doesn't mean it's as efficient both in terms of how fast it runs, but also in how the quality is at a million tokens versus 10,000 tokens. And so a lot of these models have degraded performance as they reach up to those million tokens. It is unintuitive because it's not like a hash map or a list or an array that just grows and you get some weird cache effects, but generally, you

0:57:44.760 --> 0:59:00.068
<v A>know they're kind of well understood. There are these unintuitive mechanisms, so aggressively managing what's loaded into the context and things like compaction is super important. And I think it's been very interesting to see something like taking your whole code base, understanding how to use tools to search in your code base, how to learn which pieces to extract and load up, and sort of assume—and I will say it's still not perfect. You still see sometimes I see it even using a tool like Cloud Code where we talk about hallucination; it'll hallucinate API calls, but it's in a loop now. That's the agentic part where it'll try to compile itself and realize, 'Wait, why did I call, you know, feature or you know object.foo.foo? Isn't doesn't exist?' It's dot bar. And you know, it'll figure it out, but for whatever reason like this, you know, hallucinating of foo just assumed that you have a list class, therefore you have an insert, and maybe you don't have an insert because of some esoteric reason.' And so you still see around the edges, but a lot of that is because it's trying to keep the context down in the first pass, which causes its own problems. If it loads up all your code base, if it tries to understand all of this, the speed can really get bogged down.

0:59:01.080 --> 1:00:34.120
<v B>I mean, you—the other thing that I think you touched on but just double-click on is the actual intelligence of the model goes down as you put more information in it. And so this creates a weird trap where someone will start with a completely blank slate and say, 'Build me a website for e-commerce,' and it can do that. But a lot of that is based on its sort of innate knowledge from looking at many e-commerce sites on GitHub and things like that. And so it'll build something, and you'll feel like it really understood the nature of what you built, but a lot of it might be kind of copied. And then as the context grows and you end up with more and more bespoke information about your particular product that you're selling, then it needs to keep more and more information in its short-term memory. And as it does that, the intelligence starts to go down. And so you see this trap. And so it's just something to be aware of that when you're working with bigger projects, shouldn't expect the same level of intelligence that you have at a smaller scale. But yeah, the loop thing and the tool calling super important. Okay, so we'll explain a little bit how this works under the hood.

1:00:36.280 --> 1:01:30.154
<v B>In the beginning, it takes your question, and the AI can do one of several things, right? It can answer your question just by emitting some text, or it can call a tool. And there's several tools that it can choose from. And so this technology has been around for—you know, is older than Cloud Code to Patrick's point. So when it's done calling a tool, the tool information is added to the context. So just to back up a bit: if you ask it, you know, what's the distance from here to the moon because I want to slingshot some astronauts, right? So it'll start giving you that answer, but as it's emitting those tokens, it's also using what it said to generate the next token. So it's

1:01:31.724 --> 1:02:32.200
<v B>possible. I mean, it might be hard to make it do this, but it's possible for it to generate half of an answer and then to actually say, 'Wait a minute, stop! I'm heading in the wrong direction,' and it's going to generate a totally different answer like mathematically that's plausible. It might be hard to craft a question that would cause that every time, but it can happen. So similarly when it calls a tool, the tool output—it's as if it said those tokens. In the sense of like it's now part of the context and it's used to generate the next thing that it says. So the model can either answer you or it can call a tool. When the tool comes back, that information is let's say part of the context along with your question. Then it can call another tool or it can answer you. And whether—you know—however many tools you want to allow it to call and all of that is up to the discretion of the developer, right? But at some point, it's done calling tools; it gives you an answer, and that's it.

1:02:34.414 --> 1:02:36.270
<v B>to the idea with agentic.

1:02:39.156 --> 1:02:57.128
<v B>Was what if a tool was itself like another question? Like what if we made this recursive? And so there's two basic ways you can make this recursive. One is where the tool call is actually another question that starts a session within a session. Another

1:02:58.022 --> 1:03:05.599
<v A>way is to say well, a tool call can actually give me back a list of more tools to call.

1:03:06.595 --> 1:03:09.320
<v B>Of those are implemented in Cloud Code. So

1:03:11.708 --> 1:04:55.455
<v B>you might say something like, 'I want to remove all the lint errors in my code base,' and it'll come back and say, 'Okay, I'll run a tool call, and we'll run PyWrite, and we'll look at the errors that are involved.' That might come back with you know 10,000 errors. And the model will actually say, 'Okay, you have a ton of errors here. We're not going to just fix this in one diff.' A diff, by the way, or a patch is also another tool call, you know, but we're not going to fix 10,000 errors in one patch. So I'm going to create a bunch of subtasks, and those subtasks are like isolated questions that I'm going to ask myself.' So the model will ask itself like in another process or another context: 'Fix all the lint errors where the capitalization is wrong in the variable; just focus on that.' So now the model is really like the initial model is now an orchestrator that's orchestrating all these sub-questions. And the thing Patrick was saying is, yeah, the nice thing is if anything is wrong, that's okay because the expectation is to be eventually correct. So the model, as long as it has a way to verify it, can come back and try to compile the code and say, 'Oh, I made a mistake; I'm going to call another tool call to fix it,' etc., etc. And this continues until there's some kind of stopping criteria, which again is created by the developers. At that point, the model hands the reins back over to

1:04:57.396 --> 1:06:51.480
<v A>you. There's a lot of I think maybe unintuitive to outsiders like interplay as you're describing between what—when we talked about fine-tuning before but targeting these coding benchmarks and coding applications, and code is like a way of doing sort of long-term planning, which has been something difficult for LLMs to do. But also the interplay between the harness of Cloud Code and like or you know Codex or any of them, and the underlying model—right? How how do you tell it what tool calls to use? You have like a tool call language, or do you just let it use a command line? Do you use MCPs, right, like internal for stuff? How do you sort of craft the system prompt? How do you craft like the each turn, right? Like how far do you let the model think before the night? Like there's this interplay between how the model was trained and how you prompt it, guide it, harness it, skeletonize it, structure it, and even ask separate questions about like, 'Hey, how do I like how would I plan to do this?' or 'How do I decompose this task?' So how do you decide whether to kind of have it do more of a one-shot, like, 'Here's what I'm trying to do; just do it,' and then saying, 'Let me first ask for to decompose it into tasks,' then I'm going to say, 'For the first task, like output the task as JSON,' and then for each thing in the JSON array prompt it again, right? Like you get all this interplay between how you invoke the LLM and how the LLM works. And then we—I don't think we've yet seen to be honest—you know, we might talk about in a future episode but things like Open Claw or like the various more like computer use things, which is the LLM is really starting to train on this use case, right? Sort of actually getting to where they themselves are making sure that during training they're understanding and working with this tooling better. And so for now, it's a lot of I don't want to say hackery; that's not the right

1:06:51.480 --> 1:07:49.319
<v A>word, but like a lot of sort of humans iterating or having LLMs iterate the LLM harness for the LLM. Oh my gosh, it's just LLM. But but you know I think there's a lot of nuance, and you'll see people talk about how Codex works, how you know Cloud Code works, how the Gemini thing, but there's also like OpenCode, right, which is a version of Cloud Code but with open source and you know bring your own back end as a more supported sort of method methodology. And I think all of them have nuanced difference in fact. Just last week Cloud Code had its source code leaked, and people were kind of deep diving and seeing, 'Oh, hey, there's like you know monitoring of how upset the user is.' There's like all this stuff in it. You wouldn't assume. So I don't think there's a settled out harness approach yet, and then the question is like how much does the harness need to match the individual model? It is unclear because every model is sort of slightly different too.

1:07:50.500 --> 1:08:05.400
<v B>Yeah, yeah, great points. Yeah, I think it's still pretty early days. One sort of thing that really surprised me is how well it works on things that aren't even coding related. You

1:08:08.084 --> 1:08:52.229
<v B>know, there's recently Google released a set of skills. So we didn't really talk about this, but a skill is basically a set of tools with really detailed prompts behind each of the tools. So you know, there might be a skill where it's about reading and writing Google Docs, and so you'll get a set of tools that let you sort of do the mechanics—you know, add to a Google Doc, read a Google Doc, insert characters into a doc, etc. But then you'll also get this really detailed markdown of what is the true sort of platonic nature, what is the what is the nature of a of a of a Google Doc? Like what actually is.

1:08:54.861 --> 1:09:24.730
<v B>It. Um, and yeah, so you can plug that Google G Suite skill into—or I think they're calling it a Google Workspace skill—into any of these coding agents, and then they have that power. So I think that this is going to be pretty disruptive to a lot of industries. You know, I'm thinking like finance, medical, legal, etc., in the same way that it's disruptive to coding. That's it's a very generic system. So um um.

1:09:26.485 --> 1:10:18.040
<v B>The way that you know Cloud Code worked—that's pretty different from any of its predecessors was that the tools that it had were extremely generic. You know, people were at the time building very specific tools. Um like, 'Here's a tool to access, you know, my bespoke database,' but I'm not going to give you just open SQL access because you might just drop all the tables in my database. So I'm going to give you like this tool adds a customer record and this tool deletes a customer record.' Like very, very specific tools. And Cloud Code came around and said, 'Well, we're going to give you a tool called Bash which can do anything, and we're going to give you a tool called Patch that can just patch any file anywhere, and a read file tool and a read a chunk of a file tool,' etc., but just super.

1:10:18.040 --> 1:10:18.966
<v A>Super generic. Um.

1:10:19.590 --> 1:11:53.880
<v B>And you know, in the beginning it's kind of scary, right? And so that's why by default most of the tools kind of ask for your permission if you're going to patch a file, etc. Um but then over time I feel people have just gotten more and more comfortable where I think most people probably run with like a Workspace YOLO kind of mode where hey, as long as you're in my codebase, you can read and write whatever you want, and you don't have to prompt me all the time. And so people are getting more comfortable with it. Um I was using it recently to, you know, create notes—notes where basically I wanted a midterm kind of help guide for my students. And so I said, 'Hey, you know, here's the textbook and here's the midterm. You know, come up with sort of a help guide.' And kind of what Patrick was saying in the beginning, the help guide was okay but, you know, not great. And I realized I was kind of asking it to do a lot. You know, it's a big textbook, it's a big midterm. And I said, 'Hey, break the textbook up into chapters and give me a study guide for each chapter.' And that kind of forced it to, you know, use many tasks, and each task having only to read a chapter.' You know, as we talked about, the model gets dumber if you give it more information. So because each subtask only had to think about a chapter, I actually got a much better summary just from changing the prompt.

1:11:53.880 --> 1:13:42.934
<v A>Yeah, I mean, I guess that's like a good transition to our next thing, like a set of maybe like learnings and guidelines. Um I do agree with the general use and the bright future, but I mean, I guess like the first one—and we kind of talked about this with hallucinating—with needing to break down, you know, sort of plans (we're going to talk about some of these in a minute) is at least if you're working in code. But even if you're not, like even if you're working, you know, I've been playing around maybe we'll talk about in a future episode, but like using a lot of markdown files and something like Obsidian to track stuff and to track thinking and doing basic file manipulation, but use Git. Even if you're not going to like push it up to GitHub—even it doesn't matter. Like I will say, like a lot of these tools have skills natively to do like Git worktrees, which I will not lie, I have never used a Git worktree in my life. I don't know how it works. I don't know how it works either. Okay, good. I'm not knowing, but like sophisticated ways of using Git—they know how to do it. You just have to tell them you want it done right. And then anytime it's sort of like working, check it in. If you make a change, check it in. Like you can always go back, and then you can always tell the LM like if you want to try rolling forward—which we'll talk about in a future tip—you can say things like, 'Hey, like this last change didn't work right,' and then it will be like, 'Oh, okay. Yeah, let me revert it,' or 'Let me, you know, look at the diff,' right? And so it has access. So Git is like something we talk about kind of, you know, basic technical person software engineer skill, but definitely here and again I use this for personal workflows without any remote repository. It's just a local thing; doesn't matter. It's just to have like a history that is like archived of my instructions of, you know, what I was doing.

1:13:43.238 --> 1:13:53.880
<v B>Of my code. Yeah, the agent is amazing at writing detailed Git commit messages. So you just tell it, 'Hey, do a Git commit here of all the untracked files and with a nice message.'

1:13:55.320 --> 1:15:51.240
<v A>And if you do that in your like session, it will often include notes from the session, like things you were trying to do that wouldn't necessarily be inferable from the code. Um and then I guess that'll go to my next thing is before you start, you know, attempting—I guess like the term is sort of broadly 'vibe coding,' where like to JSON's what you try not to actually write code. Um although that term is I think a little derogatory, but like whatever. And people have we can talk about the ethics or morality of it some other time, but you know it just is here. So okay. Um but I think there's like the similar term is 'one-shotting,' which is like, 'Hey, make me a website,' and you just like let it go. You don't do any that may be a way—I don't know if it works amazingly today. Um but what I will say is all the tools I've used have like a planning, you know, sort of flow. And the planning flow—the, you know, tool is not writing anything to disk; it's not making any changes. It is simply attempting to break down your request into what it's going to do to actually go ahead and extract most of the needed context. And it used to be that it was just telling it like 'Think harder' and like make a task list, and then it would go to the task list. More recently I've seen the tools get better, and basically what they'll do is, you know, 'Hey, I want to add this feature,' 'I want to build this thing.' It will go like do the research, put snippets of code, and then you can review it of varying lengths and spend time going two, three, four times. The better you get that, the better like the output will be, and you can ask questions like, 'Why did you do that if it's in a domain you don't know about?' Or like just force it to reconsider, like 'Is this the only option?' Um there's like lots of things you can do, and there's skills that will also help to, although sometimes those skills go out of date as the harnesses roll forward in terms of like do you need to do those extra steps or not. But then once you sort of get through the planning, a lot of.

1:15:51.240 --> 1:16:10.641
<v A>them will then clear the context and keep only the like I have minimized the description of the work to be done, including function calls, pertinent snippets, and all of that. And what you'll find is the execution runs much faster and is less likely to run out of context while doing the thing you described in your planning. Yeah.

1:16:11.789 --> 1:18:09.627
<v B>I mean, so we should talk about running out of context. So, you know, as we talked about earlier, your model has a context window, which means the model can only take in so many tokens at a time. Um it's not a recurrent model, so it doesn't sort of accumulate anything. And so every time it, you know, reads in a question or a tool call result, it reads in the entire context and then makes a decision. And so because of that, the context has to have a limit. Um when the model hits the limit or gets close to it, it does what's called compacting, and it sucks. And and as of today it hasn't really gotten much better. Um so basically the way compacting works is pretty simple. Um the context is made up of a list of messages. So a message might be something you asked it. A message could be part of its response because, you know, now you have this multi-question answer session, right? A message could also be a result from a tool. A message could be a request to call a tool. These are all messages—this whole ledger of things that have happened, right? And so in compaction, you know, it's going to preserve some of the things that are really important. So when you told it something and it responded, there's probably enough important information that those will get kept in their entirety. But when a tool responds with like an entire PDF and it read the whole thing, it's going to try to summarize that. So it's going to summarize that whole PDF, you know, down to maybe a paragraph or two. And so compaction is all about sort of summarization, and then, you know, continue to summarize until you have 90% of your context window open again. Um the.

1:18:11.567 --> 1:18:43.480
<v B>Problem with the summarization is that details often really matter. And so um when you compact, often you're in this really weird state. Um I've seen situations where I ask a question, it compacts, and then it actually answers like the previous question I asked again. Um That's just probably maybe a bug; I don't even know. Um But there's there's weirdness. So you really want to avoid compaction, and you know task decomposition is the best way to do it.

1:18:43.480 --> 1:18:51.400
<v A>Yes, as another pointer notes, but yes, you know decomposition is a skill that I think.

1:18:54.839 --> 1:19:49.480
<v A>Both like for the purposes of what we're talking about, like getting the mind of, but also more generally just like separation of concerns, like object-oriented things. Like things do not go out the window, but in a time when the LLM—I will say currently—tends to generate a lot of spaghetti code. It's not pure spaghetti code, but it also doesn't spend cycles is the wrong way to reason about. Like it's not true, but whatever. It doesn't spend tokens trying to think through, like you know finding duplicate. It just is happy to copy and paste stuff, right? And not think, like, oh, I should move this out to a common predecessor function or pre-pro. Like you'll a lot of that stuff won't happen. It will also happily do super inefficient things, and the more—more things for a unit. I don't like that is trying to do the worse that becomes so really letting it be step by step, even more than maybe it's step by step, is very important.

1:19:50.978 --> 1:21:48.519
<v B>Yeah, totally. Um kind of related to this, there's now the file is a little different. It hasn't standardized yet, but for most of these, the file is called agents.md, and agents has to be all caps. This is like a magic file, so whenever you—um whenever you start Claude Code in a directory, it's going to look for an agents.md file, and it's actually going to recurse back through the file tree and looking for all of them and add them all up. So um so here's where you can put things like anytime you change the code, you should—a significant change to the code, you should run unit tests, and here's how you run them, or don't use the system version of Python because we have a virtual environment. Use the virtual environment Python. That's these are all instructions that you put um in your agents.md file, and that way you don't have to say it every single time you have a question. So think of whatever you put there as being sort of pre-pended to any question that you could ask. Um and that's been extremely important. Um, you know, if you want to have a code that doesn't have all sorts of lint errors and just, you know, really kind of cruddy design, then that's where you kind of enforce all of that. Um over time, I've been making this more and more complicated. Um I have one now where, you know, if a file gets too big, it breaks it up, and it's sort of instructed to, you know, break the file up semantically, group the file—the code semantically. Uh originally, I said, 'Hey, if a file gets too big, break it up,' and I was ending up with like cube_part one.py, cube_part two.py.' So then I had to go and add, you know, no, okay, when you break it up, it has to be semantically—has to be meaning behind each of the file names. Um so these are all the kind of things.

1:21:48.519 --> 1:22:05.960
<v B>You'll you'll do as you're iterating through, but you know, you can end up with something that granted it's going to burn more tokens, but but it kind of keeps a healthy ecosystem so that if you do end up having to go back and look at that code, you could do it. Um I had an issue recently where

1:22:07.640 --> 1:22:29.400
<v B>um there was a config parameter. I wanted to configure the size of the output video, and so I said, 'Hey, you know, the size of the output video is too small. Add a parameter and by default make it eight' so that the video is eight times as large. So it's like, sure, done. And I do a run, and the video is the same size.' So I said, 'Hey,'

1:22:30.920 --> 1:23:22.557
<v B>you know, the the video is uh um you know it's only, you know, this many pixels by this many. It should have been bigger.' And um it was actually arguing to me. It was like, 'Well, you know, the domain might have gotten smaller,' and so I actually made the video bigger, but it was now smaller, and so it canceled out. I was like, 'Well, wait a minute, wait a minute. Run all the, you know, lint tests and everything.' And then come back, and it's like, 'Oh yeah, yeah, I didn't pass this variable through,' and blah blah blah. So so like having good code hygiene is important even for the AI, and so even if you're doing a side project, just keep around an agents.md file and dump it into every project you do so that your AI will have good code hygiene, otherwise it will make mistakes, just like a

1:23:23.417 --> 1:25:17.560
<v A>person. I think maybe one day this will be different, but just like your analogy earlier where we pass all our code through a compiler and maybe now we pass all of our instructions through, you know, an agent to to code all of our stuff. I do think some classic debugging stuff is kind of like just like those skills that people have learned, and I don't know, it'll be tougher to learn them in the future perhaps, but I feel like are still like you're like how you're saying when you see a problem, it's like, 'Well, what—how do I get you to agree with me that why does that come?' That comes because you are a manager and you've had to argue with, you know, individual contributors under you before. Like, 'Why don't you write a test right if you just don't go look for the bug? There's gonna be like, I didn't see one.' Like, 'Did you really try?' Like, 'Let's write a test. Write a test change of—oh oh yeah, yeah, you're right.' Okay. You know, like those skills are still, you know, understanding where the most likely source of problems are, especially in your domain and in your experience, right? Like I feel at least for the time being are going to continue to be, you know, sort of very useful. And then also there's different ways of building a system, but certain ones sort of stack together nicely, and so can we—can like the LLM training ever really be equally good across like all the different ways of developing? Right? Like imagine one person wants to develop in an Agile way and one person in a Waterfall method and one per. Like is it really going to be equally good at all? It feels unlikely. Yeah. Um and so depending on which, but you may still want to stick in your lane or your company has policy or whatever, right? Like I—it's very interesting if if it's just like your code sucks, I rewrote all of it. Well, hang on, how am I gonna like get buy-off for that? Like how am I gonna get acceptance testing? Um Right, you know, this is all very difficult difficult questions. Um It does lead to your, you know, sort of debugging example. One of the things that I hear a lot of back and forth—I try a little bit of both—but it's just something to be

1:25:20.125 --> 1:26:39.977
<v A>aware of on different levels. There's like rolling forward with the issue and then just restarting, and that restart can be like we talked about reverting to an earlier Git commit. It can be which some tools support better than others. It can be clearing your session or closing it out and starting it again, which erases your history basically unless you know force resume it because sometimes something in the context window is tripping it up or polluting it and it's just stuck in some weird loop, and literally just exiting and starting again sounds stupid and asking the same question. Now we haven't talked about it. It is still kind of slow. Like it sucks. Like it generally takes a long time even if you're asking a simple question, um which is not how it is with humans. Like if you were just having a debate with a human, like you know that they—you ask them something simple, they're going to respond very quickly. Um but you do need to be aware of the difference between sort of adjusting your questions trying to fix bugs and just like nuking it and starting back and saying, 'Like let's try this again,' but I'm going to state it a little different.' And being aware that even if you copy and pasted the same prompt, you're not actually going to get the same answer because there is randomness in the LLM that's inserted on purpose through things like temperature and other things. So I don't—I haven't tested it, but I would not be so—I would assume that if you paste in the same question twice from the same base state, you aren't guaranteed.

1:26:39.977 --> 1:26:42.053
<v B>To get the same? Oh, totally. Okay. I'm like,

1:26:42.964 --> 1:26:49.343
<v A>Not crazy. I—I would. I know that's true in the chatbots, but I haven't tried it in Claude Code or Codex or

1:26:49.343 --> 1:27:26.080
<v B>anything. Even if you set the temperature to zero and you fix the random seed and all that, you're still not going to get a deterministic answer because your question is batched with other people's questions, and the low-level CUDA operators will give slightly different answers depending on what data is around you. And so—so yeah, there's no way. Zero way you can guarantee determinism. Ah, I learned, yeah, unless you're running on your own hardware with a batch size of one in a very efficient way, and you're waiting

1:27:27.531 --> 1:27:29.000
<v A>a day for every question.

1:27:30.400 --> 1:27:55.544
<v B>Yeah, yeah, exactly. Um okay, so we're just want to wrap up on this. You know, I get asked this a lot. You know, is software engineering dead? How is this going to work? Well, I have a job after we graduate. Uh, you know, I teach at UT now, and so tomorrow I have office hours, and I guarantee you tomorrow someone's going to ask me this. I know this because they emailed me today saying they're

1:27:55.544 --> 1:27:57.872
<v A>Going to ask me that? That's not fair, but

1:27:58.429 --> 1:29:17.800
<v B>But I get this every single week. And so um, you know, my—I guess I'll give my overall take, and I'd love to hear your take, Patrick. Um, I think that, you know, using a systematic way to solve problems that is always going to be in demand, right? The demand for doing that has done nothing but go up. You know, I'd like to think that my great-great ancestors were also engineers, like building catapults and stuff. Maybe they weren't; maybe they were tyrants. I don't know. Maybe they were probably just died of dysentery, dude. But before they died of dysentery, they were making some pretty kick-ass trebuchets. That's what I'd like to believe. But we actually work more hours than they did, right? Like we work more hours than medieval peasants. And so, clearly the demand for the—the demand for like systematic thinking and turning, uh, you know, abstract problems into like really concrete problems—that's not going away. What what is going away is like the very rote, you know, like I have this project in Java and I need to port it to C++ and that's going to take two years. Like that's going away. And so

1:29:19.640 --> 1:29:53.382
<v B>Um, and so you know, maybe the supply of software engineering jobs actually does take a permanent hit. Um, but that was never really the point. The point wasn't really to write software. The point was to build something cool. And if that as long as we stay true to that, then there's always going to be a ton of opportunity. Um and if if you know, we have thousands of years of history showing that, you know, it's just going to get more busy for everybody. So

1:29:54.934 --> 1:30:08.603
<v A>Oof, yeah. I my—we sort of talked about it. I think like curiosity building. I mean, I think these are things you see throughout history, like certain people had um and the way they applied it was maybe different opportunities for applying it were maybe different. Um but

1:30:11.219 --> 1:32:10.920
<v A>I mean, I think of all the things in the world that you can build having—maybe call it style—like having the persnickety-ness to keep through working and enforcing that it doesn't work the first time or didn't match what you expected. Like the question would be, I guess we were talking about Sora earlier. Someone was saying, 'You know, oh, Sora will kill Hollywood,' right? Like everybody will just watch their own movie. I don't actually think that's true. I think for me the most likely outcome is a new—maybe order of magnitude, maybe two orders of magnitude. Literally 100 more creators and movies get made. But there are still ones that just like—like you were talking about Drip Boards. I wouldn't have had that idea, but I enjoyed it. Yeah. And so therefore, like somebody had taste, like someone to call it whatever you want. I don't know. They had ambition to like prompt it to do that thing that I never would have. And I think the same will be true in software. Deciding what to build now may not be that—it may not be. I don't even say like fair in the same way that you just go to college, like get a good job and you're guaranteed in the same way that like, you know, you can hear the criticisms about, you know, 40 years ago, you did A, B, and C, and you like, you know, worked X years and then retired, and it was a good life. And that may not be true anymore. Maybe not. I'm not here to say—I'm not trying to get on politics. It's like there's no guarantee. Like you're saying, 'You know, medieval peasants,' it's completely different than us today. You know, like the things change. And so my take on software engineering is the skill set feels valuable: understanding how systems work, how they get built, how to debug stuff, right? Like this feels useful. But maybe there's less of a discrimination between learning it as a Mechanical Engineer, an Electrical Engineer, a Software Engineer—maybe a lot like the lines blur because you can reach further outside your domain. But maybe there's a big difference between how those people

1:32:10.920 --> 1:32:43.313
<v A>interact with tools? Like Cloud Code vs. difference between them than working with an IDE? But then maybe someone who's a Journalism student works with Cloud Code as well, but in a fundamentally sort of different interaction model, right? They're building something very different. They're building smaller things. They're building what you know, and maybe they don't want to buy the thing you're building, but maybe you're building for an end product. Or we—I think there's probably going to be seismic shifts, but it's very difficult to predict a specific shift. Yeah, that was a lot of words to

1:32:43.313 --> 1:32:44.520
<v B>say. Very little. Boom.

1:32:46.280 --> 1:34:16.024
<v B>Done. Throw the gauntlet down. Um yeah, no. I think that makes a lot of sense. I mean, I think that you know for teachers, for all these folks to use Cloud Code for their specific task. Oh, here's another thing. You know, we seem to have forgotten the importance of data and particularly data around user experience. You know, like your people say, 'Oh, I'll just make Salesforce myself for my company,' and I'll save my company $100k a year.' It's like, yeah, but Salesforce isn't just like a database and some front-end code and back-end code. It's also like a decade plus of user research. Like, oh, you know, we—we allowed everyone to just add and delete people in Salesforce, and then some rogue employee just got pissed and wiped the whole Salesforce database for one of our clients. And we learned not to do that anymore. We learned, 'Oh, we actually need our back—we need role-based access control,' so that Joe Schmo, who you know is like entry level, who's an intern for the summer, you know, when his internship is over, that the last day he can't just go and download like all the personal information out of Salesforce, right? So like that's a lesson they learned. And there's probably—there's definitely like, you know, they probably learned like 10 lessons a day for 10 years. They learned like tons of tons of lessons. And the AI is not going to have all those lessons because

1:34:16.024 --> 1:34:18.876
<v A>you know, it's not going to be in the GitHub comments, you know.

1:34:19.163 --> 1:34:51.867
<v B>It's like—like Joe Schmo did wipe the database for ExxonMobil on his last intern day, and that's why this code is here. Like that's not there. So so I don't think SaaS companies are dead. I don't think yet. I think I know companies are like we're going to cancel our Salesforce contract and in-house it—I expect that to be a complete and utter disaster all sorts of weird security problems over the next 12 months, and ultimately for SaaS to make a comeback. And I'm not even in SaaS, but but I have no horse in the race, but that's where I see

1:34:52.390 --> 1:35:50.811
<v A>it going. I also feel like it's a focus thing. Like do you Salesforce is expensive, but I think people are maybe too worried about some of the expenses. It's like if you can have focus and just outperform it, right? Like is that really what you want people spending time on? Like even if it's easy—I don't know. I mean, I don't know how much it costs, but it's like maybe it is for a certain size company or people like doing very low margin work, but growing your margins is probably like finding a way to be more effective in some new area rather than just like reduce costs. There's like always that balance, that tension in business. Yeah, and you're right. I feel like it's overblown that suddenly it'll just make sense for everyone to do it. Some people will—like the total addressable market will probably go down when there's lots of companies who are just paying too much per seat, right? Like it's it's the total amount they spend. They can employ a small group of engineers to do this, but then in a bunch it just—it takes away.

1:35:51.469 --> 1:36:07.095
<v B>Focus right. Yeah, because when you say re-implement Salesforce, that's not good enough. You know what I mean? Like you can't just go to an engineer and say, 'Do that,' um, because Salesforce is enormous. So you probably don't need all of it. And even if you did need all of it, how would you faithfully reproduce?

1:36:08.614 --> 1:36:24.360
<v B>it um so chances are they say redo salesforce and you end up with some janky thing that can't handle load and just doesn't behave the way you expect and guess what has no documentation either or if it does it's ai generated and it hasn't been looked at and reviewed

1:36:25.590 --> 1:36:33.555
<v A>by a person this is broken talk to claude talk to dr gpt there we go open open ai ring ring ring sam altman can you fix my

1:36:35.952 --> 1:37:19.894
<v B>Oh my gosh. So yeah, I mean, I'm not a betting person, but I would probably try to buy the dip on SaaS. I don't—again, I don't know where it is. I wouldn't try to convince anyone that they could time the market, but it just feels like feels like SaaS is a little underrated at the moment. All right. So yeah, we will definitely cover OpenClause. Someone requested that. There's been a lot of requests for things adjacent to this. We wanted to really set the foundation, talk about the history behind this and build the sort of first floor of this tower that we're building on this topic. So hope you all appreciated it. And yeah, we'll catch y'all later. All right.

1:37:20.063 --> 1:37:20.840
<v A>See you next time.

1:37:38.001 --> 1:37:40.009
<v A>Music by Eric Barndorler.

1:37:41.612 --> 1:38:00.040
<v B>Programming Throwdown is distributed under a Creative Commons Attribution-ShareAlike 2.0 license. You're free to share, copy, distribute, transmit the work, to remix, adapt the work, but you must provide attribution to Patrick and I, and ShareAlike, and Kind.

