WEBVTT

0:00:15.184 --> 0:00:21.512
<v A>Welcome everyone to another.

0:00:22.457 --> 0:00:33.879
<v B>Episode. We have a great topic today. I'm excited to learn. It's legitimately something that I hear all about, but I don't know too much about. So we're going to have teacher Jason joining us here in a minute.

0:00:34.019 --> 0:00:40.820
<v A>That's right. We had teacher Patrick for Compilers and Interpreters. And then we have teacher Jason for VectorDBs.

0:00:41.200 --> 0:00:44.200
<v B>I feel like we should have called ourselves professors. We have missed opportunity there.

0:00:44.420 --> 0:00:50.200
<v A>Oh, that's right. Is there anything higher? Distinguished, Emeritus, Professor, Emeritus?

0:00:52.340 --> 0:00:58.047
<v A>Okay. What's the highest level professor title? ChatGPT. Can

0:00:58.047 --> 0:00:58.114
<v B>you?

0:01:00.240 --> 0:01:02.760
<v B>Soul Emperor of the Knowledge Universe.

0:01:05.200 --> 0:01:13.020
<v B>Oh, but what if it thinks it's the smartest? It's going to tell you a lie because like, it doesn't want you to be superior. Okay. Anyway, no, sorry.

0:01:13.020 --> 0:01:34.280
<v A>Oh, you know, I have this plan now when, you know, when I get these calls from people telling me that like my taxes are overdue or I should buy car insurance or whatever. What I've started doing is asking it questions about particle physics. And when it gives me really good answers, I'm like, 'Aha, this is ChatGPT. This is not even a real person.'

0:01:34.640 --> 0:01:36.860
<v B>A particle physicist could work in a call center.

0:01:41.000 --> 0:02:10.103
<v B>Okay. Yeah. I mean, I had to go looking for a car and they legitimately like, you know, want your phone number or whatever. So I just have one of the, like, you know, a voiceover IP number, but it doesn't actually like ring through or whatever. And the place, one of the places I went has called me twice a day, every day for going on two weeks now. And I've picked up zero of the calls and I'm just like, they really want to sell me a car, but it turns me off. I'm just like, why are you pestering me this much? Like, I feel like just

0:02:10.103 --> 0:03:06.640
<v A>Text me. I will. So, you know, my dad worked for a few different car dealerships over the span of about 30 years. And so I learned the trick to buying a car and I'll share it with everybody. If you're a Programming Throwdown listener for free, you don't have to be a patron, although we really appreciate it. If you want to buy a car, here's what you do: You show up, you test drive the car in person so they know you're a real person. You're not another dealer trying to buy a car and then resell it. You know, you're there as an individual to buy a car. You prove it to them, right? And then you leave and you don't come back until you have everything finalized. So everything is over email or phone. You negotiate the price. You compare with other dealers and you should expect this to take about a month. But if you do that, you will save anywhere from like 10 to 30% of the price.

0:03:07.740 --> 0:03:42.600
<v B>Wow. The last time I bought my car, I felt like it was pretty different than now. And that like the inventories are much worse. But yeah, in general, I did something similar to what you were saying for the last time about a car. It was super nice. I just showed up and all the paperwork was done, you know? Just basically it was still took, you know, close to an hour, which is ridiculous. But you know, pretty much it was all ready to go. So yeah, that I agree. I'm going to attempt to do this one again. I haven't gotten to the negotiating stage yet. So, but showing up, I guess they're trying to, you know, do the cap show with me. So they're trying to call me back and make sure the phone number is real.

0:03:42.600 --> 0:03:51.840
<v A>So yeah, I haven't bought a car since 2020. And so I've, I've missed this shortage, but is the shortage still a thing or is it over?

0:03:52.320 --> 0:04:38.140
<v B>I don't think it's a thing in the same way, but the inventories on like desirable stuff is still not great. And at least at the places I was looking, I think it varies dealer to dealer how much they've, you know, or brand to brand, I guess how much they've recovered. So yeah, I think it's, it's still on the pricing is tighter, but I, I don't know. I will still die on the Hill that I think we are stupid for having gotten ourselves. And I guess this is us centric. People tell me in other countries, it's not like this, but I have no idea why you go in and the only place in life remaining is car dealership where you have to like negotiate. I'm going to bring in my manager. I'm going to bring in the finance. We're going to change everything on you. Like it's a total like scammy thing the whole way through.

0:04:38.520 --> 0:05:42.580
<v A>Oh yeah. That's another thing. So the thing about finance, you ever play these mobile games where there's like 24 different currencies, you know, like I play FIFA on mobile and there's like, there's literally five to 10 different currencies at any given time. And some of them get phased out and replaced with other currencies and all that. But there's one currency that like you could pay real money to get. And that's ultimately the only currency that really matters because that's the one that's tied to USD, right? And so if you start going down the finance route early, you end up in this boat where they're like, Oh, your monthly payments are the same, but like they really, they added a year of payments, you know, but each month is the same. And there's just, there's just like too many variables and it's like pretty easy to get trick. And so that's, that's another thing that my dad told me was like always negotiate the absolute dollar amount and then come in at the very end and do financing after like you're ready to pay the amount.

0:05:42.580 --> 0:06:11.800
<v B>So this has turned into a little bit of advice from Jason and a lot of whining from Patrick. But yeah, buying a car. I look forward to, I will say from what I understand, Tesla does it a little different. I look forward to, you know, some future for my children or whatever, where like you literally just figure it out online. Maybe you show up to test drive once and then you just buy. And it doesn't matter if you buy from the one on the West side of town or East side of town, it's all the same, like every other thing we go to and do so.

0:06:13.280 --> 0:06:14.300
<v A>Yeah, exactly.

0:06:14.840 --> 0:06:24.780
<v B>Can't come soon enough. All right. And I'm sorry for all the car dealership people that I probably, you know, very frustrating to or putting out of business, but you know.

0:06:25.260 --> 0:06:32.780
<v A>Yeah. If you're a car dealer, message us. We'll take we'll take your hate and read it on the air for everyone to hear.

0:06:33.099 --> 0:08:29.380
<v B>Oh no. Oh no. All right. We're jumping into news. All right. So my first article talking about programming things here, cognitive load is what matters. So this is actually a GitHub article, which is a kind of interesting way to do it. So the person can update it over time. And this is something that I picked, I will say a little selfishly because I say this all the time at work and there'll be a link in the show notes. But just talking about in development that one of the most overlooked things is how much cognitive load you're putting on the engineers. And this can come in a couple of forms and they kind of call them out. Like sometimes that could be the business surrounding the code. I work at a place that does lots of software engineering. So it's a pretty respected part of the business, but a lot of businesses, you may only have one, two, three, four software engineers building something. And so the bulk of the business isn't really the development or the software. And so there's all sorts of like, Hey, you really need to work with the accounting team and this team and that team. And who talks to whom, and this is using this database and this person is using this other one. All of that has to be loaded into your head. The other thing I always bring up that people forget is, you go to college and you learn all this stuff about compilers and programming and you show up on the job and they give you like an unconfigured Linux box. And maybe you've only ever used Windows or you've only ever used Mac OS. And now it's like, Hey, here's this Linux box. Congratulations. Like configure it, install your packages. Oh, you need to like, you know, move your PC to a different location and plug in all the cables. But wait a minute, like, is that really like a requirement listed for the job that like, you know, Bash or ZSH or any of these shell scripting? And so the amount of cognitive load that goes into programming and then, I think kind of the bulk of the article is when you're selecting paradigms.

0:08:29.380 --> 0:10:20.500
<v B>abstractions, where to divide stuff, how much to copy and paste, there's a balance to be reached. If you make stuff too small and modular or too monolithic, both introduce a lot of cognitive overload. So there's the balance there, but actively thinking about how do we build this in a way that makes it easy to reason about. So things that are similar are kept together, and not abstractions for abstraction's sake is an important thing. The one that always gets up, like we program in C++ and it's like everybody, aha, look, I wrote the ternary operator, which is you write a conditional expression, like an if statement, you know, a greater than B, and then you do a question mark. And then the first thing is what happens if it's true. And then you do a colon and then you put the second thing. Okay. Just write an if else statement or an if then's like, like you are, I have seen, I've been programming for a long time. I see it. Yes, I can parse it, but I don't parse it as fast. And I always have to double check. Like, wait a minute. It's the true is the first one, false is the second one. And secondly, when we bring on a new person, maybe you say, Hey, we're all experienced people. And this comes up in a later news article, but when new people come on, how like, maybe they aren't used to that. So using and this is matters for some languages more than other. But using like the subset of the language that is most comprehensible and most widely used gives you benefits. And that it's very easy to bring new people on. It's very easy to teach. And we were working with a code base where they've like put in a ton of macros, all of the, you know, variable names are single line, single character variable names, all this stuff. And they may be efficient in programming, but you cannot work with them. You can't work with that code base. And all of this is about that, you know, thinking about cognitive load and really programming in a way that reduces it.

0:10:21.040 --> 0:11:23.040
<v A>Yeah. I've seen so much like a kind of bad behavior with macros and metaprogramming and C++. It just makes it so difficult. It's like, Oh, you know, I'm using like the double hashtag thing to like concatenate different things in the preprocessor and like, Oh, I'm auto generating classes. And now I have this, this call to a preprocessor macro to create a class. And I'm doing that call like 30 times. And so none of these classes actually exist anywhere. So like, you know, it messes up the IDE and everything. And at some point it's like, look, if you were to copy paste this class 30 times, well, first of all, there should be a better way of doing it. But let's say you just couldn't figure it out. Just copy paste the class 30 times. At least that way someone can see kind of the point and the pattern and all of that. If you hide it behind a preprocessor macro, it just becomes like a tar pit that people get stuck in.

0:11:23.660 --> 0:12:02.000
<v B>And I think there are other languages that the same thing happens, or even, you know, the way of doing object modeling. And I've seen stuff in Java with, you know, getting into dependency injection and reflection at runtime, and just like, it makes it really hard to like, 'Hey, I have a call site here. What code is being executed?' Like how many hoops do you have to jump through if you don't know off the top of your head to find the actual line of code invoked, and how confident are you that you found the actual line that is being invoked? And if that is not a trivial amount of time, you should be really, really thoughtful about trying to reduce that as much as you can.

0:12:02.900 --> 0:12:39.680
<v A>Yep. Yep. Yeah, totally. Yeah. Dependency injection is in Java. I remember—I forgot what the tool was called—but there was this tool where like you would, maybe it was part of the Spring Framework, but you would pass in like a config and it would create classes from this config and feed them into your main function. And so it's like, 'Oh, like this object is already filled in. What the heck? Like, how do I change it?' Oh, it's actually in an XML file. It's like, what, you know, it just became really hard to use.

0:12:42.820 --> 0:14:41.175
<v A>Moving on to diffusion models are real-time game engines. I don't know, I don't necessarily agree with the premise of this, but the thing about this that blew my mind is the video. So basically what people did is they took these like image generators and text-to-image, you know, that whole like a Vision Transformer kind of technology. And they basically used it to say, 'If I have this image of a screenshot of a video game,' in this case, they use Doom because it's open source and people use it on everything. So given the screenshot of a Doom game, or maybe even like a handful of screenshots of the last second of this Doom game, and given a command, like shoot a gun or go forward or whatever, then generate the next screenshot of the game. And so they basically are trying to encode the entire Doom game, like the art assets, the engine, everything into a diffusion model. And so then they proceed to play this game and it just gets—it just gets kind of weird. Like, you know, points of discontinuity are really difficult for these models. For example, when you fire the weapon, sometimes it doesn't fire. You get the fire animation, but like, you know, the world isn't affected. And other times it is. Sometimes like you walk in a door is just kind of like flickering in and out of existence. The whole thing is wild. We're not doing it just this year over audio, but go to the show notes or watch the video. It's just pretty—I've seen a lot of weird things with Doom. Like they passed it through DreamDayem and

0:14:41.530 --> 0:15:07.939
<v A>you have this like really weird, you know, kind of like trippy kind of view of it. But this is really taking it to the next level. Like you're defying the laws of the game are not, are not, are not first-order anymore. So it's like sometimes the world works the way you expect. Sometimes it doesn't. And it's just kind of—it's just kind of insane to watch. So there's

0:15:09.188 --> 0:15:15.880
<v B>a couple games that kind of do something that reminds me of. So I actually have seen this one before.

0:15:18.098 --> 0:15:21.490
<v B>And like you're saying it's kind of weird because you—it

0:15:23.869 --> 0:16:32.000
<v B>looks like you're playing the game 90% of the time or whatever, but then occasionally something happens that is different. So like you're looking one way, you turn around and face the other way. And then when you turn back, the world is completely different. Like you're not in the same place or like the world has shifted because it's just—it's generative. So it's just doing the whole thing sort of like on the fly. So you almost convince yourself it's right. And there's a game, I think it's Superliminal. I don't know if you've played this that has stuff like this, like you can go through a door and then when you come back through the door, like you're not in the same place. It's sort of—I don't know what to call—like non-Euclidean or whatever, like, you know, it, and it kind of plays with the sort of like perspective. And so like if something changes size in your perspective, it changes size in real life as well in the game as well. And so there's all these things you can do where it sort of subverts your expectation of a game modeling reality. And it doesn't mean it's not fun to play. It's just—it can be very fun. I don't know in these cases, maybe it probably wouldn't be very fun. But it is; this is like this weird twist to your brain.

0:16:32.979 --> 0:16:58.680
<v A>Yeah, totally. I'm pretty sure you can't get much beyond the first or second level in this. I mean, this is like some academic exercise. But yeah, I just think there's something—there's something just kind of wild. I mean, it's definitely something I've never seen before where someone tries to like imitate an entire engine of a game with a neural network and then see what happens.

0:17:00.360 --> 0:17:00.968
<v B>People with too much

0:17:00.968 --> 0:17:13.119
<v A>time on their hands. No, I've just—well, you know, people with too much money in their hands, right? Because of the AI hype cycle. Oh, there we go. Well, yeah, I mean, that's—yeah, we should, maybe that should be episode 178.

0:17:13.119 --> 0:17:16.839
<v B>So we'll get sent some our way. We'll try some of this stuff.

0:17:17.500 --> 0:17:31.459
<v A>Yeah, that's right. Yeah. We're currently raising—we have a new model. It's called Claude, C-L-A-W-E-D. It absolutely does not just call Claude under the hood. Please give us $50 million.

0:17:31.719 --> 0:19:30.560
<v B>Oh, I—yeah, I've seen some of these accusations going around about people just calling other models and saying, 'Yeah, okay, anyway, we're moving on.' This is the AI gossip, I guess. Yeah, that's right. My next article is kind of related to the previous one. Your company needs junior developers. And so again, I guess a bit bias is something I push for all the time. It's a strongly held personal belief. And so, the idea here, a lot of places, I think you can fall into the trap of: 'Hey, we want to hire a senior person.' The inside joke, I guess it goes around, is like, you read resumes and it's like, 'Hey, you need 10 years of experience with a tool that came out two years ago.' And it was like everybody needs to hire the utmost expert in the field. In practice, that's not really true. Most of the time that isn't—it's not to say that there's never a good idea to hire senior people, but I feel by default most folks have this belief that like, 'Hey, we're high velocity. The way to keep this up is just keep adding people at the same level of the team or above the level of the team to bring them up.' And in some cases that may be true, but I have found that often there's a number of first and second-order effects that come from hiring junior developers. And I'm not even talking about like budget, you know, money reasons. I leave all that aside: hiring junior people and being able to train them helps them sort of—they are more plastic towards how stuff is being done and learning. If they're inquisitive, it's not to say that they won't express their own opinions or bring changes. They often will, and sort of say, 'Hey, I learned it this way,' or 'Why not this?' Or, and they ask a lot of questions, and through those questions I think can come learning. But they don't come in with as many preconceived notions of 'This is the only way to do things,' or 'Every time I've done it is this way,' or 'Hey, we first need to reconfirm everything you've done to my way of doing it.'

0:19:30.560 --> 0:21:19.860
<v B>And then we'll start making progress, but also putting the more senior developers into the mode of teaching. And by teaching they're learning; they're having to synthesize concepts and simplify stuff. Like we were talking about before: cognitive burden. If you're the only one coding, maybe it doesn't matter so much because the context is always loaded in your head. But the minute you have to share that context with someone else who's also making changes, right? Now you start to help each other and improve. Also, I think making sure that the expectation is known to senior engineers—that there will be junior engineers and that they will bear some of the responsibility for helping to onboard, teach, and show them—not that they have to become managers, but that everyone on the team has a responsibility to help those who are trying to learn or asking questions. And I think that again, these first and second-order effects—that you have someone who is contributing, who can grow and can sort of become the team member that you need—also gives like a sense of a narrative arc to the team, right? Like that we were here and now we're there, and look at the growth. And so for just a host of reasons, I think it's a really important thing to be hiring junior engineers, to bring them on, to have them grow and learn. And also sometimes projects go through phases, and the project of your phase may be: 'Hey, it's not insane growth right now. We're not able to hire a ton of people because of that.' Some senior engineers may need to or want to move on to other companies or other projects or other things because there isn't a lot left for them to sort of grow, but there's still a lot to be done. And this is when junior engineers can come in, and for them, that work is super suitable. Whereas for a senior engineer, it may be difficult to justify the continued sort of growth of their career by doing that work.

0:21:20.640 --> 0:23:20.571
<v A>Yeah, totally. I mean, like clearly Patrick and I are really invested in growing engineers. I mean, the whole point of this show when we started it so many years ago was that, and it's the point still today. So, this is something that's extremely important. And personally, I worry greatly about the sort of junior engineers and the future of the field. I think that yours have been hit with some really big whammies. The first thing they've been hit by is just exploding college tuition costs, right? Then you get hit with the pandemic and all of your mentors not being around. And now there's all the layoffs and companies saying—I was in a meeting. There was an Austin engineering leadership meeting that I was at. And this person was saying that like, 'Hey, you know, we just need to get rid of like 80% of our company,' like Twitter did, and just keep like the 20% super productive people.' The thing is that will work in the short term for your company, but there's a tragedy of the commons there where you're kind of—it's like when you buy roses from HEV or from your grocery store and they look great, but they don't have roots. And so, these roses will not live very long. They're already dead. Yeah, that's right. Like, the cat is dead in the box. So, and so like, yeah, you're without roots; it's just not—it's just a temporary facade. And I do worry a lot about junior devs just not having a safe place to land. What do you think is going to happen, Patrick? Do you think that if people keep focusing on senior devs, do you think that there's going to be?

0:23:21.027 --> 0:23:26.760
<v A>Just—people are going to get dissatisfied? There's going to be a generational gap or what?

0:23:27.675 --> 0:25:26.080
<v B>I mean, leaving aside fundamental pivotal shifts from AI or whatever. I think it's a slightly different topic. I think this is where—and this will get somewhat, I guess a little maybe controversial, but I think this is one of the benefits of having capitalist markets: companies that will hire junior devs and invest in training them will have a market advantage. So, you know, two companies participating in the same... So that company that you mentioned, just as an example, like, 'Oh, and lay off all of these junior people. They're not really pulling their own weight right now. We overhired our fault,' whatever. Those people are going to go somewhere. And if not, they might self-organize and kind of learn their own thing and build up a competitor or do something better because they kind of see what wasn't working at the previous company. So I—this is where I'm a believer that having the competition companies that do this will end up winning out over companies that don't, companies are willing to teach and hire. Again, like I said, I think they have better practices. And that's not to say in some extreme example that, like you said, somebody could do this and in the short term it works amazing, but then—and hopefully it doesn't happen—but what if you have a team of super high-functioning three people who are working at 110% capacity? That doesn't make sense. I know, but they're working at a hundred percent capacity, super optimal. You think you're a genius. And then one of them—I mean, I work with people get we're getting not that old, but up in age, have a heart attack, have like, you know, a family emergency. And they literally stopped working for six months and there's no redundancy. There's no backup. Like, is your company going to close shop? Because you were—that's kind of, I guess the equivalent in financing is people do leverage. They use all this leverage, and it works great until it doesn't. And so you seem really genius. I guess is the equivalent of what you're saying about the rose is already cut. Yeah. Seems really genius.

0:25:26.180 --> 0:25:34.080
<v B>You're making all this money on leverage, but as soon as the tide turns even a little bit, you crash really fast because everything is magnified.

0:25:35.340 --> 0:26:02.560
<v A>Right. Right. And even if like you might say, 'Well, I will let other companies hire and skill up junior engineers, and then I will just hire those people when they become senior engineers.' But I think the problem with that is you need to transfer knowledge within your company. I don't—I don't think it's as fungible as people would hope. And so, I don't think that strategy will work either.

0:26:04.180 --> 0:26:23.760
<v B>Yeah. I mean, you still have to teach them the specific; most companies still have quite specific things. So even if you say, 'I want to let other people train them and then we hire them,' you don't have a culture of training people because you don't do it very often. You buy either than back, and then you're going to have a difficulty getting them to adapt. And so you're going to go through a lot more pain and suffering.

0:26:24.520 --> 0:27:41.620
<v A>Yep. Yeah, exactly. All right. So, onto mine. Mine are a couple of libraries that let you do real-time text-to-speech, which is now becoming basically an open-source commodity, which is kind of amazing. So, so you can actually and I'll add a third one to it that just came out this morning, but you can actually have a conversation with an AI—an open-source AI in real time. And that just blows my mind. Like you can talk to it. It'll talk back. You can interrupt it and say, 'Hey, I want to talk about something else.' Like all these things that we saw in that OpenAI demo about maybe six months ago, a year ago, are now becoming a commodity. And here's—here's a few examples of libraries that do this. But yeah, I think the progress is just absolutely astounding. I mean, it's actually really hard to keep up with the technology; it's just moving so fast. But this is a really big milestone, and a few different libraries have kind of reached it in more or less the same time.

0:27:44.160 --> 0:28:39.340
<v B>With the speed of some of this, and like you were joking about earlier, but I kind of find myself thinking about it sometimes: companies, there's a lot of untapped potential and sort of developing companies around some of these things and just replacing stuff within bounds. But I think defining those bounds is really important. Like, you know, making sure that there are some—I don't want to call them guardrails or whatever—around how some of this stuff is done, but also it's got to be somewhat difficult to say you're going to build and do an integration with one of these models that came out. And then to just like three or four months in potentially a completely different way of doing or a different architecture just moves the goalposts. And then you got to decide, like, 'Do we ship on this old thing?' Or do we like completely redo and move to the new thing? And then you end up in this constant cycle of—'Oh man, I just got completely blown away by whatever the next release from whomever is.'

0:28:40.740 --> 0:29:28.520
<v A>Yeah. Right. I mean, you know, a good example of this is like the text-to-image stuff where it wasn't really that commercially practical because the images were just not that good. I mean, they were fun to look at, but you'd obviously know that this was computer generated, and basically within the span of 12, maybe 18 months, it went from that to like, you literally—maybe, okay. So I take that experts can tell, but like you and I can't tell if you see a photo of a person, like a profile photo, we can't tell if that's a real person or not. And that is just wild. I mean, if your product was kind of based on one thing, like you, it just opens up a ton of product opportunities to your point.

0:29:29.080 --> 0:30:23.780
<v B>Yeah. There was all that stuff like, 'Oh, they're not good at text.' Earrings won't be symmetrical, you know, too many fingers. And you're right. They were kind of funny to look at and they looked plausible. And then now you're right. I see pictures posted in some of the sites like, 'Oh, look at this AI trash.' And I'm like, I—you know, if I really stare at it, it looks off, but there's a lot of commercial photography that's airbrushed or whatever that also can look off or not natural. So I think it's like you said, it's a sort of—whatever, maybe some experts saying these aren't good, but they actually are very convincing a lot of times. And it's really only the fact that the situation is so improbable, like, 'Oh, a cell phone snapshot of an elephant parachuting into my backyard.' Like, okay, probably that didn't actually happen. So it must be fake, but not—it's not actually something in the scene that tells you that's not realistic. Right. Right.

0:30:24.260 --> 0:31:11.960
<v A>And the thing too is, yeah, I mean, if you're in the mindset of, 'I wonder if this is real or not because someone told me it's fake,' well, then you're going to have a very discerning eye. Right? But if you're just on the internet browsing, you're not, you're not, you know, you're not in that type two mode of thinking. You're just consuming content. And I think many people will just look at fake images and just think they're real. Like all these stock images; you go to like some random startup that just paid someone on Fiverr to build a website. Right? And there are like stock images of people just having fun around a computer. I bet if you replace those with fake images, most people wouldn't know this. Yeah.

0:31:12.720 --> 0:31:37.962
<v B>Yeah. And there are some ethics problems. I will say as well, looking at everything with the discerning eye causes its own problem because there are plenty of actual images taken that look implausible or seem unrealistic as well. And there are totally legit for a variety of things. So you could enter cons conspiratorial, whatever, like, 'Yeah, I saw a video of this bad thing happening, but it's just fake. Like it doesn't look real. There's no way that really happened.' Oh,

0:31:37.962 --> 0:32:01.139
<v A>Interesting. Yeah. I mean, that's a different direction. I hadn't thought about is if there's a real event, like we're at 9/12 today. Yesterday was 9/11. And it's like, yeah, I mean, if not, if I'm reading Twitter and I see like some big terrorist attack, my first reaction might be like, 'That's a fake picture or something.' I know, it's kind of wild.

0:32:03.520 --> 0:32:10.480
<v A>All right. So moving on to Book of the Show, Patrick, what is your book of the show?

0:32:10.520 --> 0:32:11.543
<v B>Book of the Show. It's not.

0:32:11.999 --> 0:32:13.315
<v A>a book, cheater warning.

0:32:13.315 --> 0:34:09.139
<v B>Until the show is not a tool either. So, yeah, we got to rename these sections, man. All right. I thought I must've talked about this before, but I looked, I Googled through the search, the show notes and I didn't see it. So if I have, I apologize. I have never heard of this. So I think you're good. So this is one of those, I know it's like fairly popular now, but I run across fairly popular, like million-subscriber YouTube channels all the time that I've never heard about before. It's just a million isn't what it used to be or something. Yeah. It's not like an old man now. All right. Anyways. So, Thought Emporium is a guy who does, I guess, like just call like hobbyist science research. He's my hero. Like, I wish, I wish I could do this stuff. Like I wish I could make money just poking around like random stuff in my lab. And like, I want to say close to probably half of the stuff he does. Like I have at one time or another done played with or sort of, so it's very fascinating. There's a bunch of sort of one-off projects and long-term projects that he's working on, including topics like genetic engineering. He was recently trying to grow neurons to play Doom. So how do you like grow neurons in a Petri dish, hook them up to electrodes, hook that up to a computer and make them play Doom? And not your garage. That's not exactly right. Cause 'it's garage' is kind of fancy, but not in a multimillion-dollar biomedical research lab either. The one I linked is just him and his friends, like sometimes come up with memes. I think they released one yesterday. It was like, you know, everyone always jokes, 'What did you take this photo with? A potato?' And so he actually built cameras out of potatoes. Not like a potato hooked up to a camera, like the whole camera is a potato. And so like the potato is a pinhole, and you use the flesh to sort of get a negative image. Okay.

0:34:09.400 --> 0:35:03.800
<v B>Anyways, the one I linked to in the show notes, you can check it out is a single-use emergency thermite hot dog cooker. So what would you need to do if you needed to cook exactly one hot dog, like in an emergency situation in a single-use container? And it's just like him figuring out like in a moderately serious engineering, like 'How would I,' 'What is the best way to do this?' Like, 'How would I use heat to cook the hot dog? Should I cook it like by itself? Should I cook it in a liquid? You know, is this actually going to work? How safe is this? You know, could I make the, could you make a product out of it?' And so if you've never checked them out, definitely do. There was a lot of biohacking initially, like, 'You know, I'm going to inject the genetically modifying materials into myself and do crazy stuff.' Nowadays a little less, but another one that my kids really liked was is Mummy Yummy.

0:35:05.625 --> 0:35:32.080
<v B>So let's go through the full process of making like a mummified chicken from the store grocery store. Like legit, we're going to go into reading hieroglyphics and making our own hieroglyphic statements, do the full mummification process with authentic old recipes, and then taste various parts of the mummification fluids and things. And determine if mummy is yummy.

0:35:32.319 --> 0:35:34.460
<v A>Is that healthy? I mean, or is that sanitary?

0:35:34.460 --> 0:35:51.960
<v B>So nowadays it would be formaldehyde and you definitely eat formaldehyde. But it turns out in ancient Egypt, they didn't have formaldehyde. So you actually could use most of the stuff. So I won't spoil, but if you think it is intriguing to find out if mummy is yummy, you're in the right place. Okay.

0:35:52.060 --> 0:36:00.840
<v A>Yeah, this is awesome. I think this is one of these things where, like, 40 hours of my life later, I'm like, 'How did I get on this YouTube channel?'

0:36:01.360 --> 0:36:11.080
<v B>Mine's the opposite. So I watched this and then I found myself on eBay. Like, how could I replicate what this guy is doing? Like, I have so many questions.

0:36:11.680 --> 0:37:58.180
<v A>Yeah. The YouTube to Amazon pipeline is strong. My book of the show is a project that I started with a friend of mine, called Novel Minds. And the idea was there's all these public domain books. A lot of them are required reading for kids. Even the ones that aren't, a lot of them are really good. Like War of the Worlds or classics, right? So I thought, could I use AI to turn them into sort of, I think the word is graphic novels. I mean, not, not literally a comic book, but just a picture book. So there's just a lot of pictures, you know, the cover is totally redone. And so we did this. By the time folks listen to this show, there'll be 30 books on this site. I think at the time of the recording, there's maybe seven or so. But you can basically grab them all for free. If you're on Apple, it will just open an Apple Books. Naturally. If you're on Android, you'll have to get an EPUB reader, but there's plenty of them for free. And you can read all these classic books with tons of pictures. We tried to do roughly a picture every like 400 words or so. And there's a bunch of technology. I actually gave a talk on how this works, that's recorded. And I posted the talk on LinkedIn. If folks want to watch that for more technical depth. But you know, if you have kids who are in school who have to read these books, you could give your kids or your friends, your friends, kids, copies of these versions. All the text is exactly the same. We haven't changed any of the words. We just added a ton of pictures.

0:37:59.660 --> 0:38:27.700
<v B>I did not know about this until right now. So I'm actually scrolling through one at present. This is really cool. The style is very consistent. The main character doesn't seem to be—I'm looking at Around the World in 80 Days—but the main character doesn't seem to be super consistent between them. But there used to be books when I was younger. They were like, there's like every other page. There was like a hand-drawn illustration. And then they were the like synopsis. Like they were kind of, I don't know, 10% of the full book or whatever. They were like concise.

0:38:28.660 --> 0:40:26.860
<v A>Oh yeah. Yeah. So, okay. A couple of things. One, the version that Patrick you're looking at right now is an older version. Oh, okay. We did fix the consistency in the newer version. We basically use something like a vector database for foreshadowing. But we kind of encoded a representation of each of the characters. And then when a character is referenced, we added that representation back in. And so now, the version that'll be out when we publish the show actually is pretty consistent. But yeah, you're right. The first, the very first version was like thematically completely inconsistent. Like one picture would be black and white and next picture looked like a cartoon. Next picture would be like hyper-real, you know? And so getting consistency, the second phase was where we got a pretty consistent art style, but the era was all over the place. So like, you know, it'd be 80 Days Around the World, which was recorded, which was set around the 1890s, I think. But there would be like high-speed rail pictures or a guy with like a rocket launcher, you know? And so what we did there was, we also injected the era. We figured out the era by scanning the book and then we inject that into every single image prompt. So yeah, it's been a really, really fun exercise. My goal is ultimately to get these into the Kindle store, although it's pretty difficult because there's a lot of scammers who try to push either fake AI books or public domain books onto the store. And so we're running over a couple of barriers, just trying to prove to Amazon that we're legit.

0:40:27.415 --> 0:40:40.480
<v A>But yeah, that's it's a fun little book project and I've had an awesome time actually reading these books. Some of them are a little difficult to read because they're rather old. But a lot of them are awesome.

0:40:41.480 --> 0:40:58.397
<v B>This is cool. I'm trying to determine how many of these I've read before. So can I have another request? Can you like run the generator to give me summaries? So it's like all the pictures, but only 1% of the words so that I could claim to have read all of these books much more quickly than actually

0:40:58.397 --> 0:41:24.580
<v A>read all of these. Oh, that's an interesting idea. You know, I wonder, the reason why I didn't change the words is I felt like if it is required reading through someone, there's a test or something. Bad idea. But yeah, but you're right. I mean, for people who aren't reading it for a grade or anything, we could definitely summarize every paragraph or do something like that.

0:41:24.580 --> 0:42:37.977
<v B>There was a—so what I want. Okay. So this is really dumb. Okay. We can move on. But growing up, there was this show Wishbone where it was a little dog. Did you watch the show? No, I never heard of it. Okay. Oh my gosh. All right. This is good. All right. I'll make it short. So just this dog who goes to same—same idea, like various famous, you know, plays or books and the dog is like the, you know, he goes into the story and then he, you know, it's like a dog is the main character or whatever in the book, but they're just acting out the book in a 30-minute show. And so you get the arc of, you know, the book, the main idea, the gist in 30 minutes for kids just like teach them about different worlds and different, you know, a moral story or whatever, but I—okay. That won't make sense if you know, but that's what I want is like the Wishbone, like the Hey, I could hold a conversation with someone. And unless we were doing like a book club, they would believe that I had read the book and I will have benefited from the, you know, analogy or the, you know, whatever the, you know, book is trying to get without all of the dialogue that happens between, you know, characters that, like you said, if it's required reading, it may be very important to read. I'm not trying to denigrate old books, but I also know that I won't commit that time. Yeah. It's a

0:42:37.977 --> 0:42:58.919
<v A>great idea. Maybe by the time we publish a show, I'll be able to do that. I mean, one simple thing would be, you know, take each chunk, which is around, I think it's 4,000 characters, take each chunk that's being used to make an image and summarize that chunk down to like 200 characters.

0:43:01.120 --> 0:43:10.461
<v B>All right. I'm excited. This is cool. I don't know how you're going to become rich and become the financier of the biomedical research lab, but we'll

0:43:10.950 --> 0:43:13.800
<v A>work on that. I'm talking to was it Tony Stark?

0:43:16.140 --> 0:43:26.480
<v A>All right. Time for tool of the show. Yeah. If you want to help us, Tony Stark, you can subscribe to us on Patreon, patreon.com/programmingthrowdown.

0:43:26.940 --> 0:43:34.600
<v B>Okay. That was actually a jet. Sorry. I was skipping over, but that is really cool. I'm excited. I have to move on. Cause I feel bad about myself for not contributing to society.

0:43:36.180 --> 0:45:02.629
<v B>Speaking of which, my tool of the show—not being a tool—and a way to not waste time to improve your brain is Escape Simulator. I think I mentioned before playing some escape games. I recently actually picked this one up in a Humble Bundle. I didn't know kind of think too much of it, but it's actually like, it was pretty cool and just so much content in this game. So Escape Simulator, it's kind of what it says on the box. It's an escape room simulator. And they themselves have a lot of content for levels that you can play, but they also have a big fan base that is making games of various complexities for you to play. And I just really enjoyed doing it. I've been playing. I've not come anywhere close to beating it because I refuse to use hints. And my daughter sometimes sits and we play together, and you know, my kids are pretty smart, but like way to be impressed with your kids is have them shout out an answer to something, and you can't figure out why that is the answer. You try it and it works, and then have them try to explain to you like how the puzzle had that answer. And they can't; like, they just intuit it, but it happens way too much. And then you kind of feel bad and you're like trying to look at the hints to understand why anyways. So Escape Simulator, definitely check it out. It works on macOS, Linux, at least in Steam, Steam Deck, it worked, and Windows. Yeah, this is

0:45:02.629 --> 0:45:13.024
<v A>like super fun. I'm definitely going to grab, grab this. Is it like a split screen or is it internet co-op? Internet, internet.

0:45:13.024 --> 0:45:27.940
<v B>thing. So yeah, that's the only thing I couldn't figure out. I just have one Steam account. I do have two computers, but once you're playing on one Steam library, won't you play in another one? So we just sort of sit together and like collaborate on a one-player game.

0:45:28.220 --> 0:45:30.557
<v A>Got it. Yeah, that makes sense. Very

0:45:30.557 --> 0:45:34.000
<v B>cool. But I could also just buy it under two Steam accounts and that would probably work fine.

0:45:34.742 --> 0:46:03.960
<v A>You know, there's a new thing not to digress too much, but Steam has like a new feature now called Family Mode where you can actually say like this computer is a family member. And then what'll happen is you can both play different games at the same time. If you try to play the same game at the same time, Steam says, 'Hey, you're gonna have to pay, you know, to buy a second copy.' But then when you do that, that you can play together. So like, it kind of handles all of it.

0:46:04.680 --> 0:46:06.265
<v B>This is good tips. Okay. All

0:46:06.265 --> 0:46:07.615
<v A>right. I'll be back in

0:46:07.615 --> 0:46:08.100
<v B>a few minutes.

0:46:09.700 --> 0:48:08.779
<v A>My tool of the show is Cursor IDE, and you can go to cursor.com to grab this. So this is basically GitHub Copilot and all of these things just on steroids at the next level. When you start up Cursor, it asks if you want to import all of your VS Code extensions. So it's literally a fork of VS Code. I imported all my extensions. They all imported just fine. The thing about Cursor that really blew my mind was it has like auto-complete that spans the entire file. For example, I had this app that was a WSGI app; it's basically a kind of REST API app, and I was converting it to Flask. And so one pattern that I had to go from is the old app: you would return a JSON object and it would have a status key and then the HTTP status code—it's like 200 if it was good, 400 if it was now forward, et cetera, et cetera. And then your payload. So most of the time in these little handlers, I was returning an object where JSON object were like status 200, payload is data. Data was some object I created, right? In Flask, you don't return anything from your handler. The handler gets past this response object, and you do `response.status = 200` and then `response.data.payload`. Right? And I changed it. There was a file that had a bunch of these handlers in it. So I changed it in one. And as I was changing it, it let me tab complete to do that.

0:48:08.779 --> 0:49:40.799
<v A>Like I was putting response.status—like about halfway through the word status—and it figured it out. It's like what you want to do is `status = 200`, `data.payload`, and then delete the return thing. And I hit tab. That was cool. Then it like zipped down like a third of the way through the file, and it's like, 'You want to do it here too.' Don't you hit tab. And like, I just hit tab, tab, tab, tab, tab. And it just did all of it. And that blew my freaking mind. Like I've used Copilot and Amazon CodeWhisper and stuff. And they do have some of those magical moments, like within pretty close to where your cursor is. But I've never seen it like this where it's like, 'Oh, you also want to do this thing at this file over there.' It just kept chaining these things to do. And he got it all perfect. Like, you know, there was weirdness around the brackets and all of that; you know, is kind of complicated because I was adding this handler. So it was actually adding a bracket but then deleting the return bracket. And it figured it all out flawlessly. It was really impressive. Again, I think we've talked about this in the past. If you're at a company, make sure you get the right approval. We were just talking about junior engineers. We don't want you to get fired. If you're a junior engineer and you get fired, that will make us sad. Get the right approval, but this is an amazing tool.

0:49:41.580 --> 0:49:58.566
<v B>That is very cool. Is it in the works for it to be able to do it across files in your project as well? Like if you wanted to refactor all of your files to, you know, use Flask or whatever, like, is there a way to invoke it to do that or not yet?

0:49:58.988 --> 0:50:12.080
<v A>Don't think so. I'm trying to remember. It definitely jumps around in the same file. I don't think it has a way to skip other files yet, but it might; I've just started using it a few days ago.

0:50:12.240 --> 0:50:20.240
<v B>Okay. And what is the zip down? So like it zips down and then like, if you don't want to do it, it just returns you back to where you are. Like, I feel like I'll get a whiplash.

0:50:20.240 --> 0:50:27.372
<v A>Oh, interesting. That's a good point. You know,

0:50:28.975 --> 0:50:43.520
<v A>Yeah, I guess I'm not totally sure. So the times where it did this, I took its suggestion every time. I think maybe if you don't take the suggestion, maybe you could hit the up arrow on the key and get teleported back.

0:50:43.520 --> 0:50:50.400
<v B>I sense a new developer metric, which is how often do you accept the AI suggestion?

0:50:51.120 --> 0:50:57.860
<v A>Oh, definitely. Yeah. I'm sure if they're not measuring that, they should be. And yeah, I guess you're right. They should pass that measurement along to the company.

0:50:58.844 --> 0:51:24.180
<v B>No, I'm saying you're like, as an engineer, you get a score, which is like, the more you accept the AI thing, the worse off. You're not adding anything. They could just replace you with the AI. So you need to like, then there's this whole gamesmanship, right? Because then we'll figure it out and we'll just write what the AI said, but change one character so it doesn't count. And then our metric for, you know, novelness goes up.

0:51:25.339 --> 0:51:29.472
<v A>Yeah. Well, you know, it's interesting. I mean, I think

0:51:31.851 --> 0:52:25.299
<v A>That, yeah, I mean, you know, the next level up from this would be for me to say, you know, convert this kind of raw. And then it would have just figured out. Cause the thing that I had to provide was the knowledge that when you return things, Flask doesn't care. Like it's not going to use that return value. And so like, just like running this code on Flask means that you would just error 500 everything. Right? And so like, I had to figure that out. And then once I started typing the solution, then it was done. But yeah, I mean, at some point maybe it will—that's almost like a segue, or another thing is, you know, junior engineers now also have to compete with AI. Or maybe not; maybe it makes them on a more level playing field. That's still kind of an unknown.

0:52:29.140 --> 0:52:37.479
<v B>I don't. Yeah, I don't have a—I think this is a watch this space. No, no, that's not how they say that. This is like 'watch the future.' I don't know what will happen.

0:52:40.279 --> 0:53:43.560
<v B>And I think where I'm sitting now, but you know, this has turned out to be wrong before, but like you said, it's sort of the intent still needs to be provided and sort of guided. But people who lean in on it probably see a big improvement before everyone else. And then everyone else sort of adapts or not. And we call it the same thing, but it's not really the same thing. So, you know, it's like programming a website is technically programming, and programming an operating system is also technically programming. They're not really the same thing. Like some parts are similar, but lots of parts are different. And I think we'll just end up with more divergence. So there's some things you're doing, like you're saying, 'Hey, I'm looking for another web framework that may be faster.' And it suggests to you to go to, you know, whatever. But then there's also times where you might just be straight up, like this is something that's not been done before and it needs to be written for the first time. Yeah.

0:53:43.900 --> 0:54:02.984
<v A>Yeah. That makes sense. Yeah, it's going to be wild time. I think these tools will save people a lot of busy work, which is good, but it also means that you will have to memorize the code base faster because you're not going to be sitting there staring at it while you do busy work. So,

0:54:05.464 --> 0:54:16.080
<v A>All right. On to the topic of the show: vector databases. So Patrick, have you ever used a vector database?

0:54:16.640 --> 0:54:17.180
<v B>No.

0:54:17.880 --> 0:54:20.820
<v A>All right. But you have used a database.

0:54:21.800 --> 0:54:22.100
<v B>Yes.

0:54:23.166 --> 0:54:55.339
<v A>Do you have ever put anything other than text in a database? You know, like, have you ever used a database for images or anything? Yes. What did you use for that? Do you just stuff an image into MySQL or what? Protobuf. Oh, you put Protobufs in the database? Oh, yeah. Really? So, but that's just like garbage, right? Like, you know, it's unintelligible. Wow, bro. What? You're coming at me. No, no. I don't mean it's garbage quality. I mean, when you look, when you run a query.

0:54:55.720 --> 0:55:00.180
<v B>Yes. Yes. It's unintelligible. Yeah. You can't index it properly. Yeah. You're right.

0:55:01.040 --> 0:55:15.520
<v A>Yeah. Because like when you query, I think if certain databases let you have a column be a JPEG and the database kind of knows it's a JPEG. And so when you query it, it will actually show you pictures.

0:55:16.160 --> 0:55:23.779
<v B>Oh, that's cool. Can you query? Oh, no, no, it's not. I was gonna say, can you query stuff about it? Like I'm looking for JPEGs where the subject is read.

0:55:25.160 --> 0:55:28.709
<v A>You could with a vector database. Oh, okay. Here we go. I'm ready. I'm

0:55:28.709 --> 0:55:30.020
<v B>here. I'm leaning in.

0:55:30.540 --> 0:55:39.593
<v A>All right. So yeah, my mind is still blown by the, so, okay, hang on. I got to dive in

0:55:39.593 --> 0:55:40.040
<v B>a little bit.

0:55:40.520 --> 0:56:00.434
<v A>So if you put a Protobuf in a database, right? Yeah. Then I guess most languages have Protobuf support, but you would have to like, you know, interpret that Protobuf before you do anything with that data. So like, you know, your SQL query goes straight to something that like converts the Protobuf to like. Yeah.

0:56:00.580 --> 0:56:10.100
<v B>I mean, you can use just more like a key-value store where the value rather than a relational database, right? So it becomes more like, 'Hey, I have these keys and then just have the LOBs.'

0:56:10.860 --> 0:58:10.740
<v A>Got it. Okay, cool. Uh, okay. So let's build a scaffold up to vector databases. So, um, okay. People like starting from the beginning, how do computers represent things? So for example, you open up Notepad, you start typing some letters and Notepad, you hit save and it saves as a text file. And so, not dealing with Unicode or anything like that, just putting that aside, a text file is basically a bunch of bytes where, you know, each byte represents one letter. So the letter, there's an ASCII table you can look up. And so, you know, I don't remember off the top of my head, I think isn't like 97, a lowercase 'a' or something. It's been a long time. Oh, but, but, but, but, you know, each letter and each symbol, you know, equal sign, apostrophe—they all map to numbers. And in at least the ASCII format, there's less than 256 of these, or different things. And so that's one byte in a computer. And so your document is, you know, a byte for each of these letters, including the spaces and all of that. So that's how, you know, traditionally how you could represent text in a computer. So a computer obviously doesn't have a concept really of the letter 'A', you know, in hardware. That's the way it works. And for images, the same kind of thing, you know, there isn't like any lithography—is that the right word? But, you know, there's no like dark room in your computer. Like there's no physical images in your computer. The way it works is, you know, each picture is broken up into pixels, which is like a really tiny square of an image, hopefully so tiny you can't even tell it's a square. And then for each pixel, the color could be represented as the amount of red, green,

0:58:10.740 --> 0:58:40.160
<v A>and blue representation in that pixel. And so you can imagine, you know, if you have a picture, a bunch of these triples—your red, green, blue triples—and you have one for every pixel. And you know, the width and the height of the image. So, you know, like what that whole rectangle looks like, and that's how pictures are represented in the computer. Does that make sense? Anything you want to add to that, Patrick?

0:58:41.620 --> 0:59:00.200
<v B>Yeah, I mean, I think you're right. I think what you're trying to point out is the computer doesn't have a native understanding of most of these concepts. So it just has operation bytes and like data bytes, and the data bytes need an encoding. They need a scheme, and humans are the ones that imbue meaning to those.

0:59:01.140 --> 0:59:51.060
<v A>Yeah, that makes sense. So, okay. So then we'll get into like, let's say compression. So for example, you know, you might have a document that has the same word a bunch of times. It's like 'Chapter One,' 'Chapter Two,' 'Chapter Three.' And so every time you have the word 'chapter,' that's going to be what seven bytes, right? To store all of those characters. But if it's always like 'chapter,' 'chapter,' 'chapter,' 'section,' 'section,' 'section,' there's like a lot of repetition there. And so you could actually have a smaller file by taking advantage of all of this repetition. And so one example that most people are taught in undergrad, and I've long since forgotten is Huffman encoding. Do you remember Huffman coding?

0:59:52.020 --> 0:59:52.920
<v B>Yes. Yeah.

0:59:53.300 --> 0:59:56.540
<v A>It's something to do with like trees, right? Yes. Build a tree where.

0:59:56.860 --> 0:59:57.720
<v B>And the prefixes.

0:59:58.680 --> 1:00:00.720
<v A>Yeah. Maybe you explain it. Cause I've totally.

1:00:00.780 --> 1:00:34.060
<v B>No, no, no, no, no. That was good. Yeah. Yeah. So I, and you're like, look in your file and you have to look at the whole file at once and basically decide like, how am I going to assign bits to the most common prefixes of numbers? And each branch of the tree represents sort of like going down that way. And so it can represent a bucket of the prefixes that get concatenated together. And ultimately that way, most, the most common pieces use the fewest bits. And then the longer, deeper parts of the tree encode the less common pieces.

1:00:35.180 --> 1:01:22.881
<v A>Yep. That makes sense. And so, you could do something similar for images, right? You could have some kind of Delta encoding or other kinds of things that take advantage of the nature of the image, but you know, for images really, you want to capture the sort of nature of the image, right? So it might not matter that like this single pixel in this one part of the car is like a little bit more blue in the image than it was in real life. Like it doesn't—that doesn't really change the nature of the image. And so now you start getting into what are called lossy compressions or lossy encodings. So for example, one of the most common is where you take

1:01:24.434 --> 1:01:44.900
<v A>these various kind of repeating patterns. So such as like a Fourier cycles or cosine cycles, you know, the cosine wave, right? You take these kinds of waves and you compose them at different frequencies and amplitudes on top of each other.

1:01:48.059 --> 1:02:52.540
<v A>So for example, if you imagine like a zebra. Zebra is this like really sharp kind of wave where it's white and then it's black and then it's white and it's black. And so if the camera zooms in on the zebra, now the sort of frequency is getting smaller, right? Because those stripes are getting bigger. The camera zooms out or the zebra is running away or something. Now the frequency starts going up. And so what you do is you have a ton of different waves of different frequencies and you have there's an algorithm that tries to figure out how you can compose many, like add many of these waves together to faithfully reconstruct the image. And if you can do that, then you only need to store like the definitions of those waves. Like what were the amplitudes and the frequencies of those waves? Instead of storing every single pixel.

1:02:52.540 --> 1:03:13.180
<v B>I mean, I think just to completeness there, you need infinitely many, but if you have some loss criteria, like some amount of thing you're willing to give up, then you can basically chop and only take the top X most important frequencies and then you'll get back most of the image.

1:03:13.540 --> 1:03:25.529
<v A>Yep. Yeah, exactly. That's a really good segue to the next part. So now we'll jump into embeddings, right? So we talked

1:03:29.309 --> 1:05:25.037
<v A>about ways to kind of explicitly represent something. So if you have a document and you want to store it, you explicitly represent your each individual character. But many times what we want is to know like the overall essence of a document, like is this document about cars, or is this a long document or does this, is this document at a really high reading level or at a kindergarten reading level? So, you know, these kinds of like soft attributes, those sound like kind of yes/no questions. But really there's a lot of ambiguity there, right? It's like, there's certain degrees of how much about cars is this or how much of this document is at the kindergarten reading level. So, you know, these kinds of like soft attributes and then doing things like give me all the documents that are roughly a yes for this question. So it's kind of like a lossy attribute system. And the best way that we have today to be able to answer questions like that is through embeddings. So the goal of an embedding, it's not directly to recover the original document as we talked about before. So, you know, with Huffman encoding, the goal is to compress something so that you can later decompress it and get back the original thing. In the case of embeddings, the goal is not to get back the original thing. The goal is to be able to answer questions around nearness. You know, if I have a query or a concept or the query might even be another document, right? What are the things near that query? Semantic nearness—that's the goal. And so because

1:05:25.189 --> 1:07:06.920
<v A>of that embeddings can be very, very lossy because they don't really have to reconstruct the original content. And so there are two kind of key ways that embeddings are created. So one way is through contrastive methods. The idea here is I might say these two documents were written by kindergartners. So, I know I'm going to be asking a lot of questions about reading level and reading skill for these documents. So I have documents created by kindergartners. I have documents created by adults and I randomly pick two documents. If they were both created by the same age group, then I want their embeddings to be closer together. If they were created by a mix of a kindergartner and an adult—you know, if one document was a kindergartner and the other document was an adult—then I want to pull those two documents apart in this space. And so at the very beginning of your embedding process, all the documents are just randomly projected. So they're just scattered all over the place. But after you ask yourself and answer this question many, many, many times you end up with, hopefully you'll end up with sort of two, you'll end up with many spaces, many clusters, but each of those clusters will be either a kindergartner document cluster or an adult document cluster because of all of this pulling and pushing. Does that make sense? Yeah.

1:07:07.680 --> 1:07:18.680
<v B>I mean, I think the pulling and pushing now, does that happen in like a subset of the dimensions or are they guaranteed that even in another dimension, they may be far apart?

1:07:19.640 --> 1:08:02.020
<v A>Yeah. So when you pull or push, you do it in all the dimensions. So it's almost like there's little thrusters on these two documents. There are points in embedded space and those thrusters are pulling them together in every dimension simultaneously. Or pushing them apart. And so, and so when you do this embedding, a single dimension has no real meaning because the only meaning was constructed from the distances, the pairwise distances. So like, you can't say like dimension three is how many times they said the word dog or something. It's because it's all like a composition of all these dimensions.

1:08:02.020 --> 1:08:08.200
<v B>So you need all of the things you're trying to do this with upfront though?

1:08:09.687 --> 1:08:57.899
<v A>You need okay. So when you're training, you provide a bunch of training examples and those have to be labeled, right? So you have to know this is kindergartener or adult. Once you've trained the model, then you can give it new documents and you could even ask it like, is this new document a kindergartener or an adult document? And then it would use a vector database to look at what are the nearest neighbors? And if most of those are kid documents, then that's one way that you can answer that question. Got it. So that's one way is you give it a bunch of pairs of things that are should be close together or should be far apart.

1:09:01.307 --> 1:09:09.863
<v A>There's another way to do it where it's kind of like a fill-in-the-blank, right? So for example, you might,

1:09:16.495 --> 1:10:57.000
<v A>Here's an example. Let's say I have three documents written by a kindergartner and I give a model the first one, and I give a model the third one. And then I say, 'Hey, generate the second one.' Maybe you ask a kindergartner to write a book about dogs, write a book about goldfish, write a book about cars. And you present the AI that you're training two of those three books and you say, 'Generate the third one.' In the beginning, when you start training, the AI is just going to generate random words, just like random characters. It's not going to make any sense, but you have the third book actually written by the kindergartner. And you go to the AI and say, 'Hey, you were wrong. Like, here's the actual answer.' And they do the same thing with adult books, right? And it turns out that if all you're doing is kind of filling in the middle—whether it's the middle of a sentence, the middle of a collection of work, or whatever it is—because you're just filling out the middle and you already have the middle, like, you know, the right answer, you don't actually need any humans in the loop. It's like I could just take all of Wikipedia and I could give every sentence to some AI that I'm training and say, 'Hey, fill in the beginning of the sentence or the end of the sentence or the middle.' And I know the right answer. And so when it doesn't give me the right answer, I kind of push it. You have the right answer. And so this is called self-supervised models.

1:11:01.845 --> 1:12:51.260
<v A>Now, if I give it the end of something and tell it to reconstruct the beginning, then when I actually try to use this model in real life, I have to give it the end of something. But if I'm using this to do like ChatGPT, I can't do that. Like you can't say, 'All right, you're given the end of what ChatGPT wants to say. Give me the beginning.' It doesn't work that way. So usually you have to go in the other direction. You purposely hide the ending completely. You artificially hide the middle and you give the beginning. When the model tells you the middle, you correct the model. Because you're never able to look at the future. These are called forward models because they can only take the past and predict the future. So, GPT is an example without the 'chat' part is an example of just a pure forward model. Given like a bunch of things you said—this is your context—what's the next thing that's going to be said? It's trained on terabytes and terabytes of text. Then now it generates that token. But along the way of generating that token, it constructs an embedding. So it constructs an embedding first, and then it uses that embedding to predict the next token. If you're not interested or don't need that second part, you can just take the first part. And now you have an embedding of the words that have been said so far.

1:12:53.280 --> 1:13:18.720
<v B>And so this is why, as a by-product of making these chatbots, the companies can also offer an embedding service because they need that internally. So you can offer it externally, and it's valuable for the reasons you're describing, which is now if I have text—maybe it's not a chat, it's just I have text—and I want a representation of it that I can ask questions about, then I can get the embedding and then do whatever I want with it.

1:13:19.660 --> 1:13:20.960
<v A>Yep. Yeah, exactly right.

1:13:23.292 --> 1:13:44.520
<v A>Um, and there are vision transformers that work very similarly where you take a picture, you hide part of the picture from the AI and you ask the AI to recreate that part of the picture. When it does, you have the real answer because you were the one who hit it yourself and you compare the AI

1:13:46.039 --> 1:14:45.940
<v A>And similar to language, a Vision Transformer has a step where there is an embedding representation. So you can take a picture and run like half of the Vision Transformer. And now you just have this embedding where the picture is going to be kind of semantically close to other pictures that need to get reconstructed the same way. So all the pictures where you deleted a car and asked the AI to redraw the car, all those pictures are going to have very similar embeddings, but the picture where like you deleted a person from who is in the background of your family portrait, and you wanted the AI to fill that in with something. All those pictures will occupy like a different space in the embedding.

1:14:52.965 --> 1:15:03.620
<v A>So yeah, I mean, this gets pretty difficult to explain over the air, but any questions about that part of it?

1:15:04.640 --> 1:15:08.997
<v B>No. I mean, I think it makes sense. So the idea is you're getting a vector,

1:15:10.515 --> 1:15:25.500
<v B>which I guess we talk about like it's a list of numbers and those numbers represent where in this space it is. And the hope is that for these processes that you learn the ways in which things can be similar or different.

1:15:26.800 --> 1:15:58.160
<v A>Right. Right. So now let's talk about how to correct the AI. Let's say the sentence was, 'The quick brown fox jumps over the lazy,' what is it? Lazy dog or something? Yes. Yeah. So let's say the AI generates 'The brown fox jumps over the lazy dog.' So it got everything right except for the word 'quick,' right? But because it skipped to the word 'quick,' every word after that is kind of wrong, right? It's like shifted, right.

1:15:59.993 --> 1:16:17.678
<v A>So you have to develop all sorts of different, uh, metrics to be able to say like how wrong the AI was in a way that like helps it learn, learn well. Um, that's at the training side. Um, it turns out the embeddings, you

1:16:19.230 --> 1:18:13.420
<v A>know, you're often trying to use embeddings for a different purpose. Like for example, you might want to fill in the missing image, the missing part of the image when you're training. But then when you're actually using the model, you want to group all of your images together. So someone can say, Hey, this is a picture of me at the beach. Find more like this one, which is different than when you're training it. So we call this training serving skew. It's when you use a model for a different use than it's, it's, uh, what it was trained, how it was trained. Um, and so what that means is the similarity metric on the embedding. Like there's a lot of unknowns there. Like, you know, like it might be that the model was just trained to fill in parts of the image. And so it actually does a terrible job of, of finding similar images. It might just be a coincidence, right? That it does a good job, right? So, so there's not a lot of theory at this point. And when there's not a lot of theory, the best thing you can do is try a lot of things. And so you want to have a variety of different similarity metrics on the embeddings. So a common one is just how close something is like in Euclidean space. But then there are a bunch of others too. We don't need to dive into all of them, but, uh, there's a ton of different similarity metrics. And so, you know, once you've chosen a similarity metric, then you can say, given a point in the space where the point can have a document on it or not, uh, find the nearest neighbors, find the nearest documents to that point. Um, and, uh, um, and then you can look at those results and see, okay, does that match what I was hoping from a product perspective?

1:18:14.200 --> 1:18:36.458
<v B>So the point that you're making and getting neighbors like a query, but the query is not doesn't have to be a text. Like you said, it could be a picture. It could be whatever it gets to embedding. And then the goal, maybe I'm jumping the gun here, but the goal of the vector database is to help you efficiently find the things closest to your query. Yep.

1:18:37.380 --> 1:20:35.480
<v A>Yeah, exactly. Right. Yeah. So, just like we talked about with the Fourier Transform, as Patrick said, you know, you'd need an infinite number of waves to get a perfect resolution, but you're often willing to tolerate error. Similarly, if you're willing to tolerate error in finding out which documents are the closest, then you can do what's called sublinear approximate nearest neighbors. So the idea is if I have a million documents, I don't need to check the embedding of all million of them to get the nearest neighbors to a point. I can use a variety of different data structures to say, okay, I'm extremely confident that I have the nearest neighbors, but I'm not one hundred percent confident. So for example, maybe there's a partitioning system and if you fall like exactly just to the left or right of a partition boundary, you might miss what's on the other side of the partition. And that might happen one out of every 10 million times and only affect one result. And so you're more than happy taking that penalty in exchange for just massive, massive speed up. So imagine like, we're talking like going from like a million seconds to search to six, you know, or something like that. So, so that's basically what the vector database does. So the vector database has it. So where you provide your own embeddings, it doesn't do that part, but you jam a bunch of embeddings into this vector database. It creates all of these data structures and then you can give it vector queries and it will tell you the nearest neighbors.

1:20:36.929 --> 1:21:42.019
<v A>And the vector database is responsible for managing sort of the dynamics there. So for example, if you start removing documents and adding different ones, at some point, the vector database might be like the data structures might be kind of stale. So I'll give you an example, like an extreme example. Let's say you add a ton of vectors, but only the first dimension isn't zero. So the vector database is basically going to create a bunch of partitions around that first dimension and ignore the rest, right? If you take those documents out and insert a bunch of documents where only the second dimension isn't zero, but you still use the old partitioning scheme, they're all going to fall in the same bucket. And now you don't have that speed up anymore. And so the vector database is responsible for like knowing when to rebalance and re-index, just like your hard drive has to occasionally rebalance the B trees on the folders of your hard drive.

1:21:43.500 --> 1:21:44.759
<v B>What's old is new again.

1:21:45.299 --> 1:22:03.900
<v A>Yeah, exactly. So, yeah. And then beyond that, just all of the traditional database things, you know, backups and restores, all of all of that stuff that we've come to appreciate from databases, vector databases have to provide that as well.

1:22:04.560 --> 1:22:20.460
<v B>So if you have like, you mentioned, so you get like Euclidean distance or you hear like cosine similarity or whatever, if you want to change the metric, do you need, or does per metric, is there a needed a different clustering? Or is it like one is good for all of them?

1:22:21.400 --> 1:23:10.180
<v A>Yeah, it's a good question. You basically need a different, well, the answer is always going to be, it depends on the data, right? But definitely, you can construct data sets where, where each clustering, each similarity metric needs its own clustering. And you can show that there is a data set where any clustering can only be good in one of these three different similarity metrics. So it's possible. Now in practice, I wouldn't be surprised if these database systems just use Euclidean distance for the clustering. And then you pay a slight penalty if you're using something else. I think for 99% of cases, that would be fine.

1:23:10.180 --> 1:23:11.800
<v B>Makes sense.

1:23:13.660 --> 1:25:13.240
<v A>Yeah. And so there's, this is where, this is definitely in the category kind of like authentication. It's in the category of things you don't want to write yourself. Oh, come on. It's going to be pretty gnarly. You know, even some of the top vector databases are only now starting to become mature. So for example, you know, I've used in the past Milvus, which is an open source vector database. And, we had issues where when we tried to back up the database, it would crash and we would lose everything. And so, you know, definitely I would say it's still at the phase where you probably want to store your embeddings just in a regular database as well. It's still kind of early. But you know, every month it gets massively more mature because so many people are using it. The one that I, so Milvus is good. Pinecone is an enterprise alternative. Pinecone is a lot better, but you're going to pay for it. You get what you pay for there. Another one that I really appreciate is PGVector, which is an extension to Postgres that lets you have vector columns in an existing table. And that's really nice, right? Cause you could have, like a row have, imagine you're storing houses, your Zillow or something. Right. So you could have a row and the row has like an ID, the address of the house, the number of bedrooms, all of this explicit information. And then in that same table, you just add a column that's like, you know, embedding. And that column is backed by PGVector. And it works pretty naturally. So you could, this is actually another thing, you

1:25:15.248 --> 1:25:16.940
<v A>know, the approximate nearest neighbors,

1:25:18.741 --> 1:25:57.639
<v A>it assumes that they're all valid. Right. But like, for example, let's say you only want three bedroom houses that are approximate nearest neighbors, right? That's actually a really hard thing to do because you might pull like the nearest thousand neighbors, but they're all two bedroom. Now I have to pull another thousand, another thousand, another thousand. So finally you have enough to fill up the three bedroom limit that you set. So, there's a lot of complexity there. But, but yeah, you know, these database systems are really good at handling all of that for you.

1:25:59.919 --> 1:26:10.360
<v B>Yeah. I would love to like find an excuse to use this to always sounds really interesting and cool. And I know I'm obsessed with tinkering, so love to find a use, but just haven't had one yet.

1:26:10.360 --> 1:26:48.100
<v A>Yeah, I mean, for me, the one use case that I've found outside of work has been nearest neighbors on my photos. It's not too hard to use one of these open source models, embed all of your photos. And then, given one of your family photos, you find all the nearest ones. Google Photos, these other things will do it for you. So it's really more of a fun exercise for you to do than something that's providing a lot of unique utility. But that's been a fun thing that I've been able to do with it.

1:26:53.040 --> 1:27:17.680
<v A>Cool. All right. That is a ramp up on vector databases. If you're doing anything with vector databases, let us know, give us a shout out, go on the Discord. The Discord is starting to pick up some steam, a lot of really interesting discussions about career changing jobs. Should I be a consultant after I leave my company? A lot of really interesting discussions there. So check it out.

1:27:19.300 --> 1:27:22.440
<v B>Okay. I will take that as a personal message.

1:27:24.375 --> 1:27:26.670
<v A>Well, I got you covered Patrick, but I will

1:27:26.670 --> 1:27:28.320
<v B>be there, but it'll be a bot.

1:27:29.240 --> 1:27:33.200
<v A>That's right. Patrick's AI will be there and you can interact with it.

1:27:33.720 --> 1:27:39.968
<v B>So this is awesome. I learned a lot. So thanks Jason. And thanks everyone for listening. Cool. All

1:27:39.968 --> 1:27:47.680
<v A>Right, everyone. Thanks so much for supporting us on Patreon. We really appreciate it. And we will catch you all later. Have a good one.

1:28:10.520 --> 1:28:40.500
<v A>Thank you.

