Episode 161 · Programming Throwdown
161: Leveraging Generative AI Models with Hagay Lupesko
10 Jul 2023 · 1 hr 33 min
Episode 161 · Programming Throwdown
10 Jul 2023 · 1 hr 33 min
MosaicML’s VP Of Engineering, Hagay Lupesko, joins us today to discuss generative AI! We talk about how to use existing models as well as ways to finetune these models to a particular task or domain.
Tap a chapter to play from there.
A:Hey everybody! So we have seen so many AI hype cycles around so many different areas. We've seen self-driving cars was a big deal in 2009, if people remember that. At the time, Ray Kurzweil has been talking about the Singularity forever. Oh, even beyond AI, there was Bitcoin and Web3 and all of that. And Patrick and I, we've had folks on, but ourselves, I've kept a little bit of an arm's reach from the latest shiny object syndrome. But I think generative AI is amazing. I'll just put it out there. I don't think it's Singularity AGI type stuff, but I do think that there's a tremendous opportunity to create value with generative AI. I've been really excited about it. I've been diving deep into the literature and also applications. And I know a lot of other folks have too. It's a really exciting area. It's an area that I'm pretty excited about as well. And I'm super excited to have Hagai Lebesco on the show. He's the VP of Engineering at Mosaic ML. So thanks for coming on the show, Hagai.
B:Hey, Jason. Hey, Patrick. Thanks for having me.
A:Cool. So we'll definitely dive into generative AI and how folks can use it at home or at their business. But let's start off by talking a little bit about you. What's your background? What was the path that you took that brought you to Mosaic?
B:Yeah. So I'm currently the VP of Engineering at Mosaic ML. And I guess we'll probably touch on Mosaic ML a bit later on. But I really started my career a while back now. If there's been video, you could have seen all my gray hair. So I started my career as an engineer, you know, back in Israel where I was born and raised. And really earlier in my career, I did a bunch of things around computer vision, medical imaging, you know, vision for factory automation. I even spent a couple of years living in China, working on a startup there.
A:Wow. So let's dive into that a little bit. So you were in Israel and then the U.S. and then China or straight from Israel to China?
C:No. Yeah. So straight from Israel to China.
A:And so what was that like? That must be a huge culture shock.
B:It was definitely initially a shock and then really a fantastic experience. Because, you know, as we all know, China is even today, actually, you know, kind of growing rapidly. Back then, it was really super, you know, moving super, super quickly. So just the story is, you know, I was a young engineer back then, had some experience, expertise in computer vision. And, you know, this was actually, so just to put things on kind of on the timeline, this was pre the deep learning revolution. So I'm talking about 2007, '08-ish, you know, neural networks were not working well. So computer vision actually was completely different, like the way you apply computer vision to a problem.
A:Yeah, just to put context, I think, you know, there's a bunch of hand-coded things, right? Like there was these, I remember, Patrick probably knows this way better than I do, but there was a whole bunch of filters, right? Like Sobel filter and these like directional filters. And you would basically try to build your own deep learning system by just stacking all of these filters as an expert. And then at the end, you would have some shallow model that, you know, is stacked on top of all these other things.
B:Exactly. That was exactly the way you'd apply, you know, define different filters. You would hand-tune them. I mean, today from, you know, computer vision neural networks, the convolution kernels are kind of, you know, figured out during the training process. Back then, we would use convolution quite a bit and you would hand-tune the convolution to work for your problem. Yeah. That was actually a lot of fun. It was a really interesting process. Of course, it made kind of the solutions not super scalable where for different customers, different problems, you'd have to sit down and tweak things. You know, field engineers, that's a lot of what they would do. They would sit down with the systems and tweak the parameters, including the convolution kernels by hand.
A:Oh, wow. That's wild. Yeah, because, you know, when you're, when the convolution kernel is not, doesn't know anything about your objective. Like, it's trying to, like, find edges, but that's not your objective. Your objective is to say, like, is there a face in this picture? And edges just happen to be kind of, like, tangentially, you know, interesting to that objective. And so then it's like, can you come up with a filter that's even more interesting? And then, yeah, to your point, deep learning now just does everything for us, which is pretty wild.
B:Yeah. And it's even more than that. Like, you had to, usually, typically, in a typical computer vision pipeline, you know, you'd start by taking the input image and then preparing it to be kind of ready for the convolution operator. So you'd have to do different tricks. It was like a whole toolbox of tricks you do to, like, clean up the image, you know, normalize it manually and then start scrubbing the image with different morphological operators.
B:Yeah. So it was quite a ride. But going back to kind of the experience in China, so, you know, I was just married back then. I asked my wife, 'Hey, do you want to go on an adventure in China?' And she initially said no. And then I was able to convince her that it's going to be kind of an experience of a lifetime. And we just hopped on the plane. It was a small company, startup, like, I don't know, maybe five people. Well, they brought me in as sort of the computer vision expert, although definitely there was tons of, you know, I wasn't that of an expert. But, you know, I said, 'What the heck?' And we built a whole system, including hardware. Of course, the differentiator of the system was the software. But it was hardware with, you know, robotics to, you know, from conveyor belts controlled by, you know, by different actuators through imaging system, lighting, cameras through integration with the automation in a typical factory. And that product was for the PCB industry, the printed circuit board industry.
A:Oh, cool. Yeah.
B:And, yeah, you know, we built a product from scratch. We're able to, you know, sell it to a few companies. I spent a lot of time on factory floors in China, which is by itself an experience. Oh, I bet. I heard they're massive. They're massive. It's like a city. Now, you know what's funny? I mean, I'm from Israel. So when I was growing up as a kid, Israel was about like 6 million people. Okay. And by the way, just for context, I think Israel is often in the news, but many people don't realize that it's tiny, both in terms of population and geographical size. It's actually smaller than the state of New Jersey in terms of the size.
A:Oh, I never knew that.
B:Yeah. Yeah. So, I mean, you know, here I am coming from
C:Israel, you know, 6 million people country moving to China, going to a suburb of Shanghai. And that suburb was, you know, a small suburb, 6 million people.
B:So, yeah, just the size
A:of China
B:is massive. And, yeah, so, you know, it was fun.
A:And are there tours? Like, let's say, I've never been to China, I'll be honest. I would love to go. I just never had the opportunity. But if I went, could I tour a factory? Like, is that a thing that tourists do or not really?
C:I think it's a great idea for a startup, Jason.
B:But no, I don't think it's an option. But these factories are really interesting because they're like little cities—literally little cities. Like, one of our first customers was actually a factory owned by a Taiwanese company and wasn't considered a very big factory, but it had 50,000 workers.
A:Wow. Oh, my goodness.
B:So, you know, 5, you know, times 10 to the power of 4. And what you realize after you go there is, first of all, most of the workers are fairly young, meaning, you know, 18-ish. And they actually live in the factory. Oh. And they have dorms there. They have everything they need, like, you know, food, social activities, you know, places to work out. It's literally almost like a student dorm, only you work, you don't study. So I found that really interesting.
A:Yeah. One of the things that blew my mind was—this is a long time ago—but they interviewed Tim Cook. And they were talking about manufacturing, like what it's like, because I think this is the time where they're making the Mac Pro in America. And they were talking about the difference there. And what he said was, in China, if you need a million people, literally, like you need one million people to show up to, like, you know, boost the iPhone production, you can get a million people. And when he said that, and he wasn't hyperbolic—I mean, he literally meant it. That really hit home on this kind of scale we're talking about.
B:Yeah. And China, by the way, is not done with that process. There are still, the majority of the Chinese population is still in villages and looking to go to the city where they can find, you know, have work, get proper, you know, wages, you know, and start their lives. And this is part of what I think many people don't understand about the Chinese—you know, China and the Chinese government is that they are under immense pressure to sustain growth so that their masses actually kind of have a path to a better life. And that's part of why they're so aggressive on growth. They just have to grow very quickly to kind of, you know, serve that need of their population.
A:That makes sense. So what happened to the startup? Did the startup grow very quickly or no?
B:So it started well. And then, you know, I don't know how much, how many of the listeners know, but 2008/9 was a pretty massive financial crisis. And we were hit—the startup was hit very significantly by that crisis. Started as a, you know, the mortgage crisis in the US and then very quickly kind of expanded globally. As often is the case, right? When there is a crisis, people start, you know, cutting back on their purchases. And then, you know, the PCB industry was hit significantly. So did the chip industry just because the demand for devices went significantly down. So the startup didn't shut down, but it definitely kind of grew on a good trajectory. And then it just kind of, you know, most of the orders were cut back. Budgets were cut back. Yeah. So, but we were able to still work through things. However, at some point I had some family issues. I had to go back to Israel. That was about, you know, two years later. So I went back to Israel and, you know, for a while I was flying to China every month, but it's really unsustainable, especially when you have a young family that needs, you know, needs you to be there for them. And I had my first son who was born. So at some point I just parted ways with the startup.
A:Yep. That makes sense. And so then after that you were—oh, so at that point you were back in Israel. At some point you were at the US. What happened there?
B:Yeah. So I went back to Israel and then kind of, I went back to working in an area that I had some experience on before, medical imaging. So, yeah, I actually went to work for GE Healthcare in Israel and we built a cardiovascular imaging system, which was, you know, really a lot of fun. And I think for those that have worked in healthcare, you know, there are definitely some downsides. Like it's a very slow moving field because of a lot of regulation. And in general, the audience is very, customer base is very conservative. But then you really feel, on The plus side, really feel that you're, you know, changing the world for the better, right? Because if you can develop systems that give better care, help detect diseases earlier, help treat diseases. It's really something that, you know, you feel really good about doing. So I did that for a while. And then, you know, Amazon reached out and they didn't have an R&D office in Israel back then. I mean, now they have tons of R&D offices in Israel, but back then they did not. And they, you know, interviewed me and then asked to relocate me to the U.S.
A:Ah, so you went to Seattle?
C:No. So they wanted to relocate me to Seattle. But, so again, you know, I mentioned my wife earlier and how she had to, you know, approve the, you know, moving to China. So, again,
B:my wife is the decision maker on these things. And after thinking about that together, you know, Seattle was not the right place for us in terms of, you know, just the weather, family. So we moved to the Bay Area.
A:Ah, okay. Got it. So I know Amazon has this, like, Lab 126 or that's where the Kindle came out of and some of these things. Is that where you went or was there a separate office? No.
B:So the opportunity, at least, you know, Amazon kind of offered me back then was actually to join Amazon Music in SF. Ah, okay. Amazon Music back then was a relatively small team. It was a very basic product back then. You know, they were kind of following the iTunes model of, and again, like, for folks that are a bit older like me, I mean, you'd know that digital music actually started by selling songs. So you would buy a song, you would buy an album, you would pay for it, you would have the rights to a digital copy. And, you know, you can deploy it on your, you know, whatever players you had. Audio streaming was not hardly a thing back then. I mean, now it's how all of us consume, you know, music and more than music. But back then, the technology was not there. The business terms were not there. So it was a very different world. But I joined Amazon Music at a really amazing time where streaming just started picking up. So I actually helped ship Amazon Music starting in the U.S. And then we expanded globally. And that was really super cool experience because it was part of, you know, we're participating in that revolution of kind of taking music from being, you know, download digital content to streaming digital content. And that was, you know, a huge revolution for the entire industry.
A:Yeah. I feel like this is, you know, obviously out of my depth, but I do feel like just thinking about it economically, it better aligns the incentives. Right. Because I remember I definitely remember I was huge into, you know, bands and going to concerts, you know, in high school and even in middle school. I think there was one year in high school where I went to something like 100 shows in one year and I still had all the tickets and everything. I mean, not big bands because that would, you know, that would be a break the bank, but a bunch of local shows and everything. And there were times where, you know, I saw, you know, an album and the cover, you know, this is this is back when we were buying CDs. The cover looked awesome and I'd never heard the band before, but the artwork looked really great. So I bought it and then the songs were terrible. And so it's like, OK, well, I just lost 10 bucks. And so now, you know, because it's streaming, the songs that you enjoy, that you listen to again and again, that time is logged. And then that credit is assigned, you know, to the appropriate musician. And so and so now, I mean, the sad part of it is no one cares about album art anymore. But the good part of it is that people are just laser focused on the music and the message.
B:Yeah. Yeah. I think it definitely really revolutionized the entire music industry. It also increased the pie. And that's, I think. Right. It's a good lesson, by the way, because I think whenever there is a new technology that comes by, there is always the kind of the pushback. Right. Especially when it's a fundamental technology that that changes how people, you know, interact with content, for example, interact with technology. There's always a pushback because, you know, very naturally people are concerned. And we'll get to AI, I guess, later. But I think we all can see similar patterns. Mm-hmm. But in reality, first of all, these technology changes are something, you know, usually you cannot block. You can, you know, slow them down a little bit if you really try hard, but you can't block. But second of all, they're usually opening up really new opportunities, business opportunities, consumption opportunities, education opportunities, and whatnot. They typically tend to be for the best or at least have a path that is, you know, for the best. And in the music industry, yeah, there was a lot of pushback from the big labels, the companies that control the rights to most of the content, at least on the Western world. But eventually they kind of, you know, went along with it. And now if you look at the revenue of the music industry from streaming, obviously it's much bigger than, you know, CD sales. But it's also in just if you look at the entire streaming revenue versus what CD sales were at its peak, streaming is now a bigger business. So, and it's not surprising, right? We all have now phones in our pockets, which is also audio streaming devices. And just the reach of content is much more broad today.
A:Yeah, that's right. And actually kind of a little foreshadowing here, but one of the most popular trending folks on Spotify was AI Drake, which is an AI version of Drake. And I was, I listened to some of the tracks and I was blown away. I think they eventually got banned from Spotify because they were using Drake's face as their face. And so that you can't do that. But they were, I want to say in the top 10 of trending for Spotify, which got my attention. And I listened to it. It sounded amazing. I actually was really shocked, even with everything we've seen so far, which I was really shocked at the quality of it. Come to find out that actually a person wrote the lyrics. So I kind of, you know, I thought that it was all the way AI where someone just pressed a button. But no, a person did write the lyrics. But the text to speech, you know, and getting the music and getting it all to match the rhythm and everything, it was just flawless. I mean, if you haven't yet, I don't think it's on Spotify anymore. I'm not sure what happened there. But you can definitely go on YouTube and look up AI Drake and listen to these songs. It's pretty wild. Amazing.
B:Yeah. Anyway, so going back to kind of my story. So I spent some time in Amazon Music. It was a lot of fun. But it's also very new to me. I mean, it wasn't about computer vision, obviously. It was also on machine learning, which I was anyway, you know, didn't do a lot of things on, you know, outside of my graduate studies. You know, it applies to computer vision. I was focusing over there more on algorithms for audio streaming, web applications, you know, scaling this from, you know, millions of customers in the US to tens of millions and later even more globally. And also, I think for me, I would just, you know, relocated from Israel to the Bay Area. The culture was very different. The way technology is developed, like the culture within the companies were different. And I was, you know, to a large degree, really adapting to that.
A:Why don't you double click on that? Like what is, you know, because Patrick and I have basically been in the US our whole lives. I lived in Italy for two months. Other than that, I've lived in the US or Canada my whole life. And so what really struck you about like maybe, you know, culture and then corporate culture over here?
B:Yeah. So, wow, I don't even know where to start because the changes are, you know, the differences are pretty significant. So, you know, I started by, you know, in Israel. Israel's culture is, you know, very casual and also very direct. For better or worse, you know, people, you know, would often, you know, not beat around the bushes when they have something to say. You know, they'll tell it to you in your face, even if it may be a bit offending. And in Israel, it's not considered offending. It's just, you know, people tell it for what it is. I think in the US, you know, people tend to be, I don't have to use the word respectful. It just kind of be more, you know, have more tact around saying things. So when they have something, you know, difficult to say or have some significant feedback, you
C:Know they would share it in a way that is very processed. So for me, initially, I had to really adjust, right, my noise cancellation.
B:So, you know, to really learn that, if people say something, even if they say it in a really nice way, I have to read a little bit more into it just because I'm used to, hey, if someone has something important to say and if it's critical of something that is going on or something that I have done, it would come from my experience in Israel and the way I grew up, it would come very directly. The US had to learn to understand the nuances a little bit better.
A:That definitely resonates with me. You know, like Patrick and I grew up on the East Coast or I guess maybe you'd call it the South, I mean, Southeast. But in moving to California, I think the way I kind of expressed it, I didn't really tell people this because it also lacks tact. But just to explain it, I kind of felt like the people around me were like passive aggressive and I was actively aggressive. So, but yeah, I just I felt like similar to what you were saying. I would just say stuff and then other people, that would meet, especially in the corporate world, I would realize kind of like you're saying a day later, oh, this person actually they were actually really happy with this or really upset with that. It's like there's intentionally a bit of noise in the signal to try to. Yeah, I don't know what it is. Maybe it's like there's always plausible deniability of everything. You know, it's just it's like a politeness thing. But yeah, even though I've grown up here the whole time, I had to go through the same experience. Yeah.
B:And I think to your point, it's, you know, the US is a very big place. So I guess my experience has been based mostly on the culture in California and other areas in the US. Right. Like you said, are probably somewhat different. But yeah, this is one of the differences. I think the other thing which I actually think that is kind of aligned, actually, just done a bit differently. You know, Israel versus California is taking initiative and thinking out of the box. You know, Israel is a small country. I grew up, Israel kind of was developed in an area with a lot of security problems. So Israelis tend to be very creative, out-of-the-box thinkers and also don't have too much respect for the way things are done. Right. It's like you always think about ways you can do something better. And, not surprisingly, per capita, Israel is the number one country in terms of startup. Right. Founding startups. A lot of it comes from that Israeli mindset and culture. I do think this is, California is actually kind of very similar, maybe for different reasons.
B:But in California, also kind of independent thinking, thinking out of the box, taking initiative, not conforming with status quo is, I feel, is kind of encouraged and even maybe something that is highlighted.
A:Yeah, that makes sense. Yeah, it's interesting. I think I do feel like there's a real independent spirit. You know, if you visit places like I've never been to Israel, but if you visit India, for example, it felt like like a libertarian paradise because there's so many small companies. If a policeman arrests you, you just give him money one to one. You know, you don't have to go to court. And so it just kind of felt like, yeah, like if the libertarian folks, if you kind of take it to the limit, that's what you would get. I do feel like in the U.S. there is just a new health care, particularly. There's just so much structure. And there's pros and cons to that, but it is different for sure. You were at Amazon and then at some point, like you got into. So that kind of was sort of like an intro to A.I. Your sort of introduction to like recommender systems and some of these kind of large-scale A.I., you know, and kind of augmenting what you did with computer vision. And then what's the path from that to being kind of like all in on A.I.? What happened next?
B:Yeah. So I spent a few years in Amazon working in Amazon Music and then kind of decided, hey, I need a change. And then I moved within Amazon to AWS, Amazon Web Services. Ah. And back then, AWS was already kind of, rapidly growing business that was already fairly large. So today, AWS is a business in Amazon that generates about $85 billion in revenue every year, which is just massive. Right. Like if it would be its own company, it would have been one of the five biggest software companies out there in terms of software revenue. But back then it was not that big, but still fairly big. But their machine learning offering was very limited back then. And then they doubled down on it. And I thought that was a really interesting area to be part of. And, you know, I was fortunate enough that they accepted to take me in. In my master's degree at Tel Aviv University, I studied a little bit about machine learning, among other things, just like many other folks did the CS back then. Machine learning was not what it is today in terms of kind of the dominance. And, back then we were building AWS SageMaker. That's the, you know, they like a very successful machine learning platform offered by AWS. It's a big business. From what I hear, it's, you know, the fastest service in AWS's history in terms of growth. So joined that team and then worked on, yeah, contributed to SageMaker, worked on deep running frameworks. Back then, AWS tried to double down on a framework called MXNet, which is kind of similar to TensorFlow or PyTorch, only it wasn't successful as both of these frameworks.
A:It's tough. I mean, I'm amazed PyTorch was able to take the lead.
B:Yeah. And I had, yeah, I think that was really interesting. I definitely took a lot of lessons from that because, you know, I was on the team that lost. I was on the MXNet team. And I think, you learn a lot from things that don't work according to plan. Typically, you actually learn more from the things that don't work according to plan or fail than from your success. Because success, you tend to attribute it to yourself and yourself and your team and that's it. But failure is kind of forced to think harder, right, about why things didn't work out.
A:Yeah, I would love your take on this because I don't know how that all played out. It's a little bit—I'm definitely a user of TensorFlow and PyTorch, but sort of how did PyTorch kind of take the lead and leapfrog over everybody? You know, and I guess like what were maybe some of the mistakes MXNet did or some of the gaps that PyTorch was able to fill that allowed them to do that?
B:Yeah. I think some of the things I observed, and again, I think there's definitely many more angles to it. But the first thing I think is usability is the number one thing, right? And I think especially for us as engineers, we tend to sometime underrate usability, thinking, 'Oh,'
C:no, usability, you know, it's similar, right? Like people can achieve the same goal in different way, one way maybe more complex than the other, but it's fine. Performance matters more. That's like a very
B:common pitfall. And I think definitely, I think on MXNet side, we definitely fell into that pitfall where we optimize for performance rather than optimizing for usability. So I think that is one key learning. And I think every tool developer, platform developer, framework developer out there, I recommend always put usability as the most important thing. Performance, you can catch up later on.
C:And actually for, you know, for people to get started, they actually don't look, usually, they don't look too
B:much into the performance. They would look more about the usability, how easy it is to onboard, how easy it is to learn, how easy it is to extend, how easy it is to apply to kind of core problems. Because, you know, at the end of the day, usability is what allows, first of all, people to move quickly solving a problem. And your tool exists to solve a problem. It doesn't exist for the sort of, you know, existing as a tool. And moving quickly actually saves tons of time and money. So I definitely say usability over performance. That's one key learning to keep in mind.
A:Yeah. You know, I think I actually, you know, if we follow that trail, follow that breadcrumb trail. And one of the things that Facebook did really well was having a lot of different roles inside the company. You know, it wasn't just everyone was software engineer. And I think that that, you know, although it seems esoteric, if you think about it, that really plays into this where, you know, if everyone's a software engineer and it's software engineers building things for other software engineers, then, then, of course, like, why can't you use this really convoluted API? I did, and I'm a software engineer, right? So, but if you have, you know, research scientists, machine learning engineer, and then embedded engineer and software engineer, then, you know, it's more clear that, like, the machine learning engineer or the research scientist is the customer of something like PyTorch. And that you can't really expect, even though they have engineer in the title, you can't really expect them to figure out some weird C++ error. And so, I think, you know, setting up that distinction early on, you know, kind of, you know, caused all these sort of downstream effects. Because I think if Amazon had treated, you know, the folks using MXNet as true customers instead of engineers, Amazon is amazing, you know, at, you know, customer satisfaction, right? So, it's almost like maybe it's, that's where the issue kind of started, right?
B:Yeah, yeah. That was definitely, definitely part of it. I think the other thing, which you also alluded to, is the importance of building a community. And actually building a community is definitely not trivial, requires, you know, deep thought. I'd say at the, you know, equivalent level to thinking through your, you know, software design, for example. You want to think about how do you design your community? How do you design it so people, wherever they are, they want to use the tool, they're well-supported, they have resources, they have people to follow. You know, that also requires a lot of deep thought. And I look at the PyTorch, I think they definitely, I'm not sure if they did that from the beginning, but at some point they started investing a lot in that. And I think they did it fantastically well. And I think, yeah, there is a real PyTorch community. And I actually consider myself now part of the PyTorch community. You know, I spoke at their last, you know, PyTorch developer conference. I met with a lot of people. At Mosaic ML, you know, we are part of that community. And that community is what helps PyTorch be successful and, more importantly, really be used by so many people in a really productive and constructive way.
A:Yeah, that's a really good call out. Yeah, I agree 100%. I feel like I'm also part of that community. I think they did an amazing job with SEO. When you search Google for PyTorch issues, it'll take you to the PyTorch forum. I don't even know if there's a PyTorch stack overflow. I mean, I'm sure there is, but I don't know if there's a significant one. But they've done an amazing job of being the place where you go for issues and solutions. Okay, so you're working on AWS and you're going to these really large companies working on their ML platform. And now you're at Mosaic, which is a startup, kind of full circle. It's a startup here in the US, but a startup nonetheless. And can you kind of describe that? I mean, it feels a little scary. We talked about Amazon, the enormity of the business. AWS is one of the biggest businesses in the world. And so how do you kind of take that leap to Mosaic, knowing kind of what you're up against?
B:Yeah, so I think we missed another kind of station along the way, which is after AWS ML, I joined Facebook. Back then it was Facebook, today it's Meta. And I joined the Meta's AI team. And then I did a bunch of things over there that were a lot of fun, starting with the recommendation platform called Deeper back then. And then expanding also into foundation model services there for language understanding, image understanding, video understanding.
A:I still remember, you know, Hagai and I worked together. And I still remember when Hagai first joined. And I remember thinking, wow, this person has long hair. The guy had this really long hair. And but total genius. A ton of respect. You've helped me a ton along the way. So I really do appreciate it. It's been a pleasure. It was a pleasure working with you. See, it was awesome. We did a bunch of cool AI stuff together. And I think you left after I did, I think. Or maybe it was before. It was around the same time, though.
C:A little bit after you did. Yeah. Okay.
A:Yeah.
C:I really kind of
B:really wanted to, after so many years in big tech companies, and there's definitely nothing wrong with big tech companies. I think you learn a lot. You do a lot. You have really kind of your impact just, you know, propagates to, you know, these immense customer bases, right, that these companies have. You know, in a smaller company, you
C:have more there. You feel like your impact is more direct. And you do have more bandwidth and time to kind
B:do a zero-to-one thing, right? Like build something from scratch.
C:Solve a core problem with very kind of where it's more easy for you to see kind of full ownership or work with others. Right.
B:So I was really tingling for that. And then I, you know, so the opportunity with Mosaic ML, really loved the team, the folks there, really kind of felt good about the business problem. And I can tell you about that. And then just kind of, you know, decided to make the leap and join Mosaic ML.
A:Got it. Cool. And that, is that the first time you started really diving into generative AI or did you do some of that Facebook and Amazon?
B:It was, no, you know, it was more at Mosaic ML. I think even when I joined Mosaic ML, I don't even know, like the term generative AI was probably used back then, but
C:not as often as it is used today. Right. Yep. So, yeah, you know, at Mosaic ML, the mission of Mosaic
B:ML was really to, back then when I joined, it was, let's make machine learning more efficient. The reason for making it more efficient is that, you know, anybody can see the pace at which the complexity of training deep learning models is increasing. And I think, by the way, that trend is actually toning down now. We can get to it in a minute. But, you know, if you look at even the last four years, you know, going from BERT, I think it was 2018, to GPT-3, 175 million parameters, a couple of years later. There's actually, there has been a growth in the number of model
C:parameters of an order of magnitude every year, which is just insane. Right. And obviously, it requires much more compute
B:And, you know, transformer architecture because of the transformer blocks, it's quadratic in terms of the number of parameters. So, you know, that growth just kind of limited the Number of companies, organizations out there that can actually leverage these advanced models, just because it became much more expensive to train these models and, of course, also to deploy them. So, Mosaic tried to initially to just make this more efficient, so it's more accessible. As we built our product, which is the Mosaic ML platform, it's a platform for training and deploying these
C:models, kind of realized that the problem space is more than just efficiency. I would even say efficiency, right, is a feature.
B:But then there's a lot of other things that make these models less accessible. You know, it's the complexity of setting up the infrastructure.
C:It's the complexity of getting started with some baseline model. You know, again, going back to ease of use, right?
B:How can this be made as easy as possible so as many companies as possible can leverage this technology?
C:And this is our focus now at Mosaic. It's just making state-of-the-art AI with a focus on generative AI accessible to any company out there. You know, not just kind of the usual
B:suspects of the big technology companies or big labs like OpenAI or Google Research.
A:Yeah, that makes sense. So, you know, I think generative AI might be at that point where an average person has heard the word but has no idea what it is. Like, it's not defined for them. And so it just occupies this sort of space, this soup of different things that they've seen and read about. And so this is a great time for us to really define it. Like, what is generative AI? And, you know, what's kind of the brief history there?
C:Yeah, so I would say generative AI refers to AI technology and more specifically deep learning
B:models that do a really incredible job generating media such as text, images, videos, or audio through very simple prompts. And I think typically what we see today in something like, you know, model like ChatGPT is, you put in text phrasing a request or a question and the model does a really incredible job following through on your request. And then, of course, there's also the kind of another poster child in Stable Diffusion where it's a text-to-image or text plus image-to-image model that, you know, just takes a simple natural language text prompt with a request to generate some visual and does an incredible job of generating that visual. So those are, you know, that's what generative AI is at its core. And I think we're just seeing the beginning of it, meaning these models will be much better at following through on your requests. Plus, they'll be able to generate very impressive additional, you know, mediums, right? And I think video is one such example where we're still early on in video generation and we'll see much more impressive things come along. But you can even expand this further, right? I recently did a keynote at a conference in Boston. I spent a lot of time creating the slides. Like I had the idea of what I want to talk about, but then a lot of time was spent creating the slides. I can definitely see generative AI sooner than later actually generating the slides for me, doing a pretty
C:good job at it.
A:Yeah. Yeah, that totally makes sense. You know, one of the things that always really inspired me, but I didn't know where it was going was, you know, unsupervised and self-supervised learning. I thought, and this goes way, way back. I had this idea where, and feel free to in the audience steal this idea. I'd love for someone to actually build this. But it was an idea where you would have sort of like a zombie game. So you'd be, you know, it's very typical. You fight the zombies. There's an infestation. You need to get the medical supplies or whatever. But you'd start in your own house. So the idea is, you know, I would plug into Google Maps or one of these map services. And so I would somehow render the game across the entire planet. And so whenever someone played the game, it would get their location from their phone. Actually, I guess Pokémon Go is kind of like this, isn't it? But you'd play in your own house. The thing that I ran into was, you know, how could I figure out which buildings should have which supplies? And so I thought, well, I could, you know, I could scrape Wikipedia and scrape the internet. And I could try to figure out what I wanted ultimately was for not me, for like just the computer to do the work to figure out, oh, if I sneak into this hospital, which is like a real hospital on Google Maps, that I would find medical supplies there. And if I sneak into a car dealership, I wouldn't find medical supplies. I would find gasoline or something, right? And so, you know, rather than having some content, you know, human in the loop there, I wanted it all to just get rendered, right? And then that kind of led to learning about these embeddings where somebody has, you know, scanned all of the internet through Common Crawl or Wikipedia or these things, and they figured out the similarity between words.
A:So you could actually see what's the similarity between hospital and medical kit. And that would be more similar than gas station and medical kit. And so you would use that to sort of generate your game here. And I found that to be, I mean, I never finished the project, but I found it to be just really inspiring how I created like the entire planet worth of supplies, like in these buildings. And it just was one of these kind of really satisfying moments. And so ever since then, I've always been into that. And there's just something magical about that. Maybe you could talk a little bit about, you know, how does that actually work? So, you know, when someone's on Mosaic or on, you know, SageMaker, which you also worked on, how do they build these giant models? Yeah. So
B:MosaicML offers a platform, right? Which is a platform for training and then deploying these models. I can start by maybe quickly describing, you know, when you are, you know, when someone wants to train such a model, let's say a large language model (LLM), what they need to do. And then, you know, how a platform can actually help them achieve that. So typically the first thing is it always starts with defining the task, right? You're trying to solve. And, you know, there's definitely a lot of general-purpose LLMs out there, right?
B:That just are, you know, their task is to basically be able to, on the business level, kind of be able to follow through on instructions, requests, questions, and do a good job responding to what a human is asking or prompting. And then when you look at the machine learning task, it's basically, you know, completion of the next word. So when you get an input sequence of words, which is a human sentence, complete the next word and then complete the next word after that and the next word after that. And when you do this a bunch of times, you get a coherent response. The first thing is, of course, you know, you need to figure out your data set, training data set. And I kind of breeze through some things. Of course they're pretty complex, but the data set and there's the model architecture, which covers things like, you know, both the architecture of the neural network. And, you know, we always knew our networks for these things today, as well as the scale of the model. Because you can, with the same architecture, you can, you can scale it, meaning the number of parameters across the different layers can be larger or smaller. And it has implications on the compute you'll need and the amount of data you'll need and get to that in a minute. Then you want to set up your training regime, meaning the hyperparameters for training, as well as your evaluation. How are you going to evaluate your model versus the original task that you had? The next step after that is going to be deploying the model once you have a good model. And that's almost like a, you know, related problem, but almost separate. Now, what's important to note about all of these things, as I said, is the scale of these models tend to be, you know, very large. And what we've, what
B:I think the community, the industry has found is that when you scale these transformer-based architectures up, you get what's called emergent behavior.
C:Meaning the model
B:is suddenly, it's like a step function chain where the model is suddenly able to handle new problems that, I mean, it wasn't explicitly optimized for. And they just emerge with bigger model size and more training data. One example for, you know, for that is like the ability to solve math problems. You know, I think both OpenAI and their, you know, GPT work and paper, Google with their Lambda paper, called out some of these emergent behaviors, including, you know, solving math problems, but other things as well.
A:Yeah, it's important to double-click on the scale. You know, when I was getting interested in large language models, I thought, well, you have a pretty decent GPU. I mean, it's, I don't know, three or four years old. It has, I don't know, one gigabyte of GPU memory or something. I don't know actually how much it has on the order of one or two gigabytes. And I thought, oh, I could just download the data set and train a model myself. And the answer is, I'll save everyone in the audience some time. You can't do this. So the data set is enormous. The models are enormous. Even if you want to fine-tune the model, you still have to load it into memory. And I think they said you needed like 60. Basically, you need a GPU that costs $2,000 if you want to do this yourself, which is out of my budget. So, yeah, you kind of have to use a service. I think this might be the, you know, maybe some of the computer vision models were like this. But for me, at least, this is the first time where you just, you can't try this at home. You can do it yourself, but you can't do it in your own house.
B:Yeah, exactly. So, I mean, just to give a few examples, I mean, you know, Meta published results of training a model called OPT-175, which is a 175 billion parameter model. They trained—they haven't published the weights, but they did publish a logbook and other details. It was trained for, you know, over a month on thousands of GPUs. And budget for that, you know, that kind of operation can be in the millions of dollars. And I'm not even talking about preparation before, deploying after, just the training.
A:Right. Is a parameter and a weight the same thing? When people say there's 7 billion parameters, is that 7 billion weights?
C:Yeah, usually that's how people refer to it. So, it's definitely immense. Although, I think what we're learning in an industry is that, you know, that has been a few things. So, I think, first of all, a model like OPT or even GPT-3, it was actually under-trained for its size. Meaning, you can take an actually a smaller model with less parameters, train it on more data,
B:and it will actually perform just as well or even better than a bigger model trained on less data.
A:Now, how do they know that? How do they measure that?
C:Yeah. So, there's a paper published by OpenMind. I think people tend to refer to it as the Chinchilla paper. I just don't remember the exact name of the paper. I think our community is, you know, having a really good sense of humor when choosing model names and paper names.
A:Right. It's all animals, right? There's Llama, Alpaca, Koala.
C:Yeah. But yeah, so the Chinchilla paper basically talks about the scaling laws. Meaning, you know, for a given, it's all about, you know, all referring to very similar architecture. So, transformer-based architecture and then different model
B:sizes. What's the amount of data? And typically, that's counted as number of tokens.
C:that are required to train it to its full capacity. Now, the way they—there is no fancy math kind of analysis. Unfortunately, I think, you know, machine learning is somewhat still feels like more like alchemy than science. So, what they do is just ran a bunch
B:of experiments and just took the same architecture, train it in different, you know, in different data sizes, data set sizes. And then measured the various evaluation metrics. And then kind of came up with their analysis. And they were able to train, I think it was a 60 billion parameter model on, I don't remember how many tokens. And it outperformed the valuation metrics of GPT-3 with 175 billion parameters.
C:So, although the number of parameters was, you know, a third of OpenAI's big model, the smaller model actually outperformed the bigger one.
A:Oh, interesting. Yeah.
C:So, and I think there's new kind of things being discovered by the day. I can tell you that; I'll give an example of two of our customers at Mosaic ML. One is Replit. So,
B:Replit is kind of a very popular online IDE. I'm sure the listeners, some of them at least are familiar. And for those that don't, definitely check it out: Replit.
C:is a fantastic tool for software developers. They built their code assistant, right, called Ghostwriter. And, you know, it does things that are pretty cool, making developers much more productive, like, you know, code completion. You know, it can create functions from comments. It can explain your code for you, et cetera. So, it's really a nice tool. The model behind it was trained on Mosaic ML platform. It's a 3 billion parameter model. You know, only, you quote, quote, unquote, only 3 billion. It's funny today, 3 billion considered a small model. Just a couple of years ago, it was considered huge. Yeah. But 3 billion parameter model trained on, I think, about 500 billion tokens of just code, open source code.
A:Is there a common place where I could get a scrape of all of GitHub or something? Like, how do you get that many tokens of code?
C:Yeah, there are a lot of data sets. There is, I think it's called Code, or sorry, it's called The Stack. That's an open source data set that you can access. Replit itself, obviously, because, right, people have, you know, using them for,
B:you know, they store a lot of code.
A:Right. Right. They
B:have access to some of that, of course, when the writers of that code allowed, right, to use it. Correct. Yeah. So, yeah. So, there's
C:definitely a lot of kind of specific data sets. Plus, they also tend to mix, right? So, usually, and that's where the alchemy part comes in, right? Usually, you want to mix your training data set so it's a bit balanced. So, you know, you want to mix a little bit of kind of natural
B:language from, you know, Wikipedia or other websites. Because, you know, even code comments, for example, they're, you know, they're written in plain English. They're not written in C++ or whatever other
C:programming language. So, people tend to do mixing. By the way, Replit published a fantastic blog post talking about how they build that model. And there's a lot of details there, including both the modeling side, data set management side, as well as the infrastructure. But just going back to kind of the TLDR, so it was a relatively small model specialized for, you know, being a code assistant. And it actually outperformed the, I think, two or three times larger OpenAI Codex. And that's the model. OpenAI Codex is the model behind GitHub Copilot.
C:So, I think what's interesting there is, first of all, you know, a relatively small company. I mean, Replit is a startup. It's a big startup. It's still a startup. Was able to train, able to train, you know, a model smaller than, you know, than another model, as well as outperforming it on quite a few valuation metrics. And we're able to do it with actually quite a small team.
A:Yeah, that's amazing. I think you touched on so many different things there. So, one thing is, you know, folks should definitely get familiar with doing things on the cloud. And I mean, we've talked about this for many shows. We've had folks, we literally just had a show on Kubernetes a few episodes ago. And so, you know, you kind of, you'll definitely have a lot of tools at your disposal, which will abstract away, you know, layer upon layer of this. But it's good to kind of get familiar with running things on the cloud. Because storing, you know, a 500 billion token data set on your desktop is probably out of the question. And definitely the models, yeah, it was just the capital cost would get out of control. And a lot of students, you know, if you're in college or high school, you know, often there's a whole bunch of different Amazon credits that you can get and all sorts of services there. Yeah, totally. Yeah, and then it sounds like the process for, you know, if you say, you know, I'm a musician and I want to train a model on, you know, all of these. Actually, music we talked about. Here's an even different one. You know, I'm really into theater and I want all of the, you know, English plays, you know, around Shakespeare's era, you know, all ingested into some large language model. You know, step one is to find that data set. And so, you know, it sounds like, you know, what I usually do, and Hagai, I'd love to get your advice on this too, but I usually just type what I want and then add the word 'data set' at the end into Google and try to see if someone's already done this. Do you have any tips for getting access to data?
B:I think a great place to start would be Hugging Face Hub. We have, you know, data set repositories there, and a lot of them, actually.
C:Actually, the problem is choosing the right one out of, you know, so many available. But Hugging Face Hub is a great place to start. Similarly, by the way, for starting with the model architecture. So the other nice thing is at this point, you don't have to do anything from scratch. There are data sets available, there are models available, and then there's a lot of training recipes available. And the best way to get started is just to start with something that is working and then hacking it, right? To fit your, you know, your specific needs, etc. You know, tweaking the data set mix, tweaking the task you're training your model for. That actually kind of brings me to a point where, you know, even taking a step back, you know, how can people leverage generative AI or even being more specific, large language models, for example. We've been talking about kind of training your own model a little bit now, and I think it's definitely, it's gotten much easier today. But it's still, you know, even when I look at, you know, Replit, which we just discussed, right? Training that model, Replit's model took about 500 GPUs running for about 10 days. Oh, wow. Yeah, so, you know, for those of us that are familiar with efforts at Google and Facebook, it sounds, you know, like something relatively small and fast. And definitely it is comparing to the bigger things that have been happening. But then if you approach it from the perspective of, you know, maybe a much smaller company or even just someone who just wants to do a cool project—a student or just someone doing a cool project on the side—that's definitely still big and requires a lot of, right, monetary investment. Yeah, yeah.
B:But there are other ways to actually get started with LLMs that are much faster and cheaper, and we can maybe talk about those as well.
A:Yeah, I think we'll definitely dive into that. Going back to something you said earlier, I do feel like it's very alchemic at the moment. And I think the reason for that, if you think about what actually, I want to say standardized or what took us to the next level in chemical alchemy was just the reproducibility and the affordability of experiments. So people could run just thousands of experiments, do them in parallel in very sanitized environments. You know, if we go to that factory in China, we'll have to wear those suits where we can't get any dust anywhere. And so everything has been extremely sanitized and as a result, just very reproducible. And so that's ultimately what turned alchemy from alchemy into chemistry. And so here, you're totally right. It's not only that it took those 500 machines for four days, but it's that it's probably their 20th or 30th model. So it's their 20th time dropping, you know, $2,000 to train this model. And they're constantly altering the data and mixing with it and hyperparameter tuning and all of that. So totally agree. I think, you know, very, very hard. You know, it's a big investment to train one of these from scratch. And so that's, I guess, where fine-tuning and other things come in. And what have you seen kind of on that front? Yeah.
C:Yeah. So, you know, I think two other alternatives that people approaches, people are taking that are slower, you know, have a lower barrier of entry or easier to get started. One is just using a model behind an API. And just, you know, OpenAI is, I think, is one service that is, you know, very broadly used already today, where basically it's very easy to get started, right? You just sign up with the service, get an API key, and then you have access to a really powerful general purpose model, you know. And what's nice is to have access to that capacity. All you need to do is kind of just write an API call and, you know, whatever is your favorite programming language. But it's fairly simple, and you don't need to know anything about machine learning. But still, you have that power, that capacity. You know, that's one good way to get started, especially to
B:create prototypes, right? Or to play around with the technology and understand what it's capable of.
A:Yeah, to that point, there's a lot you can do with engineering the prompt. I have a project that I'm working on with OpenAI. And, you know, it was giving me answers that were not unreasonable, but didn't fit the product that I was trying to build. And I kind of found that by, you know, playing around with the prompt, one of the tricks I found is you can kind of, if you know the beginning of the answer. So if you know, for example, it should start with 'Answer:', or a person's name: or the answer is. If you actually write that, it makes a huge, huge difference because it massively narrows the scope. So, for example, you know, I would ask a question. And this is actually, I wasn't using OpenAI. At this point, I was using Llama, which is Facebook's open-source LLM. So I, you know, I asked a question and then it generated another question and then another question and another question. And I was like, no, I want you to answer the question. And so I found it's as simple as putting, you know, 'Answer:' at the end of my question, you know, sentence—told it that, you know, it's expecting an answer. And so to your point, you know, even before you try anything with gradients and loss functions and all of that, just playing around with something like OpenAI's model or any model as a service can teach you about the problem you're trying to solve. Exactly.
B:Yeah. And I think
C:the prompt engineering is definitely another kind of field of alchemy, feels like it. But it today does have a really massive impact on the quality of responses you get from text completion. And I think an important thing to remember, I do think that folks who,
B:You know, understand how the sausage is made are the best prompt engineers out there. Although there's definitely, you know, I think if you just Google 'prompt engineering' today, you already see a lot of interesting kind of examples in Google to get started with. But one
C:An important thing to remember when you are creating your prompt is, remember, these models are typically trained with next word completion, right? So it's the autoregressive transformer models. They just try to predict the next word. So if you are giving them the beginning of the answer, for example, you already really made the problem much simpler for them, right? Because they don't have to guess your intent and get it right. You are indicating your intent to them by giving them the first few words of their answer. So that's a great way to squeeze better results out of them. I would say that personally, I expect these things to matter less and less because models will be just much better at understanding your intent.
B:Maybe even better than we are at some point. Yep. It's definitely getting there.
C:And the other interesting trend that is also pushing things in that way is what's called instruction fine tuning. So my guess would be that maybe with the Lama model you played with, it was the base Lama and not an instruction fine tuned version of Lama.
A:Right. It was just the naive stock vanilla.
C:Yeah. And with instruction fine-tuning, what people do is they take a base model. You know, Llama is pretty good. Seven. They have multiple sizes, but it's a pretty good model overall.
A:Yeah, I think I could only fit the seven billion on my computer.
C:Yeah. Yeah, that would have been my guess. But then they fine-tune that model to follow instructions. And then this means this model, you know, yeah, just has seen a lot of examples of an instruction and a response to that instruction.
B:And it's now just can do a better job following instructions. And then, you know, assuming...
A:How does that work? Like, how do they, how does the system know that the question has finished? Like, how do they actually do that fine-tuning?
C:Yeah. So typically, again, there's the art of how you format your data set. So typically, if you look at instruction fine-tune—most instruction fine-tune data sets, they'll have sort of a structure of 'instruction: [some instruction text]' and then 'response: [some text]'. Sometimes people use also, you know, like hashes to kind of highlight to the model the instruction and the response. And then if you follow, so when you create your prompt, if you follow that pattern of how instruction/response are written, the model kind of has an easier time, right, to follow your instruction.
C:Does?
A:This make sense? Yeah, this makes sense. Actually, you know, this—I don't want to take this on too much of a tangent, but how do you deal with if most of the data is just crawled off the internet? How do people deal with all the HTML and the markup? I mean, if you're reading the New York Times, and they italicize something, how do you—how does that get into the model?
C:Yeah, so there are different approaches. I mean, some models, you actually want them to be able to generate HTML, right? I'll put that aside for a minute. Let's assume for a minute you want your model only to be able to, you know, write text. So when you curate your training data set, you filter out things like HTML tags, you know, Markdown formats, and stuff like that. So your model only gets the text data and, you know, it doesn't see anything else. That makes sense. Yeah. For some models, you do want them to create HTML. In that case, you do want to preserve that, right? But again, your data set should not only understand HTML but also understand kind of the context of—you know, now you're asked to generate HTML or now you want to generate Python code. And then instruction fine-tuning is really helpful at explaining to the model that, hey, for a given response, it's expected to generate the
B:distributions that are more, you know, text distributions or Python code or whatnot. Got
A:it. And so I've seen this thing called LoRA. Is that—that seems pretty pivotal? Like the low-rank stuff seems pretty pivotal to the fine-tuning. What's the sort of connection there? Is instruction fine-tuning? How does that actually work?
C:Yeah. So I'm definitely not an expert in LoRA, and I think it's also still pretty early days. But with LoRA, the idea is that you can do fine-tuning much more efficiently by like decomposing matrices. And then your fine-tuning is more efficient, but then you can also apply—you can take a base model and apply fine-tuning by just, you know, applying the factorized matrices you got from LoRA.
A:But is that—is that the common thing? So if someone, let's say someone out there wants to fine-tune a model, let's continue with the screenwriting example. So someone takes Llama off the Internet and they want to adapt it to screenwriting. And let's say they've found the screenwriting data set and somehow they've converted it to Markdown or they've stripped out all the HTML. So they have the screenwriting data, they have the base model. How would they, you know, either using Mosaic or using something else, like how would they actually fine-tune that? Is there a module that everyone uses or something?
C:Yeah. So what you would typically do, first of all, you know, curate that data set with, you know, basically with just text. So in this case, let's say it's screenwriting. So what people would typically do is they'll curate a data set that includes a lot of examples of screenplays, text, and then they would take a base model that was pre-trained on general purpose language. So that model should be pretty good at, you know, English, you know, grammar, syntax, and understanding various concepts and all of that. But then that pre-trained model, they will just continue a training regime with that data set that they have. So they would fine-tune it on that data. Now we're not even getting into LoRA. LoRA is more like a way to do this in a more optimal manner, both for the fine-tuning and applying that fine-tune. I'll put that aside for a minute. There's a much more simple thing to do is just to take that data set you created and then just continue training the pre-trained model with that data. And what it will—it will kind of force the model's parameter to be better tuned for that kind of text, that kind of language. At Mosaic, we recently—the other thing I would say is that it's also much cheaper and faster than the pre-training. Because for the pre-training, you need to train it right on billions or even trillions of tokens. At Mosaic ML, we recently open-sourced a model called MPT-7B, so 7 billion parameters. It was trained on 1 trillion tokens of text, of language, which is huge. And this, you know, it cost us about $200,000 to train this model, this size on this number of tokens. But then to fine
C:tune it on, you know, we did an instruction fine-tuned version, a chat fine-tuned version, as well as a model that is able to actually write books or write stories, write, you know, fiction. That was much, much cheaper and much faster. Like, just to give you some data around that, it's all published in our blog post. But the base model took us about 10 days on 4x A100 GPUs. It's, you know, almost the best GPUs out there, except for the H100s just coming up. So it cost us about $200,000.
A:So those 400 GPUs for four days... For 10 days. Oh, 10 days. Okay, so 400—so that's 4,000 GPU days, cost $200,000. Yeah, it's not cheap.
C:Yeah, yeah, it's definitely not cheap. But luckily, we've open-sourced it with the wait, so anyone can build on top of it.
A:Now, how does that work? Do they need your PyTorch code? They would, right? Yeah.
C:So the PyTorch code is defining the model architecture. So that code has been open-sourced, obviously. But there is, you know... And it is... There are a bunch of optimizations we'd put in there. But, you know, it's PyTorch code. A little bit of C++ for some of the optimized operators. But that's it. And then there's the weight itself, which is typically stored in a separate file. But then, you know, it's just PyTorch weight. So we have example code, but basically, once you instantiate the class for your model, you
B:just use the standard PyTorch interface to load the parameters into the model.
A:Yeah, getting... Kind of going full circle. You know, computer vision, we've been doing this for a while, where, you know, you have a trunk model, and then you have a bunch of heads for that model. One head detects, you know, traffic lights. Another head detects pedestrians, stuff like that. And so it's well-traveled ground there. I wonder how data efficient it is. I guess there's no way to really know, right? You try to amp up the learning rate. But there's not really a scientific way to say, okay, this is how many play data scripts you need to have a model that's reasonable. It's one of these things that's like, really hard to calculate.
C:It's really hard. Yes, it's still more of an empirical trial and error. But what's interesting is, you know, so the version of MPT-7B that model we open source, that version that is instruction fine-tuned, we took the MPT-7B, we took, you know, a data set. I think we combined a couple of data sets that are just out there for, you know, instruction. I think it was the Dolly data set from Databricks. So about 10 million tokens of data for instruction fine-tuning. And basically within, like, a couple of hours on one node with eight GPUs, we fine-tuned that model. So just to put things in perspective, the base model took us 10 days of hundreds of GPUs costing $200,000 to train.
C:But to take that model and then instruction fine-tuning for following instructions, like we discussed earlier, that took us two hours with just eight GPUs costing us about $40. That's it. Wow. So this is definitely within reach, right, for anyone out there, right?
A:It's like the difference between, you know, buying a tractor or buying, you know, the seeds to plant, right? Yeah. Yeah.
C:Yeah. It's a huge difference. And that actually kind of is a segue to, you know, we spoke about the first way to leverage LLMs: just call an API. The second way is take an open-source model and either use it as is or fine-tune it for your needs. Either way, you know, it's fairly accessible today, fairly cheap and available. And
A:So what about serving? I want to use the time we have left to talk about that. Let's say you try Llama on Hugging Face, and, you know, a lot of these Hugging Face sites have a web UI where you can ask questions. You know, it's not as sophisticated as OpenAI's site, but it's good enough. You can type in your question; it'll generate an answer. And you say, 'Yep, this is good enough.' You fine-tune a model. And now you want to build a website or some service for people. You had a—because the model can't even fit on your GPU. Like, how do you even, how do you serve the model? Do people use CPUs to serve the model? Is that a thing? What's the story there?
C:So if you take, you know, for example, MPT-7B or the Llama-7B—so it's a 7 billion parameter model, right? Every parameter, let's say, is, you know, 2 to 4 bytes, depending if you're using FP16 or FP32. Typically serving today is done with FP16 or more specifically BF16, you know. So 7 billion parameters times 2 bytes, that's, you know, 14 gigabytes. That actually does fit on kind of good GPUs today, like the NVIDIA A100 40 gigs or even the A10s, 24 gigs or 32 gigs of memory. So these production-grade GPUs can, you know, one GPU can hold such a model. But then there's other complexity there. I mean, you know, first of all, when you're talking about text generation, you want it to be fast and efficient. Now, remember, the way these models work is they generate one word after the other, or one token, actually, after the other. So actually the latency of inference matters a lot, especially for interactive applications. Because, you know, a typical response to a model is definitely not just one word, right? It typically has, you know, I'd say tens, some questions, even hundreds, right, of tokens. So we want the inference to be as optimal as possible. And that's definitely one, I think, area of development. Even if you play today with models like ChatGPT, it's streaming the output word by word, but you can still see that it, you know, it takes a while. But so setting up optimized inference is, you know, one thing, one area where there's definitely more and more tooling. And I think there's more room for the machine learning community to invest in.
A:Now, one thing about that, you know, with regular deep learning models, like predicting the probability of an event, you would want to serve on the CPU because you don't have a batch versus in training, you know, you have a batch of data. Is that true here or is even just generating one word better on a GPU?
C:Yeah, so GPUs definitely can, I think, where they become really cost efficient there is, like you said, handling a batch. Now, the tricky thing is when you want to generate output for an input sequence, and let's say you want to generate, you know, 50 tokens, you have to first calculate token number, you know, response token number zero, and then you feed it in to generate token response number one, right? And then, and so on and so forth. So there is a sequential angle to this. Where you can do batch even for inference is when you have, you know, you have a service and you have multiple requests, different requests at the same time, then you can batch.
A:Right. Yeah.
C:But then you need scale to be able to handle something like that, or it's an offline process, like, you know, batch inference, which is typically offline, and then you can do those things. Going back to the question of a CPU. So I think that the main advantage of a CPU is just cost, right? Because GPUs are very expensive. I know Intel has been doing a lot of work to get their new CPU generation to be pretty good at handling transformer architecture, so people can use it. You know, I have yet to see kind of inference of these kind of models work well on CPU. But I know it is an area, actually, Intel is working on. I even saw a demo. They did something which looked pretty promising. But then when you look at the details, it was their newest generation of CPUs. And actually, the cost of that CPU, at least on AWS, was actually the same as the cost of a low-end GPU.
C:So performance was good, but then cost, there was no difference. So, yeah. I think, you know, if we look at the trend of, you know, computing and processors, is that the cost of running complex workloads, you know, always goes down, right? Right, right. And I expect this to happen. So there's been a lot of interesting work by the community of folks kind of allowing you to run, you know, these models on commodity hardware. There's something called Llama CPP, I think, that someone hacked together, where it's a, you know, super efficient, you know, low-level implementation of inference for Llama on a commodity CPU. So I think it will definitely get there, although we're not there now. It actually brings me to, there's another important angle of inference. And again, that's like a differentiating factor, I think, between using an API, a model behind an API versus using either your own model or an open-source model. And that issue is a huge issue, actually, of data privacy. You know, when you are doing, leveraging a model behind an API, you have to send your data outside of a premise into another service.
A:Yeah, this was in the news. I think, I don't want to, I hesitate to get the company wrong here, but I'm pretty sure it was Samsung. The employees were using OpenAI. And then, yeah, somehow, I don't know what actually happened there, but somehow OpenAI got their data or their schematics or something.
C:Yeah. So, yeah, that was a big story in the news where, I think, engineers at Samsung, that was the report, they were using ChatGPT to kind of write down some of their plans. And then that data somehow leaked. It's not clear if it was leaked because OpenAI is using data people send their service for inference. They're using it to retrain the model. And then the model memorizes some of what it's seeing, and then it leaked in the response somewhere else.
A:Yeah, I think somebody else searched for, like, the model number. You know, a competitor was like, 'Tell me more about the, you know, Samsung S4000.' And OpenAI was like, 'Sure, here's what I know.'
C:Yeah, so I don't know if OpenAI, I don't know if they changed it yet, but if you're using the free version of ChatGPT, the default is opt-in, meaning you're—many people are not aware of it—but you're by default opt-in to share your data with OpenAI and then use it for training their model and whatnot. And that's really something to pay attention to. And I think the industry needs to mature a little bit. And I also, my personal take is that I think there should be legislation that governs how these models are used and, you know, the privacy of data and all that stuff. But that's just something for everyone to remember. Like, there are a lot of advantages of using a model behind an API, and we went through that—those, you know, those advantages. But one drawback is definitely, if you care about data privacy, if you're like in finance or healthcare or similar industries, you probably don't want to send your data over the wire somewhere, or, you know, you want to get very strong, right, guarantees from your service provider about how this data is going to be used or is not going to be used.
A:Yeah, I mean, I think the old adage, 'you get what you pay for,' applies here. You know, if you're using a free API and OpenAI is spending—you know, we talked about hundreds of thousands of dollars, you know, on keeping these GPU machines up, even to do inference, you know, you're giving something back, right? And so, you know, using Hugging Face or Mosaic or one of these services where you're paying for the service, you could probably get much better privacy guarantees.
C:Yeah. The other thing, by the way, is also cost. So what people are finding out is when they use models behind APIs, then I think at small scale and prototyping, it's very cheap, cost-effective. But when, if and when this becomes a core use case for your application, it just—it's becoming very expensive, especially if you have a kind of a large-scale operation. So that's also something, you know, I think people are realizing sometimes a bit too late, and that's also something to factor in because you can get much better cost efficiency if you are serving your own model or either an open-source model or a model you train. If you serve it on your own infrastructure, of course, you need to set this up. And there are some services that help you do that, including Mosaic ML. But it's much more cost-efficient than actually using an external service that, you know, has a margin and whatnot. So cost is becoming a thing at scale.
A:Yeah, I think, you know, if everyone uses OpenAI, that it's driving towards a monopoly, which just putting my economist hat on results in infinite profitability for OpenAI. You know, conversely, if everyone's using their own model and it's just a matter of who can host your model, that's driving to zero profitability or like infinite competition, which is good for you as a person who wants to use the model. So it sounds like, you know, it doesn't take a lot of money or time. It really takes you kind of out there building those skills to, you know, grab the right data set, grab the right model, try a bunch of different fine-tuning and learn how that system works. And in the end, end up with a model that can create some unique value for you or for an addressable market that you have. So one thing about—I want to dive into Mosaic, the company here. So, you know, there's a ton of folks out there, listeners who are really interested in this technology, just like there are people across the world and all disciplines interested. And they would love to get, you know, their foot in the door, work more with AI and machine learning and generative AI. And so talk a little bit about, you know, what's it like at Mosaic and what kind of folks you're trying to hire for and just general kind of job-seeking advice.
C:Yeah. So let's start with Mosaic. So we're still pretty early stage startup. We're now about 60 employees. Most of us are in SF, but we also have a couple of other offices, including in New York and even Copenhagen. And we're really, you know, kind of relatively small team, just trying to do a good thing by making state-of-the-art AI with the focus of generative AI, just more accessible. So any organization out there—that's what we're out to do. Any organization out there should be able to, you know, leverage these models in whatever way works for them, you know, model behind an API, which we offer or open-source models that we open source make available to the community or pre-training and fine-tuning your own model. And we think there is, you know, great business opportunity with that. And it also—it's going to kind of really help kind of the next generation of startups as well as big enterprises to use AI. Yeah. So, and then we're hiring actually. So, you know, the business has been growing well. You know, we're seeing good traction with customers. We're seeing good traction with the community. And we are growing the team and we're hiring across, you know, software engineers, both for our cloud platform as well as for machine learning runtimes. We hire researchers who are a fantastic research team that is using our platform to, you know, build these amazing models like MPT that we've open-sourced. So there's researchers. We are hiring interns across these both teams, both the research team and the engineering team. And then we hire for other functions, you know, we hire across product, technical program management, recruiting. So it's really kind of, you know, I feel that the team is kind of hitting on all cylinders. And then as part of that, we're also continuing our growth.
C:Yeah, it's really exciting. And, you know, I guess I'm biased, but I'm really excited about both the mission of what we're trying to do, as well as kind of the culture and team at Mosaic.
A:Cool. That makes sense. And so, as we talked about, for a relatively small sum, you can take the MPT model and you can fine-tune it to do playwriting. If someone does playwriting of Shakespeare, let me know. I would love to just add me on Twitter. I would love to see that. But, I think the best way to get noticed at a company like Mosaic is to use the product, right? And to build something and have a portfolio of accomplishments that you could do relatively low cost. Kind of adjacent to that, so if someone's a student—you know, a college student, even high school student—is Mosaic a tool for them? Is it something that they should know about for when they go into industry? Is there sort of a free tier? Like, what is the story like there?
C:Yeah, great question. So at Mosaic, we do have a few open-source components that anyone can use. So there's the models I mentioned earlier, the MPT series of models, but there's also a training library called Composer. It's a PyTorch training library, which just helps train PyTorch models faster and better. And there's also a streaming dataset library, which is really useful for training models when you need to stream all the training data from cloud buckets. However, the product itself, so far it's been really geared towards enterprises, meaning there's no free tier or community tier where people can just easily get started with the platform. And the reason it was designed this way is just kind of how the company evolved, right? You know, at the end of the day, it is a business, and we were going after enterprises initially to establish the business. And that has gone really well. And the next thing on our plate is offering some sort of a community tier where a broader set of practitioners out there can get started using what we have to offer. And this will come soon. And I think at that point, definitely it's going to be very easy for anyone to just get started, try us out—either use our models as APIs or fine-tune either our models or any model out there that is available on the Hugging Face Hub or GitHub or anywhere else, as well as, of course, kind of pre-training your own model. Although this tends to usually cater more to the enterprises that have enough data and have the budgets to pre-train these models. So stay tuned. It will come. And at that point, you know, it's going to be amazing. I'm really looking forward to that moment where we kind of open the floodgates and allow the community to really engage fully with us.
A:Cool. That makes sense. I mean, in the meantime, folks can get the MPT model. They can get all of the weights, the PyTorch code so that they can continue training on their own data set. And there's a whole myriad of different services out there. So if this sounds cool to you, you should put in the sort of sweat equity here to build something neat. Definitely email us, tag us on social media with anything you build. We've actually, inside baseball here, we've been really good at placing people. I've gotten emails lately from people who have been on the show representing a variety of different companies saying, 'Oh, you know, we have our first intern who found out about us from the show.' So I think that's a real testament to the audience out there. You folks are super motivated, highly technical, which is really great to see that we're able to sort of Patrick and I can kind of connect to interested parties here, which is awesome. So we'll put the links to Mosaic ML and their careers page, all of that on the site, if that's something that interests you folks out there. Hi, Guy. Thank you so much for coming on the show. I think we did an awesome job kind of covering in the audio format how this whole system is evolving, how it works technically. There'll be tons of resources in the show notes for people to follow up. And I want to just really appreciate you spending time with us today.
C:Thank you, Jason. Thanks for having me. And I really enjoyed this chat.
A:Cool. Thanks, everyone out there. Have a good one. Programming Throwdown is distributed under a Creative Commons Attribution-ShareAlike 2.0 license. You're free to share, copy, distribute, transmit the work, to remix, adapt the work, but you must provide attribution to Patrick and I and ShareAlike.
Transcript supplied by the publisher with the episode.
by Patrick Wheeler and Jason Gauci · English · Tech & Science
Programming Throwdown educates Computer Scientists and Software Engineers on a cavalcade of programming and tech topics. Every show will cover a new programming language, so listeners will be able to speak intelligently about any programming language.
E164 · 11 Sep 2023 · 1 hr 31 min
Things to consider when choosing a database Speed & Latency Consistency, ACID Compliance Scalability Language support & Developer Experience Relational vs. NoSQL) Data types Security Database environment Client vs Server access Info on Kris & Harper: Website: harperdb.io Twitter: @harperdbio, @kriszyp Github: @HarperDB, @kriszyp.
E163 · 14 Aug 2023 · 1 hr 29 min
Patrick and Jason break down recursion as a practical problem-solving technique rather than a classroom trick. They cover base cases, recursive steps, common pitfalls such as nontermination and stack limits, and real applications in trees, graphs, and divide-and-conquer algorithms.
E162 · 24 Jul 2023 · 1 hr 8 min
In the latest episode of Programming Throwdown, we delve into the captivating world of interactive fiction. We explore: Wordnet, Inform, and how games in the past have been the forerunners of today’s NLP challenges.
E160 · 26 Jun 2023 · 1 hr 30 min
It’s a question that may seem easy to answer on the surface, but in truth hides more complexity than people expect. In today’s episode, we tackle the latest on AI, creative endeavors, and more before diving into the meaty discussion of position localization.
E189 · 24 Aug 2026 · 1 hr 23 min
E188 · 9 Jul 2026 · 1 hr 36 min
E187 · 2 May 2026 · 1 hr 38 min
E186 · 3 Feb 2026 · 1 hr 28 min
Patrick and Jason discuss what it means to become a manager and how the role differs from individual engineering work. They cover hiring, coaching, performance management, team goals, and when moving into management is or is not the right choice.
E185 · 4 Nov 2025 · 1 hr 32 min
Patrick and Jason break down workflow orchestrators and why they matter for batch jobs, long-running tasks, and resumable distributed systems. They compare tools such as Airflow, Dagster, Temporal, Ray, and Kubeflow while explaining the infrastructure patterns behind them.
E184 · 23 Sep 2025 · 1 hr 31 min
Patrick and Jason explain asynchronous programming and how it differs from traditional multithreading and multiprocessing. They cover coroutines, blocking versus non-blocking operations, promises, callbacks, async/await, and the tradeoffs behind each approach.