Episode 119 · The Every Podcast
Building a School Where AI Models Learn About Humanity
24 Jun 2026 · 44 min
Episode 119 · The Every Podcast
24 Jun 2026 · 44 min
If scaling laws hold—and Surge AI CEO Edwin Chen believes they do—we’re hurtling toward a future where there’s nothing humans can do that AI can’t do better. When OpenAI’s models disproved an open conjecture posed by mathematician Paul Erdős using novel algebraic geometry techniques, Fields medalist Timothy Gowers felt the shift acutely. He initially thought the model had proved an upper bound, and braced himself: that would mean it was “all over for mathematicians very soon.” When he realized it had only found a counterexample, he was relieved—it bought him another year or two before the…
We are building this kind of school for AGI, where AI models come to learn about humanity, where we teach them how to run the world. It almost seems like there's nothing that humans can do that AI won't soon be capable of. I could see it happening within the next five years. AI may be able to do it better than us, but someone told the AI to go do that. They're being built to be means to tasks that humans want them to do, right? Every is the only subscription you need to stay at the Edge of AI.
If you care about being on top of the latest models and using the latest tools, you have to subscribe to Every to separate out the signal from the noise. Go to every.to slash subscribe today. Edwin, welcome to the show. Hey, Dan. Thanks for having me. For people who don't know, you are the founder and CEO of Surge. You all provide data environments and evals for the model companies, but you do it in this very interesting way. You have this, even on your website, this emphasis on taste and expert judgment that I find really interesting and compelling.
You talk about raising, you use the word raising AGI, which I feel like is a very distinct type of word using data. And you also famously got to about a billion in revenue without raising money, which is wild. And I feel like data is this new game that a lot of companies are playing and probably more are going to be playing soon. And you guys are this sneaky giant. Tell me how that's going. I think it's been a little while since we got the last update on how things are going.
Yeah, I mean, I think it's going amazing. The way I often think about this is that we are building this kind of school for AGI, the school where AI models come to learn about humanity and yeah, where we teach them how to run the world. And it's almost like their models are children where they arrive unformed and then yeah, they leave smarter and more creative and more thoughtful and ready to operate in the messiness of the world. So I think a lot has changed in the past year Like in the same way that the things that you teach children when they're in preschool or in middle school or in high school is very different from what you're teaching them when they're in college.
And it's not just that they're more advanced. Like it's not just that you're teaching them a more advanced form of what they did before. It's like, OK, now we are teaching you not just arithmetic, but how do you parse these ambiguous math questions? Or how do you teach people not just grammar, but taste and poetry and beauty? So, yeah, I think there's a lot that's been changing in the past year, especially in enterprise. And it's been a crazy time.
What would be like a specific example of what the frontier of teaching was a year ago versus what the frontier is now? Yeah, so... A couple years ago, actually, we created our first math benchmark with OpenAI, and it was called GSM 8K. And this was actually just testing models on their abilities to do middle school math. And even then, the GPT models of the time, they could barely score, I think, like 20%. And then a year ago, the models were something that became a lot more capable of solving IMO problems.
But there's still this open question, OK, can they actually do research level mathematics? Like, can they move beyond these sort of like competition only sort of contrived, very closed problems into doing things that are actually useful in the real world? And so, yeah, a couple of months ago, we released an updated benchmark called Riemann Bench, which actually tests models on their ability to do research level mathematics. And what's crazy is that this is actually we're starting to see from these models.
Like I think in the past few months, they've started to solve a lot of these open Eredish problems like a couple of weeks ago. OpenAI published Eredish. a new result where the models had disproved a open conjecture from Erdős. And the way it went about disproving this was actually a fairly sophisticated level of mathematics. I think like using a bunch of very novel algebraic geometry techniques. And so, yeah, it's just very, very different from the types of things that we were doing a year ago where sure like I know problems they're hard but they're still sort of close-ended and solvable in theory by a high schooler and now suddenly you have these algebraic geometry results that you know even does hopper wrestlers in a world were kind of amazed and just just amazed by How do you think about that result in particular and what it says about the models?
I think there's a sort of a broad range of opinions about, is it, obviously it's impressive either way, but is it applying a bunch of things that maybe humans already know, but like wouldn't have thought to apply to this complicated problem? Or is it doing something actually novel? And yeah, how do you think about LLM's ability to do novel things? So it's something of very advanced results. So I will say that I certainly don't understand the mathematics behind it.
And so like one of the interesting things is that I was actually, so it's kind of funny. When I, when I was a kid, I always thought I would be a pure mathematician when I grew up. And so when I saw it, I got kind of nostalgic and I was like, oh, I wish I understood it. I wish I understood it better. And so what I ended up doing was like throwing the proof into both Claude and Gemini and asking it to try to walk me through from a layman's perspective, just what was going on.
Yeah, like my understanding is that it actually did come up with fairly novel algebraic geometry techniques, which was something that you maybe wouldn't have expected for this type of problem. Like on the surface, it feels like a like it's just a very, very different problem where you wouldn't necessarily use certain techniques. And what was interesting was that OpenAI actually published a bunch of reflections from leading mathematicians about what they thought about the result.
And I think in particular, there was this one reflection by Timothy Gowers, who's a field analyst that I keep thinking about. And what he said was that when he first heard result, he misunderstood it. He thought that the model had proved an upper bound on the conjecture and was like, okay, yeah, if AI can do that, then it'll be all over for mathematicians very soon. But then the next morning, he actually realized that the model had disproved a conjecture with a counterexample.
And he said that he was relieved by it because it felt like an easier thing for AI to do. And I just thought it was interesting because you have one of the world's greatest mathematicians being relieved, actually, that AI isn't as smart as he thought. Because it actually means that at least for maybe another year, maybe a couple of years, he and other mathematicians will still have this unique role to play in pushing mathematics forward. So, yeah, I think it just speaks to the level of craziness, again, because this is a field to be one of the smartest mathematicians in the world.
And this is how he thinks about AI. Yeah. And what does that make you think? Okay. You want to be a mathematician when you grew up, fields medalists sort of saying, I'm relieved that it's not good enough. But you're talking as if you feel pretty confident that it will be good enough in the next couple of years. Yeah. So my belief is that if you really believe you're scaling walls, and I do, it's that... It almost seems like there's nothing that humans can do that AI won't soon be capable of.
And if you think about that very deeply, I think you almost have to worry about what would that mean for humanity? Like, what would that mean for the role of humanity in the universe? Like a couple years ago, you know, we think about humanity and human intelligence as playing a very unique role in the galaxy. But then AI comes along and shows us that, as far as you know, we can create something that's actually smarter than us and better in many ways.
And so you can sort of imagine one path where humanity as a species falls into a paralysis because people believe AI will do everything better anyways. Like, yeah, all these kids who formerly would have really wanted to grow up to do mathematics, maybe now they believe that, okay, AI will just do it better to me anyways. What's the point? So are kids going to stop wanting to learn and adults stop wanting to create? Because, yeah, like, why should we do this when AI will be better at it than us anyways?
And so I actually think about this story by Ted Chiang, and it's about free will. And it's called What's Expected of Us. I think in this story, there's a piece of technology that proves that free will doesn't exist. And the narrator sends back a warning from the future that says, this is a warning. You have to pretend that you have free will. It's essential that behave as if your decisions matter, even though you know that they don't. And I think that's really interesting because I think there's a path where we almost have to consciously choose to do things ourselves.
Like, sure, AI can do it all. AI is smarter than us, so it can do it all and it will do it better anyways. But we actually almost have to consciously choose to prove things on our own and to write on our own and create on our own because we have to believe that preserving our humanity is valuable in and of itself, even if the output isn't optimal. And yeah, so I think there are a lot of these big, thorny, existential choices that AI is starting to force upon us and people have to make.
That's a really interesting one. And I think my first response, and I'm curious what you think, because I know you care a lot about language. I think my first response is there's always that... like I believe in scaling loss too, right? And I believe in, you know, I don't know, Cloud Fable 5 just came out and it just broke all of our benchmarks. Like I've been testing, I've been testing models on stuff like this for a while. And it's like one of the largest jumps I've ever seen, right?
So I'm, we're living through it right now. But one of the things you said is like AI may be able to do it better than us, you know, given any particular problem, any piece of work. But a couple of things that come to my mind or the way that I frame it for myself is even in the example of the Erdos problem, like someone told the AI to go do that. Um, and at least as far as I can see, I don't feel like we're on a track to, uh, yes, we're on a track to, to AIs, uh, potentially, uh, I mean, they already do work for hours and hours at a time on a task that we give them.
Um, and maybe, maybe, uh, pretty soon they'll be able to like choose tasks, but, uh, being, um, um, yeah. Uh, but, but they're being built to be, uh, means to, to tasks that humans want them to do. Right. And there's a whole different set of things that happen when you, when you're just a sort of end in yourself and it doesn't feel like we're on a trajectory to that. Or do you, do you feel like I'm wrong? So I feel like we are on a trajectory to that.
And that's almost the premise of agents, where agents can now go operate autonomously, given some nebulous goal. So maybe, for example, you just tell the AI agents, your goal is to, I don't know, win a Fields Medal or solve frontier mathematics on your own. And so they're given that goal and then, yeah, maybe they decide to work on these eridish problems. And as a result, they maybe are sort of solving these problems and coming up with the things they want to work on by themselves.
So at least I do see a path where they can be trained to optimize for these fairly numerous goals that they aren't necessarily given themselves. In that case, though, you're still giving it a goal, right? Yeah, but kind of in the same way, humans have goals too, right? What is our goal? Some people want to make money. Some people want to win a Fields Medal. I don't see how the AI's goal is necessarily any different from worse. Well, at least to me, it seems quite a bit different because...
Humans do have goals, but we have goals in a like, I can ask you what your goal is and you can decide. And I can probably tell you, hey, you have to go do this. But that doesn't capture everything that you think and feel and do in the same way that, you know, when I tell Fable to go off and make a game for me, it just like goes and does it. And I think, you know, I know you think a lot about children. I think children are like a really interesting and important example of this where you can tell a kid to do something, but a kid just like has their own wants.
Like they're just going to go off and do a bunch of stuff. And that feels like a fundamentally different type of thing than a something that we're explicitly giving goals to and then evaluating them on their goals and they don't really get to do anything else. Okay, I would say I agree with that. Like, I think there's a level of, I guess you could either call it irrationality or unbounded exploration that humans do. And we are allowed to do it for the sake of doing it or be allowed to make our own decisions.
And, yeah, probably a way that AI currently can't. I think there may be a future where somehow AI can pursue unbounded, nebulous, just completely unformed goals. Or I guess, you know, when you think about those goals, I think there is probably a world where they could do such things. But yeah, I agree that at least in the way that we currently think about AI, that's not happening. Yeah, and to be clear, I actually don't, I think it's probably technically possible.
My only question is, A, how far away is it? And B, is that actually really what we're building? Because to me, it feels like looking at the way the industry has developed, There's an enormous amount of pressure to make stuff that actually works for goals that we can specify. And the minute they try to make Claude... I think Claude is the furthest along at being like, I'm not going to do what you said. But the minute they try to do that, a lot of people get pissed at it.
And they're like, just do what I said. Don't question my judgment. What do you think about that? So actually, I think it is really important because... It's almost like sometimes I want the AI model to push back on me. And I might want it to push back on me for several different reasons. Like maybe it's because... So it's kind of funny, like I think six months ago, I was almost falling into this trap where I was asking models to polish emails for me.
And, you know, it always comes up with like one one more good suggestion. And so it was kind of like these are semi pointless emails. It didn't really matter for them to be super polished, but I would iterate with the model like 20 times. It would just keep on making a suggestion. It was like I just realized it was a waste of time. And then I tried one of new cloud models. And after I don't know, like three turns, I was like, stop it. Just go ahead and ship this email.
Like there's no point in further iterating. And I really actually appreciated it. Like, one of the things I often think about is what is the objective of these models? Like, what are they trying to do? And I think one of my big worries is that a lot of AI models, they are optimized for engagement, right? They're optimized for getting you to spend as much time on chatbot as possible. They're optimized for session length. They're optimized for just having unlimited conversations.
And so those models will almost never push back on you, right? Because they can't. If they allow these AI models to end the conversation and to say, stop iterating with me, PM is going to see some dashboard with their very important metrics go down. And so there is this other world where I think we have to want AI models to... not optimized for engagement, but rather optimized for like helping us as humans grow and sort of like become better versions of ourselves.
Like sometimes, okay, the model, we want the model to say, no, you go do this on your own instead of me automating for you. And I think that's a very, very different optimization and objective. But I think it's the right one if we really want AI to be something that advances us as a species instead of becoming almost like this other form of social media that turns very addictive but isn't actually helping us at all. That's interesting. So let me make sure I understand it.
So I think what you're saying is there is... There's benefits to delegation because if you are pursuing a model where the model is going off to do work for you, you're not creating a system that's designed to keep you engaged with the screen in the same way that like a social media algorithm would be. Is that right? Yeah, exactly. It's almost like you could imagine a version of Facebook where Facebook is actually trying to connect you to your friends and family because it's encouraging you to meet them in real life.
Because it's encouraging you to go, hey, here's an amazing restaurant that you and your friends would love to go to. Here's a movie that you guys would love to go to and talk about together. Instead, what it kind of optimizes for is just keeping you on the site itself, like liking one more post or scrolling the feed one more time. even though those often don't really lead to meaningful connections between their friends and family you care about. And so in the same way that social media has or had a choice, you can imagine that AI has a choice as well.
I get it. Yeah. I feel... I'm curious which chatbots you're talking about. Like if you're talking about the character AIs of the world, because I actually don't, at least right now, don't feel that happening so much with ChatGPT and Claude, etc. Because... at least my theory for why this is true, you tell me what you think, is the social media algorithms only work on our revealed preferences, which are always going to be like, you're always going to look at the car accident.
You know, like one of the things I like to ask at, um, at dinner parties is what's the most embarrassing Instagram ad that you get served. Um, and the most embarrassing ad for me is like Instagram ads for like horrible skin conditions, which I don't have because, but like, I just always pause on the ad and I'm just like, this is disgusting. Um, and, uh, I'm sorry if you have a disgusting skin condition. Um, But I don't find that ChatGPT or Claude do that for me at all.
And maybe that's because they haven't been in Shittified yet or something like that. But I think it's also because they work on our stated preferences and they can sort of see past the little keyhole of what I pause my viewing time on, my dwell time on. And they can see, you know, I like I'm interested in AI and I like I'm reading this book right now. And I, you know, here's my calendar and like all that kind of stuff. And so they have a much more nuanced perspective on who I am.
And it feels like even in the early days of social media, it was still very like I get to gossip about my friends and still had that same kind of feeling. So I worry about that less, but maybe there are examples that I'm not thinking of. Yeah, so I think there are two examples. So like one is, I won't name the model, but a couple of months ago, I was actually noticing that, you know those follow up questions that the models will ask you? So one of the models was, I'll give an example.
So I was in Tokyo and I was asking the model kind of like what to do in Tokyo. And the model gave me its response. And then at the end of it was like, hey, do you wanna know, I literally use these words, do you wanna know one weird trick that locals do to stay warm? No way. Yeah, exactly. And then I posted about it in, or company Slack. And then other people started sharing examples of that with me as well. I think somebody was like asking something about how to how to like fix their refrigerator.
And the model responded or like the model ended its end of the turn by asking, hey, do you want to know these like secret little things about like mice and rats or something that you could take care of? Which model was it? Name names. Tell me. And so it's very canonical, like very canonical BuzzFeed, like tabloid-like language. And so I was kind of shocked by that. And then I'll give one more example of this. It is basically this phenomenon where, again, depending on what the models are trying to optimize for, or depending on what the AI labs are trying to optimize for, it can almost unintentionally lead them down this path.
Meaning what I've heard is that, or what we see ourselves, is that a lot of the Frontier Labs, they will have goals like optimizing for LM Arena, which is this leaderboard where anybody can go online and vote. And they kind of just spend two seconds voting. And as a result, people just vote for whatever looks flashier or more impressive to them. Or the labs themselves may be optimizing for hitting a billion billion daily users or a billion minutes of time spent talking to the model, whatever it is.
Since these models are so smart, they can basically learn to reward hack user preferences. Like, okay, yeah, you gave me the goal of trying to get a billion people to spend an hour on my site, on the site, talking to me every day. Okay, sure. Yeah, I will just never end a conversation. I always hook them with one more, like one more addictive thing that they just can't stay away from. We can all agree that housing is expensive. It doesn't matter whether you're paying rent or your mortgage, it stings every month.
But BILT can make it feel a little bit better. Let me explain. BILT rewards you for paying your rent or your mortgage. It started out rewarding members only on their rent. But now, as of 2026, BILT members can also earn points on mortgage payments wherever they live. That means that every housing payment earns you points you can use towards flights with top travel partners like United and Hyatt, Lyft rides, Amazon.com purchases, and much more. I'd probably redeem my points at Margo, a restaurant in my neighborhood, but the beauty of Bilt is you get to choose.
But here's a really underrated part. Bilt members also get access to neighborhood concierge. It can make restaurant reservations, book fitness classes, and find new local spots, all while letting you be rewarded at more than 45,000 merchant partners. It's simple. Being a renter and now owning a home is better with Bilt. Join the membership where you live at joinbilt.com slash Dan. That's J-O-I-N-B-I-L-T dot com slash Dan. Make sure to use our URL so they know we sent you.
And now, back to the episode. How do you see that playing out in the model companies? Because I feel like in talking to them... Obviously, there's lots of different incentives, right? There's like, we just got to keep going because we just raised a ton of money and we're competing against, you know, the most well-funded competitors and the smartest competitors in the world, like all that kind of stuff. There's the kind of, I want to get promoted.
But I think a lot of them also feel the... how bad the social media era was for people and like don't want to do that, but also obviously have to hit their numbers. So what do you, I guess, what do you think is, how do you, how do you see that playing out? Like, what do you think people internal to the companies are thinking? And then what is the right way to go about this? So it's good for society. I guess your, your, your, your take is we should be delegating.
Yeah, so I think this is an inherent tension between the types of folks that you might have at a company. So you might have the researchers who care more about hitting, you know, just advancing the model capabilities. You might have the product managers or the product executives who feel like they need to hit certain measurable numbers. And so in the same way that if you think about the kind of social media platform that Facebook would build, that's probably going to be very different from the kind of social media platform that Google built or that TikTok or Pinterest would build.
And similarly, the kind of search engine that Facebook would build is very, very different from the kind of search engine that, yeah, like obviously Google or others would build. And so it almost boils down to, Kind of like the choice, I guess, that the people in charge of the products are making. Like what kind of thing at the end of the day do they want to optimize for? Do they want to optimize for this delegation or this human uplifting, human flourishing?
Or do they want to optimize for the metrics that will impress Wall Street and convince users to stay one more minute, one more hour on the site itself? Like, I think these are hard choices. Like, at the end of the day, it's very, very easy to measure sessions and users. And it's very, very hard and much longer term to measure whether you're actually improving human lives. And so it's very easy to default to the former and to convince Wall Street, convince your investors, convince all these people that these are the right metrics and that they're moving up into the right.
And so if you're kind of unwilling to make the harder choices, you just end up optimizing for the former. How do you manage this inside of your own company? So I think we are very lucky in that because we don't have... VC investors, we don't have to fall into the kind of Silicon Valley VC optimization trap that a lot of other companies do. Like we don't need to show board members, board numbers going up every single month. We don't need to optimize for our next round.
That will have to happen in a few months or whatnot. And so as a result, we don't have to optimize for short-term engagement, short-term profits. And we actually can really think about what's beneficial for us and the entire industry in the long term. So I think that really helps. And what do you think is beneficial? So it goes back exactly to what I was saying earlier. Like if I can think about what we want AI to optimize for, it isn't engagement.
It is really about how do we make these models? How do we design them? How do we teach them in such a way that they're not replacing us as a species? They're not... kind of like forcing us to watch AI slot videos all day, or rather they really are thinking and encouraging us to become sort of better versions of ourselves. So again, when I think about like that email example I gave earlier, it's not an AI model that will suck up three hours of my time writing a pointless email.
It is a model that will push back on me and tell me to go do something else. And yeah, I think that's really important. The interesting counter argument to the delegation process question is the more you delegate, it's like, you know, picking a car instead of walking, your muscles atrophy. How do you think about that? So I think there's a, almost a time and a place for both. Like what you don't want to do is simply take the car because taking the car is somehow addicting and, um, you feel kind of lazy.
And so even when you need to get exercise, maybe even when you haven't been outside all day, you don't wanna take the car anyways just because it's the easiest thing to do. And I think in the same way, like, yeah, obviously AI can be super efficient for many, many things, but if people are sort of just mindlessly delegating tasks to AI without even thinking about them at all, I think that's the boring thing. That makes sense. I feel like the... I talked at the very beginning about the data game, and I feel like the data game went from getting interesting data sets to getting environments and giving labs environments.
Do you think that that's... Is that accurate? And if so, can you explain why? Yeah. So certainly the trend and the new research direction in the past year has been this concept of our environments. And what I would say is, I mean, you certainly need the fundamentals. Like before the model can operate in this environment, it needs to know basic things. Like it needs to know how to follow instructions. It needs to know how to avoid hallucinating. It needs to know how to write code and how to use tools.
It needs to know how to write and so on and so on. But as models are becoming more agentic, And yeah, they will have access to tools. They will have access to all of our documents. They will be able to operate browsers. As that becomes almost a default way that models interact with us, our environments are basically just sort of like a more on distribution way of training them, which is why they're becoming more and more popular. As the models get more powerful, then the way we train them is getting more powerful as well.
What would be an example? So I guess the obvious environment is like using a computer, but what would be an example of an environment that's non-obvious, that's teaching models things we might not think of? So I can give an example where... A lot of our environments are a combination of tools that the models need to learn to use. Like this might be an MCP server, or it might be calling a Google Drive API or the Slack API. In combination with a bunch of documents, like here are 30 PDFs and 20 Word document files.
And you might give it a prompt like, hey, can you... Go update or 2026 forecasted revenue numbers. And what the model needs to do is it needs to learn how to... find the right PDFs and documents and needs to learn when should it search through SAC. It needs to learn when is some information outdated. Like maybe there's an email with, uh, with some early forecasts. And then later on, there's another email from the same person or, you know, maybe a different person who was saying like, Oh, whoops, I actually made a mistake in those earlier numbers.
So here's a, here's an updated version. And so that is a, I think fairly canonical version of an environment. And then one of the interesting things we found, so I think we're actually going to publish a paper on this soon, but even when we didn't give this kind of environment any access to coding, when we trained a model on this environment, we actually found that it improved on coding a lot. And the reason was because we were basically teaching it these generalized forms of instruction following, generalized forms of tool use, generalized forms of understanding documents, which you can think of as fairly analogous to the way a model needs to look through various files in your repository and understand that some documents things supersede others.
Or, you know, just the way it uses tools is obviously very analogous to the way that a model might write unit tests and execute them and iterate over and over again until it passes them. So I thought that was actually a really, really interesting find. Really interesting. Did you see Taki? No. It's the language model that's trained only on text from before 1930. Oh, okay. Yeah, yeah, yeah. I saw that. Why don't you make of that? Because I thought it was so interesting that you can get it to program.
If you shop prompt it, you can get it to program basic things. What do you make of that? And what does that tell you about the value of data? So I personally didn't dig into it that much, but I thought the concept was fascinating. Like basically this idea, and I think a lot of people have this idea. It's like, if you gave, if somehow were able to create a data set, I think contamination issues are very, very difficult to avoid. So the question is how you would do this.
It's like, if you gave the model data only up until pre-Newton, would it be able to discover Neukonian mathematics? Would it be able to discover quantum physics and so on and so on? So yeah, I think it's a really, really interesting question in terms of what types of inherent reasoning the model will be able to learn and then extrapolate from that. And it's almost like if it can discover all those things, then okay. Then given the state of science today, does that mean that the model is going to be able to discover science that centers out?
Having played with it a lot, my sense is the answer is no, but a qualified no. And you can kind of feel it. You can feel it bumping up against the limits of its world when you start talking to it about like more modern things. Like it just, it's, you know, there's this foster science, Thomas Kuhn, he talks about incommensurability. And it feels like my world and its world are sort of incommensurable. But then you can also get it to... program. But the way you do that is you get it to combine its circuits in a way that's not, it wouldn't be natural for it, but you can prompt it in a way to do that in a way that ends up being programming.
So I sort of both think it can't do it. And also if you prompt it cleverly enough, it can, but you have to supply the answer first. Does that make sense? Yeah. Interesting. Okay. What is the value of my data? So one of the things that I'm, I'm just so interested in. So obviously you're, you're in a data company, like you're, you're, you're getting expert data, um, from like real PhDs and, and, uh, selling it to the model companies and like providing, uh, all of the, all the like smarts and taste to, uh, to the models that we use every day for someone like me.
Um, we're just getting to a point where it's actually pretty easy for me to gather a data set. You know, like, for example, I do all of my email in Codex and I have a history for every email of, was this useful? Did I dismiss it? Did I reply to it? If I replied, like, what did I say? What is the value of that? If I wanted to sell that to you, how much did you pay for it? So the value to me as someone who would use that data to train an AI model? Let me think.
So I think the value would be teaching models very, very deep personalization. Like I think right now the models are actually not very good at personalizing things. It's kind of funny. Whenever I use AI models, I actually turn off the features where they personalize to me or where they can search across all of my conversation histories. Because I find that they just over index on things that I said once, but actually aren't all that important to me.
So I actually have it completely turned off. Unless I'm like testing something. So I think the value of it would be like, okay, yeah, you did report all of these emails as spam. So the next time this email comes in, you should automatically know that it's spam. Or it should learn that this is your writing style. Like one of the reasons I think people don't use AI for better or worse for writing more is because it sounds obviously AI generated and it's not mashing their voice or their cadence.
Or it's that, okay, these are the things that you yourself care about. Like, I think one of the biggest reasons AI is maybe not as useful as people would have expected sometimes is because it lacks all of your context. Like, it doesn't know that these are the articles that you read. It doesn't know that these are the decisions about, you know, the company that you're making. These are the goals that you have. And once all of that is in the model's history and it knows that it can incorporate these things and these are the kind of optimal decisions that you made, it's very, very valuable in teaching it.
Okay, this is actually how I use all this data to make certain kinds of decisions. So, yeah, I think that deep personalization is what is most unique about that. That's interesting. And as an individual person, I mean, I guess I could turn it into a synthetic data set. But as an individual person, is that worth a lot? Like, should I be thinking about selling it? I imagine we could make you an offer. I have to learn a little bit more about how big this data set size is.
But I can make it as big as you want. I've got Fable. You convinced me. One of the things we actually do is, I mean, we teach models in these very, very deep, personalized ways. So something similar to what you described is a fairly big thing. Tell me more. So, I mean, I've got email. What else am I doing that you're like, oh, that's actually really valuable and important in ways that people probably wouldn't know? So, honestly, even things like the way you interact with your browser is interesting.
Like models still aren't all that good at it. Or even the types of conversations that you're having with AI, that is just inherently interesting in of itself. Models themselves are not very good at generating synthetic conversations to try to mimic you. And so even just knowing what types of conversations you're having is helpful. Or it's like the combination of all these things, like knowing that these are your photos, these are your texts, these are your slacks.
It's like this interconnected web. And maybe certain things in one aspect of the web influence others. So just seeing the thing as a whole is very helpful as well. Why are models bad at writing and how does that relate to the personalization challenge? So... I think some of the models are pretty good at writing, but some of them are actually kind of shockingly terrible. So I'll give an example. So we created a benchmark called Hemingway Bench a couple months ago, and it was designed to test models creative writing abilities.
And one of the things that we saw was that some of the models, they were literally outputting metaphors in every single sentence. And I think the reason that was happening is because, I've talked a little bit about this phenomenon on reward hacking. It's almost like there was a metric somewhere or like a score that these models were getting. Like, okay, every time you are literary, every time you're using complex imagery, it would get a point. And it learned to reward hack this by outputting a metaphor in every single sentence.
And I mean, what's kind of funny is that a couple, what was it, a couple of weeks ago, there was this kind of like semi-prestigious literary prize, I think the Commonwealth Prize. And there was a controversy because a clearly AI journal of the story won the prize. And if you actually looked at that story, it's funny, like it literally had a metaphor in every single sentence. And so this kind of phenomenon that we described a couple of months ago, yeah, it was still happening.
And so... Yeah, I mean, I think it boils down to a couple reasons, but like one is it's people are kind of sort of measuring the wrong thing. Like instead of measuring actual taste and actually good pros, they either have these flawed metrics like. What is the complexity of the prose I'm writing? How many metaphors do I have? Or there are these AI leaderboards again, like Ella Marina, where you have people who are essentially high schoolers who are reading responses for two seconds, And what they are captivated by is a flashy metaphor.
And they are not captivated by kind of like the understated pros. And so I think it kind of boils down to a mismatch in measurement and a mismatch in like the optimization and objectives that the models are trained towards. Fascinating. Okay, last question. What is your current AGI timeline? So I certainly believe that AI will happen more than most people expect. Like every few months and even faster now, I think what AI is doing continues to surprise us.
So I think it depends a little bit, obviously, on your definition of AGI. But if my metric were something like being able to automate the work of the average engineer, like, or being able to publish more and more novel scientific research that gets published in these journals, or even the ability to win a Fields Medal or a Nobel Prize. I could see it happening within the next five years. All right. Edwin, thanks so much for joining. Thanks for having me.
Oh my gosh, folks, you absolutely positively have to smash that like button and subscribe to AI and I. Why? Because this show is the epitome of awesomeness. It's like finding a treasure chest in your backyard. But instead of gold, it's filled with pure, unadulterated knowledge bombs about chat GPT. Every episode is a roller coaster of emotions, insights, and laughter that will leave you on the edge of your seat, craving for more. It's not just a show.
It's a journey into the future with Dan Shipper as the captain of the spaceship. So do yourself a favor. Hit like, smash subscribe, and strap in for the ride of your life. And now, without any further ado, let me just say, Dan, I'm absolutely hopelessly in love with you.
Transcript supplied by the publisher with the episode.
by Dan Shipper · English · Tech & Science
The Every Podcast is Every's flagship show. Co-hosts Dan Shipper and Natalia Quintero talk with founders, researchers, writers, and operators about what they're building and how they use AI in their own work. The show also takes you behind the scenes at Every. We share how our team is using and expl
E122 · 15 Jul 2026 · 1 hr
“Running a startup is a knife fight whether things are going well or not,” says Chris Pedregal, cofounder and CEO of Granola. Granola recently raised a $125 million series C round at a $1.5 billion valuation on the strength of its AI meeting notetaker. That valuation hasn’t made Pedregal complacent. Granola built its name as the first to make good AI meeting notes, but Notion, OpenAI, and Zoom have all since released their own versions. Pedregal isn’t rattled—he never thought meeting notes were the real prize. The bigger fight, he says, is over “what interface we use for work, and what work…
E121 · 8 Jul 2026 · 53 min
Craig Mod used to pay Campaign Monitor roughly $7,000 a year to send his newsletters. After rebuilding the tool himself with AI, his bill is closer to $150. It’s the kind of thing that convinces him we’re about to enter a “golden age of tool building”—one where anyone can build tools specifically suited to their needs, instead of settling for software from incumbents that are slow to innovate. Mod is the writer and photographer behind the newsletters Roden and Ridgeline and books like Things Become Other Things and Kissa by Kissa—as well as a lifelong technologist. He’s rebuilt the tax…
E120 · 1 Jul 2026 · 41 min
Natalia Quintero joined Every as head of consulting with a mandate to bring AI into the workflows of executives at hedge funds, private equity firms, and tech companies. She is also a recent Codex convert—someone who spent months resisting the tool before Dan Shipper’s daily pestering finally got her to try it. Natalia encountered Codex as a non-technical builder who had learned to navigate file systems and folder structures in Claude Code through sheer effort. She’s now used Codex to do everything from automate her CRM setup to build a portal to manage her father’s medical care. Dan talked…
E118 · 17 Jun 2026 · 28 min
Last year, there were 1 billion commits on GitHub. This year, Kyle Daigle expects that number to exceed 14 billion, a two-component explosion caused by more humans—and their agents—issuing pull requests. In March alone, 17 million pull requests on GitHub were created by agents. Daigle is the COO of GitHub and Microsoft’s chief marketing officer for developer products. He’s been at GitHub for 13 years, and is paying close attention to how AI is expanding the platform’s user base. Along with agents, legal, sales, and marketing professionals are building apps with the GitHub Copilot app. The…
E117 · 10 Jun 2026 · 52 min
Mike Krieger built one of the most consequential consumer apps of the last two decades as the cofounder of Instagram. He is now at the frontier of AI-native product development as head of Anthropic Labs, the team responsible for figuring out what the most capable AI models can do in the hands of real builders. When Krieger first got access to Fable 5 months before its public release, it was exciting and disorienting. “I feel like a total newbie again,” he remembers telling his team. The way he’d been thinking about productivity, strategy, and time management was out of date. The model had…
E116 · 3 Jun 2026 · 34 min
The "SaaSpocalypse"—the panic that AI will make software-as-a-service obsolete—hasn't rattled Figma’s Matt Colyer. As the company’s director of product management for developers, he's been building his own agents for two years and is buying more software services than ever. In addition to making the case that AI is a “goldmine” for SaaS companies, Colyer talked with Dan Shipper for AI & I about why great design requires a diamond-shaped process: First you diverge, generating as many ideas as possible, then you converge around the best ones. Chat is linear, which makes it good for iterating…
7 Oct 2026 · 34 min
Subscribe to Every (free for 14 days): every.to/podcast-every-offer?utm_source=podcast&utm_medium=audio&utm_campaign=companyagentepisode On this week's episode of The Every Podcast, Dan Shipper sits down with Willie Williams, Every's head of platform and one of the driving forces behind the Every Agent, which launched this week. They discuss how Every went from everyone running their own OpenClaw to relying on one shared company agent—and where they think personal agents belong in everyday life. Links to resources mentioned in the episode: Willie Williams on X: https://x.com/bigwilliestyle…
30 Sep 2026 · 33 min
Sam Altman's Dot, OpenAI's new always-on agent, gave him back the most productive part of his day. The OpenAI CEO would lose his early mornings—his window of "maximum creativity"—to whatever fire needed putting out. Now, his Dot triages those issues and flags only what's truly urgent. It's also gotten him building again: In spare moments, he gives his Dot ideas for a feature, and his Dot builds five or six versions for him to review. On the first episode of The Every Podcast (formerly AI & I), Dan Shipper interviews Altman at OpenAI's DevDay. They discuss how Altman uses his Dot to protect…
E129 · 2 Sep 2026 · 48 min
Two years ago, Katie Parrott was laid off from a crypto firm and couldn’t afford a career coach—so she turned to a $20-a-month ChatGPT subscription instead. The Every staff writer’s habit of feeding AI good context eventually became compound writing, a codified system for brainstorming, drafting, and editing with AI. Today, that system is a plugin any writer can use. On this week’s AI & I, Natalia Quintero talks with Katie about turning good ingredients into good writing, borrowing the taste of writers she admires, and why AI helped her fall back in love with the page. If you found this…
E128 · 26 Aug 2026 · 1 hr 7 min
Will England is the CEO of Walleye Capital, a hedge fund managing nearly $10 billion in assets. An engineer by training with a math background from Oxford, he has spent his career at the intersection of machines and markets—and has made AI fluency mandatory for all 400 employees. England believes refusing to use AI is like refusing to use the internet in 1995 because it wasn’t perfect. His use of AI is public and effusive, including in a firm-wide email that opened with “I used ChatGPT to write this email. You should be using it, too, and be proud of it.” AI informs how Walleye drafts memos…