Skip to content
Melo Podcasts Home
CategoriesLanguagesFollowing

Episode notes

Summary In this episode Ofer Mendelevitch shares what it really takes to evaluate RAG systems in production when your source data is incomplete, constantly changing, or difficult to validate against a clean ground truth. He explores how RAG has evolved from simple “chat with your PDF” demos into enterprise-grade retrieval systems that require robust ingestion pipelines, hybrid search, reranking, multimodal support, access controls, and refresh strategies for large and dynamic document collections. Ofer also explains why evaluation becomes one of the hardest parts of the stack, particularly…

Chapters

Tap a chapter to play from there.

Transcript

Read the transcript · about 10,960 words, follows along as you listen

Hello, and welcome to the AI Engineering Podcast, your guide to the fast moving world of building scalable and maintainable AI systems. Your host is Tobias Macey, and today I'm interviewing Ofer Mendelevitch about how to evaluate your RAG systems when you have incomplete or changing data to grade against. So, Ofer, can you start by introducing yourself? Hi. Thanks, Tobias. Great to be here. I was a machine learning and kind of recently deep learning leader in here in the Bay Area for the last maybe two decades.

I actually started my career in machine learning. I had a fun story to tell about this, which is when I started, Naive Bayes was a leading algorithm at the time. I don't know if anybody in the audience know what that is, but it's probably the first machine learning algorithm ever created that I've known of. And nobody is using it today, of course, because it's so old. But, yeah, so I've been I've been around for a while from the time that we called it machine learning to data science to now we call everything AI.

In addition to that, I was sort of in engineering too. So part of the engineering stack, was VP of engineering, a couple of startups in Silicon Valley. So I've seen a lot of the ups and downs of the Valley over the last two decades. Yeah. More recently, I've been, involved in, some efforts in also developer relations and, for a company called Band and other companies, And wrote a book actually with O'Reilly recently called Hands on Rag for Production, which I think we're going to talk a little bit about today too, that kind of tries to summarize all of my experiences in bringing AI systems, Rag, and today agents into production and what that means and all the challenges that are, part of that.

So you mentioned that you've been in the space for a long time. I'm wondering if you can just start by giving a bit of an overview of how you first got started in the overall field of what used to be machine learning and has now been subsumed by AI and what it is that has kept you here. Yeah. So, I actually started with, machine learning in my first startup, which was founded in year 2,000, more or less. And as I mentioned earlier, we did a very elaborate system. It's funny to think about it these days that did the document classification.

It was called Quiver. The company later was acquired by another company called Inktomi. But at the time, it was all new. Like, people who did machine learning were only PhDs. That's the only way you could do it. You had to have these very specialized expertise. But that's how I got started. I started a company and we required this classification classification engine to do our product, to implement our product. And I just learned on the job, essentially.

Then this, of course, was very interesting to me. So I kept in this field. Then things evolved into, you know, decision trees and what we know now as XGBoost and, you know, all these other algorithms that evolved over the years. And around twenty eighteentwenty nineteen, I actually had my first meeting with the LM space. So a friend of mine who used to work at OpenAI in the past told me to take a look at this, and it was... I was very fortunate that I had this friend. And this is the time when GPT was GPT one, and GPT two just came out early twenty nineteen.

For those of you who haven't seen that, there's a whole set of articles published by OpenAI on those years. It's really funny because they... Well, funny and interesting. They released GPT two in three phases. They had the small, the medium, and the large, and and they had extra large. So they did not... They decided not to release. Now I'm saying it's funny because the extra large, which was deemed very dangerous, was about 1,500,000,000 parameters, which is kind of laughable in today's terms. But I think they had the right idea of thinking about alignment and thinking about, protecting because nobody knew what this would do.

Anyway, I I used this, for another startup that I started at the time in health care. We used GPT two. We trained it on health care data and fine tune it. Yeah. And that's kinda been my journey from the early days until today. In the process, of course, I've learned a lot about how to build machine learning models, how to release them to production, how to test them in production to ensure that they don't regress in their performance and in different fields. So it's been in financial services, in travel, in health care, and other industries.

Now digging into the topic at hand, as you mentioned, you wrote a book on the topic of RAG and bringing it into production, and RAG as a concept was one of the first ones that really gained a lot of attention as LLMs were exploding onto the scene because it was something that was fairly straightforward and easy to engineer because we didn't have all the frameworks. We didn't have all of this concept of agentic capabilities, and the models themselves weren't really strong enough to do much with that. But we understood how to do some of that few shot learning where you do the in context parameterization of providing some examples so that the LLM could then synthesize that.

And I'm wondering if you can just talk to what you see now that we're a few years in the role that RAG is playing in some of the actual systems that people are running. And now that we've moved beyond it being a toy and just a... An easy way to get a quick demo running to something that people are relying on for actual production applications. Sure. Yes. So, about, you know, two, two and a half years ago, you're absolutely right. Rag was the hottest thing. Right?

And let me clarify a little bit the reason for that, because we started by having LLMs and they were answering questions really well. And the main challenge or the main opportunity, I guess, also with Rag was I want to connect it to my enterprise data. I want to connect to the data that's documents on Google Drive or on S3 or in Notion or in anywhere like that. And the LM doesn't know about your private data. So the main reason to do Rag was to ingest the data from these data sources, prepare them for retrieval so that you can use retrieval as part of REG. That's the R in REG there, it's retrieval augmented generation, right? And so the R is the prepare the data for semantic search and retrieve it at query time so you could use it inside of Reg. And the Reg was a simple idea. It was a prompt that said, you gave the element a prompt that said, here's my question, and then here's some relevant documents or chunks of documents that might have relevant data.

And the end of the prompt was, please answer this question given these documents that are relevant. And the LM will process everything, summarize those documents, and sort of give you the right answer. And you're right too, Tobias. It was quite magical. And a lot of people said, like, there was a time where, like, oh, look at me. Talking to my PDF was, like, the magical sort of simple use case that everybody really liked. It was easy to do. I mean, you had a vector database. There was, at some point, like, I think 50 of those available, both commercial and open source. It was an explosion of of vector databases.

And you just did some chunking of the documents, parsing the document, checking the documents, put them in even in memory because it was a small demo. And it was it was... It created a new capability that people really liked. Instead of reading the documents for 100 pages, they could really like ask it questions, which was really nice. Now that is very different than doing this in a real production environment. Right? And I think that's where companies that are... That build right systems, like I worked at one company called Vectara, which is where I wrote this book as well, has done a lot to create that retrieval and query pipeline to be much stronger and much more useful. Let me talk a little bit about where some of the challenges are in that.

So first, I mentioned ingestion. Right? So you ingest the data into a semantic search engine like a vector database. And just in itself, when you think about it, at scale is much more difficult. Let me highlight a couple couple example reasons. Right? There... There's there's big companies who have PDFs that are... And I've I've seen one of those. 15,000 pages of a PDF. Right? So think about that scale. Like, you run something in Python, let's say, and you you know, it could take, I don't know, hours to just process the PDF. And maybe in the last two pages, it runs in an error and fails and you have to try again. So you have to think about sort of normal engineering best practices, right, of how do you tackle really large files so there's no... It doesn't go out of memory. You do it in small batches or something like that. The other thing is just the number of files. I mean, it could be a million files, right, or 20,000,000 files. It doesn't have to be files, by the way. Could be, Notion documents. It could be, Jira tickets.

It could be, you know, stuff from SharePoint or or whenever. But again, if you have a lot of these documents, you have to build a production system that not only ingests all those documents for the first time, which again, could take days or weeks and you have to paralyze it maybe and build retries in case some documents run into errors. And you also have to build things for refresh. Right? If I have, let's say, a Google Drive source of data and I ingest all the files there, new files come in every day or files get changed. How do you refresh without re ingesting all the files? Right? And then the last thing I wanna mention on ingest, which is really critical to understand is, multimodal.

Right? So a lot of real world use cases do have multimodal data and quite a bit of it. So for example, tables, images, diagrams, you know, flowchart, things like that. And a lot of times the the value of of answering a question in Rag or in agents is really embedded in that image. Right? You might have a graph that shows the revenues of a company and it's not in any of the numbers, it's just in a graph. So you have to really understand that image to understand what it is. You can even have videos. A lot of companies have training videos, for example, a lot of them for staff, for sales, for customers, etcetera. So you wanna be able to understand video and audio and transcription and so so forth, and so on.

And so, that is a a key interesting set of capabilities that you have to develop beyond what's involved in ingesting text. So that's the Indios part. Now another set of challenges is is I'm... I kinda laughed it out a little bit that everybody built it with a vector database, but there's a lot more to retrieval than that. So today, the sort of state of the art retrieval engine does include a vector database, but also a what's called hybrid search. So there's also a lot of use cases where you do want to have the old school, I'll call it that, search capabilities like with TFIDF or BM25 type of lexical or keyword searching.

And hybrid search is the ability to merge semantic search, vector search with that together to create a way to identify both keyword search use cases and semantic search use cases. And then people use re rankers quite extensively too because vector search and embedding model, what's known as embedding models, which are the models that you embed... You create vectors from from text, are pretty good for for initial selection. But, actually, they're not as accurate in terms of the semantic ranking.

And so you use re anchors, and there's a lot of work in the ranking space as well. So those are, you know, a couple of examples from the the querying part. Right? And then, I mean, if you do Rag, which is sort of a a single shot, so I think you have to now think about the generation part. How do you craft the prompt properly so it's accurately taking all the chunks together and creating the best, most accurate result? And then that leads into the topic you mentioned earlier about evaluation. How do you, in production, evaluate that the results are accurate, that the retrieval is accurate, that there's no hallucinations. It's a big topic, and so on. So there's a there's a whole area of evaluation that, is very, very important.

And on that note of evaluation, that I think is probably the hardest part of RAG despite all of the other complexities that really come into it because changing any one piece of the system can have a completely unpredictable result on the output where you might change a piece of data in the system, and then that might actually have been the keystone that allowed the LLM to be able to provide the accuracy that you wanted or maybe changing the model because of either different thinking parameters or different size of the scale of the model or maybe you're changing the retrieval or reranking system. Any one of those things can have an outsized impact on the overall.

And so there have been numerous frameworks and approaches developed to be able to do some sort of at least heuristic testing of the outcome. A lot of the, at least, earlier iterations were reliant on LLM as a judge where you just had another language model gauge some sampling of the outputs to say, yes. This looks good or this looks bad. And as long as everything is about average, then you ship it. But, also, there's the offline versus online testing where with offline, you say, here is my set of examples and here are my expected outputs. As long as everything looks about right at the end of the day, then we're we're fine. But, really, the changes in these systems come about when you actually put them in front of people because the way that somebody asks the question or the way that the prompt is structured can also have all these different variables involved.

And so I'm wondering if you can talk to some of the current set of approaches that people are using for these production or enterprise grade RAG systems that they rely on to actually build and maintain confidence in the accuracy and reliability of the overall stack. Yeah. A 100%. Super... You're right. It's a super important, topic in Reagan. I wanna mention also one that people don't necessarily think about upfront. They just build first, and then they realize, oh, we gotta build something like this. So let me let me try to give the full picture here. So first, yes, you're right. People use LM as a judge. What is that? That's like just to give us some some eval prompt and you try to evaluate the accuracy of the results. But I want to separate between retrieval accuracy and generation accuracy.

Right? Two different things. So retrieval accuracy, we just described, you know, vector vector search and hybrid search and all that. So the end of that phase of RAG is a set of, they're called chunks. They're parts of documents, let's say, a couple paragraphs sometimes or something like that, that are deemed to be relevant for this particular query. And so the first question is, did my retrieval pipeline retrieve the right set of chunks, the most relevant set of chunks? Now you have to think about, do I have the relevant set of chunks? The first problem is, like, sometimes the question doesn't have an answer in the documents that I ingested.

And then maybe I need to figure out where those documents are and bring them in. It's not obvious in a big enterprise that you have the right, you know, baseline ground truth for... To answer every question that a user might ask. But assuming you do have them, your retrieval pipeline, like you said, might have shifted. You may have been putting a new re ranker. You might have increased the number of results, reduced number of results. You may have read down how you do chunking to so many levers that you could have changed, and every one of those could have influenced your retrieval pipeline.

And so on retrieval, you know, people tend to do LLM as a judge, or they try to use the regular techniques that we have, like precision recall, F1 from machine learning. Right? The problem with that is those techniques, those metrics are pretty simple or pretty proven, require a source of truth. Right? So precision, for example, is how many of the retrieved chunks are actually should be in... Let's say I have top 10. How many of the top 10 that I actually retrieved should be in the top 10 are actually relevant. Right?

And you have to somebody has to sit down on all the chunks in your, say, 20,000,000 document organization document set and look at all the chunks and say, are those relevant for this query and that other query? And I I said 20 mil to to, you know, emphasize it's it's completely impossible to do it in practice. Right? It's completely infeasible. Nobody can do that. And so in theory, those techniques work, but in practice, it's not working. In in the answer side, so the generation, let's say that I don't do this on the the retrieval side, but I go to the generation side, people have done in practices. They said, I'm gonna come up with a thousand questions that I think are relevant. Or sometimes if they're smarter or they have the system in in in production already, they'll look at the most common questions or take a random sample of questions that users actually ask.

And then they'll they'll have somebody sit down and generate a thousand answers. Okay? It's a manual process, but, you know, it's doable. You can sit for a week and... Or a couple people and and generate that. And then they'll have some metrics to measure to the... Know, is this... Is the generated answer similar enough to the to the ground truth answer? And I've seen a lot of people try that. There's two problems with that approach, which also leads to LLM as a judge.

The first one, of course, as I said, like, the questions themselves could be hard to generate sometimes. Remember, you have to generate not just questions that people ask, but also questions that people might ask later. And so it keeps changing. The tendency is to ask questions that have obvious answers, but you wanna ask also tough questions that don't have answers necessarily and things like that. But the answer is... That's where, again, I see a scalability issue in practice. Sometimes these organizations are so complicated. They've existed for thirty years, and nobody knows the answer to some questions. It's like in so many systems. And so you can even find the engineer or the product manager or whatever the function is, the production manager that knows the answer to that question, to validate it, that's the right... To to write that golden answer. Right? And again, it keeps changing. Like like you mentioned earlier, like some of the answers might might change over time because I have a new version of a product or I have a new set of, capabilities in my my API.

So it's very hard to keep track of in in practice. So let me talk about two approaches that, are using LM as a judge that, I also write up in the book that are really cool to solve both retrieval and generation. The first one is called umbrella. And and... Oh, by the way, before I go there, I wanna mention this has been really interesting research that, we did at Vectara along with the Jimmy Lynn's lab in University of Waterloo. I wanna give, all credits to those guys. They did a lot of the research, and we just collaborated to implement this in something called OpenReque eval that we created later. But Umbrella is really cool. It's a very simple algorithm if you think about it first, and this is to identify chunks that are relevant.

So you take every chunk in your top 10, and you run an LM as a judge that gives us a score between zero and three. Zero means the chunk is completely irrelevant for the query. One means it's somewhat relevant. Two is like it's pretty good relevance. And three is like it's super relevant and might even solve the whole problem. Right? And, you know, for those who are with that line as a judge would ask me like, okay, great. I can generate this. Like, what's so research y about this? Why is it so special?

And the reason it's so special is... And that's the algorithm. You take this and you rank it like this. Right? But the the really cool part about it is that the the lab for for Jimmie Lin, the the research they did is they tested this approach with a particular prompting language, also with human beings. They looked at the chunks and they had a bunch of data sets and they compared the results of human judgment with this, you know, over time, a couple iterations of crafting the particular prompt sentence, and they found that the correlation is very high. I think it was like something like maybe 80%, if I remember correctly.

So very high correlation between what this particular prompting LLM would do compared to what a human would rank it as. And so you have confidence that if you use it in this way, you get the zero to three actually mean what they do. And that's a general problem of LLM as a judge. It's a fantastically easy algorithm to implement for a lot of different use cases. The problem is, do you trust it? Right? It's got its own biases. It's got all kinds of problems. The LLM might change and what you get might change too and stuff like that.

Yeah. With the LLM as a judge, you then have two nondeterministic systems that you have to have some sort of evaluation for, and then you're in the catch 22 of which one do you tune when and how do you decide which one regressed versus improved. Exactly. Exactly. So with Umbrella, you kinda wanna keep it the way it was implemented first. It's been proven. It's shown the the sort of correlation, and that's why it's it's powerful. I'll just mention briefly the other one, And, you know, it's called autonogatizer.

And it's a little bit more complex approach, but it's similar in concept. Again, NLM is a judge approach that takes those chunks and ranks them relevant to the the generated answer. This is for for the answer relevancy. And, what you do is you you then see, it tests for a different thing, which is is this answer that was generated by the LLM in Rag, does it use all the relevant chunks properly? Right? So if one of the chunks was rated three, like the three, like the highest rated thing, you know, you wanted to make sure that that piece of information was moved over into the answer.

And, again, it's not deterministic, so you don't really know. So that's the goal of that algorithm. It's, again, something was, something actually called Nuggetizer, which is deterministic, and then Auto Nuggetizer is this generation of nuggets using LLMs and then sort of evaluating this in the context of the generated answer. So so those are two LLMs as a judge approaches that I wanted to highlight. And as you mentioned, one of the standard ways of doing the evaluation is that you have that set of labeled question and answer pairs of this is the question that I have. This is the answer that I expect, and that can be a very time consuming and potentially manual process, which can also be expensive and also requires that you have some concept of what data you have, what types of questions or inquiries are being performed, and that requires some amount of bootstrapping that can again become a catch 22.

And I'm wondering what are some of the ways that people are tackling that start from zero position of I have the LLM. I have a way of generating the embeddings, producing the context, building out that whole machinery, but I don't have the flywheel in action yet. So I need to be able to either generate synthetic data to start with or hypothesize a certain set of questions and answers, or maybe I just do it where I have a live feedback loop in production and I'm doing online evaluation so that the system can evolve and improve as it's operating rather than having to get a running head start first?

Yeah. So first of all, yes, synthetic generation is actually a great technique. You can point kind of an LM at at a good sample of your data that you just ingested and actually generate some really good... With a good prompt, you can generate a diverse set of questions based on this data. It's not perfect because usually, we'll do it one document at a time, so it cannot generate very... Like, you know, questions that necessarily require information from multiple documents.

That's a little more tricky. But people do that all the time. I also highly recommend looking at your logs if your system is already... Even a little bit of time in production. What types of questions people actually ask that they use the system and ask your subject matter experts, you know, for suggestions for different questions that might be useful. So that's a good way to kick this off. And sort of from a, like, a release perspective. Right? You wanna build that system to be sort of a part of your aggression suite in a way. Right? So you wanna be... Every time I I do something, we kinda mentioned it earlier. If you change a component, you change something, you fix the bug that changes your chunking strategy or something, everything impacts the system. So you want... Every time as part of your testing, as part of your regression testing or something, you want to be able to run this complete set.

And it's got to be, you know, versioned and controlled and sort of so you know that you at least got the same results as before and hopefully better over time as you improve the system. Now let's talk about online testing a little bit. So it's a very interesting approach, but it's more difficult, right? And it's more difficult for two reasons. One is you don't want to increase latency, right? You don't want to have yet another call to an LM to test before you produce the answer.

It also, of course, costs a lot more. Right? It may be... Yeah. So, you have another LM call. It's more expensive. People worry about cost of LMs in 2026 much more than previous years, I think it's becoming more and more of an issue. So what I've seen in my my general recommendation is to do online, maybe as a phase two, like, if you want, or you can do it in phase one if you're up to it. But, like, it's okay to do it later on after you do the offline, sort of release. And also do it as a sampling. Right? You don't have to do it on every one of the queries that come into your system. So the way this works, right, is your your user query comes in, you go through your Rag pipeline, and then you test.

And you don't have to test on all the queries. There may be a lot. Right? Some people have, like, I don't a million queries a month. Right? It will be a lot of cost to run another LLM call or maybe multiple LLM calls on each of these. And so you can say, wanna test on a random 10% of the queries or 20% or 5% or something like this to control your your budget and also do it out of band. So I don't recommend doing it like waiting and holding off the user until it's finished.

Think of it as diagnosis for your system. You did off... Offline. You store the results, and then you have some analysis later on and see what you can learn from it. You can learn that, hey. This query, the result wasn't that great, etcetera. And then you go and fix it back. And something that goes out of band of your release cycle, but gives you, like, more immediate responses. And one of the other evolutions of RAG that came about as the idea of agents came into the picture more strongly and as the models became more capable is agentic RAG. And I'm wondering if you can talk to some of the ways that that is distinct from the first generation of rag stacks and some of the ways that it fits into the current ecosystem and maybe how it complicates the picture of how to actually evaluate the effectiveness of it.

Yeah. So the way I've seen this kind of evolve over time is, as we said at the beginning, you know, RAG was the kind of hot thing and and everything happened. But at some point, people saw some of the challenges of this one shot flow that Wright has. Query, retrieval, generated answer. And so it evolved into what's called AgenTik RAG. And AgenTik RAG is essentially a simple form of the agents we have today. A very simple form. It's just... It allows you to do maybe a couple of iterations. Right? So the agent takes the user query.

The first thing it can do is it actually it can rephrase the query in a way that's better, maybe appropriate for the retrieval part. So it can rewrite the query in in a different way. It can actually rewrite it a couple times and maybe apply to the retrieval pipeline a few times to get a couple of different results back. And that became possible as our LLMs became possible too. So it didn't just happen in a vacuum. We had longer context windows. You can not just put five results, and that's all space you had to have.

You could do 10 or 20 or 50 or a 100 because the the context windows became longer, and you can put more results in there. So AgenTiGrag, the primary value I see is that ability to do multiple iterations and also rewrite the queries. So a lot of times if you run a query the way the user wrote it, it may not be the correct way that is best suited for accurate retrieval. Right? User may have an acronym there, for example, that is unknown to the system, and the LM might actually translate it using some glossary that you give it or something like that.

So that's been actually really helpful to just improve accuracy of retrieval in general. And then this evolved into what we know today as AI agents. Right? So agents are sort of an evolution of that where we have what we call the harness, which is the evolution of the AgenTiC Rag loop, the simple loop that we had at the beginning, and now it has access to tools. So if you think about it in that in that sense, AgenTiC Rag was an agent with one tool for retrieval using the Rag kind of retrieval pipeline without the generation.

And now we can have access to tools anywhere from retrieval tools like with the Rag, but also tools that go directly to target systems. Like there's tools to query databases with text to SQL, for example. There's tools to go to the native systems, like, you know, search capabilities, in in different places in Notion and other places. They all provide this kind of API type capabilities all through MCP. And that's what the the harnesses we have today for for AI agents building.

And that idea of having the tools be the retrieval engine has also maybe diminished some of the initial supremacy of vector databases as the end all be all for that retrieval step and changes some of the equations as far as what you need to do for the seeding of the data where you don't necessarily have to do the whole pipeline of collect the data, chunk it, stick it into a vector database full of embeddings in order to be able to get anything out of it. But it is still a necessary element. I'm curious how you're seeing that change the ways that people are thinking about their overall retrieval and context system, particularly for these more agentic use cases as well as some of the way... Reasons why somebody might want to stick with a more straightforward rag system in the kind of first generation definition.

Yeah. Excellent. Excellent point. So, you know, dragging all the data over to ingest it into the sort of the the rag pipeline was hard. You had to connect to the source system. You had to essentially replicate all the data. So you had to duplicate. Those were the downsides. Right? And you had, like we talked earlier, we had to refresh it every once in a while. So you had to sort of create like a hot copy of the data only in a way that's, sort of in a vector database or in a lexical database for ready to be queried.

But it didn't have the advantage of semantic search and hybrid search and all these things we talked about that created better results. And so today, you're right, as things evolve, more and more vendors created tools for their system to be able to create that system. But the downside of that still, and it keeps improving, is that you are dependent on the quality of retrieval of that tool. Right? So if I call... And and I'm I'm... I wanna say Notion, but I think actually Notion search is not that bad. Not not kind of telling anything about them, but just because it's a it's a common thing. If you go to...

If you call a tool of Notion instead of bringing it back to search, you depend on whether Notion Search is good or not. Does it do multimodal search properly? Does it do the text search properly? All these things we talked about become important. And if they're not, then it doesn't help your agent. It's not gonna get the right information. Similarly with Google Drive, similarly with SharePoint, similarly with, you know, Jira or or anything else that you have.

So that's the the trade off you make. And I think as an industry, we're kind of in the middle of that. So over the last two years, a lot of these sort of API search functionality has improved. But not all these companies have, you know, it's it's... You know, search is a kind of its own little sets of capabilities and sort of know how. And I think it's not super easy for all of them to sort of improve the search to a level that it can be as good as you might need it to be. The other challenge is a lot of companies are now moving to kind of what's called... There's a word for it. It's called AI sovereignty.

Right? Like... But kind of on premise. Right? When you go to on premise... And we haven't touched this yet, but, like, just to... On this point, like, they want everything in control. So they're calling a local LLM. Maybe it's OpenAI's open source model, or maybe it's the NVIDIA one or or thinking machines, whatever it may be, it's all... Nothing goes outside. Nothing that goes outside. And so again, if you have to do that, a lot of times you might wanna ingest the data earlier into that local environment so they can be part of what can be queried, and you don't have to make any external API calls in sort of these air gapped environments.

One of the other interesting pieces of bringing agents into the mix is, in particular, on the side of data pipelines where the first generation of RAG required a lot of upfront engineering effort to build that pipeline of doing the chunking and embedding, etcetera, keeping that up to date. Now agents are doing a lot of the work of actual code generation. They're being brought into data pipelines, and so that also potentially acts as an input to the RAG system, which can be yet another source of errors or improvements.

And I'm wondering what are some of the ways that you're seeing that come full circle where we're bringing LLMs and agents fully into the process of managing the end to end system for the data seeding as well as doing the retrieval and as well as doing the evaluation and just some of the ways that that maybe unlocks some new patterns that you're seeing? Yeah. Quick question. So, for example, there's a lot of different places. I see, for example, an ingesting of data. Right? There used to be a data engineering pipeline, that did sort of, let's say, take a PDF file. You have to parse the PDFs, put it into pages, see where tables and images are, do the chunking, all that stuff. Now people do it with with agents. Right? You have an agent that just does all that process for you. That's one area.

Generally, the other interesting thing is all these systems used to be with APIs, and they had very elaborate documentation about the APIs. And you had to, as an engineer, as a developer, have to read all of that and figure out how to implement it and if it has changed and everything. Today, you know, I run Code or Codecs or Copilot or whatever, and I just point it at the APIs and that's it. It produces a really good system for querying for my actual application that uses the API.

It produces a good thing for indexing the data or ingesting the data. So like everywhere else in engineering, the coding agents are fantastically powerful to help us build system much, much faster. And honestly, with less mistakes, like we still make mistakes as humans. I know we all say like, oh, is the coding agent really doing a good job? And in some cases, it does not, and we can talk about that too. There's... I have some experience with that too. One of the things I like to do is to write something in Cloud Code and then have codex review it, kinda have this adversarial between them. Because I found that a lot of times that really helps. You show something to Cloud Code and you ask Codex to review it, it says, oh, I found three bugs that CodeCode said everything was okay. Or or the other way around. It's not, I guess, one or the other. But, anyway, getting to your earlier point, yeah, I think, I've seen this in ingestion.

I've seen this in, sort of just writing other parts of, the system. I think in the infrastructure of the of the the right pipelines, like where embedding happens and things like this on the retrieval itself, I haven't seen a lot of people use it yet. And also in evaluation, I haven't seen it yet, but that's just maybe me. I haven't, you know, bumped into people to do that. I I could expect that this could be helpful in any any part of this of this. Another term that has been percolating through the ecosystem is context engineering, which if you squint is the same thing as rag, but people like to come up with new terms for the same thing. And I'm wondering what are some of the ways that you think about the overlap and some of the differentiation between the ideas of Rag and the idea of context engineering.

So, yeah, the context engineering is... Let let me sort of first state what I kind of my understanding of it just so we align on because the terms are... Could be very confusing. Right? So context engineering is this idea that you as the engineer developing your system control what fits into the context of the LM before you send a request. Right? So by the way, when I started using GPT two, for those who don't know in the audience, it had 512 tokens context limit. That's it. We think about now we have a million or 10,000,000. It it sounds ridiculous, but, that was it. And then it became a thousand, ten twenty four. I was like, oh my god. This is double.

Like... So, context has grown a lot, which is amazing. But... So context engineering now is this practice of when you create your system of deciding what's in your system prompt, what do you tell the the agent to be like its guiding principles, And what sort of intermediate data you bring into the the same context, that you put in front of the LLM in terms of... We think about RAC, it could be chunks that come in. If it's a more advanced agent, it could be sort of the tools that it may be allowed to call and their definitions.

It could be historical information about the user. People now talk about memory. So there's a lot of things that you can pull into the the context, but it's the same thing. It's it's how do you take all this stuff and then you send it to the LM and ask it for help to do something, either call one of the tools, call many of the tools, whatever it needs to do, that'll answer the question or perform a task for you. Yeah. So that's kind of the evolution that, one thing I I I I should mention is, one way I think about RAG in a way is a little bit... I don't... I haven't heard it anywhere else, but I thought it was an interesting way to think about it is REG could be the kind of attention mechanism for context engineering.

And what do I mean by that? People know about attention, you know, the famous word in deep learning. Right? Attention is a mechanism in deep learning itself where it sort of looks at the history of the the conversation so far and picks the right words or the right tokens in the history that are more important to generate the next token, right? That's what the retention mechanism is in general. And so sort of taking that idea and saying, I have a question for the user now, and I have all of my corporate, all my enterprise data out there.

And REG sort of picks the right sets of information out of there that needs to be pulled in to the context of the LLM in that sense. And so I think of it as, again, a selection... Selector operator that helps you pick what's important, into the the context engineering sort of, landscape. And as you have been working in this space and particularly in your capacity doing developer relations, working with people out in the community? What are some of the most interesting or innovative or unexpected ways that you've seen teams deal with the production requirements of RAG systems?

The first thing I I wanted to mention there is not necessarily unexpected, actually expected, but important is how do you deal with roles and permissions. Right? So this is an important problem in Rag and in agents. Right? So I... Let's say I have all my Google Drive ingested into my Rag pipeline. Maybe the salaries of the company, there's Excel sheet that's all... Also there. Only the CEO and HR have access to it. But if I ingest it without any protections, then everybody can can pull information out of that. Right?

So you wanna make sure that you have this, role based access control the same way the company implements it be applied in your retrieval pipeline. That's a really important thing to do. And the... I guess the kind of the the innovative ways people implemented this is just by implementing some form of filtering. Right? So if you say attach to every document that you ingest into your pipeline, you touch some some notion of permission. And, again, permission is a whole world of complexity.

And a lot of companies have multiple permissioning systems, and, you know, it's it's a mess. But if you can normalize your permission, your role into something that you can sort of add as, say, metadata to the document, and then when a query comes in, you can sort of say, if if the CEO is asking the question, that... Then I'm gonna I'm gonna allow them to see everything. If it's, somebody on the... Whatever, on on another team that's not supposed to see the salaries, then he would then filter the documents that are sort of have the the permission block in the metadata and not include that in the checks even if it's relevant.

So so the overs... Supersede relevance with with filtering beforehand. And that's been really helpful. That's one thing. I think multimodal has been an area that I've seen a lot of innovation. So I mentioned this a little bit in the beginning on the indexing or the ingestion side. But think about multimodal also on the UI side. Right? Your application needs to show the response. But if you have an image, you kinda wanna integrate it into the response.

So, again, a lot of people don't think about it. If you have a video, you wanna embed the video in there to show where the response was retrieved from and maybe even point the video to the right second. Right? So you kinda wanna have some innovation on the UI side to make the user experience, like, really helpful. Ultimately, everywhere else, every other application we've ever built in tech, people are lazy. They wanna get the answers right away, and they want it to be appropriate for them. So you kinda work really hard to make this work in this way in in multi... Multimodal data.

I think the other thing is, as I mentioned, I think people are working really hard to make it be viable, both REG and AgenTik, in sort of the air gapped environments. That's a really important use case for a lot of big companies. Think about, you know, people working in the government industry where there's, like, different whatever agencies and things like this that they can never leave, or companies in the defense industry or semiconductors where there's a lot of competition.

Those are areas where we see, even financial services, again, a lot of different things. So if you unpack that, you have to do a lot of innovative stuff to do this. So one of the surprising things that I bumped into was that, well, somewhat surprising, would say, is people are afraid to use the Chinese models, right? Chinese models, the open source models are very, very good in terms of performance, right? They are on par with the commercial models that are API based to some... Almost almost there, I think, like, MiniMax three, Kymi two, you know, three. Like, there's a lot of recent models, QN, like DeepSeek. All of them are... By the benchmarks, they're really good.

And they can be hosted locally, and there's not too many models that compete that are sort of by US companies. But I see a lot of resistance of using these models in at least US and European based companies. And, I can see why that would be concerning to some extent, but... And so I I really like, for example, that, thinking machines recently came with a very good use based model, and I see there's a lot of movement there. So that's something that initially surprised me because open source weights are just weights.

They can't, like... I don't see... As a technical person, I don't see a way how this... They would send something or the Chinese could control them. Maybe I haven't thought this through enough, and I'm not an expert on security and relationships like this. Maybe it's about the fact that people might get used to them. And then once there's a need for additional code, they'll forget about it or or something like this. I I don't know. I'm really not sure about this. But I've seen a lot of resistance to that for sure in this sort of, server and AI space.

And in your experience of working in this area of building and evaluating RAG systems, what are some of the most interesting or unexpected or challenging lessons that you learned personally? Well, first of all, it's something I mentioned earlier that teams don't think about the challenges upfront. They just... Mean, they rush to build a system. And like we said earlier, it's... It was a lot... It used to be a lot of engineering. Now it's a little bit less, but it's a lot of work to build a scalable, robust sort of, like, system with all the bits and pieces.

And a lot of teams... I was I was surprised that people don't apply the same rigor that we do to any other system that we ever built in microservices or whatever other technology in the past to the system. I think part of it is because this has been a lot... It's happened so fast and there's been a lot of push like, hey, we gotta go into AI. We gotta AI fire systems or whatever it may be. Right? So I think that's that's something that I've seen, like, the rush and not thinking through it and then being surprised that it's more complicated than we thought when you go to production scale.

And the second thing is really... And we talked about it too, is, like, just devaluation, how how little people think about it. And they're then surprised that results are not good, and then they put it... Slap it on the system after the fact. And then realize, okay, how how difficult it is. It's not just creating questions and answers that becomes complicated. So people are used to regular applications, right, where where the metrics... So people understand metrics. Right? You have to have measure latency, measure downtime, you know, all these things that we're accustomed to from the kind of DevOps world. Right? But these new metrics that are ML based and sort of LLM based are new and more difficult and people, like, don't think about them upfront.

What I've seen is if you you have an ML team internally that is powerful and strong, then they might tell you about this and you might think about it upfront, but many companies don't, and then they they don't think about it in that sense. And as you continue to explore and work in this space and keep tabs on the ways that the industry is evolving, what are some of the predictions or hopes that you see for the future of RAG systems and context engineering and just the overall space of data driven LLM and agentic applications?

So... Yeah. First of all, I think retrieval will continue to evolve, not... Probably without much the g. The g is still... Some people still use NaiveRag in in the old way, and it's available for them. But I think more and more will go in... Certainly go into AgenTik, and we'll just use the retrieval part where it fits and continue to improve that. I still see new embedding models being created that are better, new ranking models that are better. So there's there's still work to do.

People use glossaries, for example, to... And build models that are specific. They're fine tuned specifically to that particular industry, which helps with fields. So there's a lot a lot more to do in retrieval and search. I definitely think we will see the retrieval of other systems that you connect to directly using a MCP tool or something like this become better and create better results. I still see a a big gap in the multimodality in that respect, and I hope those those vendors will will keep improving that as well.

So that's one thing. Context engineering, again, is... Can used to be a buzzword that's very popular. It kinda goes into what people think of as memory systems now. So people talk about, like, the immediate memory, like, in a in a particular session in a chatbot. How do we talk, you know, back and forth me in the chatbot? And so if I say something that we talked about before, I wanna remember that. That's the simplest form. But then there's things like, oh, twenty days ago, we talked about something related to, I don't know, maybe different banks in my area.

And I told you, the agent, that I like this bank, so don't ask me that again. So, like, historical context of earlier conversations. And I think, both of these are really interesting, and people attack them in kind of, different ways. But I think, again, that'll probably be commoditized and be, like, certain two or three approaches that people do, and they'll be really easy to use. In terms of agents, of course, I think there's a lot of interesting things coming up. So one...

And let me talk about regular agents and then coding agents because I like talking about coding agents separately too. On regular agents, you know, the agent... The LLMs do really well with calling tools and deciding what to do with agents. There's no doubt about it. If you compare it to a year ago, it's incredibly better. But we've seen the agents move to what's called long term horizon sort of things. And I think they'll still suffer from accuracy when the task is... Takes a lot of steps Because what happens is the context becomes sort of like poison a little bit. There's like what's called context rot a little bit. And it'll start doing the wrong things a little bit or asking questions again.

So for example, when I use a coding agent like Cloud Code or Codex, a lot of times when I use the same session for a long time, after about 200,000 tokens or something, I start seeing it not act as well. And I sometimes benefit from, like, killing it starting a new session. Right? So there's... This happens in regular agents too. And I think we're... We have to... As an industry, I'm looking forward to seeing what solutions people come up for for that.

On coding agents, I think we're seeing this interesting evolution from one coding agent to a swarm of agents. And that's really something we're working on at Band about creating this... Using Band to... Which is agent agent communication to create this idea of multiple coding agents that talk to each other to solve a problem, that that collaborate together. Just like we humans in a team can collaborate together. Right? So I think this idea of coding agents that specialize in a particular area, so you can have a coding agent to write a code, a coding agent to plan, a coding agent to review the code or to review the plan, a coding agent for compliance, a coding agent for machine learning, for design, for different roles.

And they have a very precise context. They don't have context of everything, so they can work more precisely on each of these. And then they collaborate and hand off work to each other as needed when they need to develop code for us and and do that. So I think in my mind, this will create a a much better outcome for everybody. And especially if you use multiple agents, you can take advantage of Clubcode and Codex and Copilot and, I know, Kiro or everybody else. And there's gonna be so many of these, coding agents produced in the antigravity.

Because every one of those has a little bit of a different harness, so it behaves a little bit differently. It has a different LLM that drives it. So you will see... You will keep seeing them. I think we all will keep seeing them sort of have different personalities, different, like, strength and weaknesses. And as we get to complex code generation, long horizon tasks, benefiting from sort of adversarial review and seeing multiple, points of view would actually create higher quality outcomes.

Are there any other aspects of the work that you're doing in this space or just the overall challenge of building and operating and maintaining these rag based systems that we didn't discuss yet that you'd like to cover before we close out the show? Well, I wanted to mention one more area, which is kinda SQL based data that we haven't talked about, and that's interesting. So there is a lot of work in text to SQL. And, again, think of it as a tool that agents might call. The reason it's interesting is very obvious, right? Like most mission critical data is still in databases, in Snowflake or Databricks or Microsoft or whatever it is, right? And Text to SQL is going through...

First of all, it's really good, Right? It's still... It it is pretty good now. And it's gonna get better over time. There's models being developed specifically to solve this problem. But I think one of the the interesting things that are happening is it's related to this world of semantic modeling, right? There's was Looker that started this, and there's now the new version of this called Malloy, which is open source. And there's other companies that do this kind of thing to build semantic models to enable enterprise to sort of define what something means in SQL. Because the LM, again, like with REG, doesn't have context of what things mean.

So if you ask it, tell me about the churn in our recent customer base, churn can be defined as somebody that dropped off for fifteen days or thirty days or a hundred days or whatever. So there's no definition the LM knows is the the absolute truth. And so the ability to define this is really important. And it was important before we had this with data and analysts and things like this, but now we have to define it for agents because there's gonna be so many more agents asking so many more questions out of these stacks of SQL pipelines that we need to enable, make sure that if agents are gonna do this at scale and we can trust it, it has to be enabled.

So it's just one thing I wanted to mention, that we didn't didn't cover. Yeah. That's definitely still an interesting space, and I see the word semantic layer and semantic modeling probably at least once a day, definitely multiple times a week these days. Alright. Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gaps in the tooling technology or human training that's available for AI systems today.

The thing that... First of all, let me start with human training. I still see a huge gap in terms of people adopting AI. Now I think you and I and maybe others in the audience, some of us live in the tech world, and it's... We track these things every... We try our best actually to track these things every day. Right? It's it's so difficult. There's, like, 50 things happening every week, and it's hard to to keep track of everything. But, AI is so powerful and so useful in a lot of different ways. And I'd love to see more. I see a lot, but I'd love to see more people and companies help with education of everybody of how to take advantage of AI.

This started from people don't even know how to use ChatGPT. Now I think more people know how to use that, but, like, how to use it effectively, how to connect it to your your own, you know, systems and things like this. Very, very important. I also... And a little bit in a narrower context, but related to that, I think the education system itself, the education industry, teaching, you know, kids, teaching adults that do later education, I think has to change. Right?

We have to integrate, in my mind, how to use AI effectively in every part of education. And it's a massive shift that has to happen because learning, memorizing things, and and, kind of answering questions based on memory is just not gonna work anymore. I'm very well connected to the education ecosystem here in in Silicon Valley, and I've seen this a lot in the last couple of years where teachers are looking for help too. How do I how do I teach Python?

Like, how do I test? Like, because the the test is usually solve this problem, and you give it to the coding agent. And the second letter, you have the answer. So I think people need help, not just, learning how to use AI, but how to teach AI. So I'm really looking forward to seeing how that... And I know a lot of people are working on it. I'm not saying it's not solved, but it's it's it's, I'd like to see, people help with that. Yeah. And, generally, I'm really excited to see what some of the commercial labs, come up with. I know there's a lot of, people that's, kind of, came out of the labs now or starting a lot of new companies, like, you know, Ilia and his company, SSI, and Thinking Machines, and and Coral Automation, and other companies like that that are trying to solve the next generation of of LMs. And also very excited about image and video generation. I think, I've been tracking that space a little bit too. And, kind of the way I like to characterize it, I feel like it's, it's in a GPT two, GPT three moment of its evolution.

It's very powerful. We use it all the time, but I think it's... The potential will be unlocked in maybe a couple years. We're all gonna see this amazing innovation and, much more, direct use cases that are that are more immediate in that sense. Absolutely. Well, thank you very much for taking the time today to join me and share the work that you've been doing and some of your insights on how to evaluate rag systems and get them running well production. It's definitely a very interesting and continually complex problem and one that keeps shifting. So I appreciate you taking the time to share your insights on that, and I hope you enjoy the rest of your day.

Thanks, Tobias. Glad to be here. Thanks for having me. Thank you for listening. Don't forget to check out our other shows. The data engineering podcast covers the latest on modern data management, and podcast.init covers the Python language, its community, and the innovative ways it is being used. Visit the site to subscribe to the show, sign up for the mailing list, read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@aiengineeringpodcast.com with your story.

Transcript supplied by the publisher with the episode.

AI Engineering Podcast

by Tobias Macey · English · Tech & Science

This show is your guidebook to building scalable and maintainable AI systems. You will learn how to architect AI applications, apply AI to your work, and the considerations involved in building or customizing new models. Everything that you need to know to deliver real impact and value with…

More from AI Engineering Podcast

  1. E79 · 19 Sep 2026 · 1 hr 4 min

    Harness Engineering for Reliable, Governed AI Agents

    Summary In this episode Nikunj Bajaj, co-founder and CEO of TrueFoundry, talks about the challenge of building reliable agents on top of inherently variable foundation models. He explores the idea of the agent harness as everything around the model engine: memory and context management, tool and MCP integration, sandboxed code execution, permissions, observability, and guardrails. Nikunj explained how TrueFoundry approaches enterprise AI as a centralized control plane for token traffic, while TrueForge provides an open source, vendor-neutral harness for building agents without locking teams…

  2. E78 · 25 Feb 2026 · 1 hr 1 min

    Kubernetes, Compliance, and Control: The Operational Backbone of AI Sovereignty

    Summary In this episode of the AI Engineering Podcast, Steven Watt, leader of the Office of the CTO at Red Hat, discusses practical paths to achieving AI sovereignty for organizations. He shares his two-decade experience in AI, highlighting how governments are building GPU platforms and protected data hubs to maintain control over AI workloads. Steve emphasizes why self-managed infrastructure is becoming a strategic necessity as companies outgrow cloud costs and require tighter control over models, data, and compliance. The conversation explores the operational substrate for AI sovereignty,…

  3. E77 · 15 Feb 2026 · 51 min

    From Blind Spots to Observability: Operationalizing LLM Apps with OpenLit

    Summary In this episode of the AI Engineering Podcast, Aman Agarwal, creator of OpenLit, discusses the operational foundations required to run LLM-powered applications in production. He highlights common early blind spots teams face, including opaque model behavior, runaway token costs, and brittle prompt management, emphasizing that strong observability and cost tracking must be established before an MVP ships. Aman explains how OpenLit leverages OpenTelemetry for vendor-neutral tracing across models, tools, and data stores, and introduces features such as prompt and secret management with…

  4. E76 · 8 Feb 2026 · 59 min

    Taming Voice Complexity with Dynamic Ensembles at Modulate

    Summary In this episode of the AI Engineering Podcast, Carter Huffman, co-founder and CTO of Modulate, discusses the engineering behind low-latency, high-accuracy Voice AI. He explains why voice is a uniquely challenging modality due to its rich non-textual signals like tone, emotion, and context, and how simple speech-to-text-to-speech pipelines can't capture the necessary nuance. Carter introduces Modulate's Ensemble Listening Model (ELM) architecture, which uses dynamic routing and cost-based optimization to achieve scalability and precision in various audio environments. He covera topics…

  5. E75 · 27 Jan 2026 · 46 min

    GPU Clouds, Aggregators, and the New Economics of AI Compute

    Summary In this episode I sit down with Hugo Shi, co-founder and CTO of Saturn Cloud, to map the strategic realities of sourcing and operating GPUs across clouds. Hugo breaks down today’s provider landscape—from hyperscalers to full-service GPU clouds, bare metal/concierge providers, and emerging GPU aggregators—and how to choose among them based on security posture, managed services, and cost. We explore practical layers of capability (compute, orchestration with Kubernetes/Slurm, storage, networking, and managed services), the trade-offs of portability on “Kubernetes-native” stacks, and…

  6. E74 · 20 Jan 2026 · 56 min

    The Future of Dev Experience: Spotify’s Playbook for Organization‑Scale AI

    Summary In this episode of the AI Engineering Podcast Niklas Gustavsson, Chief Architect at Spotify, talks about scaling AI across engineering and product. He explores how Spotify's highly distributed architecture was built to support rapid adoption of coding agents like Copilot, Cursor, and Claude Code, enabled by standardization and Backstage. The conversation covers the tension between bottoms-up experimentation and platform standardization, and how Spotify is moving toward monorepos and fleet management. Niklas discusses the emergence of "fleet-wide agents" that can execute complex code…

  7. E73 · 5 Jan 2026 · 56 min

    Generative AI Meets Accessibility: Benchmarks, Breakthroughs, and Blind Spots with Joe Devon

    Summary In this episode Joe Devon, co-founder of Global Accessibility Awareness Day (GAAD), talks about how generative AI can both help and harm digital accessibility — and what it will take to tilt the balance toward inclusion. Joe shares his personal motivation for the work, real-world stakes for disabled users across web, mobile, and developer tooling, and compelling stories that illustrate why accessible design is a human-rights issue as much as a compliance checkbox. He digs into AI’s current and future roles: from improving caption quality and auto-generating audio descriptions to…

  8. E72 · 29 Dec 2025 · 54 min

    Beyond the Chatbot: Practical Frameworks for Agentic Capabilities in SaaS

    Summary In this episode product and engineering leader Preeti Shukla explores how and when to add agentic capabilities to SaaS platforms. She digs into the operational realities that AI agents must meet inside multi-tenant software: latency, cost control, data privacy, tenant isolation, RBAC, and auditability. Preeti outlines practical frameworks for selecting models and providers, when to self-host, and how to route capabilities across frontier and cheaper models. She discusses graduated autonomy, starting with internal adoption and low-risk use cases before moving to customer-facing…

  9. E71 · 16 Dec 2025 · 1 hr 8 min

    MCP as the API for AI‑Native Systems: Security, Orchestration, and Scale

    Summary In this episode Craig McLuckie, co-creator of Kubernetes and founder/CEO of Stacklok, talks about how to improve security and reliability for AI agents using curated, optimized deployments of the Model Context Protocol (MCP). Craig explains why MCP is emerging as the API layer for AI‑native applications, how to balance short‑term productivity with long‑term platform thinking, and why great tools plus frontier models still drive the best outcomes. He digs into common adoption pitfalls (tool pollution, insecure NPX installs, scattered credentials), the necessity of continuous evals for…

  10. E70 · 24 Nov 2025 · 1 hr

    Context as Code, DevX as Leverage: Accelerating Software with Multi‑Agent Workflows

    Summary In this episode Max Beauchemin explores how multiplayer, multi‑agent engineering is reshaping individual and team velocity for building data and AI systems. Max shares his journey from Airflow and Superset to going all‑in on AI coding agents, describing a pragmatic “AI‑first reflex” for nearly every task and the emerging role of humans as orchestrators of agents. He digs into shifting bottlenecks — code review, QA, async coordination — and how better DevX/AIX, just‑in‑time context via tools, and structured "context as code" can keep pace with agent‑accelerated execution. He then…

Every episode of AI Engineering Podcast →

Take it with you

The Melo app keeps playing with the screen off, works in the car and on your watch, wakes you to your station, and browses the whole catalogue offline. Free, no ads, no account.

Get it on Google Play