Episode 516 · Data Engineering Podcast
What Context Really Means in Data Engineering and AI
15 Sep 2026 · 53 min
Episode 516 · Data Engineering Podcast
15 Sep 2026 · 53 min
Summary In this episode Soham Mazumdar, co-founder and CEO of Wisdom.ai, talks about what “context” really means in data engineering and AI systems. He explores why context has become such an overloaded term, spanning everything from semantic layers and data catalogs to tribal knowledge, query logs, dashboards, and even agent memory. Soham explained that the big shift is that context is no longer being prepared primarily for human analysts, but for LLMs and agents that can’t reliably fill in missing gaps on their own. That change raises the bar for how context is represented, validated,…
Tap a chapter to play from there.
Tobias Macey:Hello, and welcome to the Data Engineering Podcast, the show about modern data management. Your host is Tobias Maci, and today I'm interviewing Soham Azumdar about what context actually means in the context of data engineering and AI systems. So Soham, can you start by introducing yourself?
Soham Mazumdar:Yeah, sure. Tobias, thanks for having me on the show. Yeah, so we, I'm co founder CEO of Wisdom dot ai. We are approximately three year old company looking to bring AI into data analytics. We are solving the problem of how do you make sure that organizations can get super trustworthy insights from data. And again, as part of that, the big challenge that we're in fact solving for organizations is how do you take enterprise context and pin it down, make it usable in, should I call it, trustable, benchmarkable manner so that you can truly hang your hat on it. That's what we do.
Tobias Macey:And do you remember how you first got started working in the overall space of data and AI?
Soham Mazumdar:Yeah, so I have a I guess I started my career with Google a long, long, long time back. And as such, getting interesting insights from data has sort of forever been in my DNA. But a more proximal reason for System AI was my prior role where I was running engineering at Rubrik. It's a cybersecurity company. And, you know, I had made myself the data analyst for the company for a little bit because we were trying to kind of diagnose some pipeline related issues and later engineering productivity issues. And my main observation was that, heck, this is like a super, super hard thing to do. You know, despite having a fairly technical background myself, unless I made myself data engineer full time, there was no way I was be going to be able to get the insights I needed because it was essentially inherently fragmented.
Context was not there. There were no LLMs to help me do anything fast. Tools, I guess, Snowflake or Tableau at that point for data visualization for us, they were designed for a very different point in the you know, for a very different audience, I would say. You know, like, for somebody to get quick insights from data without spending a whole amount of time setting things up, those tools were simply not the ones for us to use.
Tobias Macey:One of the ongoing challenges that has only been exacerbated recently is the overall complexity of having a shared understanding about what any given word or phrasing actually means for a particular audience and use case. And context is one of the ones that has become increasingly overloaded with the advent of agentic engineering and AI engineering. And I'm wondering if you can just start by sharing some of the ways that the overload of that phrase has contributed to problems in your own experience.
Soham Mazumdar:Yeah. I mean, yeah, one would never think two years back that context would be the hottest buzzword in the world of AI, would we? Yeah, so I think the biggest problem with context is that it's extraordinarily hard to really define what exactly is context, right? I think, I mean, ultimately it's like a big bag of words that help guide the LLM to do something well, Right? So in that spirit, almost anything can be context. And then when people try to pin it down, there are a few relatively well understood ways. One of them would be like, let's bring in a semantic layer. So semantic layer is what I would describe as extraordinarily formal context. You know, like you have precise definitions of KPIs, precise definitions of, you know, dimensions and so on. Another flavor of context would be what, let's say data catalogs and so on bring in, which is like, you know, you bring in things like lineage, you would have your schemas documented and so on. So these are what I would still describe as formal ways of talking about context.
But I think the biggest challenge with context is there is so much that's very hard to pin down. You know, you can call it tribal knowledge. You know, like I've seen data analysts who almost maintain a cheat sheet pasted right next to their monitors because that is sort of what they are consulting when they do their analysis. So this is kind of tribal knowledge. And then you try to scale it across the organization. You try to get knowledge that sits in the heads of business users. And that's like a whole world of information that's extremely hard to capture. So I think, again, the main challenge with context is very hard to pin down. There is formal context, there is informal context. And the other thing is context keeps changing, right? So problems with like, how do you pin it down?
Tobias Macey:Within that framing of context, when you're talking about data engineering, one could argue that data engineering, since its inception, has really been the overall exercise of context engineering. What is it that this information is actually trying to tell me? Why do I care? How do I use it? True. And I'm wondering what you're seeing as some of the ways that that overall conception of the role of context at the data engineering space has shifted when AI and agentic consumers are the predominant consumers.
Soham Mazumdar:Amazing. So I think the, you know, all the work on data engineering that has happened, let's say whatever happened maybe a before a year from now, they were all designed for non humans to They designed for humans to consume in a certain way. They were not designed for agents to consume. Humans essentially have a ton of judgment that they use. You simply do not know whether it's even if you give like a semantic layer to an LLM, whether the LLM is going to do the right thing with it or not. So this whole aspect that the end consumer of, you know, like the catalog or the end consumer of the semantic layer is no longer the human data analyst, but it's in fact an LLM. That is the big shift that has happened. So this whole aspect of whether the context is good enough or whether LLMs understand the context correctly or correctly enough to take the right actions, this was never considered, right? So you just, you know, you can build documentation in a vacuum and it all sounds good, but is it going to drive the right outcomes or not? So I think the biggest shift that has happened is now you are talking about feeding it into an LLM. LLM,
you need to prove that this context is good enough for the LLM to work. So this whole thing around benchmarking context for accuracy, benchmarking context for lack of ambiguity or like, you know, for the amount of ambiguity it contains, making sure that the LLMs do not take like 20 terms to understand your context, make sure that the tokens get efficiently used because your context is not too much garbage, right? So I think those are like problems that were simply did not exist a year back, right? So when you built like a semantic layer on a catalog, it was like, you know, somebody, some human is going to look at it and they will fill in the gaps. But with LLMs, there is nobody to fill in the gaps, right? So it needs to be this like, you know, almost much more higher fidelity, pristine context. And that's where I think the problem really comes.
Tobias Macey:The other interesting element of this idea of context, what is it? What does it mean? Is being able to go from, I have a maybe implicit and contextual understanding of what this is supposed to mean, but how do you then turn that into something that is a useful and reusable and concrete artifact that can actually be verified and validated and understood as a component in a broader system.
Soham Mazumdar:100%. I think you mentioned this word right now that something that is validated. I think the moment you talk about validation, first of all, you are even forced to define what does good look like. That simple thing of what does good look like was never a constraint earlier. So now before, let's say, again, timeless plug for our product, when we go into any organization, we almost work backwards from the outcomes that you're looking for, because it is very important for us to not just drop off a system and say, Hey, go, you use it. It's always, Let's first define what good looks like, which is create a benchmark. And this benchmark is like a precise set of answers that must be delivered through agents.
And that should be, you know, like the questions themselves have to be in a colloquial language. So this is not like, you know, this is not about like dragging and dropping anything into like a visual portal, right? So this is like truly a natural language. So we start from a benchmark first. And then there is a ton of tooling that we have built. It is essentially finding flaws in people's context. So when people say, Hey, I have some context, they will give you some sort of information. It could be a catalog, could be some documentation.
But the test is whether I can give you like this 100 benchmark. We actually say 95%. Can we hit that 95%? And in order to get to that 95%, we often discover, well, there is missing context. There is overlapping contexts which conflict with each other. You have three definitions, which there are some variations which makes sense to use them in one context or the other. This is a problem. The other thing that ends up happening is when you have an agentic system live in production, your users go off script all the time, which means what you thought you have trained the system for versus what it eventually gets used for, there is often a big drift.
And we call it like context drift. You can essentially say that humans can just go off script. That's the best way to describe it. And which means that your context, I mean, now what the data engineer needs feedback on is, okay, are the gaps? Like, where is the agent trying to take it, which I did not anticipate? Where is it in fact making mistakes? Because, you know, like we did the basic benchmarking, but then over time it's kind of going in its own ways. So you kind of, there's a whole new class of tools that you need to build to essentially understand how agents are using the context, where are the gaps, how do you incrementally improve it. So this is, think, I would say a class of engineering problems that were simply things that you never solved earlier. And that's kind of where a lot of our own core IP kind of comes in, which is how do you build trustable context and how do you maintain it over time?
Tobias Macey:The other piece of it too is what is the format of context as an artifact where for a long time it was you had the dimensional warehouse. That was the business context. It was intended to be a digital twin of the organization. Obviously, it was never a perfect replica, and it has lots of its own challenges associated. But what are some of the ways that the data warehouse needs to change or supplemental systems that need to exist when you're talking about an agent being the primary either consumer or creator of that warehouse environment.
Soham Mazumdar:That's a great point. So, you know, I as part of running my company, I have to come up with some marketing friendly lines of how to describe context, right? And we have tried calling it knowledge graph. We have tried, I don't know, like ontology is a buzzword right now. I one of the problems with context is that it is a collection of data structures that all have different strengths. And if you use exactly one data structure, you are going to be in trouble. So I think, again, at a very conceptual level, I would certainly say that there is some sort of relationship graph is a fairly sound way of describing context, right? So every business has concepts and concepts have a linkage between them.
Again, that has its representation on the data side. You know, you have tables and tables have relationships and so on. So that's what I would describe as one form of like relatively graphical context. Then again, is experiential context, which is like, this is what the agent did. In fact, like agent's memory itself is a form of context that should help the agents do better in the future. And that's a context that is hard to pin down as a context graph, but the best way to kind of describe it is as a trajectory catalog that you can kind of mine over time. That's a different data structure. I think somewhere in between you have like query logs, for example, they are context. Or like old BI dashboards, they are context.
That's, you know, I mean, best way to think of them is, you know, it's this unstructured pool of data from which you can extract meaning, but like it kind of, it starts its life as a big, big unstructured data dump. So that's another flavor of context. I think what our At least I have a very pragmatic view on this. Even though I would like to market it as some beautiful knowledge graph, the reality is it's a combination of about four or five different data structures that need to work together.
And finally, the proof is in the validation routine. So ultimately, whichever way you represent context, it only matters if it gets you the right outcomes. Again, there are people who have tried to standardize this, like this OSI thing that Snowflake came up with. I think that's a good attempt. I mean, in our experience, it's insufficient to drive pure accuracy. Or in fact, if you have to bring every organization to look like OSI, what OSI model support, you know, you are up for, I mean, it works for, but it only works for a small subset of cases, right? So it's a
Tobias Macey:good attempt. But I think at this point, my view is it's a far bigger composite of different structures through which you're. And as you're talking about the different formats and shapes and requirements around these projections of context, it just makes me think about the perennial challenge of every data engineer everywhere of having to battle the siloization and fragmentation of information where the data warehouse, one of its primary goals is get all the data in one place so we could do something with it, which is never sufficient. And so you end up having to build all these different views and projections over that and reflect it into another system if you wanted to be able to do OLAP versus just historical querying. And I'm curious how you're seeing some of the complexities grow as you're bringing these fast moving machine operators into the mix and some of the ways that you can have a shared ground truth that everything else is built and projected from rather than having to do bespoke implementations in multiple different ways.
Soham Mazumdar:That's a great point. And I think one of the problems that also introduces is, I think this maybe goes a little bit beyond data engineering. General spirit of data engineering for always, I general, I would say the central fundamental premise has always been, let's bring data into one place within that one place, let's manipulate it to get to get it to a shape where it becomes really easy to consume. And then let us hand it over. I think what is happening with all agentic systems now is that particular invariant is sort of breaking because almost every agentic system has the ability to plug into a diversity of sources. You know, there is, you know, sure they're talking to your Snowflake, but they're probably also talking to some transactional database.
They're potentially also talking to like a SaaS source, which has a very different contract. Like there is an MCP or there is an API, that's a different way of kind of using it. And then how do you even express the linkages across these different systems, which are not even inside one homogenous snowflake, right? So first of all, again, that brings up another problem, right? I mean, you can think about, let's say, DBT style way of expressing your context. But the moment the data is not in its snowflake and spans multiple sources, simplicity breaks down. The moment you pull in MCP or API, then that further breaks down. So first of all, again, you need to think about ways to express context in different sources. Ultimately, the only, I would say, context expression technique we have come up with is the only truly universal context expression system we have come up with is one which is entirely built around trajectories and entirely built around observing agent traces and then extracting, finding ways to optimize it for repeat usage. That probably is the one universal, but for everything else, you actually need a collection of different techniques that are more suitable for
a different vector. I would again say like, and the other thing that you need to need is a way to express relationships across different silos. And that's yet another, I mean, the way we do it is it's a combination of, again, like agent trajectories and, you know, and again, like pure unstructured ways by which we can express it. Our knowledge graph also kind of allows you to connect between sources, but it's not one technique that can pin it down even in the it's in fact makes it even harder as you kind of get to these multi source words.
Tobias Macey:And when you are dealing with multiple different storage systems for context where an agent might need to query multiple of them to perform a given action, what are some of the patterns that you found useful for being able to build an abstraction or a facade over that disparate storage semantics versus a single unified access plane for being able to actually pull in that relevant information to complete a given objective?
Soham Mazumdar:Yeah. So, I mean, we have, I mean, we have done a few things. I mean, first of all, having a I think you ought to think about hierarchies. You ought to think about conceptual link. I mean, like, you know, a semantic layer makes you create some very precise linkages between tables, But I think you need to kind of abstract at one level and, you know, like probably like a good analogy would be like a Palantir ontology. Like I think they call Palantir Foundry, I believe. It's a way of expressing relationships across different disparate sources, which you may not have brought into a single plane.
So you are talking about relationships between entities that exist within different systems. That is a graphical representation. I think, again, you need to break down systems into We call every distinct system, we call it a domain and we have a way by which domains can share context, the way we can create hierarchical domains. So it's a bunch of additional concepts. Linkages across domains, inheritance relationships across domains, again, like sharing collaboration concepts across domains, graph structure, and then ultimately coming back to the last failback, which is essentially agent trajectories.
So it's this which lets you create a system. Once again, will say the closest analogy I can think of is like a Palantir Foundry style representation.
Tobias Macey:In terms of the source of the information that you're working with for data warehousing, it's a fairly broadly understood set of types of data that you'll be pulling from where that might be application databases. It might be event queues. It might be third party APIs. But when you're talking about bringing an AI into the mix, you also need to have more of that organizational context implicit because the AI doesn't actually understand any of that natively where a human operator building these different pipelines is working within the organization.
They build up that knowledge that they're able to store. And so I'm wondering what are some of the other ways that bringing this agentic context into the data engineering space requires branching out into an even broader set of source systems that you need to pull from.
Soham Mazumdar:I think, you know, we we were building this, you know, trying to use Wisdom AI internally, and we were actually running into a lot of, like, challenges because, you know, we one would think that it should be pretty straightforward given that we understand the system so well. But ultimately what we realized was that our source code, and I'm not talking about DBT source code, I'm talking about the source code that the developers use for building the product. So that was like a hugely beneficial way by which you were able to kind of get stuff done. That's one interesting one. Another interesting source has been, so anytime we go into very complex domains, we like to bring SMEs, not as rather, like pure business folks who are non data team members, participate in the pilot process. And a lot of the signals we learn is ultimately by shadowing the these business users because they sort of have the you know, they are using multiple systems and they are they are the ones who will go off script, so to speak. So that's a big source of learning for us. I think, again, I mean, the standard sources, of course, you know, like query logs and BI,
we're kind of reverse engineer practically every BI system that's out there so we can we can extract context from all of them. Yeah. It's like about seven or eight independent ways in which we will build context. It all comes together, though, however, because of the measurement system that we have in place. Because, you know, ultimately, as I said, you know, context is just like a bag of words. It doesn't inherently have value unless it leads to agents doing better things. So again, anytime we deploy our product, we'll start off with some base understanding and that's maybe 50% accuracy. And then over a period of like a week or so, we keep hill climbing and getting it to a point where we can deploy it across the board. And in that meantime, we are kind of like synthesizing context from all sorts of sources.
Tobias Macey:One of the other net new challenge well, maybe not net new, but a new challenge that is becoming more broadly visible in particularly when you're talking about context is the available space for being able to stuff information into the LLMs window and different LLMs will pay different styles of attention to the information, which might require changing some of the formatting or ordering. And I'm wondering how you're dealing with some of that complexity of the different LLM consumers, the agent harnesses, the other information that needs to be present so that you're not just completely overfilling the context window with all of the organizational structure and being able to do appropriate context pruning, making sure that you're ranking the relevant details appropriately for the given objective, etcetera.
Soham Mazumdar:And I I wish I had a chance to show you some go on the whiteboard and draw like what the architecture here. So I think fantastic question. I think so the way to think, we think about our context, first of all, we call it a context engine. There are essentially three aspects. So there is one which is representation of context, which is how you store context. Then there is second, which is I think we talked about earlier, which was like, how do you learn context? How do you maintain context over time? Let's just call it tooling problems, right? So there is representation and tooling. And the third one, which is arguably the most interesting one, it is the runtime aspect of context, which is how do you activate context? How do you use it optimally?
And there are two problems that you need to Actually, it's maybe four problems that you need to solve. Number one is what you alluded to, which is how do you use just enough context so as to not blow past the budget that the LLM kind of gives you. And this requires you to search prune variety of techniques to get it down to exactly what you need. We do that because again, like our product is very much like a, you know, the end consumer product. So we absolutely do that. The second big piece that you need to think about is token efficiency.
So again, so there is part of it is like the input tokens that are going on and the input tokens. If it's too many, then you wouldn't even fit the context window. So it's kind of a bit existential that way to get it right. But the second bit is the number of tokens you pull in, number of tokens that you call out, how many LLM calls are you making, how complex is your agentic loop. Because often what happens is the agent has to work very, very hard to compensate for bad context. So essentially, if you give it bad context, you're essentially forcing the agent to make wrong turns and take those wrong turns, backtrack, and then eventually orient itself and get it right. So this has two problems. One is that your latency is kind of like dramatically increasing as the agent has to do more work, right? So for, if you have like a time sensitive operation, you know, you can have vastly different latency levels depending on how optimally your agent is able to operate with that context. So that's a second dimension that comes in. And there is a third one, of course, which is purely cost, which is, you know, you have more round trips,
you have more input tokens, more output tokens, all of that essentially ends up turning into higher token costs. And lastly, of course, the choice of the model, right? Because depending on the tasks you're doing, can certainly trade off the kinds of model you do. So, I mean, these are like our, I mean, great news is that these are all problems that need to be solved. So it just gives us the opportunity to do more. So there is like a, you know, like a, I mean, we sort of benchmark ourselves on if we took, if we did the best possible context curation and then gave this context to Claude versus we take the same context and pass it through the Wisdom AI engine, what happens, right? And we've been able to kind of essentially come to a point at this point where we can save on costs 3X latency, you know, like about two to 3X as well.
And again, the accuracy was always there. And it all comes down to learning from bad agent runs and continuously improving the system, right? So because, know, initially your context is, again, a big blob of, you know, some bites, a clot sees some bites, we see some bites. It can only be so good. What becomes very interesting is, you know, we start learning, okay, where are the problems with your context? Where do we need to prune more? What should we discard?
And how do we look at old agent runs and then ultimately make the new runs better? So where is the agent getting confused? What are the wrong turns it is making? The whole goal is of continuously maintaining context is in fact to prevent the agent from doing stupid things or prevent the agent loop to compensate for the quality of your context. So I think, you know, like a lot of the IP that we had to build is in fact around how do you make the agent do the work with least, you know, with, you know, using context in an optimal manner. Pruning, routing is a big part of it for sure. But again, like finding places where it's getting confused is another big, big place of where we get learnings.
Tobias Macey:And you mentioned that you capture the agent traces. Obviously, there are other sources of telemetry and observability that you need to factor in when you're figuring out, do I have the right context? Yes. Is it fetching any of this context, or should I just delete it because it's just sitting stagnant? I'm wondering what are some of those aspects of observability that you're investing in?
Soham Mazumdar:Yeah. So there are three categories. One is what you described, which is which part of the context makes sense and where? Where do you have problems in your context? Essentially, contexts with both seem to be saying the same thing or about the same thing in a slightly inconsistent manner. So a lot of the work we do, in fact, is finding, weeding out inconsistent context. It's extraordinarily easy to create inconsistent mutually inconsistent context. So that's a big part. The second portion is finding places where, you know, the context is clean. On the other hand, the agent is still taking some suboptimal decisions. Suboptimal decisions come in many forms. So one class of suboptimal decisions is to have queries that are very expensive, you know, because there is a warehouse cost side to it, right? When an agent produces a query, it needs to run somewhere and that could be slow. And that could be slow because, you know, the query is correct, but it is just inherently slow. So, and that could be because of variety of ways, I don't know, you're not using the partitioning information correctly.
Or, you know, like, you know, the filters are not being used correctly, or this thing should have been, you know, you have broken it down into two steps and that is kind of creating some additional information that needs to be managed. So there are a variety of like optimizations on the query side, which again, often the agent is not doing because like the context sort of has not guided them towards that. But, you know, we essentially, because we are always trying to learn from what the heck the agent is doing, we always kind of like, we'll be seeing these like, you know, you know, the telemetry coming from the warehouse, the query plans and the time it takes to execute the queries. And that also further, you know, us tune the context. There is another piece that we do, which is we mine user sessions quite a bit. So essentially, I don't know, users sometimes in all caps will say, Hey, you did this thing wrong. You should have done it this way. So that's another very interesting thing, which is like, you know, when you give it out to people to use, they give their feedback in many ways. Nobody presses thumbs up, thumbs down. I mean, that's like those days are gone. It's always like some sort of implicit feedback that is coming. So that becomes like another,
you know, like, big source of context where we can mine context to to improve the system. So, like, I don't know if I answered your question, but it's a variety of different, like, qualitatively different sources of information.
Tobias Macey:I'm also interested in digging into from the perspective of Wisdom AI and the information that you're collecting and exposing and the ways that you're making it more of a natural fit for agents to consume, what are some of those typical agentic workloads that you are powering? Is it just text to SQL where the agent needs to know, okay. This is what a customer means, so I'm gonna make sure I write the right SQL. Is it somebody generating a BI dashboard based on warehouse data? Is it somebody who's using an agent to automatically generate new organizational insights and action items? Was curious, like, what are some of the styles of agents and agent use cases that you're focused on empowering?
Soham Mazumdar:Yeah. So I think Text to SQL is what I would call as like the like one fundamental low level tool that the system kind of possesses. You know? And again, like it comes down to the fact that if the data lives in a warehouse, then you absolutely need SQL somewhere in the picture. So again, so there are a bunch of like data access tools we have. So SQL being one, hence there is a SQL. Then if it lives in a differently shaped database, like, I don't know, Elasticsearch as an example, or like a graph database, or if it's like a more document corpus. So there are a bunch of different data access mechanisms that exist and SQL being one of them. And I would say the most important one given that warehouse is a big part of our dataset.
But where it gets really interesting is what happens beyond that. And so we have like a, you know, like the, I mean, there is like a sandbox within which we can actually do a lot. So there is a bunch of like data science toolkit that sort of like exists and that's, you know, just call it some flavor of Python. There is a federation area as well, which kind of allows us to combine multiple data sets in interesting ways. In terms of what you can get out of it, you know, you can certainly like ask a question and get a summarized response.
You can create entire BI style dashboards all within very simple prompts. You can also create One of the things that we just launched is apps. So, you know, you can build a, you know, instead of a dashboard, which is very ultimately, it's like one read only interface surface that you have. I think with apps, there is like some basic reactivity that you, we also allow people to do. That's another one. And then lastly, there are a lot of things that just happen behind the scenes. Think of it as building a piece of automation where one of my pieces of automation is I'm constantly monitoring customer usage and I get this summarized view on which customers I need to pay attention to or something around escalations coming from wherever things may not be going well. I have some similar agents deployed for our GTM side around health of the sales pipeline and so on. So there are a bunch of Yes, there is a sequel somewhere down at the one end of the spectrum, but like what I am consuming is like typically like many levels of inciting and inferencing that has happened on top of it.
Tobias Macey:And the other perennial challenge of any data work, but particularly when you're dealing with agents is the complexity of understanding the required freshness and being able to invalidate information before it can cause false assumptions or bad outcomes. And I'm curious how you deal with some of those aspects of data freshness, propagation, being able to understand whether and when an agent is using stale context and what can be done about that versus just retrying whatever the objective is and some of the ways that that can have cost management repercussions because you're wasting tokens on an outcome that is suboptimal because it didn't have the right data at the right time?
Soham Mazumdar:Great question. I wish we were doing more on it. I would say some of this is a shared burden with the overall operators of the system, as in beyond the point, it's harder for us as a, you know, like a consumption layer up top to make full determination on it. On the other hand, the way, you know, what we can do certainly is, again, we sorry. I'm maybe I'm giving you a cop out answer. I mean, we give the building blocks at least when we are creating agents. You know, we can essentially set up agents the a way to sort of like back out when it detects various anomalous situations. Anomalous situations would be, you know, some pipeline has not run and as a result, the data is unfresh.
Alternatively, there is some, you know, like, let's just say there a, there are some easy variations, you know, like the schema has changed in a peculiar manner, you know, like, or, you know, those are like, I would say the easier scenarios that we can just flag and that we just surface automatically. But if it's like at the data level and the data is not fully fresh, I think we give you the tools to monitor it. On the other hand, it's a shared burden between us and the administrator of the system to make sure that they're sort of like setting it up the right way. Because like, so there are at least some customers who use us as a very focused on data quality, right? So in that universe, the whole objective of the agents is to monitor data quality. On the other hand, there are folks who kind of keep that, say that is somebody else's problem and let us focus on the end outcomes. And then it ultimately becomes like how you have set it up. Know, you can certainly set it up in a manner that you will not be cognizant of it. On the other hand, you could have done it the right way, in which case we would be very aware.
It's both use cases we have seen people do.
Tobias Macey:And so digging now a bit more into the design and implementation of Wisdom, I'm definitely interested in kind of the architectural principles and components that you're building from, but I'm also interested in some of the ways that your understanding of the problem space and your approach to resolving it has changed from when you first started building it to where you are today.
Soham Mazumdar:Let's start with what is context. Three years back in our deck that we have presented to our seed investors, we had essentially, I think I'd used the word semantic layer at least like 10 times because it was just, it just seemed to, it sounded like all you needed was a good semantic layer and it will be okay. And I think with our first customer who happened to be a Looker customer, they said, Okay, take our semantic mean, we have LookML, take it. And we learned just with that first customer that, gosh, that pretty looking semantic layer is so far from what we need for success.
So that was like our probably our first learning. And that happened in the first six months of the company, six, seven months of the company. The second one I would say is six months into the life cycle of any customer, there is like a little moment of reckoning. Because, you know, like upfront people spend a lot of effort and get one of these systems up and running. You know, like it's new and it's fun and you spend the effort and you get it right. But what ends up happening is that it's very easy not to have oversight and not see whether the agents are doing the right thing or not. So if you don't have, if your evaluation criteria is that, Hey, I tried out five things and it seems to work, done, let's move on. That is fine, you know, in the early days, but I can guarantee you at scale it's going to be a problem. Now, we didn't realize it early on. Our customers also fully didn't realize it, that they needed to do this as a program where you have observability. It's just like this whole quality.
It's almost like, you know, if you're a machine learning guy, then you always think about, you know, precision recall and like, you know, quality scores. Mean, these are things that come naturally to you. But if you don't come from that background, you don't think about this kind of measurement. I mean, you mostly have been doing very deterministic work. But most engineers do relatively deterministic work. But here, everybody has to think about, think of this probabilistic mindset.
And that requires a new kind of thinking. Our product didn't have it. Our customers didn't have it. We've sort of like learned the same way. The third one would be, it's almost like a, if you are building agents every six months, the best practices or best principles on how you create really high quality agents changes. So, you know, like our chat system, which we started building, I think right around the time, I think GPT 4.0 came along. That was what we call as like V0 of our chat system. Then we had a V1 of the system that was, that we built about a year and a half back, which we felt very modern when it kind of came along.
And then four months later, have launched the next generation of it, which we're now calling the V2 version. It's just like the technology is moving so fast. So in some ways, what you thought was truly the best principles of building something new completely changes. And that's yet another like new variable that you kind of need to contend with because everything that we have built so far, whether it's chat or apps or context mining, you know, you can just like continuously find ways of doing it way better. You know, every three to six months, you can just do it way better. So if you don't have your team oriented to, you know, think that, you know, every six months you're going to get like a quantum improvement, you know, you will essentially become obsolete very soon. That was never the case. Six months was not the cycle time of, like, obsolescence of software ever. But that that is certainly, like, something that I'm kind of realizing is is there, and our teams have to be oriented towards it.
Tobias Macey:And as you have been building the technology and the company, what are some of the most interesting or innovative or unexpected ways that you've seen the wisdom.ai platform applied?
Soham Mazumdar:I mean, we are pushing the envelope ourselves. I mean, our RevOps function, mean, we looked at like some of these standard solutions out there for like, how do you modernize your RevOps organization? And we just do it very differently using Wisdom AI internally. That's pretty cool. We have like people who use our product to figure out this is really weird. You know, you have some like red light glowing in some manufacturing console. And then what should I do when this measurement goes above this threshold? People use wisdom to decide what to do to that. That was not something that I ever thought anybody would do. We have, I think again, like in more boring terms, you know, like, like I would say twelve months back, maybe let's say twelve months back when somebody said, Hey, you know, like I have Power BI. My standard answer used to be, well, we can connect to Power BI as a, you know, context source or data source, but ultimately let's live side by side with it. I think today, I just said that I am going to transfer, you know, just take your Power BI and completely modernize it overnight. Twelve months back, nobody believed me or nobody thought that was, you know, even worth considering
as a possibility. But you would be shocked at the number of people who are truly ready to move on from whatever they have spent their long career since. So I think that's been one transformation I have seen in users, their readiness to embrace a big change, that's far more than it was earlier. I would say some of the actually some where there is one organisation where typically I expect like no more than 10% of any organization to care about data.
Ultimately, people have other things that they're doing in the organization. But we've actually been able to find some places where folks have opened up Wisdom AI to 100% of the organization. And that has been a bit like it was a bit shocking when they did that, but, you know, that's gone well. That was that was pretty surprising for me.
Tobias Macey:And in terms of your experience of building this system, particularly in such a fast moving area, what are some of the most interesting or unexpected or challenging lessons that you've learned in the process?
Soham Mazumdar:Okay. One of the lessons has been, so when you speak with people around, hey, are you ready for this change? A lot of folks, there is a lot of, I would say, existential angst that a lot of people have. And that is certainly one challenge that we see. Is a lot of desire to embrace where a lot of people get a little bit worried as well about their jobs and so on. I mean, my standard message to folks in the data space is that you probably are honestly the most important people in any organization now because now more than ever you can deliver true change.
Data teams tend to be pretty small, right? So in some ways, this is a moment of great empowerment for you where you can deliver true change as opposed to thinking from a more job loss, more negative point of view. But I think that is something that I've heard in a few places and that does slow things down. I think there is a second one is given the long years of history of doing things a certain way, A lot of people, there is also like, there's this art of the possible that you need to almost, you know, you are painting a picture of how the world could be and that looks very, very different. I mean, it can look, it'll look nothing like what it looks like today. Right? And for a lot of people, is also imagination that you need to give them. I would say every passing quarter, that problem becomes less of a thing because there is just progressively greater embrace. But I would say like twelve months back, eighteen months back, there was a little bit of like, yeah, I can do this, but like, how far can I really go? And I think people are now like, certainly realize that there is a true disruption possible. That's
another one. I would say again, that's less and lesser of an issue now what it used to be. My data is not ready. This is like a very common thing that I hear. It's like, well, what do you say? This all sounds well and good, but you know, can I it's not for me? Alright, I need like this two year long project. And only at the end of this two year long project can I even think about this thing that you're talking about? I mean, there is some truth to that. I think, you know, like our answer to that is in fact, like a lot of the tools that we have built is so that we can start with your Messier data early, right? So, and in fact, what you needed to do earlier, right? You know, when people are trained in this like, bronze, silver, gold layers of data, right? So you start with something raw and you progressively improve it to make it super clean. And then only on that super clean data do you start doing interesting analysis or the final consumable data takes all of this pipelining to get to. I mean, the reality now is you can sort of like, you know, you don't necessarily need to the goal layer,
need to get to a goal layer before you can start getting insights. And that is like a, you know, that is like a bit something that a lot of people are struggling to get their heads wrapped around. So these are like some observations. I would say on the build side, I think if you spoke to our CTO, he would give you like a long list of like, I mean, there are all sorts of fun things, you know, like actually this guy Mark Andreessen keeps talking about it. Product managers, engineers, and designers, they all think that they can do each other's job into a variety of ways that creates all sorts of fun things, challenges.
I think, again, I think I mentioned this earlier, right? I mean, how do you build a chat system? What do you expect the model to give you? That changes with every iteration. And that means that if you're not constantly on top of it, you will sort of, like, make suboptimal decisions.
Tobias Macey:Well, for people who are looking for a way to make all of their organizational data and details accessible to AI and agentic workloads? What are the cases where wisdom.ai is the wrong choice?
Soham Mazumdar:I think I think the fundamental question to ask is, are you trying to bring a tool just for yourself or are you trying to bring a tool for, like, I mean, is it for the organization at large or is it just for yourself? So we are actually not the tool. If you wanted to write your data pipelines faster, So essentially, that's a tool for yourself. And my recommendation is use the best code editor out there, you know, like use Codex, use Tracode, any of them, you're going to be, whichever is your flavor, that's what you should use. So again, so there needs to be a, I need to empower a group that is much bigger than ourselves, right? So there needs to be a, and there are ways, there are symptoms of it, right? And that the other group is like, you know, a few 100 people, let's say. If that is not there, then we are not the solution. So because we are, again, know Tobias, we are in this data engineering podcast, but I'm actually not building a tool that the data engineers consume amongst themselves. It is meant to go out and help the rest. So that's probably the cleanest framing I can give.
Tobias Macey:As you continue to dig into and explore and expand on this overall challenge of context management, context propagation, what are some of the near term to medium term areas of focus that you have or any areas that you're excited to dig into?
Soham Mazumdar:I think for us, there are like a few that I would say if I were to pick two North Star metrics, it is how quickly can we learn enterprise context. I want to move it down to zero day. We are not at zero day right now, but that's probably like the single biggest area of investment for us. And the second one actually comes down to economics. I think there is a way of doing it where you would essentially go bankrupt trying to build your context and empower the organization, and that simply doesn't make sense. And that's the second north star, which is like, how do we do this in a manner that's truly economically sustainable? So I would say these are the two vectors along which we do the most investment in. And the rest of it, again, like there are features and so on that will keep happening. I mean, ultimately, what the agentic enterprise looks like, you know, just from a look and feel and use perspective, I think that it looks, I mean, it will look very different from what it is today.
So, you know, remanaging that, I think that's hard for me to quantify it in metrics. But I think that's where a lot of the, you know, for me, the excitement is that, you know, there is going to be a new way of doing work. But again, coming back to the context itself, can I learn your context in me automatically? I think that's probably, I'll let broil it down to this one. Can I learn your context? If I can, I think we'll be able to make you successful super, super fast? Again, that's a North Star.
There
Tobias Macey:is work to do. Are there any other aspects of this overall problem space of context engineering and context delivery and the work that you're doing at wisdom.ai to make that a more tractable problem that we didn't discuss yet that you'd like to cover before we close out the show?
Soham Mazumdar:I think we have more or less touched on it. But if I were to again summarize, learn from existing artifacts, learn from humans, learn from agent runtime. I mean, those are like all the elements that need to come together. And then you're essentially always, always meeting benchmarks and hill climbing. You're not doing context for the sake of doing context. You must define a goalpost and work towards it. I mean, I think summarizes our thought process on it.
Tobias Macey:All right, well, for anybody who wants to get in touch with you and follow along with the work that you and your team are doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gaps in the tooling technology or human training that's available for data and AI systems today.
Soham Mazumdar:It's actually measurement, honestly. That's a people and tooling combined problem, but people don't think about measuring context quality. And then that's a huge problem. If we solve it through tools and people,
Tobias Macey:think we'll all be in a much better shape. All right. Well, thank you very much for taking the time today to join me and share the work that you're doing at wisdom.ai and your thoughts and experiences around this complex and evolving space of context engineering and what context even means in a given use case. So I appreciate all the time and effort that you're putting into that, and I hope you enjoy the rest of your day.
Soham Mazumdar:Awesome. Thank you so much.
Tobias Macey:Thank you for listening, and don't forget to check out our other shows. Podcast.net covers the Python language, its community, and the innovative ways it is being used. And the AI Engineering podcast is your guide to the fast moving world of building AI systems. Visit the site to subscribe to the show, sign up for the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.
Transcript supplied by the publisher with the episode.
by Tobias Macey · English · Tech & Science
This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.
E517 · 24 Sep 2026 · 50 min
Summary In this episode Christopher Doidge talks about his Agile Ledger Architecture (ALA) approach to data warehousing and how it aims to reduce data debt while shortening the path from raw data to trustworthy business insight. Christopher explained that ALA is not a replacement for existing warehouse patterns like medallion architecture, star schemas, or other modeling approaches, but a complementary discipline focused on pushing business definitions upstream, enforcing cleaner ledger-style transformations, and producing gold-layer tables that stakeholders can actually use without relying…
E515 · 27 Aug 2026 · 46 min
Summary In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers. She explored why generic coding assistants often fall short in data workflows, how Otto adds the missing context around Airflow, Astro, upgrades, and troubleshooting, and why Astronomer focused first on high-leverage use cases such as DAG authoring, investigation of pipeline failures, version migrations, and legacy scheduler modernization. She also discussed the practical realities of introducing agents into…
E514 · 2 Aug 2026 · 1 hr 2 min
Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for traversal workloads, and how that perspective shaped OmniGraph’s design on top of object storage, Lance, Arrow, and DataFusion. Ragnor explained the motivation for combining graph semantics with Git-style branching and merging so that teams can manage probabilistic writers such as AI agents with stronger governance, shared…
E513 · 6 Jul 2026 · 1 hr 1 min
Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance depends on contextual intelligence: institutional knowledge, semantic meaning, procedural know-how, and access to the right tools. She also dug into how metadata catalogs are evolving into broader context layers that serve both humans and agents, and why agentic systems are changing the economics of metadata and…
E512 · 18 Jun 2026 · 50 min
Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync at scale highlighted both the strengths and common misuses of Kafka. He digs into using events as the source of truth, materialized views with KTables, and how schema registries and type safety prevent downstream breakage. Jevin explains why teams often reach for heavyweight Kafka clusters without leveraging Streams,…
E511 · 8 Jun 2026 · 53 min
Summary In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams. He explores Kaarvi’s multi-agent architecture that runs queries across seven LLMs in parallel for reliability, its synthetic data generator that mirrors source schemas for quick testing, and “Hey Kaarvi” chat for text-to-SQL, text-to-transformations, and text-to-dashboard workflows. He also digs into on-prem versus SaaS deployments, domain-specialized agents for privacy and accuracy, code…
E510 · 1 Jun 2026 · 54 min
Summary In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources. He explores how PuppyGraph lets you run Cypher and Gremlin traversals and graph algorithms directly on data in Iceberg, Delta, Hudi, Hive, and even MongoDB—without loading into a separate graph store. Weimo explains their edge-sharded, vectorized, MPP architecture that tackles hub nodes, multi-hop traversals, and shuffle at scale, targeting sub-second to single-digit-second workloads. He digs into practical graph data…
E509 · 6 May 2026 · 59 min
Summary In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads. He explores Ray’s evolution alongside Kubernetes and PyTorch, and why consolidation at these layers has enabled a new generation of complex, heterogeneous workloads. Robert explains how data preparation has shifted to GPU- and inference-heavy, multimodal pipelines; where Ray fits compared to Spark and workflow orchestrators; and why Ray excels at composing heterogeneous pools of compute, handling failures, and scaling complex…
E508 · 7 Apr 2026 · 59 min
Summary In this episode, I sit down with Gleb Mezhanskiy, CEO and co-founder of Datafold, to explore how agentic AI is reshaping data engineering. We unpack the leap from chat-assisted coding to truly agentic workflows where AI not only writes SQL and dbt models but also executes queries, debugs, runs tests, and ships production-ready outcomes. Gleb explains why teams that master this AI-first loop can see 10–50x gains, how security/compliance concerns can be addressed with platform-native LLM endpoints, and why the role of data engineers is shifting from code authors to operators of…
E507 · 29 Mar 2026 · 50 min
Summary In this episode Himant Goyal, Senior Product Manager at Salesforce, talks about how data platform investments enable reliable, accurate metering for consumption-based business models. Himant explains why consumption turns operations into a real-time optimization problem spanning metering, cost attribution, billing, governance, and cross-functional ownership. He explores the richness required in usage data to support sophisticated pricing, the importance of treating metering like a financial system, and the architectural foundations - event schemas, durable ingestion,…