Skip to content
Melo Podcasts Home
CategoriesLanguagesFollowing

Episode notes

Summary In this episode Christopher Doidge talks about his Agile Ledger Architecture (ALA) approach to data warehousing and how it aims to reduce data debt while shortening the path from raw data to trustworthy business insight. Christopher explained that ALA is not a replacement for existing warehouse patterns like medallion architecture, star schemas, or other modeling approaches, but a complementary discipline focused on pushing business definitions upstream, enforcing cleaner ledger-style transformations, and producing gold-layer tables that stakeholders can actually use without relying…

Chapters

Tap a chapter to play from there.

Transcript

Read the transcript · about 8,650 words, follows along as you listen

Tobias Macey:Hello, and welcome to the Data Engineering Podcast, the show about modern data management. Today's episode is sponsored by Parallel, where agents find answers. Most engineers today closely follow new model releases, but don't pay attention to their agent's most important tool, web search. Parallel develops enterprise grade infrastructure for agents to retrieve high quality context from the web. Their core products are a suite of APIs for retrieving high quality information from the web with Pareto optimal quality, cost, and speed.

Whether you work on voice agents that need two hundred millisecond latency, chatbots that balance speed, depth, and quality, or long horizon agents to do thorough overnight research for you, Parallel is a single platform for all of your agentic research. Get started for free today at dataengineeringpodcast.com/parallel. Your host is Tobias Maci. And today, I'm interviewing Christopher Doidge about a data warehousing approach called agile ledger architecture that is designed to drive down data debt. So Christopher, can you start by introducing yourself?

Christopher Doidge:Absolutely. Thank you so much, Tobias, and so happy to be here. My name is Christopher Doidge, founder and principal data architect at UNI Consulting LLC and author of the book that we're gonna talk about today. So where are the insights? It's the blueprint for agile ledger architecture. And, core focus here is bridging the gap between technical data engineering and executive decision making. Essentially, we're looking to eliminate data debt and reduce overall executive time to insight.

Tobias Macey:And do you remember how you first got started working in the data space and what kept you there?

Christopher Doidge:Yeah. This goes back fifteen years ago. I think starting off as a intern to be a sales rep at a medical device company. Right? And I always thought I wanted to be an entrepreneur, and that's really where it started. But I found very quickly that that was not my true forte, And it was very much more in organization and problem solving, which really led me into the world of data itself. It grew from there. I mean, my experience ranges from working in strawberry plant distribution with Driscolls here in California, all the way over to the leading classifieds in Europe with, eBay Klein and Seigen, which is headquartered in Berlin.

So a lot of it went from being in the actual trenches of analytics and just witnessing the frustration of executives argue what is the true count of our new customers for an hour and a half into then deciding, you know, it's not a tooling problem that we're having here. It's really a question of architecture, And that's really what led me to agile ledger architecture and how to build this.

Tobias Macey:Digging now into the agile ledger architecture, I'm wondering if you can just give a little bit more color around what it is and some of the story about how you came to that design.

Christopher Doidge:Yeah. Absolutely. It's a really great question. So the easiest way to explain this is we're missing a layer when it comes to our data architecture. So, generally, when engineering and analytics works together to build a platform, we're stopping at the analyst as our end deliverable. So what I've witnessed in my career is we have these wonderful data tables. They're fantastic. But in order to get it consumable to reach the actual insight that we need, we need another step beyond that. And so what agile ledger architecture does is it enforces strict ledger accounting in the intake.

That way, we're essentially pushing business definitions all the way to the very front. And our ultimate goal is the creation of a gold layer table. So imagine this is a ready to consume table that anybody at your enterprise can connect to and pull the answers exactly how they need it. What I've typically seen is that when we think of a source of truth, we imagine one giant Excel table that just has every possible thing that you can think of. And so my whole thought around this is let's deconstruct that. The true source of truth is actually understanding where does this data point come from, what is the actual source around it, and when is it refreshed, is generally what I need to know.

Using those definitions and then a series of taxonomies to link our data together, we can essentially create end state tables that allow any stakeholder in our business to get to the same answer repeatedly over and over again. And this is what the heart of Agile Ledger is. It's serving up our data not just so that your analysts have a nice clean playground to play in, but it's also so that you have a consumable layer so that your stakeholders can get to the answers they need without having to write a nice big complex SQL query.

Tobias Macey:One of the, I guess, temptations whenever you're talking about data warehousing is to say, okay. Well, I'm running into troubles with whatever approach. I'm going to develop a new one. Obviously, there is forty years plus or minus of data warehouse, best practices and principles and designs and architectures, and everybody seems to have developed their own opinion. Some of them are an addendum to preexisting approaches such as the Kimball star style or the Inman third normal form, or there's DataVault, which is sort of going in its own direction. I've seen anchor modeling. There's the whole medallion architecture, ETL versus ELT. So there's a lot to try and figure out. And I'm wondering if you can give your summary of how you see the Agile Ledger architecture fitting in that overall ecosystem ecosystem of everybody having their own opinion and there being I don't even know how many different architectures available to choose from.

Christopher Doidge:Yeah. Absolutely. And you're a 100% right. I mean, I'm remembering back to the days when we just had an Excel file, and it was this is the column is our source of truth, so to say. And now there's these wonderful tools where we can pull data in from different disconnected systems and take them through a series of training, so to say, in order to get them ready to consume and analyze on. So, essentially, ALA doesn't try to replace any of these. It's actually built in coexistence with it. Right? It's just the way to amplify it.

Essentially, what we're talking about at the core is that you do not release a gold layer table that's ready to consume for the business without having clear definitions of how this is being put together. So this is where I said earlier where we sort of stop at the analyst when it comes to delivering up data. We think, okay. Let's get the data as clean as we can, and we'll place it here. But then the analyst still has to take it through particular steps in order to get it sort of business ready.

And I've even seen this working at, as I mentioned, in eBay in Germany. It's we had tables where I still needed to run deduplication. I still needed to filter out specific statuses, and this is never gonna end. You know, business logic constantly evolves over time. And so it's impossible to pour sort of like a concrete layer of data and assume that it's always gonna solve every need that you have in the future. So this is where ALA comes into play because what you can do is lock in your source exactly how you have it. And then as the future changes and the business starts asking different questions of the data, you can evolve your gold layer tables to answer those questions.

And this is where I introduced the idea or the notion of a clay model. So you're working with clay. That structure isn't always necessarily solidified until your absolute end. But the idea around the clay model is that you can build it up to whatever sort of architecture that you need for the current sprint or for the current need and then tear it down and rebuild it as you need it. So what ALA does is it really focuses hyperfocuses actually on specific data debt that is needed for the business right now. That way I can take to attach some context to this. I could take my orders out of, like, a system like Shopify. I can break it down to particular products. I can go down to counties that I'm shipping it to, go down to whatever level of granularity that we need to. But we build those taxonomies and those bridges as our business needs evolve instead of just trying to tackle everything at the same time, which is what I've commonly seen.

So it's a way to organize this and prioritize it with the business. That way you see what's actually working and you enable analytics to be your true compass. You know, leveraging your data to tell you you're going in the wrong direction. Let's retry and test. So it really complements those systems well because ALA works especially well with the medallion architecture as well too. So it's not about tearing any of those down. It's just about attaching a philosophy or an ideology into the way that you architect your end state data.

Tobias Macey:And one of the interesting things is that medallion architecture in particular isn't actually its own style of data warehousing. It's just a way of sequencing the transformations.

Christopher Doidge:Yeah. Absolutely. Yeah. Great point.

Tobias Macey:And so digging now into the overall question of data debt where sometimes you want to actually take on some of that debt because it's necessary to get to some learning or discovery to actually help cement what is the thing that you're actually trying to do with it. Sometimes it's just something that happens incidentally, and I'm curious if you can talk to some of the ways that that debt manifests and some of the decision structure about whether and how to be intentional about it and some of the ways that it happens accidentally.

Christopher Doidge:Yeah. No. That's excellent point. So I I think back to my experience with Driscolls. And so, Driscolls were dealing with plant operations all over the world. And large companies like this have a lot of partners that they work with along the entire supply chain. So you could have the most amazing data architecture in the world, but your partners aren't going to. And a lot of times you cannot control this. And so this is a great example of data debt. If I receive an input from a partner telling me what they have in their three p l or whatever storage concern they're giving me or data they're giving me, I can't necessarily always receive it in a format that works for me and then have to convert it over.

So this is a classic form of data that you're gonna have your tech savvy analyst jump in there and start, you know, adding helper formulas to convert this data to something that fits into your system. That way you can correct the intake. You know, not everybody has an API. Not everyone can serve up the information the way that you need it. And so this just snowballs over time. And, ideally, what really helps you to drive the decision is is this going to impact the business and move us forward in a way, or is this just some easy, simple, busy work that we can continue to do because it's not really worth the struggle to actually cover this or not? So I would say the business needs and what's really gonna drive your needle should really dictate which data that you actually cover or not because all of them are required different levels of resources.

Tobias Macey:One of the other interesting aspects of debt, whether it's technical debt, data debt, financial debt, is some people unfortunately figure out is that it's not always, I guess, financial debt's pretty straightforward. But technical debt and data debt, it's not always obvious that you even have that until you get to a point where you're actually trying to achieve some outcome and there are things standing in your way. The way that I have usually seen technical debt kind of identified is that it is the thing that adds friction to what you're trying to do. And I'm wondering if you could just talk to some of the ways that people can even identify whether and how much debt they have and then more specifically at a pinpoint where it's coming from.

Christopher Doidge:Yeah. Absolutely. So for me, I found that the easiest way to pinpoint this is if I look at two different datasets and I know that I need to combine these to get to the next level of my analysis, my first initial question is how do I join this? Right? And to me that right there is your data debt. If you cannot figure that out within just a couple seconds of looking at it because there's no columns that seem to be the same or there's no primary key, surrogate key, there's your data debt right there. So it's looking you right in the face. And so with this, it comes with, alright, are we able to do a simple proration and move it forward?

Is it something that's gonna take a lot of resources? Cause now I need to build a taxonomy to then say, okay. When you see this record in system a, convert it to this record in system b as a high level example. Right? There's some way that we gotta build up that join, and that's exactly what Datadet is here. So the way that I like to do it is I'll identify what it is that needs to be done for this analysis and, what are the pieces that we'll need to put resources to in order to tackle this. How much time do we estimate this will take?

That will also help us understand whether we should even tackle this data debt or leave it on our never ending backlog because it's not that big of a deal. What I've typically found, what happens is that you tend to have these hero analysts in the team that are fantastic, but they tend to just do this in the background with scripts on their local machines or some business context that they've heard in a meeting at some point. And the idea around Datadet is let's get this documented and actually bring life to it. A lot of times, they're just living in somebody's head, and then no one's even aware of the debt, just as you mentioned.

But at some point, we all have to pay it forward. So we need to find a way to, you know, bring this to light, bring visibility to it. And I found that the best way to do that is with the construction of these gold tables. Because as I try to make data more and more consumable to the nontechnical user, that to me quickly identifies which pieces I just cannot join without something simple.

Tobias Macey:And so digging now into the architecture itself, I'm wondering if you can give a bit of an overview about what are some of the core principles that you are using as the foundation of the architecture and what are some of the ways that people should be identifying whether and how it might fit into their overall data modeling practice and data warehouse design.

Christopher Doidge:Yeah. Absolutely. So for me, it really boils down to a very simple test. I just call it the 15 litmus test. The idea around this is your total time to insight. So if you imagine a stakeholder coming in and asking a very simple question, how many customers did we have yesterday? As simple as that. If this takes you more than fifteen minutes to answer, then you are in need of agile ledger architecture. If it's less than that, you're fine. Just move on. Why why improve something that's already working quite well?

But the idea is that whole test is if it's taking longer than fifteen minutes to answer basic business questions and each time you wanna understand, okay, do I have the data to back the answer for this? And that takes longer than fifteen minutes. Then that's when you know that you need some work in your architecture. We're built off of some particular principles as you mentioned. So obviously, the core principles are that fifteen minute litmus test.

Medallion pipeline works very, very well for this. It's built off of the foundations. So, to reiterate, that's using bronze as your loading dock to just ingest the basic data. I know the listeners probably already know this. Silver for cleaned ledgers and then ultimately gold, which would be our fluid clay model. Another core principle is the upstream integrity. So this is fixing data anomalies right at the entry instead of waiting towards downstream or trying to correct them in complex queries that live outside of our medallion pipeline.

For me, the easiest way to identify this is another core principle, the unsung hero, which is my favorite report since the first days of analytics. It's the data error report. It's the one thing no one ever wants to see. It's the continuous scan that looks through our structures and looks for all the gaps. It's essentially in layman's terms, select star where everything is wrong and shouldn't be that way. And then I could go through and scrub the list and actually correct this. That way our goal layer remains intact and accurate.

Tobias Macey:Given that the medallion approach is the presumed basis, what are some of the ways that that informs or influences the technological substrates that you are supposing for this architectural approach? So are you presuming that people have any sort of data streaming? Are you presuming that they're using a traditional data warehouse appliance or a lake house architecture and just some of the ways that the underlying technology has an impact on the way that you actually adopt and implement the principles in in the architecture that you're proposing?

Christopher Doidge:Yeah. Absolutely. So, actually, I developed this architecture with the thoughts when I was working in the trenches with very sophisticated tooling. Right? So this is where it actually came to life. However, I didn't even realize it until I was supporting a client who strictly worked off of Google Sheets. So we're talking the lowest code that you can go. There was no API here. They were downloading data from a Shopify store, exporting into a CSV, opening it up in Excel, and then starting their full on analysis.

And that's where it really dawned on me because what happened was every single time they wanted to repeat this analysis, they had to redo that entire process from beginning to end every single time. And so that's when it clicked where I thought, well, why are you doing it that way? We can take the principles that we learn from more sophisticated tools that allow something like a medallion pipeline innate with its UI and actually push it all the way back to the very basics of an Excel or a Google Sheet.

So you can use this similar philosophy to your architecture regardless of what tooling and software that you actually have. The whole idea is around, like, loading the actual data, transforming it, and then presenting it in a way that can be consumed by everyone at the business. And that's down to its basic, basic core. You can do that with absolutely any system that you have. Of course, it gets easier when you have better tools because you don't have to manually export CSVs from files. You could just let an API do it for you. So it works better when you're have more sophisticated toolings.

However, it still functions fine when you're talking a low code environment.

Tobias Macey:So for people who already have a substantial amount of investment in a given data warehouse architecture, obviously, there's a lot of inertia built in to how they're doing things, way that they have their system designed. It's not easy to immediately pivot to a new way of doing things. And I'm wondering if you can talk through what the adoption process looks like for people who do have even just a medium scale data warehouse, and they've already built out a lot of pipelines around those assumptions and just some of the upfront work that's necessary to start planning out that adoption and how you can do that incrementally.

Christopher Doidge:Yeah. Absolutely. So what I like to do when faced with this scenario is to take the quarterly sprint approach. So, obviously, when you're talking a new architecture, a new system, I think every company that I've worked at always has a new system incoming. It seems to be the buzzword that corporate likes to use all the time is, oh, there's a new system coming that'll solve all our problems, and then it never comes. So what I found is this type of architecture can be built alongside what you already have.

So you're not needing to burn the entire warehouse down and waste everything that you've already built. What you already have is fine. What I recommend that you do is you look at your upcoming next quarter and you prioritize what is the most important questions that I know stakeholders are going to have. What is sort of a KPI shield that I can erect that can let them pull the data that they need at a self-service volume, answer the questions that they have, and then my team can actually focus on the things that are gonna actually help us instead of trying to give you the kind of customers to fill for your Google slide deck. Alright?

So it's bringing to light the data that we have now and then logging any data debt for anything that we cannot do. So as a crystal example or crystal clear example to give here, if the quarterly objective is to improve picking velocity by 15% in the warehouse, let's say. And we understand, okay. One of the metrics that they're gonna wanna constantly review is what is our overall packing time on orders. So for every order, how many minutes is it taking as an example?

Perhaps I don't have the data to do that yet. So my data debt is I know where the items are located in the warehouse, but for some reason, I cannot tie the items to the actual order itself yet. Right? Unfortunately, that would be a huge problem, but let's pretend that's our problem. If that's the case, that is a data debt ticket. That's exactly what I would log, and that would be one of the first things that I would tackle in order to give the business the data that they need for this upcoming sprint.

So this is the way that I would prioritize it because, again, the beauty of it being a clay model is that at the end of our quarter or even throughout it, throughout the various feedback loops that you have, where stakeholders are gonna be very clear about what's not working with the data model that you presented to them. You can iterate on it. So you can see what is it that's missing, what is it that we need to add to it, and this can obviously illuminate further data that so that you can continue down this.

It's not a rabbit hole. It's the exact opposite of a rabbit hole. This value driver, I guess, is what I could call it. You can keep going down this path of obtaining the critical data that business is going to need to make their best decisions. And a lot of times in business that doesn't happen until you have everybody looking at the same thing and speaking the same language. So this helps align everybody completely across the board. I can see too many times where I've been in these meetings and everybody is presenting the exact same metric with totally different numbers. And it's because they all obtained it a completely different way.

So this aligns the business to the same metrics and definitions.

Tobias Macey:The other challenge for any sort of correctness campaign in an engineering organization is that there's always the issue of backsliding or blocking people getting real work done. So one of the analogies that comes to mind is something like the introduction of a type checker for something like a Python code base or the introduction of a new linting rule that has 10,000 violations at the start, but that you wanna start driving down. And so there's the overall approach of, okay. Well, I'm going to accept this baseline, but I'm going to ratchet it so that every time it gets better, I'd never allow it to regress and nobody can introduce a new violation. And so I'm curious what are some of those technical controls that are available as you are adopting this Agile Ledger architecture for being able to say, okay. This is how we want it to be at the end state. This is where we are right now, and this is how we're going to make sure that we don't have any of that regression as we make these changes.

Christopher Doidge:Yeah. Exactly. With the ALA architecture, for me, the easiest way around this is to have a strict gatekeeper on the data dictionary. So it's something we haven't spoken too much about. But with ALA, one of the deliverables is every time you present a gold layer table, you must have an update in your data dictionary. So historically, this is also known as a business lexicon. It's essentially a warehouse of knowledge for your stakeholders to be able to go and understand how are we defining this metric, what is the business definition, and what is the data definition on it. Ultimately, they wanna know that if I think of a metric as new customers, that we're all speaking the exact same thing. What does new customers actually mean to the business and having this crystal clear?

So for me, enforcing a strict rule where no gold table hits production unless its sources, transformations, and business definitions are a 100% documented helps us to stay on track with this because it will keep you from backstepping and taking shortcuts to let's just throw it into the gold table now and then figure it out later because it's gonna come back to bite you very quickly. The whole point with ALA is that since it has a logical setup and you're cleanly labeling definitions all the way through from inception, you're bringing a lot of trust to your data, which is the most important thing.

It's easy to have data and to perform an analysis. The hard thing is having trust on it, especially as you scale up an enterprise. When you're moving to larger companies, such as eBay or Driscoll's, and you're dealing with billions and billions of rows of record, there's a lot of noise. And so to really strip it away and get to the true insight, it's not about scaling and performing massive rules on absolutely everything. It's focusing on what's truly important to the business now so that we can learn our best practices along the way. And then those are what we enforce as we go through. But to me, it's about prioritization is what I would say. So stick to one path and that can easily be your quarterly objectives and then build it out from there. But don't try to sprint to everything at the same time. The other challenge,

Tobias Macey:particularly in warehouse architecture, is that every time you introduce a new principle, you have to educate everyone about how it's supposed to work, how to do it. And, I mean, the star schema approach has been around for, I think, more than thirty years at this point, and it's still shocking how many people who work in the data space who have never actually come to grips with its formal definitions or even the a casual approach of how to do it correctly. And so for people who do have an existing warehouse or even people who are saying, I'm gonna go and build a warehouse right now, what are some of the ways that they need to be thinking about sharing some of the core tenets and principles and then also layering in some of those technical constraints about how to do this properly so that you don't have one person having one interpretation and doing their own thing over on this set of tables and somebody else thinks about it differently and has a different way of managing that in a different area of the warehouse and just some of the ways of building cohesion and shared understanding?

Christopher Doidge:Yeah. Absolutely. No. That's a fantastic question. And I think that's where the simplicity of the medallion architecture comes into play because you enforce strict rules at each layer. So for example, bronze, there's no rules. It's all about speed. All I'm trying to do is get the source original source data loaded into my warehouse as fast as humanly and machinely possible. That's my ultimate goal. No rules. When it comes to silver, this is where I'm applying all of our business context.

So for example, every order ID is always duplicated, so make sure you deduplicate it. Make sure you run a window sum across this. You know? Any other rules that you can think of on it. Strip out the dollar sign that seems to be coming into the revenue field, convert it into decimal, whatever the case may be that you wanna apply to that. But those rules are documented and loaded into Silver itself. So as you're constructing that ledger and you're making a very clean, immutable ledger with the strict accounting philosophy, you're setting a core list of rules.

And so it cannot change from those. These are how it is because the if you make a change into here, you're gonna obviously break everything downstream that happens after that as well too. So that is how I found strict enforcement onto it and also by isolating them. So silver should be specific to the source that it comes into. And then as you mentioned with the it it actually fills straight into a star schema because if you think of the visual layout of it, each of those independent tables are your clean silver ledgers is what you're doing here. So it's taking the exact same principles that exist. We're just saying on top of the architecture that you're doing, you're focusing on the wrong things to start off with. So continue what you're doing. Just your focus should be aligned closer to what the business needs to move forward and bring that information visible so that everybody's on the same page.

There are so many times that I've seen us create data tables and construct a full star schema for then the business rule to completely change. And then when that happens, your immediate instinct is to go back and rebuild the ledgers, But that doesn't necessarily have to be the case. You can simply adjust the logic that goes into one of the ledgers to now exclude order IDs that come from one channel as an example. And that is how you can update and fix these without having to break anything downstream.

Tobias Macey:With the idea of ledgers, it brings up a number of different associations. So there's the ledger for an accounting ledger where you've got double entry bookkeeping. There's also the ledger where in the sort of crypto coin area of having the blockchain of an ordered sequence of auditable events. There's the ledger of sort of a decision ledger of this is why we did it this way, and I'm wondering if you can just talk through some of the ways that you want to think about the durability of that ledger. You said you don't necessarily wanna go and wipe it and rebuild it. You just wanna modify it and just some of the technical aspects of that ledger and making sure that it doesn't have too much of a oh, we'll just rebuild it whenever we want to sort of from an event stream principle, but the ledger itself actually being a durable system of record.

Christopher Doidge:Yeah. And it it stems from the movement from bronze to silver. Exactly. Because if you think about it, when your bronze stumps in and it creates a specific now set of data that's available for you to now transform into your Silver Ledger. If you think about it as a hard coded list is the way I try to think of a ledger is this is the truth. This is what it is that we want. And so if you're adding new filters, new business context that comes down into play, it's about recording that logic into here so that that can evolve over time because that will happen.

Businesses will make critical decisions and decide to change the way that we count customers, change the way that we actually want to classify our orders or categorize them. And those things can happen. But what I've commonly found is that this context doesn't live in our systems and it's not documented. What tends to happen is that it lives in private scripts. And so the core philosophy and ideology of ALA is pushing this documentation through to the inception. It's so making sure that we do not lose it over time because there's been too many times that you have that one data person leave your company, and then there goes the entire business context and logic of how this is constructed and going on.

The idea around this is constructing a system that's going to outlive you and is going to be there. Right? And in order for that to happen, you have to have the rules listed into there. It helps when you're in those conversations with the stakeholders because then it's not so much of, I don't know the source and I'm not quite sure the rules that we're applying to it to get you this list of customers. It's clearly documented. We can go back and review it together and see it.

So for me, it's really about the documentation that enables the governance for this.

Tobias Macey:One of the other aspects of the time in which we are right now is that you can't really get through any conversation about anything without AI coming up. And in particular, in the data engineering space, Text to SQL is one of the earlier examples of, oh, this can actually be really useful to help me answer questions. But as we invest more and more into agentic approaches, being able to make sure that the warehouse is navigable and scrutable by an LLM or an agent is increasingly valuable.

And I'm wondering what are the benefits that that ledger and governance as part of the warehouse adds to some of those workloads where you don't necessarily have a human with all of that organizational knowledge ready at hand. And instead, you have an agent that needs to build that context fresh every time, you know, where maybe that context lives as a durable system of record somewhere, but it's still every time the LLM launches, it's coming up fresh. It needs to make sure that it's pulling in all the right detail.

Christopher Doidge:Yeah. Absolutely. And it's a wonderful point. I mean, we can think back to the age old analogy of garbage in garbage out. And one thing that we've seen with AI is that it's phenomenal at accelerating very bad information when it can. So in order to prevent that and ensure that you're actually, you're having the AI build off of useful business context exactly to what you're trying to accomplish here is through this system. What I have seen so many times is users will take a massive dataset, hand it to an AI agent and say, go find me the insights.

And that can be incredibly useless because it's gonna dive through and perform multiple regression analysis. It's gonna try to find different links that it can. But the point is you've been in business for x years. You've already done that for a period of time. So why would you have the AI start from scratch every single time? It translates directly into the context windows because there's nothing more frustrating with AI systems than the retention windows of its context.

So you'll talk through and spend hours getting down to the core of an analysis with an AI system. You'll upload the dataset to it, you'll figure something out, and you'll get a wonderful moment. Oh, wow. All of the marketing campaigns that we've been doing have been completely wrong. We've been leaving out this little URL or this little ID in the URL. So when it comes into our warehouse, we have absolutely no way to tie it to the true orders. And you finally discovered this through digging through with the AI and figuring that all out.

Why would you want your next conversation to completely start from the beginning and now go into a different route when you've already uncovered from that? You wanna start with that as the context base and for it to load that up and then go. So what you're doing is you're enabling these tools to work off a very clean, refined data. And so now instead of co connecting AI directly to the bronze layer, which is what most people are doing, you're now letting it play with clean data that has already gone through this. To put it also into simple terms is you can make it cheaper for you. It costs a lot of tokens to load a gigantic dataset and have it find insights for you.

Instead, you could give it a much more refined list that is quite a bit smaller, already has gone through a rigor of everything that you know is just flat out wrong, eliminated that, and then it can help you to take it to the next level for that. It's an exciting time with AI where I know businesses wanna build a lot of tools. I've helped companies with building AI chatbots. You know, how do we guide our stakeholders to the right datasets into what lives out there? Well, it doesn't just know the AI agent didn't just wake up or get programmed, and now it knows where everything resides. You need to teach it and train it over time just like a standard human being.

So the purpose of this is by building it throughout these levels with your data and architecting in a certain way, you can give it clean context. You can give it the business rules. By forcing yourself to have clean documentation at your gold layer and understanding this is how we define this metric. This is how we count it. It's not gonna send you down a random rabbit hole. It won't give you an insight that you think is gonna be really exciting to share with your VP just to realize that it overinflated leads because it's not using the definition that your business uses.

Tobias Macey:And so as people are going through the process of adopting this architecture, they're either designing a new warehouse from scratch or migrating an existing warehouse. What are some of the most interesting or innovative or unexpected ways that you've seen people use some of the principles from this design in actually facilitating their overall objective of the organizational warehouse?

Christopher Doidge:Yeah. Absolutely. So I've seen this apply in something as simple as smaller consulting operations. You know, that would to me was really interesting. It went from somebody having to here here's the business scenario. They had about fifteen, twenty contractors on their team, all with independent time sheets, and and they needed to do a payroll at the end of the week and pay everybody. So you had a massive problem of data ingestion, data unification, cleansing, and then being able to take an operational step on it from there, which is typically what we're doing with our data anyway.

So using an ALA approach, they were able to take something that repeatedly was taking them hours every single week into simply running in a matter of minutes. By taking this approach and breaking it out, they could load independent bronze layers for each of the independent time sheets, unify it based on their master data that they wanted to view it as in their business, and then be able to spit out an output. I've seen this work incredibly well in very fascinating ways because it's able to take the data that exists, apply a logical base to what your business needs are, and then help you get to that deliverable much quicker than you can imagine.

The great thing is it'll help you to even use more sophisticated tooling, such as Python scripts or, other reporting software even. They can help you elevate this and take it to the next level because you're feeding it already a cleaned object. So by taking it through the rigor of ALA where it surpasses three different levels, and then it's coming out to a consumable layer that is easily ingestible by any stakeholder in the business, you can then operate on that and get to your completion.

So I've seen it used in multiple different scenarios. Mentioned the payroll. I've also seen this applied to operations where this was relating back to a hyper growth company in Berlin. They were using this to actually plan operations for their solar panel installation. So they were able to unify their data across very many different partners and systems that they had to then plan out the operation of which locations were they gonna go to in the upcoming week and, actually install solar panels based on team availability, proximity, and scheduling.

So a lot of these you can think are very manual processes that take a long time. But by tackling very specific prioritized data debt, for example, where are these teams located? Where are our customers located? What does their schedule look like? If you can bring to light this data debt, clean it up and ingest it properly, then you can enable your business to fly very quickly with the data that you actually have.

Tobias Macey:And in your own experience of working in this space and developing this overall design system, what are some of the most interesting or challenging lessons that you've learned to the process of data warehouse design and implementation?

Christopher Doidge:Yeah. So, first one is if the time to insight is already fifteen minutes, don't try to improve it beyond that. That was a hard lesson that I learned. It's if you already have a system that's working great for you, then leave it as is. The whole philosophy here is to help you to get to the true insights that you need quickly. And to easily define insight here, it's you need to also trust that data. So it's not just the fact of I can give you the answer very quickly. If you don't believe it, it's it's not gonna be an insight. It's like that simple joke that I've heard where someone says, I'm very fast at math. And they say, okay, what's 14 plus 35? And they just say a 120.

It's like, that wasn't right. Yeah. But it was fast. So that's a great example of what an insight is not. And so the whole point here is to deliver delivering credible data. So it has to also be understood at the same time. But, yeah, definitely, I would say if time to insight is already under the fifteen minutes, it's just not needed in this sort of way. Other particular challenges that I've seen along the way too is it's easy to get excited when there's a new system and wanting to tackle multiple data debt at the same time.

So my recommendation is always to stay on track, and this can easily be prioritized based on what your objectives are for the next upcoming period.

Tobias Macey:And so you touched on this a little bit, but what are the cases where the agile ledger architecture is the wrong choice?

Christopher Doidge:Yeah. Absolutely. So if you're in the early stages of still trying to figure out what you're trying to do as a business, this is not for you. If you have a simple system where everything is done in just Shopify, you're strictly running campaigns, and you can get to the answers that you need to very quickly. This is a 100% not for you. Right? So to me, it's ultimately down to that litmus test that we talked about. It's the time to insight. If you're can already get to what you need and it's quick enough for you, then this is not either.

Also, if we're talking about unstructured exploratory dumps. So if for purely exploratory unindexed datasets, like raw computer vision streams or server logs, where business definitions don't really matter, then this is also not for you. That takes a different approach, and usually that data is not trying to be consumed in order to generate business insights. So if that's not the ultimate goal, then, agile ledger architecture is also not for you. And as you continue to invest time and effort into popularizing

Tobias Macey:this and flushing out the ideas? What are some of the ideas that you have planned for new and exploratory aspects of what the architecture can enable, how it can be extended, some of the ways to maybe build tooling around its adoption or just some of the other plans you have going forward?

Christopher Doidge:Yeah, absolutely. It's it's a very exciting time for it. So I really see it, lined into three pillars for the future of this. One is really the mentality shift. So it's helping data teams to transition from reactive ticket takers into strategic business partners. So helping them to step out of that daily query grind. You know, one of the most, demoralizing tasks is if somebody comes asking for leads and every time you need to run a twenty, thirty minute SQL query, you know, that can be very weighing on you as an analyst or a data person.

And so enabling this sort of a structure to where you're already taking that logic and putting it into your consumption layer, as we mentioned, that gold layer tables, you're essentially moving that process to being done one time. You're making it incredibly repeatable. You're making it visible, and you're ensuring adoption. Everybody's doing it the same way. Along with this mentality shift, it's really viewing data as a compass, and it's reminding teams that data is a true compass. It's not ball.

You know, a good data compass shows you which direction not to go, and it empowers leaders to execute high velocity decisions by telling us, well, that path definitely did not work. Let's not do that again. Introduction of feedback loops to say this is where the data is pointing us to and then seeing if it's the right track. Let's go that way and see if it's actually helping our business in the right way or not. And if it's not, then we adjust our compass and adjust again.

So it's a bit more around mentality shifts and bringing to light that you don't always have a problem with your data. It's more about the way that you've stopped short. You've made it strictly for an analyst, the environment, and essentially created a bottleneck. We should take it a step further and enable the data for all stakeholders within the company. Last pillar for this, for the future of the architecture is, open source and low code templates. So I'd love to get to the stage where I have prepackaged modeling templates already built, maybe even DBT libraries that can help a team implement this ideology quickly, set what their bronze would be, have it suggest some silver rules to get you a clean layer, and then spin up your very first gold tables to see how your nontechnical stakeholders can self serve data.

Tobias Macey:Well, for anybody who wants to get in touch with you and follow along with the work that you're doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gap in the tooling or technology that's available for data and AI systems today.

Christopher Doidge:Yeah. Absolutely. So the gap between technical data engineering speed and operational business understanding, we have incredible tools to stream gigabytes, petabytes of data in seconds. But if data analysts are are not embedded in operational business meetings and understanding how decisions are being made, we just end up building these fast monuments to past priorities. And I always had this joke because it seemed like every data team that I worked in, we always built this monument to past priorities.

By the time we finally got it built, the business had already moved on. So all hail the monument to past priorities. So, it's exactly the biggest gap that I see. We spend so much of our time focused on tooling and ensuring that we have the flashiest software to accomplish this, but you don't need it. I mean, you can get to the true insights using a Google Sheet. It just architecting your data in the proper way. It's about convincing your stakeholders by showing them how you're applying logic to the data in order to get it ready for them to consume.

And then trusting that the number that you're giving them is coming from the sources that is agreed upon. If you can cross that bridge, I feel like you're definitely reducing that gap between tooling and technology.

Tobias Macey:Alright. Well, thank you very much for taking the time today to join me and share the work that you're doing on the Agile Ledger architecture. It's definitely a very interesting approach. It's great to see that it is an addition to, not necessarily replacement of the warehouse that I've been building all along. So I appreciate the effort that you're putting into that, and I hope you enjoy the rest of your day. Absolutely. Thank you so much for your time. Appreciate it.

Thank you for listening, and don't forget to check out our other shows. Podcast.net covers the Python language, its community, and the innovative ways it is being used. And the AI engineering podcast is your guide to the fast moving world of building AI systems. Visit the site to subscribe to the show, sign up the mailing list, and read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@dataengineeringpodcast.com with your story. Just to help other people find the show, please leave a review on Apple Podcasts and tell your friends and coworkers.

Transcript supplied by the publisher with the episode.

Data Engineering Podcast

by Tobias Macey · English · Tech & Science

This show goes behind the scenes for the tools, techniques, and difficulties associated with the discipline of data engineering. Databases, workflows, automation, and data manipulation are just some of the topics that you will find here.

More from Data Engineering Podcast

  1. E516 · 15 Sep 2026 · 53 min

    What Context Really Means in Data Engineering and AI

    Summary In this episode Soham Mazumdar, co-founder and CEO of Wisdom.ai, talks about what “context” really means in data engineering and AI systems. He explores why context has become such an overloaded term, spanning everything from semantic layers and data catalogs to tribal knowledge, query logs, dashboards, and even agent memory. Soham explained that the big shift is that context is no longer being prepared primarily for human analysts, but for LLMs and agents that can’t reliably fill in missing gaps on their own. That change raises the bar for how context is represented, validated,…

  2. E515 · 27 Aug 2026 · 46 min

    Specialized AI for Data Engineers: Inside Astronomer’s Otto

    Summary In this episode Yetunde Dada discusses Otto, Astronomer’s AI agent for Airflow, and the broader challenge of making agentic tooling actually useful for data engineers. She explored why generic coding assistants often fall short in data workflows, how Otto adds the missing context around Airflow, Astro, upgrades, and troubleshooting, and why Astronomer focused first on high-leverage use cases such as DAG authoring, investigation of pipeline failures, version migrations, and legacy scheduler modernization. She also discussed the practical realities of introducing agents into…

  3. E514 · 2 Aug 2026 · 1 hr 2 min

    Why Multi-Agent Systems Need Shared State, Graph Semantics, and Governance

    Summary In this episode Ragnor Comerford talks about OmniGraph, a lakehouse-native graph storage layer designed around the needs of agentic systems. He explores how graphs are primarily a semantic model for representing the world, rather than just a specialized engine for traversal workloads, and how that perspective shaped OmniGraph’s design on top of object storage, Lance, Arrow, and DataFusion. Ragnor explained the motivation for combining graph semantics with Git-style branching and merging so that teams can manage probabilistic writers such as AI agents with stronger governance, shared…

  4. E513 · 6 Jul 2026 · 1 hr 1 min

    Building the Context Flywheel for AI Data Agents

    Summary In this episode Prukalpa Sankar, co-founder of Atlan, talks about what it takes to build a “context flywheel” for AI agents in data-intensive organizations. She explained why model intelligence alone isn’t enough to make AI useful in production, and how real performance depends on contextual intelligence: institutional knowledge, semantic meaning, procedural know-how, and access to the right tools. She also dug into how metadata catalogs are evolving into broader context layers that serve both humans and agents, and why agentic systems are changing the economics of metadata and…

  5. E512 · 18 Jun 2026 · 50 min

    Holding Kafka Right: Product-Friendly Streaming with TypeStream

    Summary In this episode Jevin Maltais talks about the practical realities of building reliable, product-focused streaming systems with Kafka. Jevin shares lessons from roles at Zapier, Humi, and Clio, where real-time synchronization, customer data unification, and document sync at scale highlighted both the strengths and common misuses of Kafka. He digs into using events as the source of truth, materialized views with KTables, and how schema registries and type safety prevent downstream breakage. Jevin explains why teams often reach for heavyweight Kafka clusters without leveraging Streams,…

  6. E511 · 8 Jun 2026 · 53 min

    Text to Data Products: Kaarvi’s End-to-End AI for Ingestion, Quality, and Dashboards

    Summary In this episode Shravan Gunda, founder and CEO of Kaarvi AI, talks about building an AI-native, agent-driven data platform designed to eliminate the janitorial work that consumes most data teams. He explores Kaarvi’s multi-agent architecture that runs queries across seven LLMs in parallel for reliability, its synthetic data generator that mirrors source schemas for quick testing, and “Hey Kaarvi” chat for text-to-SQL, text-to-transformations, and text-to-dashboard workflows. He also digs into on-prem versus SaaS deployments, domain-specialized agents for privacy and accuracy, code…

  7. E510 · 1 Jun 2026 · 54 min

    Scaling Graph Analytics Without ETL: Inside PuppyGraph’s Architecture

    Summary In this episode Weimo Liu, co‑founder of PuppyGraph, talks about the engineering behind their “zero-copy” graph querying engine for lakehouse and database sources. He explores how PuppyGraph lets you run Cypher and Gremlin traversals and graph algorithms directly on data in Iceberg, Delta, Hudi, Hive, and even MongoDB—without loading into a separate graph store. Weimo explains their edge-sharded, vectorized, MPP architecture that tackles hub nodes, multi-hop traversals, and shuffle at scale, targeting sub-second to single-digit-second workloads. He digs into practical graph data…

  8. E509 · 6 May 2026 · 59 min

    Maximizing GPU Utilization: Heterogeneous Pipelines with Ray and Kubernetes

    Summary In this episode Robert Nishihara, co-founder of Anyscale and co-creator of Ray, talks about maximizing hardware utilization for AI and data-intensive workloads. He explores Ray’s evolution alongside Kubernetes and PyTorch, and why consolidation at these layers has enabled a new generation of complex, heterogeneous workloads. Robert explains how data preparation has shifted to GPU- and inference-heavy, multimodal pipelines; where Ray fits compared to Spark and workflow orchestrators; and why Ray excels at composing heterogeneous pools of compute, handling failures, and scaling complex…

  9. E508 · 7 Apr 2026 · 59 min

    The AI-First Data Engineer: 10–50x Productivity and What Changes Next

    Summary In this episode, I sit down with Gleb Mezhanskiy, CEO and co-founder of Datafold, to explore how agentic AI is reshaping data engineering. We unpack the leap from chat-assisted coding to truly agentic workflows where AI not only writes SQL and dbt models but also executes queries, debugs, runs tests, and ships production-ready outcomes. Gleb explains why teams that master this AI-first loop can see 10–50x gains, how security/compliance concerns can be addressed with platform-native LLM endpoints, and why the role of data engineers is shifting from code authors to operators of…

  10. E507 · 29 Mar 2026 · 50 min

    Treat Metering Like Finance: Building Data Platforms for Consumption Economics

    Summary In this episode Himant Goyal, Senior Product Manager at Salesforce, talks about how data platform investments enable reliable, accurate metering for consumption-based business models. Himant explains why consumption turns operations into a real-time optimization problem spanning metering, cost attribution, billing, governance, and cross-functional ownership. He explores the richness required in usage data to support sophisticated pricing, the importance of treating metering like a financial system, and the architectural foundations - event schemas, durable ingestion,…

Every episode of Data Engineering Podcast →

Take it with you

The Melo app keeps playing with the screen off, works in the car and on your watch, wakes you to your station, and browses the whole catalogue offline. Free, no ads, no account.

Get it on Google Play