Episode 79 · AI Engineering Podcast
Harness Engineering for Reliable, Governed AI Agents
19 Sep 2026 · 1 hr 4 min
Episode 79 · AI Engineering Podcast
19 Sep 2026 · 1 hr 4 min
Summary In this episode Nikunj Bajaj, co-founder and CEO of TrueFoundry, talks about the challenge of building reliable agents on top of inherently variable foundation models. He explores the idea of the agent harness as everything around the model engine: memory and context management, tool and MCP integration, sandboxed code execution, permissions, observability, and guardrails. Nikunj explained how TrueFoundry approaches enterprise AI as a centralized control plane for token traffic, while TrueForge provides an open source, vendor-neutral harness for building agents without locking teams…
Tap a chapter to play from there.
Hello, and welcome to the AI Engineering Podcast, your guide to the fast moving world of building scalable and maintainable AI systems. Your host is Tobias Maci, and today I'm interviewing Nakunj Bajaj about the challenges of keeping your agents reliable when the underlying models aren't and how your harness can help. So, Nikunj, can you start by introducing yourself? Thank you so much for having me here, Tobias. I'm Nikunj, co founder and CEO at TrueFoundry.
My background is in machine learning, used to lead one of the conversational AI teams at Meta. And prior to that, have worked at the intersection of building ML models and ML infra at a couple of other startups as well. And do you remember how you first got started working in the ML and AI space? Sure, I do. My my absolute first exposure to ML and AI happened during undergrad, and this was not through ML and AI, but the underlying concepts in ML and AI that was through linear algebra. Okay? At in my freshman year during undergrad, I got a chance to implement some of these matrix decomposition algorithms on top of FPGAs, which are like free field programmable gate arrays.
And when you have to really break down these complex algorithms into their unit of operation to be able to run them efficiently on FPGAs, you understand what's happening under the hood. And I think that learning really exposed me to machine learning, and then it continued through my undergrad and the rest of my career. In terms of what you're building now, you mentioned that you're at TruFoundry, you've built the TruForge harness. I'm wondering if you can just start by giving a bit of an overview of what is your current area of focus and how did you end up working in this space?
Absolutely. Yeah. So very quickly, I describe what is TrueFoundry and what is TrueForge and I will talk a little bit about how I personally got interested in this entire area. So TrueFoundry is the control plane for enterprise AI. This is one pipe through which all the enterprise token traffic is supposed to flow. And when I say token traffic I mean traffic including calls to models, to MCPs and calls to agents. Once all of the organizations traffic flows through a common pipe a common gateway then you get observability unlocked, you get the reliability unlocked, you get cost optimization unlocked, you get security completely taken care of, and lastly you get your identity management taken care of. Right? So it basically solves for these five pillars, observability, reliability, cost management, security, and identity.
Okay? So that's what true Foundry is about. It's an it's a control plane to enterprise AI implemented through an AI gateway. And TrueForge on the other hand is actually an open source agent harness. And the way we think about this is if you think about building an agent as a car, we think of the model as the engine, but everything else around the around the model, the engine is essentially the agent harness. Right? So it means how do you steer it? How do you put brakes on it? Right? How do you have a dashboard where you can see everything that's happening? Right? Those are all components of an agent harness.
And technically speaking, that you include in the agent harness are how does the model get a sandbox to execute some code that it might be writing? How does the model does context and memory management and the the token optimization? How which tools, which MCP servers, which sub agents that you call? All of these concepts are handled by the agent harness and that's what TrueForge is. It's a neutral open source harness that we actually launched recently. Just a little bit about my personal background that led to us working on this. So when I was at Meta, I got a chance to lead some of the conversational AI teams. And what stood out to me was the internally developed ML platform at Meta.
And the core design principle that really hit home with me was this notion of federated execution that I as a developer can pick the tools, the models, the frameworks that I really care about and be able to move as fast as I as I would like. So I was enabled with the right tooling to just execute, but at the same time, it never left the sight of centralized governance. So I didn't have to worry about instrumenting my code. It would always log things that I care about. And similarly a central platform team, if they had to figure out any erroneous or malicious activity, they always had the control to check what's happening and be able to take care of that. So this entire dichotomy of whether there's a trade off between federated execution and centralized governance turned out to be false by looking at Meta's internal ML platform.
And we want to take that learning and bring it to the rest of the organizations out there. So the whole space of agent gateways and agent harnesses kind of spans at least two of the major hype trends over the past year or two, where maybe a year or so ago, AI gateways were gaining a lot of attention. There were a new gateway probably every other day. Some of those have been consolidated into other systems. Some of them have ceased operation. Some of them are continuing to grow.
And then the past six to eight months has seen a lot of focus on the idea of harness engineering and that being one of the key differentiators beyond the models themselves. And I'm wondering if you can just talk through some of the ways that you're navigating some of the challenges and complexities of being in such a crowded space and what you see as some of the benefits of working across both of those and some of the ways that having a harness and a gateway that work well together can be a force multiplier.
Absolutely. Happy to zoom in. And by the way, I cannot agree more, like, you know, with the pace that everything else is changing. In in the startup ecosystem now, like, you there's this new saying. Right? Like, you know, that earlier people had to find PMF or the product market fit once. Now you are fundamentally finding product market fit every six months. Right? The space is changing that fast. But also the interesting thing around this is that this is what is creating the opportunity as well. So this is certainly a challenge for building, but this is an opportunity for operating, right, in the startup space. And the reason being, that think about what's happening with every consumer of call it like an AI gateway or a control plane layer, right. An organization builds out a centralized platform engineering team that's going to develop this AI gateway because they care about solving those five problems that I talked about earlier, right.
They start with doing this, they build a thin proxy layer and they're like okay we are sorted. And then soon enough, the model proxy layer that they have developed, all the API signatures start changing out there in the market. A couple of months later, MCP is launched. And now they're like, oh, I need to bring my MCP's in the same proxy layer. And they have to solve for authentication and authorization and all that, right? And then A2A is launched. So they have to figure out this agent to agent communication protocol through the platform, right? Then multiple agents that their teams have developed hit production, so they care about ensuring uptime, reliability, latency optimization.
And then agent identity becomes a real thing, right? So it's an ever changing space from their standpoint. New models keep getting launched, new protocols keep getting launched and the complexity and the requirement from this platform keeps on changing every single day, right. So for thousands of team out there whose main job is not building and maintaining a platform like this, it becomes a completely intractable problem to solve. Right? So we have seen that organizations went from what we call as this build, break, and then buy cycle.
Right? And that has created an opportunity for a startup like us to build a platform like this. Now, of course, at our end, the complexity does not reduce, it only increases because now we have to deal with the stack of all of those thousands of enterprises. Right? While the space is ever changing. So so we have done a lot of optimizations at our end where we have built out internal agentic solution that keeps pulling what are the new things that are getting launched in the market, what's breaking in our platform, and how do we continuously keep optimizing this besides of course like the fact that there's a team that really cares about solving this problem statement.
Same thing for agent harness, I feel like this is relatively newer but we have seen a major demand from organizations that they care about being able to take control of the agents that are being developed. And you're just not able to take the control of the agents that that you're developing if you're using a closed proprietary agent development platform that you just don't have any control over. Right? So so that's where people want more of a vendor neutral harness and it's been fun to fun to work on this.
The reason this becomes extremely interesting is vendor neutral agent harness by design is open to models from multiple providers. Right? So you are able to access frontier models from the labs, or you're able to work with frontier models that are open source. Right? And all of those models for all of the organizations that you work with are available through the gateway layer. So people can bring their own model extremely straightforward. Right? And this is extremely straightforward to bring their own models, basically.
Similarly, people care about being able to bring their own MCP servers. Now the good news is that a lot of MCP servers that people want to access are these public MCP servers remotely hosted, centrally managed. So they don't have to do a lot of put in a lot of effort to try out different MCP servers. It's all connected through the platform of the gateway. Right? But then these MCP servers can be easily plugged into the the agent harness there and help them invoke these tools. And same thing, you can extend this to agents as well. Agents are all accessible through the through the gateway there. Right?
Now across any of these models, MCP's and agents, if you wanted to apply any kind of guardrails, your guardrails are also available in your central control plane layer that gets implemented through the gateway, right. So fundamentally what's happening is when you're building your agents, all the raw building blocks that you need and the corresponding registries are available as part of your control plane. The logic to build these agents, right, the engine that operates the model and the tool and the agent loop that sits within the harness and it's able to help you orchestrate your calls better and the two together actually becomes a powerful force because now you're able to build agents that are accurate but cost efficient and centrally governed. So that's roughly the unlock that we're getting by bringing the pieces together, if that makes sense. Yeah. It definitely does. And one of the other interesting pieces of the gateway is that one of the, I think, motivating factors, at least in the initial push, was to give a unified interface to all of the different models. But as the different providers have
changed their APIs and as OpenAI itself has evolved their API, that that became a much more difficult moving target, as well as the fact that by forcing everything into that one shape, you potentially lose out on capabilities of Anthropic or DeepSeek or whomever else. So I'm wondering if you can just talk to some of the tensions there of having an easier onboarding where you just have one API contract to code against versus being able to actually take full advantage of the underlying providers.
That's that's a very interesting question. There are two aspects to it that I would like to call out. Number one, that, of course, like, you know, with the API signatures of the model providers continuously changing, this becomes, like, you know, relatively difficult to to implement this. But to that end, what we have done is, for the most part, we do the job of unifying the API signatures across, like, you know, this like, the the diverging API signatures on the on the on the right side and bring them all with the unified access for our end developers. So we don't lose the capability of different model providers, but at the same time, we try to solve for, like, you know, bringing them all into a more unified API signature for our end users. So they don't have to worry about these waiting signatures.
But at the same time, you're right that this is not possible to execute every single time. And for that, what we have done is we have also implemented what we call as a custom endpoint that you can invoke through the gateway. And in that case, like you do not have to be compliant with the gateway specific signature, and you're able to run them through through the gateway, you're able to get a lot of benefits of the gateway to the end customers. You obviously get full benefits of the model in that case, but you do lose some out of the box capability that the gateway is able to do so long as it understands the underlying model signatures. So I guess we have divided the platform into those two parts.
One where you get complete benefit of the gateway and unified API signature, and the other case, you bring a custom signature and you don't get complete benefit of the gateway. Digging more into the harness itself, as we've noted, that's I think the most recent hype cycle that we're in where everybody's building a harness. It's become a very complex market. There are a few that have gained a decent amount of traction. I think pie is probably the most notable one largely due to the success of OpenClaw.
I'm wondering if you can just talk through some of the key challenges and components of building a harness that actually works well. Absolutely. Our team has been having a fun time building out this harness. And the first thing is getting the accuracy right. Right? Getting the agent's accuracy right, getting ensuring that you are able to get the quality of responses that you get from some of these like you know out of the box harnesses, managed harnesses is actually a tough problem to solve. Right? You don't quickly implement a harness and get there. Right? But after that, once once you have gotten there, the the challenging aspect of harness is a lot of the companies that additional companies that you have to deal with to ensure that all of this is happening. Right? So for example, how do you how do you make sure that you give access to a sandbox to an environment and by the way like of course along with all the recent incidents, how do you check that the agent is truly sandboxed, right? It's not it's it's not able to get out of it. How do you check that the agent when it's when it's executing the code, it's actually executing it efficiently
and in the right environment. Right? How do you make sure that which of the tasks should actually be done through coding and like when you do the memory management around that, done through coding, but it's it's done through the tokens directly and you have to build the memory management around that. So there's some of those complexities are actually extremely difficult. Then anything around durable execution, human in the loop that you have to ensure that certain tools you get feedback from the human beings and then only you let the agent proceed forward is complex.
And lastly, the the one thing that is extremely challenging with the harness layer is how do you manage permissions and access control through agent chaining. Right? Because what what's happening now is agents are often acting on behalf of users. Right? And what I mean by that is as a user, I can send a Slack message to Tobias or an agent acting on my behalf can send a Slack message to Tobias. Right? How does Tobias know? Whether this is directly coming from Nikunj or an agent acting on Nikunj's behalf. Right?
And what happens if this agent invokes a sub agent invokes another sub agent in the chain basically. Right? And each of these different agents had their own identities and had their own access control. How do you ensure that the least privilege is managed throughout that entire chain? Right? So all of these, like, know, are actually fairly difficult to handle. And what lives in the harness layer of what has to live in a more centralized governance layer, the right thought process around that is also also something that that made things very complex to implement.
The other interesting aspect of harnesses is that that term has become a little bit conflated because most people are thinking about it as far as their Claude code or codex or PI agent harness or any of the numerous other harnesses that are out there specifically for coding agents. But then there are also harnesses that are focused more on deployed production agents that are interacting with various other systems, and you want to make sure that those are more constrained than your local terminal agent might be and have different capabilities. And so I'm wondering if you can talk to some of the aspects of how to build that harness in a way that is layered and composable so that you can have that core functionality that's shared across all of the different use cases, but be able to actually exclude functionality or exclude permissions from the harness for cases where you want it to be operating in a more constrained environment.
That makes a lot of sense. And and you're exactly right. Like, you know, first of all, you know, there's actually a lot of noise in the market. Like, people actually are not able to understand the the difference between all of these different products and the corresponding intended use cases with which they were they were developed. Right? From from our standpoint of the the stars that we have taken here is that we are focused on building more general purpose agent harness. Right? So the general productivity agent harness is what we are like, know, what we have taken a stance on building. But also like you know as you mentioned that we are developing this with an intent that we start with like this entire personal productivity type of an agent harness. Right? And eventually if we want to focus a little bit more on the on becoming a coding harness, we are able to steer the harness in that direction, essentially. The fundamental building blocks don't necessarily change between these different harnesses, but the way these components come together and, like, you know, how in one case you're always dealing with a constant context layer, let's say, from a coding side versus in the other case, like, you know, you get a lot of the context from your MCPs and stuff that you're connecting on the harness. Right? So how do you how how does the harness handle those nuances?
Is something that, like, you know, we have basically kept it configurable from our standpoint. One thing, though, that I would mention that that has been that has been a big win, by the way, in exactly differentiating these between different types of harnesses is by not trying to be all in one, you actually gain a lot of efficiency. And what I mean by that is if you try to solve all the problems in one harness from the get go, then essentially, like, you think about all of your system prompt has to become massive. Right?
And because a harness at any point in time might have to be dealing with any of the underlying, call it, tasks that you care about it's solving. And because of that and that system prompt, by the way, goes in every turn of the agent conversation. Like, you cannot avoid that. And when that starts to happen, like, start paying taxes for that. Right? So you are basically consuming more tokens in every task that you are that you are doing, and you are spending more dollars on every task that's that you're doing.
And that has been actually one of the design decisions that we have made that the way we have designed the harness that we we are able to specialize this for different tasks and, like, you you don't have to do all in one at the same time. That said, I also wanted to share that because of this being one of the reasons, and actually this is one of the reasons, not everything, we have seen that when we ran the benchmark of our harness against other tools out there like cloud managed agents and all, we were able to get over 30% cost reduction for a maintained accuracy while using this exact same model and as much as 70% cost reduction when we let the harness choose the models, like, you know, choose its own model essentially. Right? So it it also had that flexibility of being able to choose the right model from different providers and stuff. So that's how much range we have been able to achieve by, like, you know, on the task that the harness is accomplishing as well.
On that concept of model routing, there's also been a lot of focus there. The project that I'm most familiar with in that space is the VLLM router that lets you dynamically choose which model to route to for a particular request or set of requests. And I'm curious whether that's something that you see is belonging in the harness itself, in the gateway, something in between, and just some of the aspects of how to make sure that you are effectively choosing which model is going to have the right result without burning a bunch of tokens on a failed result and then having to retry it with a different model.
The way I see it, you have to have this kind of routing capability at this point in both places at least the way the industry is going right now. You have to have this in both places and to serve a different purpose in the two places. Right? A harness has to be able to figure out smartness when it's running the loop, right? That which agent I'm calling, which model am I calling, which tool am I calling, right? Around that time it also has to have certain level of smartness that I know that certain models are really good at coding tasks. When I'm getting a coding task, I route my query to those models. In this task, I have such a large token context window that I need to use a model of model accordingly. Right? So it has to have a certain level of that intelligence built into it. And that's the routing logic that I believe the harness will continue to implement and own.
That said, you can also imagine that when you're using when you're making calls to the models, many of these calls are coming from the harness, but many of these calls are directly coming from end users where there isn't an intelligence layer built into the caller to begin with. Right? So so from a and all of those calls are fielded by a gateway layer. Right? So a gateway has to be also smart by looking at the prompt, right, and figure out which model should it route it to. And of course, gateway solves a problem from a different lens as well, right? Because the gateway has the awareness, the real time awareness of what's the status of different model providers. Like is one of them serving high latency and I may want to route to a similar model but to a different region or the same model from a different provider, right? And it also of course has a lot of context around how much overall budget and rate limits have been consumed so far. So it has a lot of real time awareness and it will do the routing based on complexity but also all of this awareness as well and that layer sits rightly within the gateway.
And the task contextual awareness routing sits more within the harness, and both of them need to coexist. The other interesting aspect of harnesses is because you're working with these models so closely and because testing anything that touches a model is so challenging due to their probabilistic nature, what are some of the validation methods that you use when iterating on a harness to make sure that you're not regressing and that it is able to actually work effectively across such a wide swath of different models with different capabilities?
I think that is certainly a challenge for for for the exact reasons that you mentioned. Right? And I I wouldn't say that, like, we have we have fully solved the problem, like, where we have absolutely contained the the the test cases and regressions and ensure that, like, this is always fully solved. But, of course, we try as much as possible and the community understands this problem statement. Right? So now there are a bunch of benchmark datasets against which, like, you know, you can always keep running your harness to ensure that you are not regressing on your quality, and cost benchmarks.
Alright? So we keep running that. And then internally for ourselves, you know, think I I think of these benchmark datasets as more integration tests from the from the old software world. Right? And then you still need your functional test and unit test as well. So what we have done is for different modules that we are building within the harness, we have actually designed some of these test cases as well. Right? So the independent modules don't progress, but meaningfully with like, you know, so many variables and the stochastic nature of the models and the harness, the only true way of testing this out is against the benchmark data sets from our standpoint.
But even then so, it so happens that the end users of Harness should also actually, like depending on the use cases, they should have at least some coverage of these test cases at their end as well. That's a final litmus test to ensure that this doesn't regress. As people are designing and implementing agents of various shapes and given the wide variety of harnesses as we've discussed, what are some of the ways that people should be figuring out what are the differentiating factors across these different harnesses and how to identify which one is going to suit their particular use case?
It's very interesting question, Tobias, because the way I'm thinking about this, and I would love to brainstorm with you on this one, like hear your thoughts as well. Right? But of course, while we see that there's a bunch of harnesses out there in the market today that you're able to leverage to build some of your agents, the divide in the industry is at a very macro level. Right? And actually, we are seeing our customers choose their harnesses based on those macro divides at the moment rather than trying to get super micro with it with a bunch of low level capabilities of different harnesses or at least that's what we're seeing at the moment.
What what do I mean by that? At the first level you can imagine that you have harnesses that are SaaS like managed harnesses that you get to use out of the box. There are some of the harnesses that are more vendor neutral, more open harnesses. Right? And people are able to take a call based on their own needs that am I happy to be locked into this ecosystem and get the advantage of an integrated ecosystem, right, which could come at a price? Or do I want to do a little bit more work and get into this open source neutral type of a harness, right? So that's one type of a thing.
The the second here is that the way some of these harnesses are implemented, they're intended for certain audience. Like, some some harnesses assume that you're a developer, you understand, like, you know, you're able to set up a bunch of configurations. You're even, if need be, you're able to write some code, right, or micromanage how the harness is working. Right? So it's intended for a more technical audience, where some harnesses actually are just, like, you know, built for a knowledge worker who does not necessarily come from a coding background and does not want to manage complex configurations overall.
So today, I'm seeing that, like, you know, there's a few handful of tools that are out there that, you know, people for the most part are leveraging can be divided based on based on this itself. Right? We people don't have to go too low level that oh like in this case they harnessed us very well on memory management and in this case they harnessed us very well on a sandbox execution so and so forth. And we are not seeing people take calls based on that. Are you seeing a different trend by any chance? So, yeah, I think that in the harness space, one of the interesting pieces, just as an anecdote, is that a lot of them seem to be written in TypeScript and running on the node ecosystem, which as a Python developer, it is curious to me, but I understand that JavaScript is taking over the world.
But, yeah, the the key factors that I see when I'm looking at different harnesses are, yeah, what is the model support, which is generally fairly broad, but also how does it integrate with some of the different providers. So I think Claude and Codex are probably the more interesting ones, Claude specifically because they've been having a challenging time of figuring out how their pricing that when you're going through their Claude desktop and Claude code work versus if you're using a third party harness where you don't get the benefits of their flat rate pricing. So that's one of the considerations when I'm looking at some of the harnesses as a software engineer.
But the other pieces are the MCP and a two a support. So how quickly they're adopting some of these different protocols, particularly as the protocols themselves change. So MCP just released their latest revision maybe a month ago. And then the what is the footprint of the harness? How much resource does it take? How does that factor into deployability, particularly for if I'm using this as a deployed agent that I want to run? How does it fit into things like instrumentation and observability?
Because I wanna make sure that I have a full trace of what the agent is doing so that I can figure out when things go wrong. Why did they go wrong? How did they go wrong? Because they're moving at machine speed, you have to have other agents that are monitoring those outputs to make sure that you can catch those failure modes versus just if you just have a harness that says, yes. I can run the model. I can plug into all these protocols, but it's a black box and opaque, you have no introspection into the system.
That's definitely a nonstarter for me as an operations engineer. So those are the some of the things that I look at when I'm thinking about different harnesses. That that that makes a lot of sense. That makes a lot of sense, Tobias. For sure. Like, you know, we we definitely see our users considering these, but, like, you know, a lot of the times that, at least from our lens, what we are seeing is we work with organizations. And on an organization level, I think the the decision matrix is broader than this because people really care about being able to suit some harness to their own broad scope needs, and they choose based on that. But we'll get exposed to more and more of these nuances, like as our open source agent harness gets a broader adoption as well. So I think that would definitely be an interesting learning for us.
So digging now further into the TrueForge harness itself, could you talk through some of the key design and architectural elements of it and some of the ways that the shape and scope have evolved from when you first started working on it? So so in terms of, like, I just generally, the way we have designed first of all, like, I just let me give you little bit of a background on how this entire concept of Harness started from from our standpoint, right, and then I will describe how some of the architectural choices have evolved for us as a team as well. So there were two main pulling factors around this.
We wanted to build some capabilities on our own product like like TrueFoundry AI gateway product where people are able to consume a lot of the information through the product but in a more agentic way. I'll give an example. TrueFoundry AI gateway sees a bunch of traces, agentic traces, and we want to be able to reason on the traces, alright, that which of my agents are doing well, which ones are prone to any guardrail problem injection attacks, or where is PII leaking that I don't want want it to, how can optimize for cost and stuff like that. Right?
So we wanted our users to have that capability where they can more naturally improve the agents that they are, like, you know, invoking through our gateway. Right? So fundamentally, this means that we wanted to bake in a first class agent as part of our own product so that people are able to leverage this. And to that end, fundamentally, everything that we needed this agent to access, all the context, all the traces, all the models and MCPs that the user had permission to was available within the product. And we had all of these things available as context and as MCPs.
So from our standpoint, this felt like it should have been as simple as that we're able to connect a bunch of these, like, in a context and use a harness to quickly spin up this agent that we are able to ship with the product. Right? But when we try to use the tools out there today, like back then, right, we realized that you are either locked into a bunch of these closed harnesses, which does not allow you this capability of shipping agents as part of the product itself. Right? Or you would have to fundamentally leverage more pro code frameworks, which, like, require you to micromanage the performance of your agents and then quickly things diverge and you keep, like, upgrading that essentially. Right? So so we really wanted to strike that balance where we are able to optimize for speed, ship things fast, and they cannot get locked into any ecosystem.
So that's that was one of the motivation that, like, you know, led us to kick things off. And at the same time, we started seeing that a lot of our customers started caring about, like, you know, not being locked into a certain vendor that way. Right? So when we started building this out, the neutrality, the openness, the speed at which you are able to ship, these were some of the design principles with which we had started building out the harness.
Over a period of time, when we started actually implementing this, and this actually changed a lot in how we thought about the overall architecture, we realized that while we are able to make building of the agent relatively straightforward, being able to run this, right, and being able to run this at times and in environments that we care about is another challenge. Like, there's a lot of things that you want this agent to do that you want to be able to schedule, that you want to run every day or every few hours. Right? Or you want to run this on a trigger.
How do you expose or sometimes you want the agents that you're developing, you want this agent to run as part of another ecosystem. Right? So how do you build out that entire runtime of the agent is another major problem that we realized that we needed to solve. Right? So, like, we continue to evolve the harness from, like, you know, primarily being neutral, like, you know, to this entire runtime. Now the two different types of runtimes, basically. Right? Similarly, some of the design choices that we needed to change was how do you think about sandboxing overall? And what we thought was sandbox should be exposed to the agents of the tool. That's the other design choice that we had to make. Most of the harnesses that we know run the entire agent itself in a sandbox. We actually ended up flipping that design principle over a period of time.
So now what we do is we run the agent loop in the server itself, and we provision a sandbox only when the agent actually needs to execute that code. Right? So the advantage that we got by doing this was one server is now able to run many of our agents concurrently. And it turns out that that when you are not using writing code, basically, it makes those turns go much cheaper and much faster. And you're ensuring that things like your entire secrets and all are never leaving the layer that the hardest can touch at all. Right?
So a lot of these things that we started editing in our architecture came as evolution of what we saw as this accuracy cost and efficiency trade off that we really, really wanted to nail down. And over a period of time, these sort of things evolve. And as you have been working on the harness and using that to validate the gateway, what are some of the most complex or challenging aspects of building these two different systems and using them to play off against each other and just some of the discoveries that you've run into that were unanticipated roadblocks on your path to actually building these products?
Yeah. So in all honesty, I think one of the problems that you always have to solve for in such settings is organizational, right? Because the way you organize your teams building out the by the way, these are problems that everyone needs to solve for. Right? Like, like, people have to figure out the solution. So I would like to share this with with folks as well listening listening to the podcast. When you think about building out the the gateway itself, you are building for one intended audience, which is getting used in a certain way. Right? Organizations buy a gateway for observability, for control, for being able to access the rich ecosystem relatively fast. Right? When you're building out an open source harness instead, right, you are building for the capabilities to to empower the developers with a bunch of these capabilities. Right? So that they are able to developers or even knowledge workers to take up a harness and they're gonna build out these agents.
Now because the the user is slightly different in the two cases. Right? The team that's building it has to always continuously empathize with the end user that they're building towards. And the pace at which the team is developing different components also needs to change because one is serving more production systems. The other one is serving more building phase of the ecosystem as well. Right? So so from our standpoint, how do we allocate the bandwidth? How do we divide the team? How do we ensure that we are able to meaningfully make progress on the on both the two fronts was actually a relatively big problem for us to solve for. Right? The way we we came up with like, the the approach that we came up is to solve this problem is we actually combined the the part of our gateway that was leveraged to build a bunch of these agentic solutions directly. Like, we had a component in our gateway where people are able to register registries, like agent to a registry.
And from there, they would they would always run into some roadblocks into how they have built out an agent using another closed source product or some other framework. They want to register to the registry and then access it centrally through the gateway. They would keep running into those roadblocks. And there were a lot of learnings that were coming out about how the agent development platform or the agent development harness itself should solve the problem.
So we combined the the team that was always getting that learning from this part of the product to also work on the agent harness layer. Right? So the top of the funnel, like, the core pain points that people want to leverage an outsourced open source harness was always coming to this team. So the empathy for the user, both from a gateway standpoint and from a harness standpoint, is how we solve for. And a lot of the capabilities that we ended up building within the harness by design, the roadmap of the harness actually came from our early users of the gateway directly. So that way, a lot of the downstream challenges that we would have needed to solve for actually got front loaded by working with the right set of users using the product as a design partner, I guess. You touched a little bit on some of the cost aspect of the harness being a contributing factor where that's largely the system prompt and trimming that down can have a substantial savings. But as you have been building the system, what are some of the other key elements of the harness itself that contribute the most to the cost efficiency,
the accuracy of the outputs, and the reliability of the agent operating within that harness? So there three major factors to this entire cost optimization and the accuracy and reliability trade off. Like one we covered, which is your system prompt, right? The second one is generally your harness efficiency. And what I mean by that is the number of tool calls that you are making, the order in which you are making the tool calls, the order in which you are applying any kind of policies that you may want to do, which actually has a significant impact on how overall harness is working. Right? And to this end, there are certain more technical concepts that you also bake into the harness.
For example, do deferred loading of tools, for example. When you do not need access to certain tools at a certain point, you defer that, right? So how do you bake in that intelligence that for any given task that the harness is accomplishing, you always pick the right combination of whether you are loading a certain tool, whether you are using the actual persistence to be able to store some context and only get a part of it when you need to answer a certain question.
What do you solve for using tokens of the model directly versus where you think that being able to write code and execute this would be the right way to solve this problem? So this is what I call as the hardest efficiency. And the right choices that you make in this one is extremely critical to solve for the cost and accuracy trade off. And the last one, which is the overall model benchmarks overall. Right? Like, that's the that's the other one that's that's extremely critical.
Many times, like, for the same task, like, models can have an order or even two orders of magnitude cost difference. But their performance could be fairly comparable at that point. Right? So how do you choose the appropriate model for different types of tasks that you're doing? And how do you train the harness so that it's able to make those design choices? I think these are the three key factors that we have worked to optimize on. Obviously, a lot of effort that has gone into is problem number two, where different components of the harness, how do you optimize against those?
Like, that's where a lot of efficiencies and and citing some numbers here, there have been tasks where we have seen some other harnesses which are taking, like, more than 75,000 tokens to implement, like, know, to accomplish the task. Our harness was able to do within 23,000 tokens basically. So we have seen anywhere from 50 even 7075% of token reduction in similar kind of tasks by getting this part right. And then going back to what we touched on briefly of different models behaving differently given even the same prompts, the general wisdom that has built up over the past couple of years, and, of course, that general wisdom is frequently invalidated because everything moves so quickly, is that just changing a model given all of the same inputs can have wildly different outcomes.
And given the current generation of models that we are operating, is that assumption still accurate? And what are some of the ways that people should be testing and thinking about their preconceived notions given what they've learned from the past two years of evolution where I think the hallucination rate was one of the first ones to gain a lot of ground, and that has is it's still a factor, but the the rate has dropped a lot. I'm just wondering as the models evolve and as they all become more capable, has that assumption of switching models will dramatically change the outcome?
Does that still hold? Actually, we have seen that assumption break. It's been now a few months that we have seen that assumption break, in all honesty. And what that means is by the way, hallucination is a term. Like, I I think if there was a way to build a heat map of how frequently that term is being used, I think now over the last few months, it has become quite cold, that heat map. Right? We don't we don't hear people talk about hallucinations as much when they are talking about the model performance. Same thing. Earlier, we really needed to craft a prompt for each of the models, and that was a big thing that people needed to solve for. Now we are seeing that models are getting better, that that you don't need to micromanage your prompts as much to work with these different models. Right? Now, of course, like when you go from a frontier model to a very specialized model, the specialized model expects sometimes even if you have fine tuned some of these models and we we work with still with some of the fine tuned models, That expects prompt to come in a certain way, and that really helps to
to be able to change the prompts. But, like, so long as you're working in a, like, a beyond a certain level of, call it, the number of parameters in the model. Right? Like, once you're working with a threshold, model is larger than a certain threshold, you actually don't need to worry about switching prompts anymore for to to get similar kind of outputs now. And that, of course, has downstream implications now in terms of implementation of the harness that you you just if you you're able to make choices more independently now. Right?
But at the same time, we we in our gateway product as well, we keep this flexibility that if people are switching model, they have a scope to be able to change the the prompt as well. And it's especially useful when you're using fine tuning and stuff. Right? So we actually leverage that capability also in our harness as well. One of the other factors of model selection that I've experienced is that you might need at least a certain threshold of model to be able to effectively drive the harness where, a while ago, they I've had a conversation with somebody who had the analogy of the LLM is the engine and everything else that you put around it is the car. So you can have all the horsepower you want, but if there's no car, you're not going anywhere. And so bringing that a little further, it sort of feels like you you have sort of as you're going on the amusement park rides, you must have this many parameters to drive this harness. And I'm wondering what you're seeing as some of the cutoff threshold where if you get below a certain size model, it's not going to be able to effectively use the harness because there is too much in the context window.
So and and, like, the most the most interesting thing that we are seeing with this is people, the the users of the harness, including including ourselves. Right? We have built out a lot of systems that, like, leverage our own harness. And people actually don't necessarily need to go below a certain threshold anymore to make things work for them. Effectively, it's not like people are thinking about, should I use a 3,000,000,000 parameter open source model to run my entire harness versus should I use Fable five, for example, right? Like, it's not like people are typically operating in those extremes.
People are thinking like, like, am I doing a Fable or a GLM? Am I thinking Haiku, right? So people are still operating in that domain. And what what's what's nice about this is the overall token cost has reduced, like, you know, by such a large margin over the over the period that intelligence per token, as I think about this. Right? Intelligence per dollar that you're spending has has gone up so much that that's the ballpark in which we are we are talking about. So now you are still able to like, even if you choose a base model, something like a HYCU and above, basically, right, the the hardest works. And then for extremely complex task, you go to, like, your GLMs and the Opus and the Fable equivalent, basically. And otherwise, you're able to a bunch of tasks using high equivalent type of products.
And as you have been working on TrueForge and TrueFoundry and working with your customers, what are some of the most interesting or innovative or unexpected ways that you've seen your harness used? So first of all, we noticed that people are We saw through our product itself directly that we see queries in languages that we absolutely don't understand. We were under the assumption that most people are using the harness in English with the most common user group that we work with. They're saying queries coming in completely different languages that don't understand.
That was just a fun thing that we noticed in terms of the usage of the platform. Outside of this, what we have seen with our harness usage, people are doing a couple of things that we just did not imagine will be a use case because it's supposed to be such a deterministic use case. People have actually built out their own pricing systems, like how they price their product overall based on some data that they capture using our gateway and how they build out the end pricing dashboard for their end users using our harness.
And this entire loop came together for our end users and they were able to build out this entire thing within a few hours that their understanding was it would have taken them weeks or months to build this entire system. That kind of a usage, we had not expected people to leverage the harness for. So basically, the input comes from the gateway, the implementation happens through the harness, and that combination became very powerful. So what this has opened up for our users and also just in our mind in terms of what we were expecting versus what we are seeing is people leverage gateway to build out pretty much all of their agentic capabilities.
All of the agent capabilities by default get routed through the gateway. Now, any derivative component that they need to build from here, that's where now the harness is starting to get used very strongly. So I think that's that's been an insight that we have learned through the usage of the product. And in your experience of working in this space and building these systems, what are the most interesting or unexpected or challenging lessons that you've learned personally?
So one of the very interesting things that I've learned building these products over the last two, three, four years is when you work with different personas with their own priorities, right, they are looking at the same solution from completely different lens. And it so turns out that if you're building a system that caters to people with different goals that they're optimizing for, and this is the only system through which all of those goals get a common ground, right, to be solved for, right, then the only way that you can think about building a product like that is by having a good mix of opinion from all of these different varying voices in the room. And I'll give an example of why that has been such an important learning experience for us.
In the very beginning, I talked about different pillars that this gateway and harness product is solving for. It's solving for reliability that your platform teams, your development teams care about. The fact that you build an application, it should be always up, always running, operate at low latency. Then there's a security aspect to it that your CISO org cares a lot about. The four walls within my organization should always be secure. Then there's a CFO aspect to it that they really care about being able to do this cost optimization.
Overall, they should have an ROI, and that entire story should be clear. Which agents are we building? How do we do the reporting and stuff like that? And then there's this compliance aspect to it that if tomorrow I have to be able to answer why my agent built a certain decision, how do I do that? And turns out that the interest of the CISO, the CFO, the compliance officer, the CTO, all of these things are now being answered by this common layer.
And the most interesting way that you can solve for this entire problem statement is by hearing the voice of all of these personas. And where we have seen some of the solutions within the industry where they have tried to create a completely specialized view for the CFO in terms of the FinOps layer. And then they have basically done so much instrumentation that it doesn't get adopted by the platform engineering or the actual development layer. So yes, you have built this amazing, beautiful FinOps layer. But because there's no usage coming from the left side, there's nothing to see on the right side, and so on and so forth. We have seen multiple combinations of these break in the industry. We have seen multiple personas pull our product in different directions. And the way we are able to maintain this balance is by listening to these point of views. And I guess that's been a personal lesson that you really want to understand that, yes, who is your main buyer persona? But who are the other extremely important stakeholders in the room? And you want to make sure that you hear that holistic voice. That's been a major learning for me.
And as people are evaluating different agent harnesses, what are the cases where you would say that TrueForge is the wrong choice? Overall, number one, at the moment, like, no, this is all a point in time answer. Right? So as of whatever, 09/04/2026, right, this answer likely is valid today. But but today, we are not specializing ourselves for a coding harness. So if somebody wanted to build out the users of the coding harness, probably I find out other open source or closed source harnesses that serve that use case better.
In some other cases, if you really care about being able to basically, if you have a lot of visibility that roughly this is how my system should behave and these are the three lanes in which my agent should go on and there should not be any other fourth lane that they need to explore because I care about whatever ten millisecond latency implications, etcetera, etcetera, right? If such is your use case, then you are actually better off being able to more be in control of the code that you're writing and not let it run through a harness layer like TrueForge, I choose a more deterministic framework for building out your agents. I think those are a couple of situations where I would not pick TrueForge at the moment. If you're building anything that is more general purpose, that you care about the evolution of the agents over a period of time, you care about being able to build it quickly, you care about being able to build it efficiently, and you care about being in control of the type of models that you're able to use, you care about being able to control the cost, we feel like TrueForge is a great choice.
And as you continue to build and iterate on these systems and continue to explore the current state of the ecosystem, what are some of the things you have planned for the near to medium term or any particular projects or problem areas that you're excited to explore? Absolutely. Absolutely. This is the thing that we have been talking a lot about. And a lot of our own roadmap is shaped directly by our customers as well, the current users of the platform as well. So one of the things that we have been very excited about is going much deeper into this entire agent identity area.
And how do you manage cross app access? How do you manage this agent chaining, permission controls? And I've given an interesting problem statement here. I had mentioned briefly about how do you do least privilege across agent chaining. But when you start doing least privilege across agent chaining, one of the problems that happen is now your configuration management of which agents have access to which MCP's and which tools and which users have access to which MCP's and which agents, that becomes a very complex problem to solve.
So then how do you take this into a direction where you're not compromising for the UX, where you understand the intent as well? It's not just about theoretically what is possible, you understand the intent when you implement that without compromising the user experience while ensuring the safety and security of the system. That's the direction that we are taking the product in. And then, of course, there's constant work happening around cost optimization, sandbox security.
So some of those areas we are continuing to work on and that's part of our immediate roadmap as well. Are there any other aspects of the work that you're doing at TruFoundry or on TruForge or the overall space of Vagintic harnesses that we didn't discuss yet that you would like to cover before we close out the show? There's one thing that maybe I'll I'll call out here is, overall, the like, one of one of the learnings that we have had in this space is people have been trying to address problems around the control plane for enterprise AI or how you bring together this harness using smaller targeted point solutions.
The goal has been for a lot of folks has been that, can I do my model routing there? Can I do my MCP registry there? Can I build out my A2A optimal agent to agent communication layer? People have been thinking a lot about that. One of the learnings that we have had working with the organizations that we work with is the only meaningful way that you build this overall control plane is if you bring these systems together. Like if you bring all of your tokens together in one place is how you only meaningfully solve this problem.
So it turns out that a model proxy or a separate MCB proxy or an agent registry in isolation or a guardrail solution overall is just not able to solve the problem that organizations are facing because the true problems don't happen within these point solutions. The true problems happen at the interface of these point solutions. And to meaningfully solve for this entire thing, you need to bring these components together and do it in a very AI native way.
That's been one of the learnings that I wanted to share with the folks as well. So whoever is thinking about the architecture of how to build and scale and manage all of their agents, maybe there's a considerable consideration to be made here. Alright. Well, for anybody who wants to get in touch with you and follow along with the work that you and your team are doing, I'll have you add your preferred contact information to the show notes. And as the final question, I'd like to get your perspective on what you see as being the biggest gaps in the tooling technology or human training that's available for AI systems today. This is such a deep topic.
Right? It's it's it's very difficult to give a quick answer to this. But the way I'm seeing this is we need a holistic shift to solve for this entire problem statement. I literally talk to teenagers, college kids, and folks are asking what should we be thinking about? What's the next thing that we should be doing? What's the discipline that I should go study in? Should I still learn coding? Should I not learn coding? How do I upskill myself in this entire world of AI?
And it's such a deep question that's there in everyone's mind that I feel like a tactical solution around how you learn AI or how you learn this latest new tool or any framework out there is just not an answer to the question. I think the only way to meaningfully solve this is what are some of the more fundamental layers where we expect human beings to contribute and add value. And what are the fundamental concepts, aspects in which we expect AI to assist us? We need to think about that. And that needs to get inculcated into our education system more fundamentally as opposed to solving this problem superficially.
So I don't have a great answer to this question, but it's a problem that I resonate with very, very deeply through many of the latest conversations. Absolutely. Yeah. It's definitely a big a big thorny issue that the entire world is having to tackle right now. Even just in my own small ecosystem of working in an engineering team, it requires a lot of reimagining how you work every day because you can't human review everything the agent ships because otherwise, it'll actually never get to production. So it requires rethinking and reimagining entire processes and entire workflows, and it'll be an interesting few years to come.
Absolutely. Well, Tobias, I really enjoyed the conversation. Thank you so much for having me on the show here. Yeah. Thank you very much for taking the time today to join me and share the work that you and your team are doing on TruFoundry and TruForge and sharing some of your lessons learned on harness engineering and AI gateways. It's definitely a very interesting space and very relevant topic, so I appreciate all of the time and effort you're putting into that. I hope you enjoy the rest of your day. Thank you so much. Take care.
Thank you for listening. Don't forget to check out our other shows. The Data Engineering Podcast covers the latest on modern data management, and podcast. In it covers the Python language, its community, and the innovative ways it is being used. Visit the site to subscribe to the show, sign up for the mailing list, read the show notes. And if you've learned something or tried out a project from the show, then tell us about it. Email hosts@aiengineeringpodcast.com with your story.
Transcript supplied by the publisher with the episode.
by Tobias Macey · English · Tech & Science
This show is your guidebook to building scalable and maintainable AI systems. You will learn how to architect AI applications, apply AI to your work, and the considerations involved in building or customizing new models. Everything that you need to know to deliver real impact and value with…
E80 · 7 Oct 2026 · 1 hr 3 min
Summary In this episode Ofer Mendelevitch shares what it really takes to evaluate RAG systems in production when your source data is incomplete, constantly changing, or difficult to validate against a clean ground truth. He explores how RAG has evolved from simple “chat with your PDF” demos into enterprise-grade retrieval systems that require robust ingestion pipelines, hybrid search, reranking, multimodal support, access controls, and refresh strategies for large and dynamic document collections. Ofer also explains why evaluation becomes one of the hardest parts of the stack, particularly…
E78 · 25 Feb 2026 · 1 hr 1 min
Summary In this episode of the AI Engineering Podcast, Steven Watt, leader of the Office of the CTO at Red Hat, discusses practical paths to achieving AI sovereignty for organizations. He shares his two-decade experience in AI, highlighting how governments are building GPU platforms and protected data hubs to maintain control over AI workloads. Steve emphasizes why self-managed infrastructure is becoming a strategic necessity as companies outgrow cloud costs and require tighter control over models, data, and compliance. The conversation explores the operational substrate for AI sovereignty,…
E77 · 15 Feb 2026 · 51 min
Summary In this episode of the AI Engineering Podcast, Aman Agarwal, creator of OpenLit, discusses the operational foundations required to run LLM-powered applications in production. He highlights common early blind spots teams face, including opaque model behavior, runaway token costs, and brittle prompt management, emphasizing that strong observability and cost tracking must be established before an MVP ships. Aman explains how OpenLit leverages OpenTelemetry for vendor-neutral tracing across models, tools, and data stores, and introduces features such as prompt and secret management with…
E76 · 8 Feb 2026 · 59 min
Summary In this episode of the AI Engineering Podcast, Carter Huffman, co-founder and CTO of Modulate, discusses the engineering behind low-latency, high-accuracy Voice AI. He explains why voice is a uniquely challenging modality due to its rich non-textual signals like tone, emotion, and context, and how simple speech-to-text-to-speech pipelines can't capture the necessary nuance. Carter introduces Modulate's Ensemble Listening Model (ELM) architecture, which uses dynamic routing and cost-based optimization to achieve scalability and precision in various audio environments. He covera topics…
E75 · 27 Jan 2026 · 46 min
Summary In this episode I sit down with Hugo Shi, co-founder and CTO of Saturn Cloud, to map the strategic realities of sourcing and operating GPUs across clouds. Hugo breaks down today’s provider landscape—from hyperscalers to full-service GPU clouds, bare metal/concierge providers, and emerging GPU aggregators—and how to choose among them based on security posture, managed services, and cost. We explore practical layers of capability (compute, orchestration with Kubernetes/Slurm, storage, networking, and managed services), the trade-offs of portability on “Kubernetes-native” stacks, and…
E74 · 20 Jan 2026 · 56 min
Summary In this episode of the AI Engineering Podcast Niklas Gustavsson, Chief Architect at Spotify, talks about scaling AI across engineering and product. He explores how Spotify's highly distributed architecture was built to support rapid adoption of coding agents like Copilot, Cursor, and Claude Code, enabled by standardization and Backstage. The conversation covers the tension between bottoms-up experimentation and platform standardization, and how Spotify is moving toward monorepos and fleet management. Niklas discusses the emergence of "fleet-wide agents" that can execute complex code…
E73 · 5 Jan 2026 · 56 min
Summary In this episode Joe Devon, co-founder of Global Accessibility Awareness Day (GAAD), talks about how generative AI can both help and harm digital accessibility — and what it will take to tilt the balance toward inclusion. Joe shares his personal motivation for the work, real-world stakes for disabled users across web, mobile, and developer tooling, and compelling stories that illustrate why accessible design is a human-rights issue as much as a compliance checkbox. He digs into AI’s current and future roles: from improving caption quality and auto-generating audio descriptions to…
E72 · 29 Dec 2025 · 54 min
Summary In this episode product and engineering leader Preeti Shukla explores how and when to add agentic capabilities to SaaS platforms. She digs into the operational realities that AI agents must meet inside multi-tenant software: latency, cost control, data privacy, tenant isolation, RBAC, and auditability. Preeti outlines practical frameworks for selecting models and providers, when to self-host, and how to route capabilities across frontier and cheaper models. She discusses graduated autonomy, starting with internal adoption and low-risk use cases before moving to customer-facing…
E71 · 16 Dec 2025 · 1 hr 8 min
Summary In this episode Craig McLuckie, co-creator of Kubernetes and founder/CEO of Stacklok, talks about how to improve security and reliability for AI agents using curated, optimized deployments of the Model Context Protocol (MCP). Craig explains why MCP is emerging as the API layer for AI‑native applications, how to balance short‑term productivity with long‑term platform thinking, and why great tools plus frontier models still drive the best outcomes. He digs into common adoption pitfalls (tool pollution, insecure NPX installs, scattered credentials), the necessity of continuous evals for…
E70 · 24 Nov 2025 · 1 hr
Summary In this episode Max Beauchemin explores how multiplayer, multi‑agent engineering is reshaping individual and team velocity for building data and AI systems. Max shares his journey from Airflow and Superset to going all‑in on AI coding agents, describing a pragmatic “AI‑first reflex” for nearly every task and the emerging role of humans as orchestrators of agents. He digs into shifting bottlenecks — code review, QA, async coordination — and how better DevX/AIX, just‑in‑time context via tools, and structured "context as code" can keep pace with agent‑accelerated execution. He then…