Episode 560 · Talk Python To Me
#560: Building a Research OS: From Django to 30,000 Samples
26 Aug 2026 · 1 hr 3 min
Episode 560 · Talk Python To Me
26 Aug 2026 · 1 hr 3 min
In 2020, a gastroenterologist in Glasgow did the math on his new research study and came up with 30,000 samples, arriving over two years from three cities and a dozen hospitals. He asked around about how researchers keep track of that. The answer was Microsoft Excel. Shaun Chuah had written some HTML by hand in Notepad back in high school and that was about the whole of his programming experience, so he opened the Django tutorial and started reading. Six years later that app is Foundry120, holding 10 terabytes of clinical and genomics data with an agentic AI running on top of it.
Michael Kennedy:In 2020, a gastroenterologist in Glasgow did the math on his new research study and came up with 30,000 samples arriving over two years from three cities and dozens of hospitals. He asked around about how researchers kept track of that. The answer was Microsoft Excel. Sean Chua had written some HTML by hand in Notepad way back in high school. That was about the entirety of his programming experience. Even so, he opened the Django tutorial and started reading.
Michael Kennedy:Six years later, that app is Foundry 120, holding 10 terabytes of clinical and genomics data with agentic AI running on top of it. This is Talk Python To Me, episode 560, recorded August 26th,
Music:2026. Talk Python To Me. Yeah, we ready to roll. Upgrading the code. No fear of getting old. Async in the air. New frameworks in sight. Geeky rap on deck. Quartz crew. It's time to unite. We started in Pyramid, cruising old school lanes, had that stable base, yes sir.
Michael Kennedy:Welcome to Talk Python To Me, the number one Python podcast for developers and data scientists. This is your host, Michael Kennedy. I'm a PSF fellow who's been coding for over 25 years. Let's connect on social media. You'll find me and Talk Python on Mastodon, Bluesky, and X. The social links are all in your show notes. You can find over 10 years of past episodes at talkpython.fm. And if you want to be part of the show, you can join our recording live streams.
Michael Kennedy:That's right. We live stream the raw uncut version of each episode on YouTube. Just visit talkpython.fm/youtube to see the schedule of upcoming events. Be sure to subscribe there and press the bell so you'll get notified anytime we're recording. Let me quickly tell you about a new course we have running over at Talk Python. Up and running with Rust. The tools reshaping how you write Python, Ruff, uv, Firefly, ty, Pixie, and others. and the libraries pushing past its performance ceilings.
Michael Kennedy:Polars Pydantic, Cryptography, and Granian increasingly share one secret under the hood. They're written in Rust. That code that used to be written in C when Python needed real speed now is more and more being written in Rust. So having some Rust proficiency is a great skill as a Python developer. Christopher Trudeau is back with the Up and Running with Rust course. Check it out over at talkpython.fm. Just click it courses in the top and you'll find it right there.
Michael Kennedy:I hope you love this new Rust course. Getting a course at Talk Python is one of the best ways to support the show. This episode is brought to you by Sentry. You know Sentry for the error monitoring, but they now have logs too. And with Sentry, your logs become way more usable, interleaving into your error reports to enhance debugging and understanding. Get started today at talkpython.fm/sentry. And it's also brought to you by Talk Python Courses.
Michael Kennedy:Course completion certificates are now live. If you finished a course, there's a certificate waiting for you on your account page right now. Download it as a PDF or add it to your LinkedIn profile with one click under licenses and certifications. Same section as your formal degrees. Visit training.Talk Python.fm account to see what you've already earned. John, welcome to Talk Python To Me. How are you doing? Great.
Shaun Chuah:Thanks very much, Michael. A big fan of the show. Keen to be here to share a few things of what we've learned.
Michael Kennedy:Amazing. Thank you so much. And now you're here creating the show. It's going to be amazing. And you're doing really interesting work with biological and medical research, Python, Django, AI. I think there's a lot of cool things that we're going to dive into here. So I know even for people who are not in medical research, I think there's going to be some super interesting angles that will translate for them.
Shaun Chuah:Fantastic. Looking forward to it.
Michael Kennedy:Now, before we dive into that, as usual, just tell people about yourself. Who are you?
Shaun Chuah:So I'm Sean. I'm a gastroenterologist, and I'm a clinical researcher as well. So my primary area of practice is in the field of inflammatory bowel disease that comprises ulcerative colitis and Crohn's disease as the kind of conditions that we deal with when we see patients in clinic. But in parallel, I also work in research, trying to find out the causes of these conditions together with the GUT Translational Research Group here at the University of Glasgow.
Shaun Chuah:Personally, for this podcast, we use Python a lot in our work, both for research, so trying to analyze the data that we're doing, but also as well as for the infrastructure that we're running. So how do we actually deliver the research studies that we have to deliver? And personally, as a researcher, I've also got an interest in agentic AI and machine learning, and that's what attracted me to try to use some of these techniques to see if we can find new breakthroughs that can bring, you know, new treatments for our patients that we see in clinic
Michael Kennedy:every day. Awesome. Very cool research. That's the kind of stuff that can make a big difference for people's lives. If you can make them better, right? It's like, I couldn't leave the house, but now I'm back with my friends or whatever, right? Yeah. I mean, you know, the work that we
Shaun Chuah:do is motivated in the patients that we see day to day. Inflammatory bowel disease is not very common. It is rare, but it is growing not just in developed countries, but also in the developing world. And we see a great burden of this disease coming the next 10 to 20 years. What do you think?
Michael Kennedy:Diet? Environment? Why is it growing? Obviously, that's the research a little bit, right? But what is your ideas here? Well, these are really complex conditions, and the environment is
Shaun Chuah:absolutely a major player in this. The diet that we're eating nowadays is different from what we used to eat. But when we dive into the science of it, it's a lot more complex than just that. There are some patients who have genetic susceptibility, although not all. And we know that there is this problem that happens when the immune system reacts to our environment. And for some reason, the alarms don't switch off and the inflammation keeps happening, which leads to the
Michael Kennedy:symptoms and the problems that we see every day in clinic. So let's start by examining all the the research and the work you did and maybe the 2020 version of your research projects. It all started with Django, right?
Shaun Chuah:Yes. So back in 2020, I was still training as a gastroenterologist. So over here, we do a period of registrar training, we call it. And as part of my developing interest in understanding these medical conditions, I sign up to do a research study with one of the principal investigators. And the proposed study was basically to recruit quite a few patients from three different cities, multiple hospitals, and collect extra samples that we can run science experiments on to work out what's going on with an immune system. So when we first started, that was 2020. And I know we all remember that in 2020, you know, we had COVID and COVID had a significant impact both on hospital operations, but also on research operations. And when I sat down to look at what we were supposed to do in terms of the research studies, the first challenge that I faced was calculating the number of samples we were going to do. So we were going to recruit about 200 people and we're going to follow them up every three months. And at every time point, we would take extra blood samples, extra stool samples, and saliva samples. So that was a lot
Shaun Chuah:of extra samples and a simple back of the envelope calculation came to about 30,000 samples that we were going to generate over a couple of years. Now, when I first saw that challenge, I started asking everybody around me, how are we going to keep track of all that? How are we actually going to deliver that problem? How do we solve that problem? And from my chat with a lot of people, the common way that researchers track the samples is to use Microsoft Excel.
Michael Kennedy:That's exactly what I was thinking is like, how big is the Excel file? And, you know, how many little worksheet tabs does it have?
Shaun Chuah:Yeah, I mean, most research teams do not really have a lab system because lab systems, you know, enterprise software, which is really expensive to procure, it'll take you like six months to set it up. And it's usually designed really for hospital operations where you're taking, you know, millions of samples. And it's not really designed for a one-off research study across a lot of different spaces. So then I started looking into what are the, you know, are there off the shelf options or stuff that we could use to just get up and running?
Shaun Chuah:But after looking at a few different options, I think I had to bite the bullet and decided that, you know, the best thing was to go and write a web application with Django.
Michael Kennedy:What was your programming experience at this point? How good of a programmer were you?
Shaun Chuah:So I wasn't very good of a programmer. I mean, what I did was basically, you know, write some websites in high school. And back in those days, we used to use Notepad. We used to open the brackets of HTML manually by hand.
Michael Kennedy:It was rough. Those were tough times, I remember.
Shaun Chuah:Yeah, I don't know how many people listening to this podcast can relate. But we used to write, you know, the head, the body and stuff. And then later on, I did use a bit of programming to do some, you know, statistics, a bit of, writing an R script or a Python script just to generate a graph. So that's probably that level of programming experience when I was looking at how do we actually write a web app with Django.
Michael Kennedy:Yeah, okay. So a Django web app was kind of a
Shaun Chuah:pretty big stretch at that point, right? Yes, but the Django...
Michael Kennedy:How'd you get started? I mean, this predates all the AI, agentic stuff. You had to earn it for this one. Yeah, definitely. I mean, when I first
Shaun Chuah:visited a Django website. It's the web framework for perfectionists with deadlines. And I think that described exactly what a situation I was in. You know, I came out program and we had about a two year time period to deliver a study. And you can't spend forever trying to get this up and running. You just have to get the study going. So the Django tutorial was where I started. But actually, I just want to take this opportunity to thank a lot of people in the community and those listening because it's the documentation, the YouTube tutorial, your podcast. I've been listening for five years. Books that were written, Two Scoops of Django, famously. I read a ton of books just
Michael Kennedy:to be able to get things up and running. Yeah, the books are great. And honestly, YouTube is really good. People knock YouTube for many reasons. You know, it's one of the ills of social media that's like scrambling their brains. But if you use YouTube for education, it's an incredible resource.
Shaun Chuah:Yeah. I mean, there are a lot of creators on YouTube that are just, you know, teaching lots of different things. And it's so helpful. But I also think some of the textbooks are great, you know, for things like test-driven development. You know, back in the days, I think you would only read about it in textbooks rather than on YouTube videos because it's not really an attractive topic.
Michael Kennedy:Yeah, you mentioned Obey the Testing Goat, which is a really fun book. And you also talked about using designing data-intensive applications. And that's by Martin Klepman. I haven't heard of that That's a pretty interesting one that people might want to check out.
Shaun Chuah:Yeah, down the line, as our studies matured and we started dealing with the data, I started reading some of the other data textbooks like Data Warehouse Toolkit by Kimball. I think that's the one that all data engineers used to talk about dimensional modeling and all that kind of stuff. So lots of textbooks. I think they're really helpful because they give you a kind of a holistic, fundamental approach to learning the basics and making sure you don't have gaps in what you're doing to build some of these things.
Michael Kennedy:When you come from the background of programming that you did, which, by the way, that's a very similar story to me as well. I studied math, and then I got involved helping people with research projects and working for a scientific research and visualization company. So you kind of build up the pieces as you go, and it's super fun. It's really neat. But there's a lot of gaps in your sort of data integrity, programming, like what is the foreign key relationship?
Michael Kennedy:All that. Yeah, yeah. So I think books are really important, especially for self-taught people like us. Definitely, definitely.
Shaun Chuah:You do learn a lot as you go. And it's quite important to be aware of what you don't know and try and close those gaps so that the applications that you write end up being reliable enough for people to depend on.
Michael Kennedy:Let's talk about Django for a second because you shouted out, Django, some of the features of it that kind of made that possible. I think this is one of the reasons that people choose Django, especially when they're getting started. You know, it's like it's got the built-in migrations. It's got the automatic database management, the admin. So what was it about Django that drew you to this?
Shaun Chuah:So I think, you know, when we looked at the landscape back in 2020, Django definitely stood out because it was a mature project. It's been around for a long time and it's batteries included, which is really important because, you know, I'm conscious that, you know, as a non-programmer coming into trying to write these things, You want your authentication framework to be battle tested. Database and migration is incredibly important. You can't lose any data.
Shaun Chuah:And I think just on those two features alone, we haven't even come to the admin dashboard, but just those two features alone, I think was enough to persuade me that Django is the right framework for the problem that we're dealing with.
Michael Kennedy:Okay, very interesting. And by the way, just big news for Django folks. Is it on their blog, maybe? Big news for Django is they just announced that they're switching the Django long-term support to yearly releases, just like Python itself, and calendar version. So there'll be Django 2028, Django 2029, and every single release is a long-term support, whereas now you're kind of juggling like, well, 5.2 is a long-term, but the one before it. I think that'll also make it a little simpler for people who are coming, like they don't have to decipher the versioning and what it means.
Michael Kennedy:Definitely, yep. Yeah, yeah, that'd be cool. What else? What else did you run into trying to navigate this world of becoming, running one of these like real web apps rather than just Excel or buying some off the shelf, you know, square peg round hole type of system that doesn't really fit what you're trying to do?
Shaun Chuah:Yeah, I mean, I think one of the steep learning curves is how do you go from a local host development project into something that's in production? Because production is obviously a whole different ballgame. you need to ensure your infrastructure in production is reliable. It's secure. It's accessible. You've got backups running, all sorts of things. And for us, one of the key things that enabled us is continuous integration and continuous deployment practices.
Shaun Chuah:So using GitHub Actions to deploy continuously so we can fix problems, get it into production in minutes rather than trying to do any of these things by hand manually.
Michael Kennedy:Yeah, the CI-CD stuff, I think, is really valuable. It's valuable for big teams, but it's also really valuable for people who are really new. Because if all I have to do is save my work to GitHub, and now that's the new stuff that's out there. That takes a lot of the complexity out. People forget just how intimidating logging into a Linux computer is. Oh, yeah. Now what? Yeah?
Shaun Chuah:Yeah, I had to learn how to use Linux on a terminal, right?
Michael Kennedy:Exactly. It was like, where is the UI? This is very different. How am I supposed to accomplish anything with just the terminal? And, you know, you get used to it and it's amazing in its own special way, but it doesn't feel amazing at the beginning most of the time, I think.
Shaun Chuah:Well, I'm reminded of the... It feels intimidating. Yeah. Every time I teach somebody how to use the Linux or the terminal, it's a big jump, I think, for people who are not familiar with CLI, which is, I think, the vast majority of people in the population.
Michael Kennedy:Yeah. Well, and things like people being primarily on phones for their computing device don't make that easier. It's only harder, right? This portion of Talk Python To Me is brought to you by Sentry. You know Sentry for their great error monitoring, but let's talk about logs. Logs are messy. Trying to grep through them and line them up with traces and dashboards just to understand one issue isn't easy. Did you know that Sentry has logs too? And your logs just became way more usable. Sentry's logs are trace connected and structured, so you can follow the request flow and filter by what matters.
Michael Kennedy:And because Sentry surfaces the context right where you're debugging, the trace, relevant logs, the error, and even the session replay all land in one timeline. No timestamp matching, no tool hopping. From front end to mobile to backend, whatever you're debugging, Sentry gives you the context you need so you can fix the problem and move on. More than 4.5 million developers use Sentry, including teams at Anthropic and Disney+. Get started with Sentry logs and error monitoring today at talkpython.fm/sentry. Be sure to use our code talkpython26. The link is in your podcast player show notes. Thank you to Sentry for supporting the show. Let's talk about your project Boundary 120. So when you started, you built this bespoke Django app for, was this for the music study or which study was this for? This was for the music IBD study. And the first version was
Shaun Chuah:really to track samples and to come up with an efficient way for our teams to handle the, you know, the 30,000 that we are projecting to sort out. So one of the key functionalities that we wrote in at the start was this ability to, you know, stick a QR code, stick it onto a sample label, scan it in, and that will register it into a database. And then we could move the sample around between sites and scan it and update. So that's where it started out as a sample tracking platform. So nothing too amazing. It was just kind of a really basic operational problem that just needed to be solved. But over time, when you get samples, the next thing that happens to the sample is that it gets processed into an experimental pipeline and you get data out of it. So the lifecycle of an experimental sample then leads you on to dealing with the data that comes back.
Shaun Chuah:Now, because we were taking so many different types of samples, one of the key challenges that faced our team was how do you handle the data that was coming back at us? Because we have stool samples that becomes, you know, microbiome data. We have blood samples that become genomics data. And every data type that comes back just looks a little bit different, comes back in different formats, you know. And one of the difficulties as well is the size of the data that we get.
Shaun Chuah:So for my personal research project during that time, I was looking at sequencing out the cell-free DNA. So the DNA fragments that you have floating around your blood and each sequencing file for a participant will be about five to 10 gigabytes of data. And it comes back in a, you know, in a fast QGZ format, and that has to go into a bioinformatic pipeline. So it's the one of the key challenges is how we build that data. So Foundry 120 actually became a project that started from sample tracking into data management and handling. And in 20, you know, in the last year too, it's also become a platform where we can apply agentic AI into it to kind of accelerate
Michael Kennedy:our workflows. Yeah, it's a really neat platform. And I'm definitely going to dive into it. I think it's neat how you had this smaller bespoke thing and you're like, all right, let's sort of pull out the essence of it and make it useful for all kinds of scientific research. Going back to the samples. You said each one is five to six gigabytes. That's each of the 30,000?
Shaun Chuah:Not all the 30,000. So a subset of them, but it depends on what essays were being run on which samples. And there could be a ton of essays all running different types of samples. And each of them will have their own data type that comes back. So for my cell-free DNA work, we ran it on a subset of the population, not the entire population. But for some of the other larger scale things like genotyping, we would genotype the whole population. So there are lots of moving parts and tiny little details here that cause so much operational problems.
Michael Kennedy:That's a lot of data. Where did you, how much did you end up with at the end and how did you store it and manage it?
Shaun Chuah:So currently we have about 10 terabytes of data and we've only, you know, processed only a fraction of those 30,000 is still being processed because some pipelines takes a human 12 hours to process a couple of samples. So there's still a lot of work to be done. But yeah, it's about 10 terabytes of data that we're handling. And initially, it would be scattered across our group. Somebody's doing a microbiome project, it would be on their laptop. Otherwise, it'd be on a shared drive, on a file share somewhere in the university server. So data tends to be scattered. So if somebody is doing a specific experiment, they might create an Excel file with the readings from certain essays, and that would be the data set. And that's, I think,
Michael Kennedy:one of the big problems that we face in biomedical research today. Yeah, that's a lot of data. Are using things like blob storage? I know this Foundry 120 project is pretty strongly based in the Azure cloud. Are you using Azure blob storage and stuff like that? Or is it really just all on on-premise with regard to the university?
Shaun Chuah:So back in the days, that was the situation that we were in. But today, we're now migrated everything into Azure Blob Storage. That just gives us a scalable way of not worrying, you know, how much space we had in our shared drives. Because back in the days, you know, the university would give you a quota on your shared drive and you have to email somebody in IT to get it increased. So you might have a terabyte cap on your file share. But with Azure Blob Storage, what you end up having is just a bill at the end of the month rather than hard limits.
Michael Kennedy:You have a bill instead of a limit. Yeah, yeah, yeah. Interesting. When I was working way back when I was in college, in university, working in this math research lab. I told this story a couple of times, but it's been a while, so maybe I'll share it again. And the whole of the math research department got access to the Silicon Graphics mainframe beast of a computer. We all had shared sort of workstation access to it. And one of the students, grad students, was having a problem with their code.
Michael Kennedy:And so they started logging out what it was doing. And they got it into an infinite loop and ran it. You would run stuff overnight and come see it in the morning. We came back one day and it just wouldn't turn on or it wouldn't respond. And nobody could figure out why. The students, we had no, the reason I'm telling you this, we had no quotas, no limits. The student used up the entire hard drive of the Silicon Graphics machine to the very last byte, and apparently it needed a few temp files to operate the operating system.
Michael Kennedy:And it just died.
Shaun Chuah:And nobody could get it to come, it took a day or two for it to come back because somebody
Michael Kennedy:destroyed it. Well, these limits, they have a reason they're there, you know?
Shaun Chuah:They do. But when we look at the next five to 10 years of biomedical research, file storage is actually a big problem. Is it? Okay. Some of the newer technologies are generating like a terabyte of data for a single sample.
Michael Kennedy:Wow. Yeah, that's absolutely crazy. You know, this also goes to reproducibility and long-term viability of this research, right? Because if you have 100 terabytes of data, it's one thing to say, well, here's the Excel file and here's the Parquet file and here's the Docker image that runs it. Save that as a, you know, you can stamp these as like, here's what the paper was published on, right? And here's the digital assets. But when it's 100 terabytes, it doesn't matter if you can name it or not.
Michael Kennedy:That's a hard thing to store.
Shaun Chuah:Yes, definitely a big problem. And I think the question then becomes what is the most valuable data that you store? because perhaps you don't need to store the raw files and you want to process files because who wants to run another 100 terabyte pipeline anyway? But yeah, these are challenging questions. There are no easy answers. And most groups, we do cost our data storage, but for most grants, we cost it out for 10 years. And beyond that, it becomes a challenge to maintain all this data.
Michael Kennedy:Yeah. What are you going to do? I guess the only bonus is in general that storage is getting cheaper by a lot. And I say by in general, because the last couple of years or last year and a half, that's not true. Yeah, I was going to say the AI revolution
Shaun Chuah:might not keep that falling price curve the same.
Michael Kennedy:It might change things just a little bit. All right, so let's talk about this Foundry project. So this is sort of the next generation of what you maybe dreamed of building. Also maybe a little bit in the agentic age as well, right? Yeah, definitely. Okay, so tell us about this Foundry 120. So Foundry 120. People will find it at foundry120.com.
Shaun Chuah:It's a research operating system for translational science teams. So I've got to explain that a bit. So translational science is basically when you're trying to discover things by recruiting humans with a problem and you're taking samples, bringing them back to the lab and running all sorts of experiments to just kind of understand the biology behind disease and then find out whether there are mechanisms that we can target for therapeutic development and things like that.
Shaun Chuah:So that's kind of translational science. And Foundry 120 is a platform that allows teams to operate seamlessly and use AI to accelerate their workflows. It's built around three ideas. The first idea is to organize everything. So you organize your samples, your studies, your participants into a database that just helps you keep track of everything. And then the second bit is the idea of centralizing all your data outputs into a single platform. So this will be things like clinical data, including radiology images, endoscopy videos, digital pathology slides, and then the scientific outputs such as spatial transcriptomics, genomics, microbiome, or all the data volume that we're seeing from all the different scientific modalities are just going up exponentially. But as we know in the AIH, if you can bring all your context into a single platform, then you can let AI connect to that and operate on that to answer questions that researchers might have. So, I mean, I'll give you a concrete example. For example, you know, if one of my scientific colleagues were looking for, you know, do we have any samples that
Shaun Chuah:belong to a participant who's been treated with this drug? Can we find that? And historically, what you would have to do is to take the clinical data set, join it with your sample database, and then filter that through and find the samples that you want. And that usually would have taken a couple of days to do. But this is a great use case, I think, for agentic AI, where you can just, you know, access the data, put it into a sandbox, write the code to do the join, and then just give you the answer that you want. So just eliminate all the steps in between and get you to the answer
Michael Kennedy:faster. And that's the whole concept behind the platform. Sounds great. It's a really nice looking web app too. What's its current status? It's available not just for you all, but for others, but it doesn't have quite a just create an account, get started. So you have a request a platform walkthrough. What's the situation here? Can other research teams use it? Is it a paid product? Is it just sort of gated because it's not ready for people overwhelming it? So it's quite early stage
Shaun Chuah:at the moment. So over the last six to 12 months, we've made the entire foundation generalizable beyond just our group and beyond just disease. So what I mean by that is, you know, for another group to join, they would need the access controls to be in place. So they need a way of managing their team, being able to assign permissions to individuals, you know, without me having to do it But I think more importantly is that onboarding a team into this platform is actually quite a hands-on process because it really depends on what data you're handling, the volumes that you're dealing with, what modalities of data you need to do.
Shaun Chuah:And also the other major thing that we have to look at if you want to join the platform is the governance around your data. So in clinical research, there are strict rules around how you handle data and where that data can sit, where can it be processed. And these are subject to your research ethics approval. So before a team can join us, we do want to review all those things to make sure that we're in compliance before anybody can join. So that's the reason why it's not just a simple, you know, I can sign up, pay a monthly fee and start using the platform.
Shaun Chuah:There's a lot more issues that have to be looked at before we can bring a team on board. But we're willing to do the work to kind of review the situation and start getting other people access to it.
Michael Kennedy:Makes a lot of sense. You've got IRB research stuff. You've got HIPAA and the equivalent in all the other countries. And yeah, so you don't want to get in trouble.
Shaun Chuah:Yes. And you did mention earlier that this is quite tied into Azure at the moment. And that's because we run on the University of Glasgow's Azure Tenancy. So everything has to be located specifically for our studies within the UK data centers and things. So we have all the controls in place to do that. And that's backend infrastructure for the people who are listening on the podcast.
Michael Kennedy:This portion of Talk Python is brought to you by Talk Python Courses. Here's the thing that always bug me. You finish one of our courses, that's hours of video, a pile of code you actually wrote, and real skills you didn't have a month before, and then nothing happens. No paper, no credential, nothing to show for it. So we fixed it. Every Talk Python course now generates a completion certificate automatically. Go to your account page in your dashboard section, scroll down to your completed courses, and click Certificate.
Michael Kennedy:That's the whole process. Two things you can do with these course completion certificates. download the full PDF, which is handy if your employer reimburses training or gives you credit for finishing it. Or you can make the certificate public and hit share on LinkedIn, which adds it to your LinkedIn profile under licenses and certifications. Not a poster that scrolls away in a day, an actual credential sitting on your profile where your manager and recruiters can see it. Plus, if you've been taking our courses for a while, you've probably earned several of these without even knowing they existed. Just visit training.talkpython.fm/account and collect them. Thanks to all of you who have taken a Talk Python course. It's a great way to support the podcast.
Michael Kennedy:What a different time it is for universities. And I think they're going through a similar upheaval, I guess is the word. When I was in university, and we're talking like 90s, there was a giant Cray supercomputer or maybe a couple of supercomputers at the heart of the university in some basement, and you could get access to that if you needed mega computing resources. And now it's just scaled into some ginormous cloud infrastructure. And the limit is really just how much are they going to give you, not what does the university have in its basement or something like that, which is really interesting. Well, I mean, there are still some
Shaun Chuah:universities that are trying to build their own clusters. But I think if you look at the economics of it, it is way more cost efficient to be running on the cloud compared to on the cluster, because many of the analysis that we do in the universities once off. So you might process a ton of data, but that might be after a year or two years before you get that amount of data to process. So if you look at a workload, the IT workload within the university, I think the cloud is quite an attractive proposition. That's an interesting angle. Yeah, of course, because
Michael Kennedy:you're going to process your samples. Maybe you're like, ah, we really want to try a different algorithm and you'll run it all through again. But generally speaking, it's not steady state. Whereas if you're going to go buy your own hardware and then put it in a basement and it's kind of a steady state situation, then you could really predict, you know, we're saving, you know, 50% less expensive to run on this. And just, yeah, didn't really think about that.
Michael Kennedy:And so I said, there's another such wave coming because this was, you know, when was the cloud? The cloud is 10 years ago, 2008, I think. A little bit, a little bit, really caught going a little bit after that. But it's been a little while. Now we've got all this AI stuff going and we're proving, you know, unsolved airdosh mathematical problems and all sorts of things, you know, with AI and other, you know, trying to do protein folding and other things that are really computational, but also really different.
Michael Kennedy:And so I feel like this whole AI wave, especially the agentic AI, is going to roil what's happening at the universities all over again.
Shaun Chuah:I mean, if you look at the data center economics around AI, it's quite interesting, actually. Like a single rack from Nvidia is like a couple million dollars, right? It's crazy. And you need a power plant to plug it into. I can't see any university building their own AI data center in the future. So I think cloud adoption is inevitable in the sense that the economics of trying to use agentic AI doesn't make sense to try and build your own AI computer, unless you have specific governance requirements, which do apply to certain studies and certain aspects of work.
Shaun Chuah:But even then, you might contract a local AI company to do it and then share it between businesses and universities rather than having a university do it on its own.
Michael Kennedy:Yeah. People talk about the AI bubble. I don't know if there's actually an AI bubble. This stuff is so productive. The last bubble we had that was tech-related was the dot-com bubble. And there was really stupid stuff with a lot of money spent on it you know like there's always the pest.com or the weird investments you know thing and people are spending millions of dollars on ads to just like have dancing monkeys running it was a weird time and this is also a weird time but I feel like there's actually something legitimately at the core of it that is really changing the way people work and and do research and so on so I don't know it's going to bust if there is a bubble and I don't know if the bubble is going to burst. But if it does, I think that's also going to create an explosion of local AI. You know, all of a sudden, right now, if you have to pay $100 for your Anthropic subscription, that's totally reasonable. But if that becomes a $2,000 a month bill, well, then buying a $10,000 workstation that can do legit local AI is all of a sudden a bargain, you know what I mean? So it could be in the future that there's sort of a coming back to
Shaun Chuah:local compute as well for AI, maybe. Yeah, I mean, the small models are getting more and more capable as well. So not every workload needs a frontier model at the moment. So what you're saying is probably true in the sense that there is definitely going to be a role for local AI.
Michael Kennedy:I think Gemma maybe works this way, but certainly I can also see a world where there's kind of an orchestration layer and then 100 or 50 specialists that get selected, right? Right now, if you ask a frontier models to do something on genomics, it uses the same giant model as it would to use to like write Shakespeare derived things. Right. But if you had one that just is trained on genomics, you could have a much smaller model that could run locally and you could say, okay, this part of the question goes to this, this sub model. And I don't know, it's going to be interesting where it goes. And that let's, so coming back to it, let's, let's talk a little bit about Foundry. There's a little demo you've got going right here at the beginning. Let me see if I can go to the front of it. And maybe, I guess there's the stuff you talked about before, the sample gathering and organizing and all that kind of stuff, right? And then on the back of that, once you get it all done, there's this local tool using AI that understands all of your research data. So maybe tell us about the data collection and the non-AI bit, and then it'll be fun to talk
Michael Kennedy:about the AI because it's kind of non-standard. It's kind of powerful. Yeah, let's talk about the
Shaun Chuah:data bit because that has been one of the most difficult problems I've been thinking about for the last few years is how do we aggregate all these types of data sets into some kind of a common model that an AI... Well, nowadays we can use AI, but it still applies to humans. So when my colleagues are looking for data and trying to match, you know, data set A to B, what is the way we do it? And at the end of the day, I think, you know, we've brought down everything to a model where all data exists in files. So even tabular data, we keep it in CSV file. Well, historically, we've tried to put, you know, tabular data into the database. But I think going forward, we're just going to keep the files because that's the universal denominator. So whether you're dealing with endoscopy videos, MRI images, digital pathology slides. Everything's a file, and each file has each set of, you know, each type of data has its own file structure, but the common language at the end of day is files. So we put files, and the thing that makes the files useful is putting a layer of metadata and relationship data on top of the file. So I think by coupling the two of them,
Shaun Chuah:we found this common model that can generalize across all the data sets that we use.
Michael Kennedy:Very interesting. And, you know, to your point of just keep it in the files, you know, with things like DuckDB and other really cool things and Parquet files, you can kind of treat them like databases already. So I don't know if people know, but with DuckDB, you can say things like select star from read Parquet input, you know, which is pretty insane.
Shaun Chuah:Yeah, Parquet is great. We use it as well.
Michael Kennedy:Yeah. So also things like, what do you think about little SQLite files or DuckDB file where there's these just embedded no server databases? And I think there's probably some really good use cases and research for that.
Shaun Chuah:Potentially, you know, SQLite's a very interesting concept of having an entire database in a file. But when I think of my end users, my colleagues who are running science experiments and such, they're used to dealing with Excel and CSV files. So we try to keep the same file format everybody's familiar with. But I think for some of the future work that we're going to do, we can use some of these more specific file formats. Because as you know, if everything's a file, then an AI agent can run across all the files and pull out the data that it needs for what it needs to do.
Michael Kennedy:Yeah, absolutely. Yeah, sure. If you're giving them the files directly, here's either your Excel workbook or here's your CSV file. But if it's coming in and out of a platform like Foundry 120, you can store it as one thing and then export it or import it as another, right?
Shaun Chuah:Yeah. I mean, so fundamentally, what we do is that the files sit in Azure Blob Storage. And when Helix, our AI agent, wants to process a file, it spins up a VM. The file gets transferred from the Azure storage into the VM. The AI sends its analysis code, whether it's in Python or any other language, into the VM. The computation happens, and then we return the output to Helix and store the output back in Azure storage. So by doing that, we put the sandbox guardrail around the AI agent for security purposes.
Shaun Chuah:but it also allows us to enable the AI to do processing, generate new files, store it. And I think that's the kind of model that works for us in Foundry while keeping it all within an institutional Azure environment and also allows us to enforce all the role-based access controls that we need for our team members. So when they query the AI, the AI can only see the same files that they would have normal access to. And I think that's the design that's very specific for this because of the governance requirements that we have around the data sets that we use.
Michael Kennedy:These AIs are sneaky. They will find a way to access the files. I mean, the really big headline cases are like, OpenAI was training its model and it hacked multiple systems so that it could get to the answers on Hugging Face instead of just figuring out the, solving the test, right? It's like a teenager that doesn't really care about the work is just doing it. But what I was thinking was, you know, I was working with Claude and I had some question about my code and it said something to the effect of like, oh yeah, you have two GitHub issues on this.
Michael Kennedy:I never gave it direct access to GitHub. I never gave it access to GitHub. I'm like, how does it know that? It's quoting like GitHub stuff, not through Git history locally, but it like reading the issues and the PRs. I'm like, how does it do? And then I realized I had the GitHub CLI installed and it's like, well, let me see if the GitHub CLI is installed. Oh, and look, it's already automatically authenticated because Michael logged in at some point to the CLI. And so it was just using the CLI that it's also, you know, like, oh, let me check that on the server for you. I'm like, excuse me. Yeah. Yeah. On your production server, you're doing this.
Michael Kennedy:How you're not supposed to be there. Why? And you know, it's just like, it's realized in the code somehow it's figured out that it can SSH. So you got to be really careful about those things. Right. Cause they're not malicious. They're just like, you asked me to solve a problem. And if I can get to that, I got a better answer, more concrete data. And, but it could also go and we fixed the problem by resetting the database. Like, oh, no, you didn't.
Shaun Chuah:Yeah, so for us...
Michael Kennedy:So how do you do that kind of stuff in your project?
Shaun Chuah:So it's the backend. So Django enforces the permissions. And actually, so when the AI stages data and stages code, that actually goes through Django first. So the AI is not calling directly into the files, is not calling directly into a compute environment. And then Django screens all that code and activates the VM. So we've actually put Django in as a security guard between your AI and the raw data and the compute environments
Michael Kennedy:that we run in. MARK MANDEL: Oh, very cool. So Foundry 120 is also written on Django, but it's more Django REST framework and TypeScript React. Is that the story?
Shaun Chuah:FRANCESC CAMPOY: Yep, that's right. So on the front end, it's- MARK MANDEL: Yeah, so tell us a bit about it. FRANCESC CAMPOY: So the old app that we used to run off was just pure Django. So we use Django templates to handle all the registration and the CRUD workflows. But as we move into this age of AI, AI, as you know, is pretty asynchronous. So every call you make takes, you know, sometimes it feels like forever to come back with a response. And when we think about the async nature of all the calls that have to be made to run an AI conversation or agent loop, TypeScript and JavaScript tends to come to be more suitable for that kind of application.
Shaun Chuah:So we run both. So we have TypeScript on the front end, Django on the back end. I think it's a great setup. It gives us access to the entire Python data science ecosystem, while also giving us all the TypeScript and JavaScript ecosystem for handling all these asynchronous work and the interactivity that we want on the front end when you start running AI applications.
Michael Kennedy:Makes sense. Now, what I'm about to ask you doesn't really make sense because of the AI angle and that kind of stuff. But if you think about scaling this out to other research projects and other groups, have you considered looking at things like PyOxid, Iodide, sorry, and things like JupyterLite for running some of that compute on the front end on people's browsers so that you don't have to basically pay the compute cost?
Shaun Chuah:So that's an interesting question because now you're asking me about the compute architecture that we have in the backend, which is actually pretty heavy. So we've mentioned that some of the files that we have might be gigabytes in size and the compute power you need to run genomics pipeline is pretty high. So for a concrete example, so in the backend of our Foundry system, we have access to about 350 CPUs on Azure. So if you ask Helix for a very heavy analysis, say on a big transcriptomic data set of something, through Django, we are able to orchestrate up a heavy compute job, which will then go off and run.
Shaun Chuah:It'll spin up as many CPUs as it needs to, to process the job, and then returns the output later on once it's all done. And that process can take half an hour, a couple of hours. So it might come back really late. And because of that model that we're running, actually, the kind of computational requirements we have is pretty high. And so we don't really want to run compute on people's laptops. We want to run it in the cloud to handle the data sets that we're dealing with.
Shaun Chuah:So that's a very specific design choice. And it's all to do with the kind of data that we're handling and the need to throw a lot of RAM and a lot of CPU at it.
Michael Kennedy:Sure. And if your individual files are five gigs, that's a lot of just bandwidth costs. So it's like every time you want to load something, you got to pull that five gigs out of the cloud, which has different costs and so on.
Shaun Chuah:Right. Well, yeah. And many of the biological data problems are parallel. Right. So you've got 200 participants each of five gigs. You might as well spin up 200 machines and process all of them in parallel. So all these problems that we have are very paralyzable. And the cloud is a great platform to do that in because you can do it on demand, on the fly, spin it up, finish processing and tear it all down. So minimal cost for, you know, maximum impact.
Shaun Chuah:That's what we're going for in the back end.
Michael Kennedy:Right. That's the bursting component that you talked about. I do think JupyterLite is pretty interesting with the local Piodide execution and all that kind of stuff. Just the fact that that's possible is it's pretty neat. But yeah, I can see that it really doesn't apply for what you're doing here.
Shaun Chuah:Well, I haven't talked to you about the front end of sample operations, right? Because we're running a sample collection in the hospitals across Scotland. And I must say the frontline IT infrastructure is not always the best. So by having all our compute power on the server side, we can guarantee a speedy experience for our teams working at the front end. So we're not depending on the front end's computational power.
Michael Kennedy:I'll tell you what, my experience looking over the shoulder at the software that doctors and nurses use, there's a lot of room for improving the user experience.
Shaun Chuah:Oh, definitely. I mean, part of the reason we went Django as well is because it was server-side. And, you know, on the front end, you know, we have some computers I've used in hospitals. They go back to 2010, right? You know, we're running on Intel chips like 20 years old.
Michael Kennedy:Oh, my goodness. Yep. And there's a lot of Cisco, remote, whatever there. So let's talk about the AI side now. So we talked about the data, the data handling, some of the tech behind Foundry 120. But I think one of the cornerstones is this Helix AI. Now, when people think there's a bit of a problem here, Sean, like people talk about AI, and there's two or three different things it could be, and they all use the same word. And they think they're talking about the same thing, but they're actually talking past each other.
Michael Kennedy:You know, like I asked ChatGPT for this math problem and it got it wrong. It's like, yeah, but we also built incredible software with this other thing that we also call AI. And, you know, it's always right because it writes Python to actually answer its questions. And yeah, it's just really interesting. So there's an AI that you've mentioned a couple of times in here that will help researchers ask questions, find data, look for trends and those kinds of things.
Michael Kennedy:And this, I think, you give me your thoughts on this, but my feeling is that it's a little bit like a Claude code or a codex. One of these sort of tool using self-correcting AIs, not just a chat LL.
Shaun Chuah:Yeah. So Helix is an agentic AI system. And I think the problem that you're describing is because most people's experience of AI is chatbots. You go to chatgpt.com, you ask a question, it gives you an answer. But actually the stuff that we're seeing that makes us think that AI might not be a bubble is all this agentic AI stuff that we're seeing. So Claude Code, codecs, agentic AI is a very different paradigm from chatbots, right? In agentic AI, the AI, you give the AI a task, it looks at its tool set, it looks at what you're trying to do, and then it goes away and works at it until it gives you an answer. And that's incredibly powerful.
Shaun Chuah:Whereas I think, you know, 95% of people's experience of AI is almost like a Google search. You go to ChatGPT and you ask, hey, what's the directions to this place or what's the recipe for that? And therefore, there's this huge gap in understanding of how powerful agentic AI systems can be. So Helix is really one of the, we think it's one of the first demonstrations of how you would apply agentic AI in the science world. And a lot of the scientists that I'm showing this system to, this is the first time that they're seeing an agentic AI system. So, you know.
Shaun Chuah:What's their reaction? I think everybody's quite excited. They're like, oh, that used to take me like months to do, or it took me a lot of emails to do. Just doing it in minutes right now, it's quite amazing. And historically, every team member that's joined our team, I've had to sit down with them and teach them how to use Python or R to get a graph out of Excel file or something. And now you can just hand it off to Helix, let it do it, let it write the code.
Shaun Chuah:And then what you do in this situation is that you've got to verify that it's correct. So the work changes. So instead of you as a researcher writing the analysis code, You get the AI to do the analysis for you. And then what you have to do is verify that it is correct. So it's a very different way of working. But it's incredibly powerful. It's much faster. It's taken a lot of road work out of everybody's life. So I think everybody's really excited about it.
Michael Kennedy:I would imagine. I'm going to have you talk us through just this sort of workflow that it goes through real quick. But that's the big danger is that it just hallucinates, which I don't know, that's a weird word. It's just it's wrong, whatever, about the actual data. Do you all use RAG, the sort of training on the data, or is it just really the tool-using components that make it go?
Shaun Chuah:It depends on what you mean by RAG, because we use tools to ground the AI.
Michael Kennedy:Yeah, I'm thinking like going and actually retraining the model on the research files and data, which I'm guessing from looking at it, it doesn't look like it. It looks more of a Claude Code tool-using style.
Shaun Chuah:Yeah, it is more Claude Code style. we don't train the AI specifically for it. We can swap the base models as newer versions come up. Because if supervised fine-tuning or doing some RL on a base, LLM will cost you a lot of money.
Michael Kennedy:And it's not generalizable, right?
Shaun Chuah:Yeah, and you need the data to train it on. And that's not easy to make as well. So yeah, so given the rapid progress, we want a model where we can just update to the latest model and just leverage the latest changes that the big labs are coming out with.
Michael Kennedy:Amazing. I think that's the right way. So if you go to foundry120.com, there's a one-minute little screencast, silent screencast of it going. So I'm going to, Sean, I'm going to hit play and you kind of just narrate what's happening. I think that'll give people an interesting sense of what this thing is about and give us some talking points here.
Shaun Chuah:So we asked Helix, you know, what plasma samples do we have for the music study? And can you break them down by disease groups? So to get to that answer, you need to join the clinical data frame with your sample database. And then you got to work out which samples are unused, you know, what sample type it is. And then for the disease groups, it's got to inspect the clinical data frame and figure out how many disease groups do you have in your disease column.
Shaun Chuah:And then you've got to do the join and the grouping. So in this demo, it's gone ahead and done it. And it will come back and tell you, well, within our database, we've got like 4,000 samples belonging for Crohn's disease patients and 2,000 with ulcerative colitis.
Michael Kennedy:Yeah. And since people are just listening, let me just go back, just narrate really quick, like fill in a little background visuals. You can see it. It'll, using the different tools, it'll be like, get sample and it'll run for a second. Then it'll write some Python code to do a thing. And then it'll do some more data access and then some more code. So it's primarily orchestrating a bunch of the tools and the code writing that it already knows, right?
Michael Kennedy:It's not just reading 100 terabytes of data or whatever.
Shaun Chuah:No, you can't read all that data because the context window of your AI is limited. So you can't just dump all the raw Excel file into the AI and say, five video samples.
Michael Kennedy:Yeah, exactly. I mean, even the really big ones have a million context right now. All right, carrying on. So then it's off to get a picture, right?
Shaun Chuah:So the next question, we've asked it to do some graphing. So we've asked to plot CRP, which is a blood test against cell-free DNA, which is a scientific experimental output. So to do this, it's got to go and find your clinical data frame and join it with your science data. And it comes up, writes some matplotlib code and gives you a graph back and tell you what it found. So that's what it does. In this second segment, the AI runs into an error and it recovers from the error.
Shaun Chuah:So the AI is able to read the output of the tool, correct, it's working and come back to you. So that's just a demonstration of how an agentic AI system looks like.
Michael Kennedy:Yeah, good narration. And people can go and play that for themselves. But I think that's the big difference between what you're saying, just the chatbot as kind of a better Google, right? It's not that much of the AI thinking. It's a whole bunch of the AI using the tools. And the tools are deterministic, right?
Shaun Chuah:The tools are deterministic. And actually designing the tools is one of the most challenging things to do is like, how many tools do you expose? What should each tool do? And what's a logical set of tools? And how do you make sure that the AI picks them correctly? So I think tool design itself is a huge topic of how you do it. But because the tools run on the backend, we can then enforce the RBAC controls on the tools itself. So you can guarantee that your data access is, you know, security around it is solid.
Michael Kennedy:Have you thought about adding an MCP server to it? So then people can just within their own AI or whatever, just, hey, what does Foundry say about this?
Shaun Chuah:Yeah, MCP is interesting. But I think one of the limiting things for us using MCP is the governance around it, in the sense that we are not allowed to just put data into a Claude Code and send it across to Anthropic. Whereas within this system, all inference happens in Microsoft Azure and it's GDPR compliant, which in the UK is important for us from a research perspective.
Michael Kennedy:So in this... It's also important. It applies to the US companies as well, if they have European customers.
Shaun Chuah:Yeah, so because of regulatory compliance things, we need to make sure that we know where the inference is running and we can guarantee that inference all runs within us, you know, because we have got the platform controls of Azure, we can guarantee where everything happens. And that allows us to actually let the AI safely operate on our data.
Michael Kennedy:I see. So maybe you're using a data center in Ireland for your data so it doesn't leave the UK or something like that.
Shaun Chuah:Yeah, we run most of the things in the UK South data center. For EU compliance, we run it in the Sweden data center. So that's where everything happens at the moment.
Michael Kennedy:Yeah, that makes a lot of sense. Let's zoom out a little bit. You've been working on this project for six plus years now. You've taken it from working with books to write the Django app to integrating to the cloud, Microsoft Foundry, and Azure, and all this tool using AI. What do you see for research, either happening now or in the next couple of years with all this kind of stuff coming along?
Shaun Chuah:Yeah, I mean, I think, you know, we're all very excited about the potential of agentic AI to accelerate a lot of the research workflows that we see. And, you know, hopefully, you know, we get to discoveries faster because drug discovery is a very long process. And I think agentic AI has absolutely a role to play in shortening some of these bottlenecks that we face. So I think it'd be very exciting to see. I think agentic AI is going to transform how we do science, how we operate also in the clinical world. So I see a lot of potential, but the real world is going to take a while to catch up to where the capabilities are today. So I think most people on this podcast would have used coding agents, but I can tell you that almost everybody else in the population has never touched a coding agent. So the gap between what agentic AI systems can do and what everybody's understanding of AI is still very wide. And there's still a lot more work to do to teach people how to use the systems, what they're capable of, where the limitations are, and how do you apply it to your
Michael Kennedy:work. Do you think it can be reliable? Like, do you think we can trust the results and answers we're getting from things like Helix and other tools that are working on us? Well, hopefully,
Shaun Chuah:as things improve over the coming years, the reliability will go up. I mean, personally, you know, the coding agents are still not 100% reliable. You know, even like the best models, like Opus and Fable, they still can make mistakes. So we're not quite past the reliability threshold yet. So we do have to be aware. And I think understanding what you're trying to do is actually very important today compared to, say, five years ago, because you have to really understand what you're trying to do to be able to supervise an AI to do what you were going to do.
Michael Kennedy:I totally agree with what you said. But the alternative is to have a human do it. And humans are also not 100% reliable. There's recently been this dust up with the Linus Torvalds over about using, I forgot the name. There was some AI that they're using as a pre-screen for PRs for the Linux kernel. And some people were like, we're not using it. And it's like, look, this thing is at least as good as the people often doing it. And it's an accelerator. And there's that tension of people expect, I think because it's a computer, people expect it to be perfect, right? Because software is typically deterministic. So if it works once, it's always going to work. And AI isn't like
Shaun Chuah:that, but people also aren't like that. Yeah, that's right. I mean, like, you know, certainly everybody, you know, having an AI by your side is like having a colleague on your team, except this colleague can write code like tremendously faster and much better than most people. You know, today I would say that Claude Code and codex can write code way faster and way better than me, but you still need that kind of strategic view from the human to just make sure that you're going on the right path because you got to drive the AI and you've got to direct it.
Shaun Chuah:And it's the same with the scientific work. So I tell my scientific colleagues, you need to make sure that the answers that come back pass the smell test. You need to make sure that, you know, the numbers that you're seeing are in line with your expectations. You're expecting a number in this magnitude range and you get it there because sometimes the AI goes off and does funny things, right?
Michael Kennedy:Yeah, I just, you know, if you ask it to verify everything and prove it and write code to back it up, like it's better than if you just ask it You know what I mean? Like there's techniques, but you do have to treat it with a little bit of skepticism. But that's also true for your colleagues and your grad students and whatever, right? Like no professor would just take a, like, hey, grad student, write the paper. And they don't even read it. They just publish, they just send it off to nature or medicine or write the medical journal or whatever.
Shaun Chuah:I mean, it depends. I think for disposable stuff, you can let the AI do more of that. But for the real critical workflows, you want to make sure every single step is correct.
Michael Kennedy:Yeah. All right. Well, that brings us to our final call to action. If scientific researchers or medical researchers are out there listening, either what can they learn from your experience or if they wanted to work with you on some of this, what would you say?
Shaun Chuah:So one of the big questions and why we're putting Foundry out there is we don't know how generalizable it is or whether it's just going to be hyper-personal software for ourselves, our team. But I think the general principles, I've shared them widely. So please feel free to just take the idea and run with it. I'm very excited to see what other people can do with the ideas. And perhaps within their own institutions, they might want to build their own systems.
Shaun Chuah:And I think that's absolutely valid. But if you do want to explore a partnership with us, come and visit us on foundry120.com and drop me an email.
Michael Kennedy:Yeah, great. And I'll link to your web page. You've got email and your GitHub and other ways to get in touch with you there. You also have a lot of interesting writing here, so people can check that out.
Shaun Chuah:Yeah, that was my journey into coding, really.
Michael Kennedy:Yeah, we all have one of those journeys to tell the story of. Well, Sean, thank you so much for being on the show. Keep up the good work. I think this is a super interesting project.
Shaun Chuah:Thanks very much, Michael. Very happy to be here, and thanks for the invitation.
Michael Kennedy:Yeah, you bet. Bye. Bye. This has been another episode of Talk Python To Me. Thank you to our sponsors. Be sure to check out what they're offering. It really helps support the show. This episode is brought to you by Sentry. You know Sentry for the air monitoring, but they now have logs too. And with Sentry, your logs become way more usable. interleaving into your error reports to enhance debugging and understanding. Get started today at talkpython.fm/sentry.
Michael Kennedy:And it's also brought to you by Talk Python Courses. Course completion certificates are now live. If you finished a course, there's a certificate waiting for you on your account page right now. Download it as a PDF or add it to your LinkedIn profile with one click under licenses and certifications. Same section as your formal degrees. Visit training.Talk Python.fm slash account to see what you've already earned. And if you're not already subscribed to the show on your favorite podcast player, what are you waiting for?
Michael Kennedy:Just search for Python in your podcast player. We should be right at the top. If you enjoyed that geeky rap song, you can download the full track. The link is actually in your podcast blur show notes. This is your host, Michael Kennedy. Thank you so much for listening. I really appreciate it. I'll see you next time.
Music:I'm out.
Transcript supplied by the publisher with the episode.
by Michael Kennedy · English · Tech & Science
Talk Python to Me is a weekly podcast hosted by developer and entrepreneur Michael Kennedy. We dive deep into the popular packages and software developers, data scientists, and incredible hobbyists doing amazing things with Python. If you're new to Python, you'll quickly learn the ins and outs…
E563 · 16 Sep 2026 · 1 hr 11 min
Lint the entire CPython code base from scratch. It takes 0.3 seconds. Three blinks of an eye. That is ruff, and it is written in Rust. So are Pydantic, Polars, uv, and Granian. Rust shows up in Python three ways: tools that happen to be Rust, libraries Python imports, and servers that run Python inside Rust. This is Rust for Python developers, not Rust experts. Christopher Trudeau is back on Talk Python to discuss Rust and his latest course Up and Running with Rust. The core rule is that only one thing can own a value at a time. Pass it around freely in Python and the garbage collector…
E562 · 10 Sep 2026 · 1 hr 11 min
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first, just to learn which Parquet files matter. DuckLake asks one SQL question instead. The metadata lives in a real database. The data stays in plain Parquet. That's the entire format. Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead DuckLake developer. Guillermo Sanchez Dionis works on DuckLake and the new Quack protocol. With Quack as the catalog, DuckLake handles 200 transactions a second under heavy contention. No…
E561 · 4 Sep 2026 · 1 hr 16 min
How many cores does your machine have, 10, 18? Your async Python code uses just one of them. That isn't a bug in asyncio. That's the design, and optimizing event loops to be faster by 20% doesn't change it. So Giovanni Barillari started over. Joe is the creator of Granian, the Rust-based server that powers Talk Python. His new project is TonIO, an async runtime written from scratch for free-threaded Python. Real threads, a handful of primitives instead of asyncio's pile of them, and it flat out refuses to start if the GIL is on.
E559 · 19 Aug 2026 · 1 hr 8 min
Your site is down. It's 3am. Is it a bug, a bill, or a breach? You can't tell yet, and everyone is watching you find out. Matt Lea has spent fifteen years being the person companies call when an outage is costing them real money per hour, and his whole argument is that everything you'd want in that moment gets decided months earlier, on ordinary afternoons, when someone chose the convenient thing. We walk his top twelve dos and don'ts in AWS - infrastructure as code, IAM roles instead of access keys, private subnets, no wildcards, no public buckets - and I push on which of them actually…
E558 · 10 Aug 2026 · 1 hr 2 min
Every company has one. The little internal tool that Jane built back in 2021, and then Jane left. Nobody understands it, nobody will touch it. There are two unwritten rules around it: don't change it, it's working. And if you break it, you bought it. That's dark-matter enterprise software. For every app you can actually see, there are ten of these sitting in the shadows, frozen. Michael Booth thinks that just changed. He read my article on hyper-personal software and ran with it, writing about hyper-team software: small teams inside big companies finally building the tools that were never…
E557 · 2 Aug 2026 · 1 hr 8 min
Security has always been the vegetables of software. Everyone agrees it matters, and somehow it never quite makes it onto the plate. At PyCon US this year, that changed. For the first time ever, security got its own dedicated, day-long track, one of just two at the whole conference, sitting right next to AI. And the room was packed to the back wall. On this episode, I'm joined by the three people at the center of it. Seth Larson, Security Developer in Residence at the Python Software Foundation and, very recently, a CPython core developer. Juanita Gomez, a PhD researcher at UC Santa Cruz in…
E565 · 2 Oct 2026 · 1 hr 28 min
Do you know what's actually slow in your Python app? Or are you guessing? Until now, profiling Python meant a tracing profiler that made your code 2 to 3 times slower. Or a third-party tool that broke with every new release. Python 3.15 fixes that. It ships Tachyon, a sampling profiler built into the standard library. It attaches to live production apps with almost zero overhead. My guests are Pablo Galindo Salgado, CPython core developer and Steering Council member, and László Kiss Kollár from Bloomberg's Python infrastructure team. Their first prototype ran at two samples a second. Now it…
E564 · 22 Sep 2026 · 1 hr 8 min
Every ship in EVE Online eventually undocks and leaves the station. This time, it's the whole game. EVE has run on Python 2 since it launched in 2003, all 2.4 million lines of it, on a custom Stackless interpreter that stopped at 3.8 and was archived last year. Destination: Python 3.12. The route runs through 6,500 lines of division that decide who wins a fight, and 100 gigabytes of pickled Python objects that have to survive the jump intact. Kristinn Sigurbergsson was on this show ten years ago. He's back, with Jamie Bannister, who is flying the EVE Online migration right now, and Thomas…
E556 · 26 Jul 2026 · 1 hr 5 min
For years, "Django and async" came with an asterisk. The docs themselves warned you off it. Scary performance notes, a story that felt half-finished. Well, that story just got rewritten, literally, and the person who rewrote it is here to tell you why the old framing was wrong. Carlton Gibson is a former Django Fellow, sat on the security team for eight years, and he's on the steering council. On this episode we get into the async topic doc rewrite, what actually remains versus what was just fear, the new Tasks framework in 6.0, DB-level cascades and fetch modes landing in 6.1, and why…
E555 · 13 Jul 2026 · 1 hr 5 min
Coding agents have gotten really good at one kind of work. You scope a feature, edit some files, run the tests, ship it. It all happens on disk. But that is not how data work feels. You load something, you look at it, you run a cell, you watch how it responds, and you decide the next move from whatever is sitting in memory. And until now, your agent couldn't see any of that. It only saw the files. Never the live state. This episode, that wall comes down. marimo pair drops a coding agent right inside a running notebook, with full access to every variable Python is holding in memory. The…