Episode · How I AI
GLM 5.2: why I’m replacing Opus in Claude Code with this new model
24 Jun 2026 · 27 min
Episode · How I AI
24 Jun 2026 · 27 min
I put GLM 5.2, the open-weight coding model from Z.AI, through four real tasks inside my actual codebase: a codebase architecture audit, a UI redesign, and a 45-minute autonomous bug-hunting session pulling from Sentry and Vercel logs. Total cost: $3.36 for roughly 6 million tokens, a prioritized bug-fix dashboard I’m actually shipping from, and a landing page redesign that matched Chat PRD’s design system on the first try. What you’ll learn: What “open-weight” actually means and why it matters for cost and vendor independence How to connect GLM 5.2 to Cursor and Claude Code How it performs…
by Claire Vo · English · Tech & Science
How I AI, hosted by Claire Vo, is for anyone wondering how to actually use these magical new tools to improve the quality and efficiency of their work. In each episode, guests will share a specific, practical, and impactful way they’ve learned to use AI in their work or life. Expect 30-minute…
6 Jul 2026 · 36 min
Alessio Fanelli , founder of Kernel Labs and co-host of Latent Space podcast, walks us through two very different AI workflows: (1) a fully autonomous coding setup using OpenAI Symphony + Linear, where Linear acts as a state machine and Symphony manages agents through the whole dev lifecycle with zero babysitting; (2) Codex with browser access searching eBay for underpriced Pokémon cards—autonomously browsing, extracting PSA certificate numbers, and flagging deals on $10K–$20K cards for his San Carlos card shop, Merlin Games. What you’ll learn: Why “agent manager” is a better mental model…
30 Jun 2026 · 26 min
I’ve been testing every major frontier model release since the start of the year, and when Anthropic dropped Sonnet 5, I wanted more than a vibe check. I got tired of one-off tests I couldn’t repeat or compare over time, so I built something better: the How I AI Bench, a repeatable eval harness I constructed live using Claude Code while recording this episode. I ran Sonnet 5 blind against four other frontier models (Sonnet 4.6, Opus 4.8, GPT-5.5, and Gemini 3 Pro) across PRD quality, prototype generation, agentic task completion, and agent personality. The results were not what I expected.…
29 Jun 2026 · 52 min
Eddie Kim is the co-founder and CTO of the payroll and HR platform Gusto, which just crossed $1 billion in revenue and serves more than 500,000 small businesses. Recently he did something most CTOs don’t: he went back to writing code. With three other engineers and one designer, Eddie built Gusto Cofounder, a net-new AI product, from zero code to a tier-one launch in 10 weeks. He walks through how that team actually worked, why they threw out nearly every process, and how anyone can copy the approach. What you’ll learn: The trash-can method: how to write, review, and delete a full PR as a…
5 Oct 2026 · 36 min
Kath Korevec is a member of the Product staff at OpenAI working on Codex, and she spent over a year building and using ChatGPT Sites internally before its public launch. She’s been on the front lines of shipping Plugin Insights, MCP plugin hosting, and the connector ecosystem, which now includes around 60 integrations. What you’ll learn: What Plugin Insights actually does, and why it changes who can use a site The incident command site Kath built for her OpenAI team, and how it uses live Slack and Notion connectors The one phrase that tells Codex to wire up connectors for you The…
30 Sep 2026 · 46 min
John Lindquist created egghead.io, a developer education platform used by hundreds of thousands of working engineers. These days he’s building mega.dev, a hands-on program specifically for developers who want to do real work with AI agents, not just prototype them. What you’ll learn: Why Jev is a decision engine, not a chatbot, and what that distinction actually changes about how you build How John built a real-time voice to-do app that classifies and executes commands with no visible pause The data deduplication pattern that merges messy records in milliseconds using confidence scores Why…
30 Sep 2026 · 24 min
I spent the day at OpenAI’s DevDay in San Francisco, and I have good news and bad news: OpenAI released a lot of stuff. In this episode, I break down the announcements worth paying attention to - and show you what happened when I tested some of them early. We’ll meet my Dot, explore why Spaces and Sites could matter for how teams work, and get into the model and API updates I’m most excited about as a developer. I use the Decisions API to find podcast thumbnails where nobody looks awkward, build a collaborative sketchpad with Astra ultrafast, and let my kids redesign a 3D world in real time.…
28 Sep 2026 · 26 min
Jev is TypeSafe AI’s new decision model. It returns type-safe structured values (a choice, a score, a probability) instead of generated text, at 4 cents per million input tokens with no output charge. This week I ran it on five real projects: PR categorization, a meta-analysis of my own Claude and Codex sessions, Gmail triage, the ChatPRD product insights graph, and a live audience dashboard built from 4,500 YouTube comments. What you’ll learn: What makes Jev fundamentally different from every other model I’ve used How I analyzed 1,700 PRs for 9 cents and what I found out about where my…
22 Sep 2026 · 39 min
I got up early to record an Opus 5.5 review. Then Anthropic and OpenAI dropped new models on the same morning, and I decided to do something I’d never done before: take the How I AI bench live. I put GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, and more through the work I actually care about: emails, PRDs, frontend prototypes, backend work, long-running agents, SVGs, and video editing. I scored the outputs without knowing which model made them, so you get to watch me make predictions, change my mind, and reveal my own very inconsistent taste. Astra won my heart. Opus 5.5 won my week. Sol still…
22 Sep 2026 · 25 min
I’ve been off Claude for months. Not because it got dumb, but because it got annoying. The rambling, the hedging, the preachy little disclaimers on tasks that didn’t need them. I moved most of my daily work to Codex and I didn’t miss it. Then Anthropic shipped Opus 5.5: 40% cheaper than Opus 5, faster, and with what they’re calling a fundamentally different alignment approach. I ran it for a week across real work, including four long-running agentic tasks, a full ChatPRD homepage redesign, an SVG benchmark, and one very firm refusal, and I’m ready to give you the honest verdict. There’s a…
21 Sep 2026 · 47 min
Zach Lloyd is the co-founder and CEO of Warp, an AI-powered terminal and software factory platform used by tens of thousands of engineers. Before Warp, he spent nearly a decade at Google, including time as a principal engineer on Google Sheets. He built Warp from the ground up as a modern, AI-native alternative to legacy terminals, and the team has since expanded into software factories: a full cloud-based system that takes an idea in Slack all the way through to a merged PR. In this episode: Why a software factory is more than a coding agent The public Slack → Linear → GitHub → QA workflow…