Engineering conversations, recorded.
Practitioners on what it actually takes to run AI in production — costs, failures, governance, and the org design around the code. New episodes monthly. Streaming platforms launching soon; episode summaries below.
Inside an AI-First GCC: Operating Model, Talent, Unit Economics
How AI-first capability centers staff differently, why pod structures beat pyramids, and the levers that change cost per outcome.
RAG Post-Mortems: Five Retrieval Failures and What Fixed Them
Chunking gone wrong, embedding drift, permission-aware retrieval, and why most 'LLM bugs' are corpus bugs.
The Inference Bill: An Honest Conversation About LLM Unit Costs
Token economics, caching layers, model routing, and how to forecast COGS before finance does it for you.
Agents at Work: Orchestration Patterns That Survive Production
Planner-executor splits, tool-call budgets, human checkpoints, and the observability agents demand.
Regulated AI: Shipping ML Under GxP and Model Risk Management
Validation strategy, audit trails, and change control when your model faces an inspector, not just a user.
From POC Purgatory to Production: Why 80% of AI Pilots Stall
The missing engineering disciplines between a promising demo and a system the business trusts.
Quantum for Operators: What to Track, What to Ignore
Error-correction milestones, PQC deadlines, and a watchlist that costs one architect-hour a quarter.
Running AI in production? We want the war story.
We interview engineering leaders, architects, and operators — especially the unglamorous middle of the story where the demo met reality. If that's you, get in touch.