Home / Podcast
Audio · The Production Console

Engineering conversations, recorded.

Practitioners on what it actually takes to run AI in production — costs, failures, governance, and the org design around the code. New episodes monthly. Streaming platforms launching soon; episode summaries below.

POD/01Episode log
EP·07

Inside an AI-First GCC: Operating Model, Talent, Unit Economics

How AI-first capability centers staff differently, why pod structures beat pyramids, and the levers that change cost per outcome.

42 min · audio episode
EP·06

RAG Post-Mortems: Five Retrieval Failures and What Fixed Them

Chunking gone wrong, embedding drift, permission-aware retrieval, and why most 'LLM bugs' are corpus bugs.

38 min · audio episode
EP·05

The Inference Bill: An Honest Conversation About LLM Unit Costs

Token economics, caching layers, model routing, and how to forecast COGS before finance does it for you.

45 min · audio episode
EP·04

Agents at Work: Orchestration Patterns That Survive Production

Planner-executor splits, tool-call budgets, human checkpoints, and the observability agents demand.

41 min · audio episode
EP·03

Regulated AI: Shipping ML Under GxP and Model Risk Management

Validation strategy, audit trails, and change control when your model faces an inspector, not just a user.

47 min · audio episode
EP·02

From POC Purgatory to Production: Why 80% of AI Pilots Stall

The missing engineering disciplines between a promising demo and a system the business trusts.

36 min · audio episode
EP·01

Quantum for Operators: What to Track, What to Ignore

Error-correction milestones, PQC deadlines, and a watchlist that costs one architect-hour a quarter.

39 min · audio episode
POD/02Be a guest

Running AI in production? We want the war story.

We interview engineering leaders, architects, and operators — especially the unglamorous middle of the story where the demo met reality. If that's you, get in touch.

Prefer the conversation without the microphone?

Book a working session instead — same candor, your architecture on the table.

Average first response: under 24 hours · straight engineering answers, no pitch theatre