How to Engineer AI Inference Systems with Philip Kiely - #766
Philip Kiely, head of AI education at Baseten, discusses inference engineering as the stickiest and most critical AI workload. The conversation covers the blend of GPU programming, applied research, and distributed systems that define inference, and how teams can now go from research to production in hours rather than months. Philip explains the key knobs—batching, quantization, speculative decoding, and KV cache reuse—that let teams balance latency, throughput, and cost. He traces the maturity journey from closed APIs to dedicated deployments and in-house platforms, surveys the runtime landscape (vLLM, SGLang, TensorRT LLM), and looks ahead to the need for specialized, workload-specific runtimes as agents and multimodality demand ever more efficient inference.