Optimizing LLM Inference Serving
This event examines the internal workings of large language model inference serving, focusing on the systems layer between API endpoints and GPU clusters. It is intended for machine learning practitioners interested in understanding bottlenecks, algorithms, and architecture involved in production LLM deployment. Attendees will gain insights into performance optimization for LLM inference in production environments.