Optimizing LLM Inference Serving
EventAI & Machine LearningOnline
14:00
Online, United States
Houston Machine Learning
This event examines the internal workings of large language model inference serving, focusing on the systems layer between API endpoints and GPU clusters. It is intended for machine learning practitioners interested in understanding bottlenecks, algorithms, and architecture involved in production LLM deployment. Attendees will gain insights into performance optimization for LLM inference in production environments.
OnlineFree to AttendOperatorsStudents & Researchers