Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).