gitaskhub

Why is LLM inference slow — and how do you make it fast? A hands-on, first-principles course: roofline → KV cache → quantization → parallelism → vLLM/SGLang, with GPU labs on open models.

Stars · 19
Language · Python
License · MIT
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).