Implemented and benchmarked LLM inference compression: int4/int8 quantization, GPTQ-like calibration, int8 KV cache, pruning, distillation, speculative decoding, torch.compile, and ONNX. Every number from a logged, self-audited run.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).