speedup for lm-evaluation-harness; support tensor-parallel inference and data-parallel inference; support gptq, bitsandbytes, peft and exllamav2.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).