CPM.cu is a lightweight, high-performance CUDA implementation for LLMs, optimized for end-device inference and featuring cutting-edge techniques in sparse architecture, speculative sampling and quantization.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).