LLM inference in C/C++ (fork of PrismML fork that enables CPU (incl AVX2 and AVX512) and ROCm for AMD GPUs
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).