gitaskhub

PXQ: PXA-native low-bit MoE quants (2/3/4-bit, E16-row scales) + fused CUDA kernels for Pascal/Volta — run a real 35B on a salvaged 12-16GB card. Fork of ik_llama.cpp.

Stars · 20
Language · C++
License · MIT
Ask anything about this repo to start.
Full explanation on explaingit →

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).