gitaskhub

llama.cpp fork with TurboQuant WHT-rotated KV cache & weight compression + Gemma 4 MTP and Qwen 3.6 NextN speculative decoding (+30-50% throughput).

Stars · 1
Language · C++
License · MIT
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).