Prefill/decode benchmark of Qwen3.6-27B at bs=1 on 4x RTX 5060 Ti (Blackwell): Q8_0 + f16 KV, MTP, llama.cpp vs vLLM FP8, interconnect analysis
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).