gitaskhub

NVIDIA Qwen3.6-27B NVFP4 on SM121 (GB10) — vLLM v0.24.0 with native NVFP4 KV cache via FlashInfer FA2 JIT. 67% more KV capacity than FP8.

Stars · 8
Language · Python
License · MIT
Ask anything about this repo to start.
Full explanation on explaingit →

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).