gitaskhub

Differential-KV (DKV) is a sparse KV-cache inference runtime designed for high-efficiency, memory-bounded long-context Large Language Model (LLM) inference across Apple Silicon (MLX) and CUDA GPUs.

License · MIT
Ask anything about this repo to start.

By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).