git
ask
hub
Privacy
Terms
Sign in with GitHub
lukallm/dflash_qwen3.6_27b_llamacpp
↗
Stars ·
15
Language ·
Python
Simplify and Visualize This Repo
Scan the safety of this repo
Self Host this repo
Check out more work of Developer: lukallm
Ask anything about this repo to start.
Explain how DFlash speculative decoding works and why it can speed up text generation without changing outputs.
Summarize the speed and accuracy tradeoffs found in this DFlash benchmark on Qwen3.6-27B.
Show me how to regenerate the benchmark charts using benchmark/plot_results.py.
What hardware and llama.cpp version were used to produce these DFlash benchmark numbers?
Full explanation on explaingit →
📎
Send
By chatting or signing in you agree to the
Terms
and chat-message logging (revocable in
History
).