btw we just shipped something to make this clearer in the dedicated section for GGUF models 馃
DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 馃槈
Frontier running at 50 tok/s is wonderful!

