DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 馃槈
Frontier running at 50 tok/s is wonderful!
Loktar 馃嚭馃嚫 on X: "DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 馃槈 Frontier running at 50 tok/s is wonderful!"
- Draft acceptance came in at 53-57% I ended up dropping to Q4_K_XL to IQ4_XS to make VRAM room for the drafter, honestly couldn't tell the outputs apart. Definitely a win, now onto the next one!
- nice! i鈥檓 averaging 35 t/s with llama.cpp and deepseek-v4-flash-Q3_K_S with minimal tweaks on 11x3090s. i鈥檒l try your setup and if I get 50+ tok/s - the beer鈥檚 on me 馃嵒

