Post

Log inSign up

Post

Loktar 馃嚭馃嚫 on X: "DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 馃槈 Frontier running at 50 tok/s is wonderful!"

  • user avatar
    Loktar 馃嚭馃嚫
    @loktar00
    DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 馃槈 Frontier running at 50 tok/s is wonderful!
    7:28 PM 路 Aug 4, 202628.9KViews
  • user avatar
    Loktar 馃嚭馃嚫
    @loktar00
    Aug 4
    Draft acceptance came in at 53-57% I ended up dropping to Q4_K_XL to IQ4_XS to make VRAM room for the drafter, honestly couldn't tell the outputs apart. Definitely a win, now onto the next one!
  • user avatar
    0xNeoArch
    @0xNeoArch
    Aug 4
    nice! i鈥檓 averaging 35 t/s with llama.cpp and deepseek-v4-flash-Q3_K_S with minimal tweaks on 11x3090s. i鈥檒l try your setup and if I get 50+ tok/s - the beer鈥檚 on me 馃嵒

Log in or sign up for X

See what鈥檚 happening and join the conversation

Continue with phone
or
Log in with username or email

Relevant people

Avatar
Loktar 馃嚭馃嚫@loktar00Follow
Building with local LLMs, retro PCs, and 30 years of making things. Christian. Dad. Veteran.

Trending now

Terms路Privacy路Cookies路Accessibility路Ads Info路漏 2026 X Corp.