I stopped dumb prompts on deepseek flash and started asking things on a huge codebase and I confirm -> this person feedback is correct (using it with Pi but want to try Codex too)
not saying you should drop your frontier model assistant but for day-to-day work I really recommend having Pi on the side running DeepSeek-V4-Flash-0731-GGUF (it's nearly free, much faster and the perfect side chat :))
DeepSeek V4 Flash on my 7x3090s went from 36 tok/s to 47 tok/s.. using llama.cpp's new DSpark speculative decoding and a 10GB drafter unsloth quietly shipped right next to the main weights 😉
Frontier running at 50 tok/s is wonderful!