Run Ling 3.0 Flash locally 🌀
We released GGUF quants on Hugging Face, from lossless BF16 to 1-bit, plus NVFP4!
AD-Q5_K_M is the best fit for 128GB hardware (tested on DGX Spark). It matches the original's token choice 97.5% of the time and drifts 31% less than the llama.cpp
🚀 Today, we’re releasing INT4 and FP4 (MXFP4) variants of Ling-3.0-flash.
Both run end to end on a single NVIDIA DGX Spark via our Spark-adapted SGLang path.
For FP4, W4A16 is the stable default, while W4A8 is tuned for higher throughput.
The efficiency and accuracy of the

