llama.cpp v0.6.0 released with multi-token-prediction speculative decoding
The llama.cpp v0.6.0 release introduces multi-token-prediction speculative decoding for Qwen4Exp, reportedly achieving a 1.5x decode speedup on DGX Spark. The update also includes day-one support for Zhipu's 320B hybrid GLM-5.3-Flash, Cloudflare's Clef decision models, Nimble, and Ling 3.0 VL. New Metal and Vulkan flash-attention kernels are also claiming up to 3x faster matrix multiplication on Apple GPUs.
Want more?
Open NewsSnap.ai for the full app experience, including audio, personalization, and more news tools.