llama.cpp v0.6.0 released with multi-token-prediction speculative decoding

AI-generated NewsSnap summary based on source reporting.
Published: 2026-10-05T21:35:00Z
Category: technology
Source: AI Weekly

The llama.cpp v0.6.0 release introduces multi-token-prediction speculative decoding for Qwen4Exp, reportedly achieving a 1.5x decode speedup on DGX Spark. The update also includes day-one support for Zhipu's 320B hybrid GLM-5.3-Flash, Cloudflare's Clef decision models, Nimble, and Ling 3.0 VL. New Metal and Vulkan flash-attention kernels are also claiming up to 3x faster matrix multiplication on Apple GPUs.

Want more?

Open NewsSnap.ai for the full app experience, including audio, personalization, and more news tools.

Open NewsSnap.ai