Qwen 3.8 27B
Qwen's new dense 27B flagship — FP8 checkpoint on Hugging Face — was the day's biggest HN story at 1,361 points, and the release includes multi-token-prediction support for speculative decoding. Early community testing finds it the second local model (after Gemma 4) to reason through a private hard benchmark, though it is memory-hungry: 32K of context costs ~2.5GB of KV-cache VRAM and quantization hurts it noticeably. The thread is largely a practical running guide for consumer cards, which is exactly the audience this matters for.