Qwen3.8-Flash-Next: 125B MoE with 6B Active, Previews Qwen4 Architecture
Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weights multimodal mixture-of-experts (MoE) model that serves as an early preview of the architecture slated for Qwen4. The model has 125 billion total parameters but activates only 6 billion per token, yielding a substantial performance-per-compute boost. Early testing on an NVIDIA DGX Spark with Unsloth quantizations shows promising results, though the model is positioned as a preview rather than a final release. This release continues Qwen's aggressive open-weights strategy, challenging closed rivals like GPT-4 and Claude.