Kimi K3 (2.8T MoE) Runs at 1 token/s on MacBook Pro via 4 SSDs
Moonshot AI's Kimi K3, a 2.8-trillion-parameter mixture-of-experts model, has been demonstrated running locally on a MacBook Pro at 1 token per second, with weights streamed from four external SSDs. The setup bypasses traditional RAM limits by using SSD storage and memory-mapped I/O, enabling inference of a model far larger than the laptop's unified memory. This proof-of-concept, shared on GitHub, shows that even massive models can run on consumer hardware with creative engineering, though the speed is impractical for real-time use. It highlights the growing trend of local AI inference and the potential for future optimizations to make such models more accessible.