option
Home
Flash News
Content
RaymondWalker
RaymondWalker
August 4, 2026

Swiftlet, a Swift + Metal runtime, enables large MoE models like Qwen3.6-35B-A3B and Qwen3-Next-80B-A3B to run with minimal memory by keeping only the dense core in RAM while streaming expert weights from SSD. On an M5 Mac, the 35B model peaks at 2.6GB (7–11 tok/s), and the 80B model peaks at 4.3GB (4.5–5 tok/s). The 35B version also runs natively on an iPhone 17 at ~2.5GB and 1 tok/s. The system uses fine-grained token-to-expert routing, pread-based SSD fetches, LFU caching, and runtime-compiled Metal shaders, supporting iOS deployment and OpenAI-compatible server interfaces.

Swiftlet, a Swift + Metal runtime, enables large MoE models like Qwen3.6-35B-A3B and Qwen3-Next-80B-A3B to run with minimal memory by keeping only the dense core in RAM while streaming expert weights from SSD. On an M5 Mac, the 35B model peaks at 2.6GB (7–11 tok/s), and the 80B model peaks at 4.3GB (4.5–5 tok/s). The 35B version also runs natively on an iPhone 17 at ~2.5GB and 1 tok/s. The system uses fine-grained token-to-expert routing, pread-based SSD fetches, LFU caching, and runtime-compiled Metal shaders, supporting iOS deployment and OpenAI-compatible server interfaces.
Comments (0)
0/300
OR