option
Home
News
Swiftlet Packs 80B Qwen Into Mac With 4.3GB Memory; iPhone 17 to Run 35B Natively

Swiftlet Packs 80B Qwen Into Mac With 4.3GB Memory; iPhone 17 to Run 35B Natively

September 24, 2026
5

A Swift + Metal runtime named Swiftlet is redefining "where large models can run." It is specifically tailored for the Qwen3-Next and Qwen3.5/3.6 family of MoE (Mixture of Experts) models. The core idea is quite clever: only the small dense core of the model stays in memory, while the large part—the routing expert weights—normally resides on SSD and is streamed in as needed.

The results are clear from the numbers. On an M5 Mac, the 4-bit version of Qwen3.6-35B-A3B takes up 18GB on disk but only peaks at 2.6GB of memory, decoding at 7 to 11 tok/s; even more impressive is the 4-bit version of Qwen3-Next-80B-A3B, which requires 42GB on disk but peaks at just 4.3GB of memory, with a speed of 4.5 to 5 tok/s. The 35B version can now run on an iPhone 17, with about 2.5GB of memory and a speed of around 1 tok/s — according to the developers, this is the first time such a model runs natively on a phone without touching a server.

image.png

The project is fully end-to-end usable, and both models can produce verified correct outputs. The developer also admits a trade-off: each token actually activates about 3B parameters, so these models behave like large models during conversations and writing, but still act like small models in terms of factual memory.

Its working principle lies in the fine-grained scheduling. Each layer routes each token to 10 out of 512 experts (for the 80B version) or 8 out of 256 experts (for the 35B version). Swiftlet keeps the dense weights — attention, DeltaNet projections, routers, shared experts, and embedding vectors — firmly in memory, taking about 1.3GB (for 35B) or 2.5GB (for 80B) in 4-bit. Thousands of routing experts are repackaged into fixed-step data blocks, stored in .qpack containers. To fetch one expert, it only needs one pread operation on the SSD, no mmap, and it doesn't disturb the page cache. Popular experts are cached in a limited pool using LFU plus recentness strategy for eviction; the hit rate ranges from 43% to 70%, and the cache size has almost no impact on speed, as Apple's SSD can handle the overhead of misses.

The entire forward pass runs on Metal using runtime-compiled shaders, so there's no need for a Metal toolchain during build, and the same code can be directly deployed to iOS. More interestingly, 75% of the layers use Gated DeltaNet linear attention with a fixed-size cyclic state, meaning that regardless of the context length, it never grows an expanding KV cache — the hidden burden of long text conversations is quietly offloaded.

Swiftlet isn't just a command-line toy. It has four identities designed for itself. As a library, SwiftletCore can be embedded into any macOS or iOS app, providing chat capabilities with streaming incremental output and conversation caching; as a command-line tool, swiftlet chat and swiftlet generate handle local execution and benchmarking, while swiftlet-repack builds containers directly from MLX checkpoints and supports resuming downloads from Hugging Face. As a server, swiftlet-server provides OpenAI-compatible chat-completions interfaces on a loopback address, allowing any OpenAI-compatible chat interface to connect to the local model; as an application, the iOS version of Priv AI embeds SwiftletCore as a streaming model engine, allowing regular users to start chatting by simply downloading it.

Regarding correctness, every layer's forward pass is compared layer-by-layer against the mlx-lm reference implementation, covering f32 and int4 quantization forms. Incremental decoding is also compared with full sequence results, and the Metal kernels have been repeatedly verified against precise CPU references. The container can also perform byte-level verification against the source checkpoint — whether experts are read from cache or disk, the answers are exactly the same.

Its inspiration path is clearly explained: TurboFieldfare first validated the feasibility of streaming experts for Gemma on Mac, and Swiftlet borrowed some publicly available design experiences from it, such as using pread to stream experts into bounded slot pools, using LFU plus recentness for eviction, and using fixed-step packaging so that one read equals one fetch. However, the rest was written from scratch, with about 10,000 lines of Swift and Metal code, tackling the completely different Qwen mixed architecture: Gated DeltaNet linear attention, gated GQA, and high-sparse MoE with shared experts. It even implements MLX-style int4/int8 group quantization calculations in Metal, using byte-addressable kernels and 64-bit offsets to support gigabytes of shards.

Related article
Tencent reshuffles large model team as Yao Shunyu takes charge of basic model Tencent reshuffles large model team as Yao Shunyu takes charge of basic model Tencent’s Hunyuan multimodal team has recently completed a significant restructuring and strategic pivot. Lu Xudong, formerly leading xAI’s multimodal understanding efforts, has joined Hunyuan to head multimodal content generation algorithms. This sh
Tongyi Qianwen Open Platform Launches, Bringing Conversational-as-a-Service to Daily Life Tongyi Qianwen Open Platform Launches, Bringing Conversational-as-a-Service to Daily Life The Tongyi Qianwen Open Platform has officially launched, providing developers with access to services across mobile devices, PCs, and AI glasses. Users no longer need to switch between apps; they can complete various daily service operations directl
Cancer-stricken founder deploys AI in fightback Cancer-stricken founder deploys AI in fightback Conno Christou refuses to leave his health to chance. He monitors his sleep with a Whoop band, cross-references the data with an Oura ring, and undergoes nearly 100 biomarker tests annually. For four consecutive years, he followed the bloodwork proto
Related Special Topic Recommendations
Music composition Best AI Vocal Demo Tools for Track Prototypes
Best AI Vocal Demo Tools for Track Prototypes

2026 Latest Best Top-rated AI Vocal Demo Tools for Track Prototypes are here on XIX.AI! This curated list features powerful, game-changing tools that go through real-world tests to ensure top performance. You’ll find a free vs paid comparison, detailed rankings, and must-try options perfect for boosting creative efficiency and helping you unlock your AI edge. Explore now to discover your perfect tool!

10 tools
xix.ai
Software Development AI API Mocking Tools for Frontend-Backend Collaboration
AI API Mocking Tools for Frontend-Backend Collaboration

2026 Latest Best Top-rated AI API Mocking Tools for Frontend-Backend Collaboration are here on XIX.AI! This curated list features powerful, game-changing solutions that help teams streamline workflows, boost productivity, and create high-quality apps faster through real-world tests and strict rankings. Get a free vs paid comparison to find the must-try option that suits your needs. Explore now to unlock your AI edge!

12 tools
xix.ai
Design & Art Midjourney Style Reference Tools for Character Concepts, Posters, and Ad Visuals
Midjourney Style Reference Tools for Character Concepts, Posters, and Ad Visuals

2026 Latest Best Top-rated Midjourney Style Reference Tools for Character Concepts, Posters, and Ad Visuals! XIX.AI has curated a powerful game-changing collection that undergoes weekly updated real-world tests. You can find detailed free vs paid comparison insights and reliable rankings to help you pick the must-try option that boosts your creativity and speeds up workflow significantly. Explore now to Discover your perfect tool for all your visual design needs.

11 tools
xix.ai
Music composition AI Vocal Generation Tools: Create Demo Vocals and Harmony Layers
AI Vocal Generation Tools: Create Demo Vocals and Harmony Layers

2026 Latest Best Top-rated AI Vocal Generation Tools for Creating Demo Vocals and Harmony Layers are here on XIX.AI. This curated collection features powerful, game-changing solutions that have passed rigorous real-world tests. We provide a weekly updated free vs paid comparison along with detailed rankings to help you find the perfect tool for boosting your creativity and productivity. Explore now to Unlock your AI edge.

12 tools
xix.ai
Productivity Best AI Focus Assistants: Reduce Context Switching and Stay on Track
Best AI Focus Assistants: Reduce Context Switching and Stay on Track

2026 Latest Best Top-rated AI Focus Assistants Curated to Slash Context Switching and Keep You Productive All Day. These powerful tools undergo rigorous real-world tests, featuring a free vs paid comparison along with weekly updated rankings. Discover your perfect tool to boost writing efficiency, unlock creative ideas, and stay focused without distraction. Explore now at XIX.AI to Unlock your AI edge.

9 tools
xix.ai
writing Best AI Editing Tools for Clear Business Writing
Best AI Editing Tools for Clear Business Writing

2026 Latest Best Top-rated AI Editing Tools for Clear Business Writing are here on XIX.AI! This curated list features powerful, game-changing solutions that help you boost writing efficiency significantly through real-world tests and rigorous rankings. You’ll find a free vs paid comparison to help you choose the perfect fit. Explore now and Unlock your AI edge!

10 tools
xix.ai
Comments (0)
0/500
OR