Swiftlet, a Swift + Metal runtime, enables large MoE models like Qwen3.6-35B-A3B and Qwen3-Next-80B-A3B to run with minimal memory by keeping only the dense core in RAM while streaming expert weights from SSD. On an M5 Mac, the 35B model peaks at 2.6GB (7–11 tok/s), and the 80B model peaks at 4.3GB (4.5–5 tok/s). The 35B version also runs natively on an iPhone 17 at ~2.5GB and 1 tok/s. The system uses fine-grained token-to-expert routing, pread-based SSD fetches, LFU caching, and runtime-compiled Metal shaders, supporting iOS deployment and OpenAI-compatible server interfaces.
By clicking "Accept All Cookies", you agree to the storing of cookies on your device to enhance site navigation, analyze site usage, and assist in our marketing efforts.Privacy Policy Notice
When you visit any website, it may store or retrieve information on your browser, mostly in the form of cookies. This information might be about you, your preferences or your device and is mostly used to make the site work as you expect it to. The information does not usually directly identify you, but it can give you a more personalized web experience. Because we respect your right to privacy, you can choose not to allow some types of cookies. Click on the different category headings to find out more and change our default settings.However, blocking some types of cookies may impact your experience of the site and the services we are able to offer. Privacy PolicyStatement
Manage Preferences
Strictly Necessary Cookie
Always Active
These cookies are necessary for the website to function and cannot be switched off in our systems. They are usually only set in response to actions made by you which amount to a request for services, such as setting your privacy preferences, logging in or filling in forms. You can set your browser to block or alert you about these cookies, but some parts of the site will not then work. These cookies do not store any personally identifiable information.