Fullmoon
Chat app for Apple silicon that runs language models on device, for anyone who wants an assistant working fully offline.
Open Source Alternative to:

Fullmoon is a chat client for language models that never leave the device. Nothing is uploaded, no account is needed, and it keeps working when the network does not, which is the entire reason to install it rather than talk to a hosted assistant.
It is built on MLX Swift, Apple's array framework for machine learning research on Apple silicon, and draws through Metal 3. Builds run on iOS, iPadOS, macOS and visionOS, and chat history is stored locally on the device.
What you get is a short list of models and a few controls around them.
- Llama 3.2 1B Instruct: a 4-bit build of about 0.7 GB, the smallest option.
- Llama 3.2 3B Instruct: a 4-bit build of about 1.8 GB.
- DeepSeek R1 Distill Qwen 1.5B: available as a 1.0 GB 4-bit build or a 1.9 GB 8-bit build.
- Appearance: theme, fonts and the system prompt are all adjustable.
- Shortcuts: call a local model from the Shortcuts app and pass its output into other actions.
Fullmoon comes from Mainframe and reaches users through the App Store, with a TestFlight build carrying the newest features and models, and the iOS source published on GitHub. Apple silicon is the hard requirement, so older Intel hardware is out. Storage is the other practical limit, since every model is downloaded onto the device and the largest of them runs close to two gigabytes.
Stars
2,265Forks
219Last commit
1 year agoRepository age
2 yearsLicense
MITVersion
1.2.3Repository
mainframecomputer/fullmoon-ios
Auto-fetched from GitHub .
Open source alternatives similar to Fullmoon:
Stars
Forks
Last commit
Stars
Forks
Last commit
Stars
Forks
Last commit