Moonlet app icon

Run a 35B coding agent
on a 16 GB Mac.

Moonlet is a coding agent that runs open-weight models on your Mac. Its MoE offload keeps only the parts of a model in use in memory, so Qwen 3.6 35B-A3B and Gemma 4 26B-A4B run in about 10 GB. No cloud, no API keys, and nothing you write leaves the machine.

macOS 14+  ·  Apple Silicon  ·  free for personal use

Companies need a commercial license. By downloading you agree to the EULA.

A real session: Qwen 3.6 35B-A3B adds a --month option and its test to a small Python project, then runs the suite. Recorded in the app, sped up in the waiting parts, nothing else changed.

Memory

How a 35B model fits in 10 GB.

A mixture-of-experts model is big on disk but uses only a small slice of itself for each token: a few experts out of many, and a different few every time. Loaded the usual way, every expert has to sit in memory anyway, which is why the 35B-class models have not fit a 16 GB Mac.

Moonlet keeps the experts a model reaches for often in memory and brings the rest in from disk when a token needs them. Each token runs twice: a quick first pass to learn which experts it wants, then the real pass with those experts loaded. You get the same model, not a smaller one, at a fraction of the memory.

ModelPeak memorySpeed
Qwen 3.6 35B-A3B 10.4GB 36tokens/s
Gemma 4 26B-A4B 8.8GB 40tokens/s

Measured in July 2026 on an Apple M-series MacBook Pro with offload on, on code and tool-call work. The same Qwen weights take about 20 GB when loaded whole. Speed depends on your chip and on what else is running.

It is one switch in the prompt bar, next to the model name. Turn it on when a model does not fit your Mac. On a Mac with more memory, leave it off and the whole model loads the normal way, which is faster.

Agent

The model calls the tools. Moonlet runs them.

Ask for a change and the agent works on its own. The model decides what to read, what to search, what to edit and which commands to run, and Moonlet carries out each call on your Mac as it comes, then hands the result back. The loop keeps going until the task is done or it needs you, and every step shows in the panel as it happens.

It is the same way the cloud coding agents work, on a model that never leaves the laptop.

Assistant
The agent panel: the request at the top, then the model's tool calls (search, read, edit, run) and the diff it produced.

Supported models

Each one is tuned and tested inside Moonlet before it ships. Pick one on first launch and it downloads inside the app from Hugging Face, no accounts anywhere. They are 4.5 to 19 GB each.

Qwen 3.5 / 3.6 / 3.8 Alibaba Tuned for stability and real-world coding. 3.8 lets you dial how hard it thinks.
Gemma 4 Google Google DeepMind's open general-purpose models.
Nemotron 3.5 NVIDIA Fast agentic model, reasoning on demand.
Mellum 2 JetBrains A reasoning-augmented coding assistant.
Muse Glimmer Meta Agentic work, built for consumer hardware.
LFM 2.5 Liquid AI A hybrid model built for on-device use.

Get Moonlet

A 260 MB download. Pick a model on first launch, point it at a folder, start working.

Download for macOS
macOS 14 or later Apple Silicon (M1+) 16 GB RAM with offload for the 35B-class models
Feedback

Tell us what broke. Or what didn't.

Bug reports, model requests, rough edges — we read all of it.

Before you download

Moonlet is free for personal use. Companies and paid work need a commercial license.

It is an agent: it edits and deletes files and runs commands in the folder you open, at the model's direction. Keep your work in version control or a backup.

By downloading you agree to the End-User License Agreement.

Agree and download