Bring your own model

Bring your own model: AI Polish with Ollama or LM Studio

SayWrite can send your dictated text, never your audio, to an Ollama or LM Studio server running on your own Mac at 127.0.0.1, so AI Polish uses the model you choose and nothing goes to the cloud. Your models, your hardware.

Quick start · Ollama

# 1. install Ollama from ollama.com, then:
ollama pull qwen3.5:9b

# 2. SayWrite › AI › Use local AI: on
# 3. hold fn, talk, tap space

qwen3.5:9b is 6.6 GB. SayWrite picks it automatically once it's installed.

Overview

How it works

  1. You hold Fn and talk. SayWrite transcribes on the Neural Engine, as always.
  2. Tap Space while holding Fn (or use a per-app auto-polish rule) and the text goes to your local model.
  3. SayWrite sends it to Ollama's or LM Studio's API at 127.0.0.1 with a prompt that keeps your facts and voice: point first, rambling cut, tone matched to the app you're typing in.
  4. The rewrite is typed at your cursor. If the model fails, nothing is typed or changed.
Compare what you said with what SayWrite writes (Polish with qwen3.5:9b)
Polish with qwen3.5:9b

You said: “um so the thing about the migration is uh we tested it on staging and it's fine but like we need a rollback plan before friday”

SayWrite writes: We tested the migration on staging and it works, but we need a rollback plan before Friday.

What the microphone heard, word for word.Filler-word removal drops the ums and uhs. Plain dictation, no AI.Fn + Space: rewritten on your Mac with your Ollama model.

Engine routing

Auto, Fast or Best: which model answers

SayWrite can use Apple's on-device model and your own model side by side. The Rewrite engine setting decides which one handles each kind of request. Pick one to see the routing.

How SayWrite routes AI requests for each Rewrite engine setting
Rewrite engine
  • Apple Intelligenceon-device model · no server
  • Ollama127.0.0.1:11434
  • LM Studio127.0.0.1:1234
Polish, emails, your own instructionsOllama or LM StudioApple IntelligenceOllama or LM Studio
Grammar, concise, bullets, friendlierApple IntelligenceApple IntelligenceOllama or LM Studio

Default. Heavy rewrites go to your bigger model when its server is running; quick fixes use Apple's. If one engine fails, the other is tried once.Apple's on-device model for everything. Quickest, and good at grammar, but it restructures casual speech less.Your Ollama or LM Studio model for everything, with Apple's model only as a fallback.

Step by step

Set up Ollama

  1. Install Ollama from ollama.com and open it. Its server listens on 127.0.0.1:11434.
  2. Pull a model in Terminal:
    ollama pull qwen3.5:9b
  3. In SayWrite, open AI, turn on Use local AI, set Quality server to Ollama and leave Quality model on Automatic, or pick any installed model.
  4. The status row should read “Ollama is running with qwen3.5:9b”. Try a sentence in the Try It box, which shows the engine and how long it took.

Step by step

Set up LM Studio

  1. Install LM Studio, download a model, and load it.
  2. Open the Developer tab and start the local server. It listens on 127.0.0.1:1234.
  3. In SayWrite › AI, set Quality server to LM Studio and choose the model, or leave it on Automatic. SayWrite prefers a model that's already loaded.
  4. LM Studio unloads idle models on its own schedule (its TTL setting), so the first rewrite after a break can be slower.

Tested by us · September 2026 · M5 Pro

Which models work well

Local models for AI Polish
ModelDownloadNotes
qwen3.5:9b (default)6.6 GBBest Polish results in our tests. about 2.4–2.6 seconds warm, about 5.6 seconds when it has to load; about 16 GB of memory while loaded. Picked automatically when installed.
Gemma family (e.g. gemma3:4b)3.3 GB for gemma3:4bNext in SayWrite's automatic ranking. Smaller Gemma models are a good fit for Macs with less memory.
llama3.2:3b · qwen2.5:3b2.0 GB · 1.9 GBLight and quick, non-thinking. Lighter rewriting than a 9B model.
qwen3.5:27blargerMeasured too: not better for Polish, and about three times slower. Not recommended.
Reasoning models (gpt-oss, DeepSeek-R1, QwQ…)variesSkipped by the automatic choice because thinking adds seconds. Allowed if you pick one; SayWrite asks for low reasoning.

Timings measured by the developer on a Mac with an M5 Pro chip, with the model already loaded unless noted. Your numbers will vary with your Mac and model. Download sizes are from the Ollama library.

Requirements

What you need (honestly)

  • An Apple silicon Mac with macOS 15 or later, like SayWrite itself.
  • Enough memory for the model you pick. qwen3.5:9b used about 16 GB while loaded on our test Mac. On an 8 or 16 GB Mac, start with a 3–4B model, or use Apple Intelligence, which needs no extra memory from you.
  • Patience for the first rewrite after a break: models unload when idle. SayWrite starts loading yours the moment you tap Fn + Space, and Keep model loaded (5, 15 or 30 minutes, or until Ollama quits) keeps it warm, at the cost of that memory.
  • Plain dictation needs none of this. AI is off until you turn it on.

Privacy

Why loopback only

“Local AI” means nothing if one typo sends your text to another machine. SayWrite accepts only 127.0.0.1, ::1 or localhost (rewritten to 127.0.0.1), refuses redirects and proxies, and re-checks the address before every request. LAN servers and cloud endpoints are refused even if you type them. The full list is on the security page.

Fixes

Troubleshooting

  • “Ollama isn't running”: open the Ollama app (or run ollama serve). In Auto mode SayWrite falls back to Apple's model and the pill says so.
  • “Model missing”: run ollama pull qwen3.5:9b, or pick an installed model in SayWrite › AI.
  • LM Studio doesn't answer: make sure the server is started in the Developer tab on port 1234 and a model is loaded. If it doesn't respond within a few seconds, SayWrite moves on to the other engine instead of hanging.
  • The first rewrite is slow: that's the model loading. Turn on Keep model loaded, or choose a smaller model.

Frequently asked questions

Can I use a remote Ollama server or a cloud API?

No. SayWrite only connects to 127.0.0.1, ::1 or localhost, and refuses other addresses even if you type them. That's deliberate: it guarantees your text never leaves your Mac. If you need a cloud model, SayWrite isn't the right tool.

Does my audio go to Ollama or LM Studio?

No. SayWrite transcribes your speech on the Neural Engine first. Only the resulting text is sent to your local model for rewriting, and the reply is typed at your cursor.

Which model should I start with?

qwen3.5:9b is what SayWrite picks automatically when it's installed, and it did best in our Polish tests. On a Mac with less memory, try a 3–4B model such as gemma3:4b, llama3.2:3b or qwen2.5:3b, or use Apple Intelligence.

Why does SayWrite skip reasoning models?

Reasoning models such as gpt-oss, DeepSeek-R1 and QwQ spend seconds thinking before they write, which feels slow for dictation. The automatic choice skips them. If you pick one yourself, SayWrite asks it for low reasoning effort and strips the thinking from the reply.

Does SayWrite download models for me?

No. SayWrite never pulls models. You install them in Ollama or LM Studio, and SayWrite lists what's installed, with sizes.

Is this better than Apple Intelligence?

For heavy rewriting, usually yes: in our tests a 9B model restructured rambling speech more than Apple's on-device model, which is faster but mostly copy-edits. Auto mode uses both: your model for Polish and emails, Apple's for quick fixes.

Say it messy. We'll write it neat.

Hold Fn, talk, and let SayWrite type it. Everything stays on your Mac.

Free during early access · macOS 15+ · Apple silicon