← All posts

Splash vs Ollama vs MLX on an M4 Max: is “2× faster” real?

· Alexandru Tudorica

Splash is a new open-source inference engine for Apple silicon from Inco AI. Its pitch is “two times faster, no configuration”. I run local models for coding agents every day, so I tested that claim on my own machine against Ollama and Apple’s mlx-lm, with contexts from empty up to 200,000 tokens.

The short version: for writing code, Splash undersells itself. For reading a long prompt the first time, it’s barely faster than anything else, and on an M4 that means minutes.

What Splash is

Most inference engines try to run every model. Splash runs two: Qwen 3.8 27B and Qwen 3.6 35B-A3B. Supporting so few models lets it do three things a general engine can’t:

  • A trained draft model. Splash ships DFlash 2, a small model trained to guess a block of upcoming tokens for that specific big model. The big model checks the whole block in one pass.
  • Custom GPU kernels written for that model’s exact layer shapes.
  • A memory plan worked out at startup, so there’s nothing to tune.

Inco’s launch numbers, on a 48 GB M5 Pro: 74 tokens per second for Splash, against 38 for oMLX and 24 for Ollama. Those are the maker’s own benchmarks, so I ran mine.

The setup

  • Machine: MacBook Pro, M4 Max with the 40-core GPU and 128 GB of memory.
  • Model: Qwen 3.8 27B at 4 bits on every engine.
  • Engines:
    • Splash 1.0.2 with its DFlash 2 draft model.
    • Ollama 0.34.3 on its MLX backend.
    • mlx-lm 0.31.3, Apple’s reference library, with the same weights Splash starts from and no speculative decoding. This is the baseline.
  • Tests: greedy decoding, thinking off, 400 output tokens. Two tasks: writing a Python LRU cache class (code), and a ~300-word explanation of Python’s import system (prose). The long context is real Python standard-library source.

Round 1: short prompts

Bar chart of generation speed with a near-empty context. Writing code: mlx-lm 29, Ollama 65, Ollama with DFlash2 62, Splash 122 tokens per second. Writing prose: mlx-lm 29, Ollama 53, Ollama with DFlash2 48, Splash 64.

With an empty context, plain mlx-lm does 29 tokens per second on both tasks. That’s close to the ceiling for this chip: each token means streaming all ~15 GB of weights through the GPU, and memory bandwidth only allows so many trips per second.

Splash writing code does 122, over 4× the baseline and almost 2× Ollama’s 65. On prose the gap shrinks: Splash 64, Ollama 53, mlx-lm still 29.

The reason is speculative decoding. A small model guesses the next several tokens and the big model verifies them all in one pass, so a right guess gets you several tokens for the price of one. Code is predictable (indentation, self., closing brackets), so the guesses land far more often than on prose. mlx-lm doesn’t guess, which is why it doesn’t care what it’s writing.

Ollama isn’t the old Ollama

The launch post had Ollama at 24 tokens per second. On my Mac it did 65, which is impossible one token at a time. Ollama’s logs explain it: current Ollama also guesses ahead, using the multi-token prediction (MTP) head that ships inside Qwen 3.8. It gets about two extra tokens accepted per pass on code, and about one on prose.

So is Splash fast because of its draft model, or because of the engine around it? Inco publishes the draft separately, and there’s an open Ollama pull request that adds support for it. I built that branch and ran it with the same weights and the same draft:

Same weights, same Mac Draft Code, short (tok/s) Code, 200K (tok/s)
mlx-lm none 29 14
Ollama 0.34.3 built-in MTP head 65 11
Ollama PR build DFlash 2 62 18
Splash DFlash 2 122 55

With the same draft, the Ollama build is no faster than stock Ollama. Its logs show it drafting two or three tokens at a time, when the draft was trained to propose blocks of up to seven. The draft is half the story; the engine is the other 2×. That Ollama branch is an unmerged pull request, so this may change.

Round 2: long context

Coding agents fill the context with files, tool output and conversation history, so this is the round that matters most to me.

Line chart of code generation speed as the context grows from 0 to 200K tokens. Splash falls from 122 to 55 tokens per second. Ollama with DFlash2 ends at 18, mlx-lm at 14, and stock Ollama at 11.

Every engine slows down as the context grows, because each new token has to attend over everything before it. At 64K tokens it’s Splash 90, Ollama 34, mlx-lm 22. At 200K, Splash still does 55.

Stock Ollama ends at 11, slower than mlx-lm, which doesn’t guess at all. I re-ran it alone with nothing else loaded and got 11.8. Its draft is still being accepted at about 1.5 tokens per pass, so my guess is that its long-context attention is the slow part. Either way, at 200K Splash is 5× Ollama and nearly 4× mlx-lm. Splash’s lead grows with the context.

Prose follows the same shape at lower speeds: Splash goes from 64 to 32, Ollama from 53 to 18, and mlx-lm from 29 to 14.

The catch: reading the prompt

Everything above is about writing tokens. Before a model writes anything, it has to read the prompt (prefill), and here the story flips.

Line chart of minutes until the first token with nothing cached, by prompt size. All three engines take about 2.5 to 3 minutes at 32K tokens and 9 to 10 minutes at 100K. At 200K: Splash 23 minutes, mlx-lm 26, Ollama 29.

On the M4 Max every engine reads at a couple of hundred tokens per second, slowing as the prompt grows. A 32K-token prompt takes 2.5 to 3 minutes before the first word, whichever engine you use. 100K takes 9 to 10 minutes. 200K takes Splash 22 to 23 minutes, mlx-lm 26 and Ollama 29.

Splash is still the fastest reader, but only by 10–20%. The 2× and 4× numbers are writing speed; reading speed is roughly a tie.

What rescues Splash is caching. On the same 200K conversation, a follow-up question got its first token in 1.6 seconds instead of 22 minutes. Coding agents send the same system prompt, files and history every turn, so after the first slow read every turn is fast. Just don’t restart the server, and don’t make the agent reread everything each turn.

Does 200K actually work?

Speed doesn’t matter if the model can’t use the context. I buried three codes in a 200K-token dump of Python source, at 10%, 50% and 90% of the way through, and asked for all three. It found 3 out of 3 at 100K and at 200K.

The whole thing, model plus draft plus a 200K context, used 32 GB of memory. Qwen 3.8 keeps a growing cache in only 16 of its 64 layers, so context costs about 33 KB per token. A 48 GB Mac holds plenty of context; it just takes a long time to read in.

What about oMLX?

oMLX was the other engine in Inco’s launch comparison, so I tried to test it the same way, using an oMLX build with DFlash 2 support, the same weights and the same draft. On my machine the draft failed to load (“Received 23 parameters not in model”), and oMLX quietly fell back to plain decoding. That would have been an unfair comparison, so I left it out. If I get the draft working, I’ll update this post.

Verdict

  • Writing code: yes, it’s 2× faster, and then some. 2× today’s Ollama and 4× plain MLX with a short prompt, widening to 5× Ollama at 200K tokens. That’s exactly where coding agents spend their time.
  • Writing prose: a smaller win. About 20% over Ollama with a short prompt, about 2× at long context.
  • Reading big prompts: about the same as everything else on an M4, which means slow. Budget minutes, not seconds, for the first read. The cache hides it after that.

If you have an Apple-silicon Mac with 48 GB or more and run a local coding agent all day, Splash is the fastest way I’ve found to run this model. If you want lots of different models, or you’re on a 36 GB machine, stay with Ollama or LM Studio. Splash only runs two models.

Method notes

  • Weights: incoai/Qwen3.8-27B-Splash for Splash, qwen3.8:27b-mlx for Ollama, and mlx-community/Qwen3.8-27B-4bit for mlx-lm and the Ollama pull-request build.
  • Runs: one run per point, with one engine loaded at a time for the long-context and cold-start runs.
  • Caveat: these numbers come from one machine and will differ on yours. Some people testing reasoning-heavy maths report that the 4-bit weights fall apart on hard problems, and that 8-bit fixes it at about 60% of the speed. I tested with thinking off, on code and prose, so I didn’t hit that.

Sources: the Splash launch post, Splash on GitHub (Apache 2.0), and the DFlash 2 post.

If you’re deciding which hardware and engine to run local models on, get in touch.