A Splash contributor reports that custom M1 Max kernels raised Qwen 3.8 27B generation from 18.9 to 39.1 tokens per second. Here are the test conditions and agent delays.
A developer using the GitHub handle paperniuk ran Qwen 3.8 27B on an M1 Max MacBook Pro with a modified version of Splash. In a public test report, they compared their Metal kernels with the stock kernels on the same laptop. Average generation speed across five short prompts rose from 18.9 to 39.1 tokens per second. Those are the contributor’s measurements. The upstream Splash instructions still require an M3 or newer chip.
The test used a 4-bit Qwen 3.8 27B model on an M1 Max with 32 GPU cores and 64 GB of unified memory. The contributor started a fresh server for each run, set temperature to zero and used a five-prompt benchmark capped at 250 output tokens. So 39.1 tokens per second describes generation in that benchmark, not every conversation or agent run.
In a Reddit post, u/Erp4759 described a separate 43-minute coding-agent session in opencode. Context grew from 30,000 to 89,000 tokens. Median generation speed ranged from 19.5 to 36.5 tokens per second as context length changed. After tool calls, the agent still had to wait for the first token of the next answer; that wait accounted for roughly half the elapsed time. Agent users should measure that delay alongside generation speed.
The same Reddit post explains that Splash’s stock kernels use operations for which the M1 lacks the needed hardware support. The branch rewrites some operations in formats the chip handles faster and fixes an error in how partial results are combined. The code remains separate; it was not part of the upstream installation instructions when we checked.
The 39.1 tokens-per-second result is no promise for every M1 Max. In a public LinkedIn comment, Maximilian Moore says he got about 17 tokens per second on another M1 Max MacBook with 64 GB of memory. He did not provide his prompts or settings, so this is not a repeat of the same benchmark. The branch contributor has not tested M2 either. Anyone trying the fork should compare it on their own tasks and measure generation speed separately from the wait after tool calls.