Qwen3.8 Flash-Next on a 16 GB Mac

A local inference experiment on MacBook Air M3: the roughly 175 GB converted model stays on an external SSD, while selected expert weights are streamed into memory during generation.

Download the experimental Mference runtime, build instructions and recorded results below. No training or fine-tuning was performed.

Experimental fork of Mference. Historical 128-token runs: 2.705 / 2.775 tok/s; later repeats 1.733 / 1.773 / 1.758. The 2.7 range was not stably reproduced.

Selected historical measured results

Source archive, build instructions and Git history · All 100 timing rows

Download complete source + results ZIP · Download Git history

This page hosts code and evidence. The model runs locally on a Mac; no browser inference or weights are bundled.

Built on MIT-licensed Mference by Neel Mitra and contributors. No new training or sustained performance guarantee.