Get the latest tech news

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s


Run Qwen3.8-Flash-Next (125B MoE, 104 GB at 4-bit) on Macs with a fraction of that RAM by streaming experts from SSD. MLX + Swift, Ollama-compatible API. - carloslfu/slotstream

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of ~12

~12

Photo of GB Mac

GB Mac

Photo of tok

tok

Related news:

News photo

Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

News photo

Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

News photo

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s