Read news on tok with our app.
Read more in the app
Qwen 3.8 27B available on Cerebras at 1500 tokens/s
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone
Run Kimi K3 using 29 GB of RAM at 0.50 tok/s
Performance per dollar is getting faster and cheaper
We got 207 tok/s with Qwen3.5-27B on an RTX 3090
Qwen3.5-397B at 4.74 tok/s using 5.9GB RAM
Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5