tok

Read news on tok with our app.

Read more in the app

Qwen 3.8 27B available on Cerebras at 1500 tokens/s

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

Show HN: Maple-Preview – Ternary 20B MoE running at 120 tok/s on a iPhone

Run Kimi K3 using 29 GB of RAM at 0.50 tok/s

Performance per dollar is getting faster and cheaper

We got 207 tok/s with Qwen3.5-27B on an RTX 3090

Qwen3.5-397B at 4.74 tok/s using 5.9GB RAM

Deepseek R1 Distill 8B Q40 on 4 x Raspberry Pi 5