Get the latest tech news

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s


Qwen3.8-Flash-Next on any consumer hardware: one-click install for Windows / Linux. Strata inference engine, OpenAI/Anthropic API on localhost, optional image input. - Niko1221/Strata

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of RTX 4090

RTX 4090

Photo of Qwen

Qwen

Photo of consumer hardware

consumer hardware

Related news:

News photo

Shapelearn Qwen 3.8 27B (13.1 GB VRAM)

News photo

Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks

News photo

Qwen 3.8 27B Uncensored – Testing the Uncensored Qwen Model