inference

Read news on inference with our app.

Read more in the app

Qwen Image 2.1

A search-and-inference database from scratch in pure Zig

Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks

Google’s TPUv8s for Training and Inference at Hot Chips 2026

Kog is going deeper to squeeze more inference out of GPUs

MacPaw taps Liquid AI to offer on-device inference to devs building for its app store

Codex starts encrypting sub-agent prompts

AMD ZenDNN 6.0 Brings Many Improvements For Accelerating Inference On Ryzen/EPYC CPUs

Hot French startup ZML releases free product to speed inference across lots of AI chips

Inference cost at scale with napkin math

KV Cache Is Becoming the Memory Hierarchy of Inference

Inference is giving AI chip startups a second chance to make their mark

Train-to-Test scaling explained: How to optimize your end-to-end AI compute budget for inference

Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference

Your developers are already running AI locally: Why on-device inference is the CISO’s new blind spot

The team behind continuous batching says your idle GPUs should be running inference, not sitting dark

Launch HN: RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon

Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model

TTT-Discover optimizes GPU kernels 2x faster than human experts — by training during inference

Inference startup Inferact lands $150M to commercialize vLLM