inference

Read news on inference with our app.

Read more in the app

MacPaw taps Liquid AI to offer on-device inference to devs building for its app store

Qwen3.8-Max: A New Bar for Coding and Cowork

Codex starts encrypting sub-agent prompts

AMD ZenDNN 6.0 Brings Many Improvements For Accelerating Inference On Ryzen/EPYC CPUs

Hot French startup ZML releases free product to speed inference across lots of AI chips

Inference cost at scale with napkin math

KV Cache Is Becoming the Memory Hierarchy of Inference

Inference is giving AI chip startups a second chance to make their mark

Train-to-Test scaling explained: How to optimize your end-to-end AI compute budget for inference

Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference

Your developers are already running AI locally: Why on-device inference is the CISO’s new blind spot

The team behind continuous batching says your idle GPUs should be running inference, not sitting dark

Launch HN: RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon

Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model

TTT-Discover optimizes GPU kernels 2x faster than human experts — by training during inference

Inference startup Inferact lands $150M to commercialize vLLM

Quadric rides the shift from cloud AI to on-device inference — and it’s paying off

Nvidia just admitted the general-purpose GPU era is ending

ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference

Tensormesh raises $4.5M to squeeze more inference out of AI server loads