Read news on inference with our app.
Read more in the app
MacPaw taps Liquid AI to offer on-device inference to devs building for its app store
Qwen3.8-Max: A New Bar for Coding and Cowork
Codex starts encrypting sub-agent prompts
AMD ZenDNN 6.0 Brings Many Improvements For Accelerating Inference On Ryzen/EPYC CPUs
Hot French startup ZML releases free product to speed inference across lots of AI chips
Inference cost at scale with napkin math
KV Cache Is Becoming the Memory Hierarchy of Inference
Inference is giving AI chip startups a second chance to make their mark
Train-to-Test scaling explained: How to optimize your end-to-end AI compute budget for inference
Google Gemma 4 Runs Natively on iPhone with Full Offline AI Inference
Your developers are already running AI locally: Why on-device inference is the CISO’s new blind spot
The team behind continuous batching says your idle GPUs should be running inference, not sitting dark
Launch HN: RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon
Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model
TTT-Discover optimizes GPU kernels 2x faster than human experts — by training during inference
Inference startup Inferact lands $150M to commercialize vLLM
Quadric rides the shift from cloud AI to on-device inference — and it’s paying off
Nvidia just admitted the general-purpose GPU era is ending
ChunkLLM: A Lightweight Pluggable Framework for Accelerating LLMs Inference
Tensormesh raises $4.5M to squeeze more inference out of AI server loads