Read news on LLM Inference with our app.
Read more in the app
Apple Silicon and macOS VMs: Faster LLM Inference with llama.cpp
The CPU is back: Rethinking the CPU-GPU split for LLM inference
Predictive Speculative KV Replication for Bursty LLM Inference
Hetzner is working on LLM Inference
DeepSeek open sources DSpark, a new framework to speed up LLM inference by up to 85%
Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0
Defeating Nondeterminism in LLM Inference
PyTorch 2.8 Released With Better Intel CPU Performance For LLM Inference
Making AMD GPUs competitive for LLM inference (2023)
How We Optimize LLM Inference for AI Coding Assistant
Benchmarking LLM Inference Back Ends: VLLM, LMDeploy, MLC-LLM, TensorRT-LLM, TGI
AMD's MI300X Outperforms Nvidia's H100 for LLM Inference
How attention offloading reduces the costs of LLM inference at scale
Show HN: Speeding up LLM inference 2x times (possibly)
Effort – a possibly new algorithm for LLM Inference