Get the latest tech news

Smaller, faster, safer: running Kimi and GLM at scale


Serving frontier models like Kimi and GLM means fighting for GPU memory. Here's how we quantize KV caches, compress model weights, and add integrity checks to serve them faster, cheaper, and safely.

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of scale

scale

Photo of glm

glm

Photo of Kimi

Kimi

Related news:

News photo

Moonshot's Kimi Built With Nvidia Compute

News photo

Moonshot’s Kimi Uses 20,000 Nvidia Chip Cluster From Alibaba

News photo

Making Postgres queues scale