Get the latest tech news

Learning to solve hard problems in RL for LLMs by never giving up


This is a blog post for my recent paper on RL post-training of LLMs: introducing the Matthew Effect and proposing to solve it with Never Give Up. It is presented interactively and less formally, more like how I give the talk. For a deeper, more technical dive, check out the paper on arxiv and code on github. What is your eval actually measuring? # Every good RL practitioner has no doubt seen an eval curve go up. Here is the AIME 2025 eval during our RL training of Olmo 3.1 RL-Zero Math(1) (1)see Olmo 3.1 blog post and arxiv

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of LLMs

LLMs

Photo of hard problems

hard problems

Related news:

News photo

LLMs are real, AI is fake

News photo

Show HN: StemJSON – a language for LLMs to extend native mobile apps on the fly

News photo

LLMs are real, AI is fake