Moonshot AI releases weights for Kimi-K3, firing a shot across the bow of OpenAI and Anthropic - open-weight model performs almost as well as frontier models while being 2-3x easier to run
Drone-Bench: Tracking simple drone surveillance capabilities of frontier models
The people testing AI for danger can't keep up | The pace of AI development combined with soaring compute costs are squeezing the AI researchers responsible for evaluating frontier models — just as those models' capabilities are becoming harder to measure.
The real prices of frontier models
Trump signs AI executive order seeking 30-day government access to frontier models before release
DeepSeek previews new AI model that ‘closes the gap’ with frontier models
Frontier models are failing one in three production attempts — and getting harder to audit
PA bench: Evaluating web agents on real world personal assistant workflows
Psychometric Jailbreaks Reveal Internal Conflict in Frontier Models
TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task
Inside the US Government's Unpublished Report on AI Safety | The National Institute of Standards and Technology conducted a groundbreaking study on frontier models just before Donald Trump's second term as president—and never published the results.