Get the latest tech news

Needle: The benchmark your search engine can't memorize


Machines are now the majority of consumers of web content. AI agents will soon be the largest consumers of search. Agents need search because the information required to complete real tasks is often too fresh (”Who won the Dodgers game yesterday”) or too niche (”Can the same GM sensor be used across both 2007 and 2008 Silverado and Sierra platforms?”) to live in model weights or context. But search is not built for agents: every engine, including Keenable, falls short of what's achievable on agentic traffic. That gap is hard to measure, because overfitting and data leakage make standard search benchmarks like BrowseComp unreliable. To fix this, we introduce NEEDLE, a live, open-source benchmark for search engine quality. NEEDLE’s queries are drawn partly from real agent search logs and partly generated to reflect the search intents we observe in production. It runs continuously in public and anyone can reproduce it. All queries and metrics are on the live benchmark page, and the full evaluation code is in our GitHub repo. Contributions are welcome.

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of search engine

search engine

Photo of needle

needle

Photo of benchmark

benchmark

Related news:

News photo

Linux 7.3 Improving RAID 5/6 Benchmark-Based Algorithm Selection

News photo

GCC Patch Adjusting AMD Zen 5 Misprediction Cost Nets 12% Win In Benchmark

News photo

Firefox Announces Free VPN and 'Startpage' Search Engine Rolling Out to Android, iOS - Plus GeForce NOW Support