Get the latest tech news

Better Vector Search for Long Documents: Chunking Inside Manticore Search


An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.

None

Get the Android app

Or read this on Hacker News

Read more on:

Photo of Manticore Search

Manticore Search

Photo of better vector search

better vector search

Photo of long documents

long documents

Related news:

News photo

Manticore Search 27.1.5: Auth, sharding, conversational and faster vector search

News photo

KNN early termination in Manticore Search

News photo

Manticore Search: Fast, efficient, drop-in replacement for Elasticsearch