Get the latest tech news
Better Vector Search for Long Documents: Chunking Inside Manticore Search
An embedding model reads only the first few hundred tokens of a document and silently drops the rest. Manticore Search now splits long documents for you at INSERT time: add chunk_strategy to the vector column and pick one of five strategies. No ingest pipeline, no splitter library. On our own manual, recall@5 for deep content went from 55% to 83%.
None
Or read this on Hacker News
