The new multilingual encoders allow users to search document pages including charts and tables without performing an optical character recognition step. These models integrate with the Sentence Transformers library through a MultiVectorEncoder to process text and images simultaneously. H Company released the tool in two sizes with 260M and 800M parameters.

Day-0 compatibility with the Sentence Transformers framework facilitates immediate fine-tuning and deployment for multimodal applications. This integration provides a standardized way for developers to implement high-efficiency retrieval for a variety of text-and-image datasets in the wild.

Sign in to suggest edits

Key sources

  1. SOURCE@tomaarsen“search document pages with MultiVectorEncoder, charts & tables included, without an OCR step”x.com
  2. SUPPORT@tonywu_71“NeoMME day-0 compatible with ST thanks to Tom's help”x.com
  3. SUPPORT@tomaarsen“Their 800M model reaches 0.556”x.com
Markdown