H Company launched NeoMME, a family of 260 million and 800 million parameter multilingual encoders for text and images. The 260M model posted a 0.523 nDCG@10 score on the ViDoRe v3 benchmark, nearly matching the 0.524 score of the 3.75 billion parameter ColQwen2.5-v0.2 with 14 times fewer parameters. These single-tower encoders remove the vision tower and causal decoder, enabling the search of document pages, charts, and tables without an OCR step.
The models were trained on 524 billion packed input tokens using a masked-diffusion text denoiser and the NorMuon optimizer. The 260M model was developed using 16 H100 GPUs, while the 800M variant required 32. NeoMME supports a 16,384 token context window, sufficient for two 3840x2160 4K UHD images.
Key sources
- SOURCE@tomaarsen“search document pages with MultiVectorEncoder, charts & tables included, without an OCR step”x.com
- SUPPORT@tomaarsen“Their 800M model reaches 0.556”x.com
- SUPPORT@marktechpost“16,384-token context, enough for two 3840×2160 4K UHD images”x.com
- SUPPORT@lateinteraction“make NeoMME day-0 compatible with ST thanks to Tom's help”x.com
- SUPPORT@tonywu_71“NeoMME day-0 compatible with ST thanks to Tom's help”x.com
- SOURCEhuggingnewshuggingnews.com