---
format: "aidr-story-markdown/v1"
id: "9aa79481e521b40b7bdcd14f7278e99eb2e86651110d35800f469af11eb729d5"
canonical_url: "https://aidr.today/9aa79481?lang=en"
title: "H Company NeoMME 260M Matches 3.75B Model With 14x Fewer Parameters"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-07T14:47:24.000Z"
category: "Research"
topics: ["neo-mme","fine-tuning","benchmark","multimodal","inference"]
source_urls: ["https://huggingnews.com/ai/h-company-neomme-260m-matches-375b-model-with-14x-fewer-parameters-058fefc3","https://x.com/tomaarsen/status/2096958665974767855","https://x.com/tomaarsen/status/2096958668336173214","https://x.com/Marktechpost/status/2096708134714900817","https://x.com/lateinteraction/status/2096969884420682014","https://x.com/tonywu_71/status/2096963912172249311","https://huggingnews.com/ai/h-company-neomme-encoder-matches-375b-model-with-14x-fewer-parameters-7978b1c2"]
summary: "H Company launched NeoMME, a family of 260 million and 800 million parameter multilingual encoders for text and images. The 260M model posted a 0.523 nDCG@10 score on the ViDoRe v3 benchmark, nearly matching the 0.524 score of the 3.75 billion parameter ColQwen2.5-v0.2 with 14 times fewer parameters. These single-tower encoders remove the vision tower and causal decoder, enabling the search of document pages, charts, and tables without an OCR step. The models were trained on 524 billion packed input tokens using a masked-diffusion text denoiser and the NorMuon optimizer. The 260M model was developed using 16 H100 GPUs, while the 800M variant required 32. NeoMME supports a 16,384 token context window, sufficient for two 3840x2160 4K UHD images."
---

# H Company NeoMME 260M Matches 3\.75B Model With 14x Fewer Parameters

> [Open the canonical story](<https://aidr.today/9aa79481?lang=en>)

**Published:** 2026-09-07T14:47:24.000Z
**Category:** Research
**Topics:** neo\-mme, fine\-tuning, benchmark, multimodal, inference

## Summary

H Company launched NeoMME, a family of 260 million and 800 million parameter multilingual encoders for text and images\. The 260M model posted a 0\.523 nDCG@10 score on the ViDoRe v3 benchmark, nearly matching the 0\.524 score of the 3\.75 billion parameter ColQwen2\.5\-v0\.2 with 14 times fewer parameters\. These single\-tower encoders remove the vision tower and causal decoder, enabling the search of document pages, charts, and tables without an OCR step\. The models were trained on 524 billion packed input tokens using a masked\-diffusion text denoiser and the NorMuon optimizer\. The 260M model was developed using 16 H100 GPUs, while the 800M variant required 32\. NeoMME supports a 16,384 token context window, sufficient for two 3840x2160 4K UHD images\.

## Sources

- [Story source](<https://huggingnews.com/ai/h-company-neomme-260m-matches-375b-model-with-14x-fewer-parameters-058fefc3>)
- [Story source](<https://x.com/tomaarsen/status/2096958665974767855>)
- [Supporting source](<https://x.com/tomaarsen/status/2096958668336173214>)
- [Supporting source](<https://x.com/Marktechpost/status/2096708134714900817>)
- [Supporting source](<https://x.com/lateinteraction/status/2096969884420682014>)
- [Supporting source](<https://x.com/tonywu_71/status/2096963912172249311>)
- [Story source](<https://huggingnews.com/ai/h-company-neomme-encoder-matches-375b-model-with-14x-fewer-parameters-7978b1c2>)

