*NeoMME*: an efficient Multimodal-native and Multilingual Encoder
3.60T1.5 sourceHugging Face Blog
Source record
Published by Hugging Face Blog (T1.5 source). The original is at https://huggingface.co/blog/Hcompany/neomme.
Pipeline notes
The summary and note below are generated by the signal pipeline — they are Beyond Desk’s reading, not quotations from the source.
SummaryNeoMME is a 260M/800M multilingual multimodal encoder that processes text tokens and image patches in a single bidirectional Transformer trained from scratch with masked discrete-diffusion. Fine-tuned for visual document retrieval, the 260M model encodes ~51 pages/sec on an L40S GPU, about 2× ColModernVBERT throughput, and cuts late-interaction index storage from ~1.5 MB to 6 kB per page (255× smaller) while retaining >95% of baseline nDCG@10. Released under Apache 2.0 in Hugging Face Transformers.
Why it mattersFirst-party release with concrete throughput and storage benchmarks for visual document retrieval. The 255× index compression and 2× throughput over ColModernVBERT are specific, verifiable claims practitioners can test directly.

Cited by
No citations on record.
