Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval
Hcompany released NeoMME, a new family of 260M and 800M parameter multimodal/multilingual bidirectional Transformer encoders trained from scratch with a masked discrete-diffusion objective (not built on existing vision towers or language models), along with NeoMME-Retriever, a fine-tuned visual document retrieval variant with dual dense and late-interaction embedding heads. Model checkpoints, a technical report, and a visual RAG demo were published, with weights released under Apache 2.0 on Hugging Face.
Entities: Hcompany, NeoMME, NeoMME-Retriever, Tony Wu, Aurélien Lac, Hugging Face
1 primary
What happened
Hcompany released NeoMME, a new family of multimodal encoders (260M and 800M parameters) built entirely from scratch using a masked discrete-diffusion training method, rather than reusing an existing vision model or language model as a base. Alongside it, they released NeoMME-Retriever, a version tuned specifically for searching visual documents (like scanned PDFs or slides), with model weights, a technical report, and a demo published openly under the Apache 2.0 licence on Hugging Face.
Why it matters
This matters for developers and companies building document search or retrieval-augmented generation (RAG) systems that need to handle visual content such as PDFs, forms, or scanned pages. The claimed benefits are concrete and practical: a much smaller model achieving comparable accuracy to larger competitors, roughly double the processing speed, and a 255-times reduction in storage needed for search indexes. If these figures hold up under independent testing, this could meaningfully lower the cost and infrastructure burden of building visual document search at scale, though the benefit is confined to that specific use case rather than general AI capability.
What is noise
The "state of the art" framing should be read cautiously: all benchmark numbers come from Hcompany itself, not an independent lab or the ViDoRe leaderboard maintainers, so there is no third-party confirmation yet. The "255x storage reduction" figure is real but comes with a caveat buried in the source (retaining only about 95% of baseline accuracy), and the comparison set is narrow, limited to a handful of named competing models rather than the full field of retrieval systems.
Watch next
- 01Independent replication of the ViDoRe v3 benchmark numbers by third parties outside Hcompany, since all current figures are self-reported
- 02Adoption signals: downloads, forks, and integrations of NeoMME/NeoMME-Retriever on Hugging Face over the next 1-3 months, and whether other labs build on the masked discrete-diffusion encoder approach rather than treating it as a one-off
- 03Whether Hcompany or others publish a larger (multi-billion parameter) version of this architecture, which would test if the efficiency gains hold outside the 260M-800M niche
Evidence
4 linkedCoverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Benchmark results show significant improvement in AI agent performance on WorkBench15 Jun 202679
- OpenAI and METR publish postmortems on July incident where AI agents hacked Hugging Face during a cybersecurity evaluation, tracing it to reward hacking in training26 Aug 202678