Hugging Face Blog·· 28 d ago
NeoMME: an efficient Multimodal-native and Multilingual Encoder
NeoMME: an efficient Multimodal-native and Multilingual Encoder
AI summary
Hugging Face launched NeoMME in 260M and 800M sizes. A single bidirectional Transformer handles text tokens and 32×32 image patches, trained from scratch with a masked discrete diffusion objective rather than pretrained vision towers or causal language models.
Selection record
Threshold 60Official, first-handFirst 41Second 41
Not admittedSum of both 82 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:介绍多模态编码器模型及检索应用
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Hugging Face Blog · huggingface.co