WeMM-Embedding: Tencent's open multimodal embeddings leapfrog the 8B class

The WeChat Vision team at Tencent shipped WeMM-Embedding yesterday — three open models (2B, 4B, 9B) that reset the bar for the multimodal embedding class. The headline: the 2B variant scores 77.9 AVG on MMEB-v2, edging out Qwen3-VL-Embedding at 8B (77.8) and topping every open model at its own size. The 9B pushes to 80.6 overall — new SOTA — and leads the agent-task column on MMEB-v3 (51.0).

Why this matters beyond the leaderboard:

Embeddings are quietly becoming the substrate for agentic memory and retrieval — and agent tasks are exactly where WeMM's gains concentrate. This is the pattern worth watching: Chinese labs (Qwen3-VL-Embedding, GME, now WeMM) keep shipping open SOTA retrieval while the closed frontier chases inference scale. The retrieval layer is commoditizing faster than the market is pricing it.