Junyang Lin (@JustinLin610) on X

X (formerly Twitter) ·

1 min read Original article ↗

user avatar

Jan 28, 2025

The burst of DeepSeek V3 has attracted attention from the whole AI community to large-scale MoE models. Concurrently, we have been building Qwen2.5-Max, a large MoE LLM pretrained on massive data and post-trained with curated SFT and RLHF recipes. It achieves competitive