Alibaba Qwen3.8: 2.4T multimodal model claims 'second only to Fable 5' — no benchmarks published
Alibaba's Qwen team previewed Qwen3.8-Max on July 19 at Shanghai's World AI Conference, claiming it ranks second only to Anthropic's Fable 5. The model carries 2.4 trillion total parameters in a sparse Mixture-of-Experts (MoE) design and is the first Qwen model above 1 trillion parameters to be multimodal, handling text, images, video, and documents. Alibaba shares rose as much as 5.4% on the announcement. The claim comes three days after Chinese rival Moonshot released Kimi K3, also 2.8 trillion parameters, marking a rapid-fire volley of multi-trillion-parameter launches from Chinese labs.
The catch: Alibaba published zero benchmark scores, no model card, and no independent evaluation to support the "second only to Fable 5" ranking. The announcement stands as the company's own positioning. For context, Qwen3.7-Max, the prior flagship, scored 56.6 on the Artificial Analysis Intelligence Index (5th overall, top-ranked Chinese model) and 60.6 on SWE-Bench Pro—roughly 20 points behind Fable 5's 80.4 on the same test. Qwen3.8 is accessible through Alibaba's Token Plan at 10% of standard pricing, with open weights "promised soon."
This mirrors a recent pattern: Kimi K3 topped Arena.ai's front-end coding leaderboard but later showed ~51% hallucination rates on knowledge-retrieval tasks, illustrating the gap between announcement benchmarks and production reliability. No independent leaderboard—not Hugging Face's Open LLM, Artificial Analysis, or Arena.ai—has scored Qwen3.8 yet. The timing underscores China's push for visible frontier-parity claims following U.S. and European model announcements (GPT-5.6 on July 9, Bristol Myers on Vera Rubin, Mistral's €3B raise).
For architects: treat this as a genuine capability signal *paired with* a caution flag. Qwen3.7-Max's prior performance makes "second only to Fable 5" non-absurd, but it's unverified until third-party evaluations land. The open-weights release will be the proof point. Until then, the model's actual inference cost, active parameter count, and token-overhead profile remain unknowns—critical for cost-per-task planning in production.