Why this matters now

On August 15, 2026, Alibaba announced that its Qwen family of open models had passed 3 billion cumulative downloads, overtaking Meta’s Llama and Google’s Gemma as the most-downloaded open AI model family in the world. Bloomberg, Fortune, Global Times, and China Daily all ran the milestone within 24 hours, and Alibaba’s stock (BABA) ticked up on the news.

This is the first time a Chinese lab has held the top spot in the open-model race, and it did not happen by accident. Qwen shipped more open-weight releases in the past year than Meta and Google combined: the Qwen3 line, Qwen3.5, the MoE-heavy Qwen3.6 series, and this month’s Qwen3.8 2.4T open weights with a claimed match on Anthropic’s Fable 5. The open-weight market now has a clear volume leader, and it changes which model most builders reach for first when they need weights they can actually run.

What changed

  • 3 billion cumulative downloads across the Qwen family, per Alibaba’s announcement on August 15.
  • Passed Meta and Google — Llama and Gemma were the previous download leaders on open model platforms.
  • 462 Qwen model repositories now exist on Hugging Face alone, spanning 0.6B-parameter edge models to the 2.4T-parameter flagship (verified via the Hugging Face API on August 17).
  • ModelScope is the other half of the story — Alibaba’s own platform carries a large share of Qwen downloads, especially from China-based developers, which is where much of the 3B figure comes from.
  • The milestone lands the same week Qwen3.8 shipped under two different licenses (Apache 2.0 for most weights, a separate Qwen license for the largest ones), and a week after Claude Code added Qwen support-adjacent tooling momentum (Cloudmagazin reported Qwen models running inside Claude Code on August 3).

Qwen organization social card on Hugging Face, the platform where the family holds the top download counts in open AI.

Alibaba’s Qwen organization on Hugging Face — 462 repos and the top download counts in open AI.


The 3 billion number, checked

Vendor download milestones deserve a skeptical read, so I pulled the numbers directly from the Hugging Face API on August 17, 2026:

OrganizationRepos countedTotal downloads (HF only)
Qwen462351,032,311
meta-llama7034,348,638
google1,000+ (pagination cap)148,676,272

Three takeaways:

  1. Qwen leads on Hugging Face too — but by ~2.4x over Google, not by 10x. The 3B figure is cumulative across Hugging Face and ModelScope, where Qwen downloads in China dwarf everything else. On HF alone, Qwen sits around 351M.
  2. The lead compounds at the top. Qwen’s five most-downloaded models — Qwen3-0.6B (29.2M), Qwen3-8B (16.0M), Qwen3.5-9B (14.3M), Qwen2.5-7B-Instruct (12.3M), Qwen3.6-35B-A3B-FP8 (11.6M) — each out-download Meta’s entire catalog roughly one-for-one. Llama’s 70 repos total 34.3M; Qwen’s top repo alone is at 29.2M.
  3. Meta’s decline is structural, not statistical. Meta has kept Llama releases gated behind a custom license and a download form for years, and its cadence slowed while Qwen shipped a new generation roughly every two months. Google’s Gemma family is Apache 2.0 and healthy at ~149M, but it targets smaller sizes and never built the China distribution channel.

None of this makes the 3B number dishonest — it is cumulative, cross-platform, and company-reported, which is exactly how Meta’s own multi-billion download claims were framed at their peak. But the gap between “3B cumulative” and “351M on HF” is the kind of nuance worth keeping in mind when a vendor milestone hits your feed.


Why Qwen won the open-model race

Alibaba did three things Meta and Google did not:

1. Shipped every size, constantly. The Qwen zoo runs from 0.6B to 2.4T, with dense and MoE variants, VL (vision), ASR, TTS, and embedding models. A builder can pick a Qwen model for a phone, a laptop, and a data center from the same family, with the same tokenizer and chat template. Llama covers big and small, but Gemma is capped at 27B and Meta’s releases are quarterly at best.

2. Bet on verifiable, RL-trained coding models early. The Qwen3 generation’s hybrid thinking mode and the RLVR-heavy Qwen3.6 series made the family a default for local agent experiments. The Qwen 3.8 Max open-weight launch pushed that further with a frontier-sized claim at open weights.

3. Owned the China distribution channel. ModelScope gives Alibaba a captive, high-volume platform that Western labs cannot reach the same way. Chinese developers downloading Qwen from ModelScope is the single biggest contributor to the 3B figure — and it is a structural advantage Meta and Google simply do not have.

The strategy is working where it counts: downloads, ecosystem mindshare, and the narrative that open weights no longer mean American labs only. CNBC’s November 2025 feature on how Alibaba quietly became an AI leader aged well; this milestone is the public confirmation.


Independent testing & community response

The download race and the quality race are different races, and the community knows it:

  • The 2.4T claim is still unverified. Alibaba’s August 13 claim that Qwen3.8-2.4T “matches Fable 5” drew immediate pushback on r/LocalLLaMA, where the recurring complaint is that Alibaba’s benchmark tables compare against publicly documented numbers rather than head-to-head runs. The Qwen3.8 2.4T launch coverage lays out the claim and its caveats in detail.
  • The license split matters. SQ Magazine flagged on August 14 that Qwen3.8’s largest weights ship under a separate Qwen license, not Apache 2.0. Download counts treat both the same, but enterprises with strict license policies cannot treat the whole family as equally open. The Apache 2.0 core — Qwen3.6-35B-A3B and below — is what most builders actually deploy.
  • The positive take is real. Independent testers consistently find the smaller Qwen3.6 MoE models punch above their size on coding and agent tasks, and the AI model price war coverage shows Qwen API pricing undercutting Western equivalents by wide margins. Artificial Analysis rankings have had Qwen models as the price-performance leaders on several size bands all year.
  • The skepticism angle is the usual one: download counts include quantized variants, fine-tunes, and repeated pulls by the same users. Llama’s numbers were inflated the same way in 2024 — which is precisely the point. By the metric everyone used to crown Llama, Qwen now wins.

What this means for builders

  • Default open-weight pick is shifting. For a new project with no hard dependency on a Western vendor, Qwen is now the safest ecosystem bet: the most models, the most quant coverage (GGUF, AWQ, FP8), first-class vLLM/SGLang/Ollama support, and the largest community producing fine-tunes. The Kubernetes-moment argument for open weights applies more strongly to Qwen than to any other family today.
  • Check the license per model, not per family. Apache 2.0 covers the mid-size core; the largest 3.8 weights do not. If your legal team cares, pin the exact model card before you commit.
  • Meta is not dead. Llama 5-class models remain competitive and Meta’s distribution is still the widest in Western enterprise. But its download lead is gone, and nothing in Meta’s release cadence suggests it wants the volume crown back.
  • Expect the Western response to be political, not technical. The U.S. has already pushed open-weight export controls toward Chinese labs; the “China leads open AI” narrative will accelerate that pressure. Builders should treat open-weight supply chains the same way they treat any single-region dependency — the same lesson DeepSeek’s pricing shift taught — and keep a second source warm.

Final takeaway: 3 billion downloads is a marketing number with a real signal under it. Qwen is the most-downloaded open model family in the world, it got there on breadth, cadence, and a captive China platform, and it is now the default first pick for a large share of open-weight builders. The honest caveat: on Hugging Face alone the lead is ~2.4x over Google, and the quality gap at the frontier is still being litigated in benchmarks that Alibaba has not fully opened.


Sources


About the author

Charles Jasthyn De La Cueva is a full-stack developer and the founder of Open TechStack. He writes about AI engineering, developer tools, and practical model evaluation — grounded in real workflows, not press releases.