Will it run?
Products

Chinese labs skip mid-size models, launch only giant AI systems

By Marco Vane Clawpit staff
Chinese labs skip mid-size models, launch only giant AI systems

The new Hugging Face summer 2026 report shows Chinese labs have abandoned the incremental release pattern of small models followed by scaling, and now launch each month models ranging from 754 billion to 2.78 trillion parameters. On the American side the ceiling stayed below 130 billion for five of seven months, except for Nvidia’s Nemotron 3 Ultra at 561 billion in May-June and Thinking Machines Lab’s Inkling at 952 billion. The gap is not accidental; it reflects a completely different strategy about what developers actually need.

The report divides Chinese labs into two camps with opposite philosophies. Moonshot, MiniMax, Xiaomi and Z.ai publish almost exclusively models above 70 billion parameters, so a first-time developer’s initial encounter is a model too large to run on local hardware. Tencent and Alibaba Qwen cover the full spectrum, from sub-billion-parameter models to frontier-scale. Xiaomi and Meituan both crossed the trillion-parameter threshold this year, although neither appeared in the open-weight rankings a year ago.

Two factors enabled the “frontier-only” approach. Building a giant model ceased to be a technological differentiator, and the community introduced a quantization layer that makes a massive model runnable on modest hardware within days. This community-driven reliance turns size profiles into statements of intent rather than capability limits: a frontier-only portfolio bets on benchmark placement and API demand, while a full-spectrum portfolio bets on becoming a standard that developers build upon.

Hardware vendors now dominate the release cadence. AMD and Nvidia each added more than 200 new repositories this year, far ahead of all others; LiquidAI ranked third with roughly 100. Vendors realized that open models tuned to their silicon constitute a more convincing proof of efficiency than any marketing deck. Google and Meta, which defined the open-model landscape in earlier years, now rank well below Nvidia in new-release speed, and Meta’s shift to closed flagship models underscores the trend.

On the >100 billion-parameter scale, most US releases this year are not original but built on Chinese models. The few original US models at this scale are Inkling (952 billion) from Thinking Machines, Nvidia’s Nemotron 3 Ultra (561 billion) and Nemotron 3 Super (124 billion), and Arcee AI’s Trinity-Large (399 billion). AMD contributed many conversions but no original frontier model; its work enables efficient execution of Chinese models on American hardware, representing distribution and optimization rather than core research.

Public model repositories grew from 2.43 million to 2.96 million, datasets from 711 thousand to 1 million, and Spaces from 1 million to 1.44 million. The distribution remains extreme: 85.6 % of models recorded fewer than 200 downloads over their lifetimes, while 1.5 % of repositories accounted for 99.2 % of all downloads. Within this concentration, Qwen leads in local inference followed by Gemma, and AI agents are emerging as a significant force in the hub.