Chinese AI Models Top OpenRouter Weekly Calls; Hugging Face Downloads Reach 41% Using MoE
In 2026, Chinese artificial intelligence models have gained global traction, with weekly call volumes ranking among the highest on OpenRouter and all top five slots held by Chinese firms in a late-July statistical week. Hugging Face's spring report showed Chinese open-source models accounted for 41% of global downloads, surpassing 10 billion cumulative downloads. The cost-performance advantage stems from widespread adoption of the Mixture-of-Experts (MoE) sparse architecture and full-stack engineering optimizations. Applications now span multiple industries and regions, from Southeast Asian government services to Middle Eastern energy and Latin American e-commerce.
On May 4, 2026, the Wensan Digital Life District in Hangzhou, Zhejiang Province, hosted an AI technology market where the public experienced DeepSeek's artificial intelligence models. On July 17, at the 2026 World Artificial Intelligence Conference held in Shanghai, Moonshot AI's KIMI model booth also attracted many visitors.
The capability and price advantages of Chinese AI models are drawing widespread attention from overseas developers. Many report that using Chinese models for a full day of complex development tasks costs significantly less overall than comparable overseas products. One startup team said that after switching entirely to Chinese AI models, technology operating expenses dropped sharply, and some AI projects previously shelved due to cost were restarted. From North American tech communities to Southeast Asian startup teams, from European independent developers to Asian technology enthusiasts, acceptance of Chinese AI models continues to rise.
Public data from global AI model aggregation platform OpenRouter shows that since the start of 2026, Chinese models have consistently ranked among the global leaders in weekly call volumes, with multiple flagship models occupying the top of the call-volume charts for extended periods. In a statistical week in late July 2026, the top five on the platform's call-volume list were all from Chinese companies. On the open-source community Hugging Face, China has become the largest model provider by monthly downloads. The community's spring 2026 report showed that downloads of open-source models developed in China accounted for 41% of the global total, with cumulative downloads exceeding 10 billion. Over the past year, several Chinese open-source models have consistently ranked high on the platform's download charts.
Overseas technology figures have repeatedly publicly praised the performance of Chinese open-source AI models. Nvidia Chief Executive Jensen Huang has on multiple occasions praised Chinese open-source AI models as "excellent," and Tesla CEO Elon Musk has also expressed appreciation for Chinese AI technology.
Chinese models have entered practical application overseas. A Singaporean engineer trained a model adapted to local needs based on a Chinese open-source model, serving the local digital economy. American creators use Chinese video generation models to produce professional-grade visual content, lowering the barrier to creative production. From government services in Southeast Asia to the energy industry in the Middle East, from startups in Europe and the United States to e-commerce platforms in Latin America, Chinese models have formed an application footprint spanning multiple industries and regions.
The dual advantage of performance and cost in Chinese models is directly linked to the large-scale adoption of the Mixture-of-Experts (MoE) sparse architecture. Unlike traditional dense models that activate all parameters for each inference, the MoE architecture uses a gating network to call on a small number of expert subnetworks on demand, leaving most parameters dormant, thereby reducing computational consumption while ensuring output quality. Several leading domestic model teams have adopted this technical route and are undertaking engineering optimizations for common industry issues such as routing load imbalance and communication bottlenecks.
In full-stack engineering optimization, domestic teams use native low-precision quantization to reduce model memory footprint, employ KV cache and multi-level caching to reduce redundant computation, and optimize system-level communication and pipeline scheduling to unlock computing cluster potential. The large-model industry has thus formed a virtuous cycle of "technology iteration, cost reduction, and application explosion."
China's intelligent computing scale ranks among the world's largest, green electricity supply capacity continues to strengthen, and large-scale data center construction leverages economies of scale to lower per-unit computing costs. Domestic computing chips are continuously being adapted and entering mass production, with the domestic AI industry system building corresponding supporting capabilities.
Leading Chinese large-model teams generally adopt a strategy of open weights or partial open-sourcing. With model weights open, global developers can participate in defect identification and contribute optimization solutions, while feedback from real-world scenarios flows back to model teams, driving capability iteration. International mainstream AI development platforms have integrated Chinese models, and overseas inference platforms treat the launch of new Chinese models as industry progress. Some overseas developers conduct secondary development based on Chinese models, adapting to local languages, industry scenarios, and cultural needs.
China has a vast internet user base and a complete industrial application ecosystem. Scenarios such as e-commerce customer service, document processing, code development, content creation, industrial quality inspection, and government services continue to generate application demands for large models. Domestic enterprises' needs for processing contracts of hundreds of pages and entire codebases have driven Chinese models to be widely equipped with ultra-long context windows.
Progress in capabilities such as agents and code generation is related to the digital transformation needs of the domestic internet and manufacturing sectors. Some overseas developers' real-world tests show that Chinese models perform stably in terminal task execution, code debugging, and multi-tool collaboration, improving development efficiency and offering attractive overall usage costs. After being honed by massive user scenarios, Chinese models demonstrate more stable handling of ambiguous requirements and abnormal situations in real-world settings.
Why this event matters
The event has a measured impact on 1 industry. The strongest current signal is positive for Artificial Intelligence, with intensity 90/100 and 85% confidence over a medium term horizon.
Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.