IndustriesOther

Enterprise AI Spending Shifts to Multi-Model Allocation as Token Costs Surge 10-100 Times

Published: Updated: By 24TopNews Editorial Desk

Enterprises are shifting from indiscriminate use of frontier AI models to multi-model fine-grained allocation based on task difficulty, driven by rapid adoption of low-cost open-source models. Token costs have surged 10 to 100 times in the short term, with examples including a financial institution's commitment rising from $10 million to $60 million in three months. Chinese open-source models from Zhipu, Moonshot AI, Alibaba, and DeepSeek are widely trialed due to costs at 10% of training and 10-20% of API pricing of overseas peers, while maintaining 20-40% gross margins. OpenRouter data shows US enterprise token usage of Chinese models has stayed above 30% since February 2026, reaching 46% in some periods.

The rapid proliferation of low-cost open-source AI models is fundamentally changing how enterprises deploy AI. Companies are moving from a phase of indiscriminate use of the strongest models to a new efficiency-first stage of fine-grained allocation based on task difficulty. The focus of enterprise AI management is shifting from token maximization to value maximization, measuring the business return generated per unit of token.

Enterprise AI adoption continues to grow rapidly. Several AI-native companies have seen extremely fast revenue growth over the past six months: one company's revenue rose from $400 million in January 2026 to $1 billion; another tripled its revenue within the year to exceed $500 million; and a third grew from zero to $500 million in 14 months. The acceleration is mainly driven by improved model capabilities, expansion of agent applications that consume more tokens, and some companies shifting from per-seat pricing to token- or usage-based billing. As adoption scales, cost issues have quickly surfaced. Token cost increases are staggering: a financial institution initially committed $10 million to Claude, which grew to $60 million after three months; an AI company's spending on Anthropic rose from $20,000 in December 2025 to nearly $1 million in July 2026, a 50-fold increase in seven months; some companies' token bills have grown 10 times or even 100 times in a short period. Enterprises are not stopping AI usage but are reducing waste.

Many routine tasks were previously assigned to the most expensive frontier models. One bank spent hundreds of thousands of dollars per month on high-priced models, yet employees used them only to check weather, restaurants, and meeting locations.

Enterprises no longer view the AI market as a zero-sum competition among a few frontier model developers. Instead, they are rapidly adopting a multi-model strategy, mixing frontier models from OpenAI and Anthropic with open-source and custom models. Avoiding dependence on a single AI supplier has become a key consideration for large enterprises. Chinese open-source models are the biggest beneficiaries of this trend. Models from Zhipu, Moonshot AI, Alibaba, and DeepSeek are being increasingly trialed and even deployed by enterprises. Although some highly regulated industries remain cautious, acceptance is steadily rising in secure environments provided by cloud platforms such as AWS and Microsoft Azure. The training cost of China's leading models is about one-tenth that of overseas leaders, and inference API pricing is only 10% to 20% of comparable overseas models, while still maintaining healthy gross margins of 20% to 40%.

Data from the OpenRouter platform shows that the token share of Chinese AI models used by US enterprises through the platform has remained above 30% since February 2026, reaching 46% in some periods. NVIDIA's open-source model Nemotron is among the most frequently mentioned US open models.

Open models do not reduce GPU demand; instead, they shift computing power toward more cost-effective inference tasks. Most open models still use NVIDIA hardware for training, fine-tuning, and inference. Low-cost models expand the scope of AI applications, increasing demand for inference computing, storage, networking, and deployment infrastructure. NVIDIA itself has launched the open-source Nemotron 3 Ultra model with 550 billion parameters, offering up to five times faster inference speed and up to 30% lower usage cost. Cloud service providers are also well positioned, as their platforms already support multiple AI models. Amazon, Microsoft, and Google are in a favorable position, with platforms like AWS and Azure already capable of handling multiple models. Regardless of which model enterprises choose, they still need cloud-based inference computing power. Hyperscale cloud providers continue to face supply constraints and strong customer demand for AI computing resources.

Frontier model developers are most vulnerable to cost reductions in the short term. The outlook for software vendors is more pessimistic. With multi-model becoming the norm, business models that rely purely on model call intermediation layers are facing severe challenges.

24TOPNEWS IMPACT INTELLIGENCE

Why this event matters

The event has a measured impact on 2 industrys. The strongest current signal is positive for Cloud Services & Data Centres, with intensity 74/100 and 85% confidence over a short term horizon.

Technology · 10.3

Cloud Services & Data Centres

Direction
positive
Intensity
74
Confidence
85%
Horizon
Short term
Effective impact +45
Technology · 10.4

Artificial Intelligence

Direction
mixed
Intensity
68
Confidence
80%
Horizon
Medium term
Effective impact 0

Impact figures are analytical estimates that combine direction, intensity, confidence and event importance. They are not investment advice.