As agentic AI takes off, so has the price tag of “tokenmaxxing.” Agentic requests can consume 15 times more tokens than human-led requests, making the cost of inference an increasingly important variable in model selection, with Chinese labs emerging as attractive suppliers.
A run of open-weight releases this summer, such as Moonshot AI’s Kimi K3 and Z.ai’s GLM-5.3, landed close to the U.S. capability frontier while remaining free to download and much cheaper to run. The U.S. still leads in capability, but China is competing aggressively on cost, openness and distribution.
For investors, the key question is whether cheaper models pressure pricing power for U.S. labs or lower adoption costs enough to expand the market. So far, they have not derailed the AI investment case, but may have pushed the industry into a new phase where cost and efficiency matter as much as capability and speed.
The AI Scoreboard
- Capability: The U.S. still leads, but Chinese models have narrowed the gap. The best Chinese open-weight models now sit within roughly one generation of the frontier,1 despite U.S. chip controls. And the frontier keeps moving: the newest U.S. models are trained on Nvidia’s latest Blackwell systems, while Chinese labs continue to squeeze more performance from constrained compute.
- Price: Chinese models remain cheaper, but the price advantage often narrows when measured per task. Token prices can overstate the advantage because models differ in how many tokens they use to reach an answer2. Separately, open-weight models may be free to download, but they are not free to serve, so chips, electricity and data center capacity place a floor on inference costs. U.S. models are also moving rapidly down the cost curve.3
- Usage: Chinese models are gaining traction, especially in agentic workloads. On OpenRouter, Chinese-origin models now process more tokens than U.S.-origin models, but token share is not market share. Long-running agents generate far more tokens than short interactions, and Chinese models skew heavily toward those workloads. Counted by requests, U.S. models still handle slightly more activity, and likely far more once enterprise workloads are counted.4
- Enterprise: Two ecosystems can coexist. Large U.S. companies still predominately use U.S. models while large Chinese companies tend to use domestic ones, reflecting regulation, procurement and data governance needs. Open weights are bridging that divide, allowing U.S. companies to host a Chinese-origin model privately or through a cloud provider, but security, reliability and legal risk should keep enterprise adoption more segmented than developer usage suggests.5
What this means for investors
Few developments have tested confidence in the durability of the U.S. AI boom more than the rise of capable, low-cost Chinese models, and volatility around new developments will likely continue. The biggest implication is around the economics of intelligence, as the cost of a given level of AI capability is falling quickly.
The U.S. and China are building parallel but increasingly interconnected AI ecosystems, with the U.S. leading at the frontier and China exerting growing influence on cost and diffusion. And AI leadership cannot be taken for granted. As past technology cycles have shown, early leaders do not always prevail.
For investors, that argues for diversification across model developers, enabling infrastructure and companies positioned to turn cheaper intelligence into lasting productivity gains.