The Killing Line of China's Large Models Has Been Cut

By: www.theblockbeats.info|10/08/2026 07:16:28

Original Title: "The Killing Line of China's Large Models Has Been Cut"
Original Author: Sleepy

On October 8, Anthropic released Haiku 5.5. For short requests of no more than 100,000 tokens, the cost is $0.1 per million tokens for input and $0.5 for output, which is only one-tenth of the previous generation Haiku 4.5.

On the morning of the same day, Haseeb Qureshi, managing partner of Dragonfly, posted two scatter plots from Artificial Analysis on X, accompanied by the phrase:

"If you want the cheapest LLMs, you should now buy American."

The horizontal axis of the scatter plot represents the cost of completing a benchmark task, while the vertical axis represents the score. The dashed line connecting all optimal solutions is known as the Pareto frontier in microeconomics, where each point on the line indicates that you cannot find a smarter model at the same cost.

In June, the cheapest end of the dashed line featured names like Xiaomi MiMo, DeepSeek V4 Pro, MiniMax-M3, and Zhizhu GLM-5.2, almost all of which were Chinese large models.

By October 8, the entire line had been taken over by Anthropic and OpenAI. The Chinese models were left with only MiMo barely hanging on the edge, while Zhizhu GLM-5.3-Flash and DeepSeek V4.1 Flash were pushed down to the lower right of Haiku 5.5, indicating that they cost more and scored lower in comparison.

When the Hong Kong stock market opened that day, large model concept stocks fell sharply. At 10:26 AM, MiniMax's decline expanded to 10%, and Zhizhu dropped by 5.6%. The Hang Seng Tech Index fell by 1.93%, with the two new large model companies underperforming the market by nearly five times.

Over the past two years, China's open-source and cost-effective models had established a killing line in the global large model market. As long as prices were driven into the ground, closed-source models abroad would struggle to survive. Star companies in Silicon Valley lined up to replace their underlying pipelines with Chinese open-source large models.

Now, the table has been turned.

The killing line has been cut.

Formation of the Killing Line

When GPT-4 was first launched in March 2023, the official price for the 8K context version was $30 per million input tokens and $60 per million output tokens.

Intelligence was extremely scarce, with supply monopolized by a few Silicon Valley oligarchs, making it a complete seller's market. It was a honeymoon period where it was believed that "intelligence could be priced separately."

What broke this myth was a tech team incubated by a Chinese quantitative hedge fund. On May 6, 2024, DeepSeek launched V2 with a total of 236 billion parameters, activating only 21 billion per token; relying on the MLA architecture and extreme engineering cuts, the team reduced the KV cache by 93.3%, increasing throughput by 5.76 times. The saved computing power was directly converted into commercial pricing, charging 1 RMB per million input tokens and 2 RMB for output.

Domestic giants were almost dragged into the price war.

On May 15, ByteDance's Volcano Engine released Doubao, offering a corporate price of 0.0008 RMB per thousand tokens, claiming to be 99.3% cheaper than the industry average. On May 21, Alibaba Cloud announced a cliff-like price drop for its Tongyi Qianwen series, charging only 0.5 RMB per million tokens for the input side; hours later, Baidu Wenxin announced that its two main models would be free.

By August, DeepSeek further introduced hard disk caching technology, charging only 0.1 RMB per million for input tokens that hit the cache.

Across the ocean at that time, the giants had not fully followed suit. In July 2024, OpenAI released GPT-4o mini, reducing input and output to $0.15 and $0.6; but four months later, Anthropic launched Claude 3.5 Haiku, priced at four times that of the previous generation. Dario Amodei's logic at the time was that the model was smarter, and the pricing must reflect the intellectual advancement.

What truly sent chills down Silicon Valley's spine was January 2025.

DeepSeek-R1 emerged, with reasoning performance closely approaching OpenAI o1, while charging only $2.19 per million tokens, compared to o1's price of $60.

On January 27, Nvidia's market value evaporated by nearly $589 billion in a single day, setting a record in U.S. stock market history. Marc Andreessen called it the Sputnik moment in AI, and Microsoft CEO Satya Nadella lamented on social media that the Jevons Paradox had once again proven true.

The killing line was thus formed.

Open weights made this blade even sharper. Faced with closed-source models, customers could only do arithmetic between a few monopoly price lists; but once the weights were open, developers could deploy privately on their own clusters or run the same model on any third-party inference cloud. With multiple suppliers for the same model, customers also had the option to deploy it themselves, reducing the space for model providers to maintain high premiums.

A number of domestic American companies were the first to cross the river.

In October 2025, Airbnb CEO Brian Chesky publicly stated that the company's intelligent customer service agent heavily relied on Alibaba's Qianwen, with the entire system coordinating 13 models. Although it also integrated OpenAI's latest flagship, it was rarely called upon in the production line due to the availability of more cost-effective alternatives.

In the same month, Cognition launched SWE-1.5, claiming to be built on "a leading open-source foundation in a certain industry," which Zhizhu later confirmed to be GLM-4.6.

Shortly after, a16z partner Martin Casado revealed to The Economist that about 80% of the American AI startups pitching with open-source model architectures were using Chinese foundations.

In March this year, the code editor Cursor released Composer 2, pricing input at $0.5 per million and output at $2.5, with the team eventually admitting that the foundation was tuned from Kimi K2.5 from the Dark Side of the Moon.

Developers are voting with code and budgets. A report released by Mozilla showed that the token share of Chinese open-weight models on OpenRouter rose from less than 2% at the end of 2024 to over 45% in April 2026.

In the ecological niche of volume and structured calls, the proud closed-source models once had no power to fight back.

The Cost

Killing opponents does not necessarily mean living decently oneself.

In the same report from Mozilla, there is another set of comparative data. From May to September 2025, open models accounted for about 20% of the model usage on the OpenRouter platform, yet only garnered about 4% of the model layer revenue. This is a sample from a platform during a specific period and cannot represent the global market, but the gap between usage share and revenue share is already quite evident.

Creating value has never equated to capturing value.

In January this year, Zhizhu and MiniMax rang the bell consecutively. The subsequently disclosed mid-term performance allowed the outside world to see the costs and losses behind the revenue growth.

Zhizhu's revenue in the first half of the year was 954 million RMB, with revenue from the open platform and API business reaching 825 million RMB, accounting for 86.5%. The gross margin of this business has turned positive, rising to 24.6%; however, the company still recorded a net loss of about 2.072 billion RMB, narrowing by 12.1% year-on-year. Selling more tokens began to generate gross profit, but there is still a distance to cover all of the company's expenses.

MiniMax's revenue in the first half of the year was about $117 million, a year-on-year increase of 283.1%; the net loss was about $358 million, narrowing by 11% year-on-year. After excluding stock-based compensation, changes in the fair value of financial liabilities, and listing expenses, the adjusted net loss was about $293 million, expanding by 111.2% year-on-year.

After going public, the growth brought by low prices needs to undergo another level of scrutiny. How much gross profit can each call leave, and how much of this gross profit can cover R&D and operational expenses? Revenue growth is rapid, but the overall profitability of the company is still far away.

By late spring this year, the Agent wave pushed computing power consumption to a critical point.

The open-source project OpenClaw swept the global developer community, with cumulative token consumption on OpenRouter exceeding 8.52 trillion by early March, ranking first. In mid-March, Chinese models ran out 7.36 trillion tokens in a single week, a month-on-month surge of 56.9%.

GLM-5, launched on February 12, once topped the usage chart, and the surging demand instantly broke through the infrastructure quota. Developers both overseas and domestically began to complain about second-level delays and high-frequency throttling in the GLM Coding Plan, marking the first time that a leading Chinese AI team publicly called for more computing power.

On February 23, Zhizhu's stock price plummeted nearly 23% in a single day, evaporating over 70 billion HKD in market value.

What followed was a round of defensive price increases.

Zhizhu raised fees when launching GLM-5, and the price of GLM-5-Turbo in March increased by another 20%, averaging 83% higher than the previous generation specifications.

DeepSeek's V4-Pro, released on April 24, featured a total of 16 trillion parameters and offered a 75% peak discount upon launch. This fire first burned towards its peers, with MiniMax experiencing consecutive declines of about 9% and 10% over two trading days.

However, by mid-August, DeepSeek also shifted to a time-based billing strategy, with peak output for V4-Pro rising to $3.96 per million tokens, and cache hit costs increasing by over 12 times.

In July, the Dark Side of the Moon released Kimi K3, with API prices raised to $3 per million input and $15 per million output, more than three times the price of K2.6.

Surprisingly, the price increase did not trigger a significant loss of customers.

Zhipu AI's CEO Zhang Peng publicly disclosed that after an 83% price increase in the first quarter, the scale of API calls surged by 400%. By the end of August, the annualized revenue of Zhipu's MaaS platform reached $1.6 billion, and the gross profit margin of its API business miraculously turned positive at 24.6%.

During a few months of supply-demand imbalance in computing power, the Chinese model team first touched the precious pricing power.

However, the line that once kept overseas giants awake at night has unknowingly been pushed higher by themselves.

Blitzkrieg

Across the ocean, a strategic pivot was completed in the summer.

At the end of June, Anthropic launched Sonnet 5, marking a promotional price of $2 per million inputs and $10 per million outputs, clearly stating that it would revert to $3 and $15 in September.

By July, OpenAI released the GPT-5.6 series, and three weeks later, suddenly launched a blitzkrieg, reducing Luna's price by 80%, with input dropping to $0.2 and output to $1.2.

OpenAI attributed this price cut to a pure underlying engineering overhaul, with a rewritten GPU kernel reducing end-to-end service costs by about 20%, and a retrained speculative decoding draft model speeding up generation throughput by over 15%.

On August 10, Anthropic withdrew its price increase announcement, declaring that Sonnet 5 would permanently lock in low prices.

On September 22, Anthropic introduced Opus 5.5, reducing overall usage costs by 40%; on the same day, OpenAI launched GPT-6 Sol and Luna, locking Luna's entry price at $0.1 for input and $0.5 for output. Following this was Haiku 5.5.

Two years ago, Anthropic firmly believed that "intelligence must enjoy a separate premium," but now they have torn off their own label.

Now, the rhetoric during the release of large models has also changed. No longer merely listing slight victories in MMLU, they began to repeatedly emphasize engineering efficiency in terms of output per dollar, latency, and the consumption of each step of reasoning.

Moreover, at the unveiling of Haiku 5.5, Anthropic announced that it would fully distribute API quotas to subscription users, with Max-level users receiving $100 to $200 per month, and enterprise Team users enjoying up to $500 in deductions.

This move cuts to the core. The discounts on the price list only affect single-call decisions, but once customers' code engineering and Agent orchestration are adapted, the barriers to migration will be as high as a city wall.

The American giants dared to initiate this price war, relying on an unfathomable pool of funds and supply chain privileges.

Anthropic's annual operating revenue skyrocketed from about $9 billion at the end of 2025 to over $65 billion by the end of July this year. The $65 billion financing completed in May directly pushed its valuation to $965 billion.

In terms of computing power, Anthropic announced in April an expansion of its cooperation with Google and Broadcom, obtaining multi-gigawatt-level next-generation TPU computing capacity, expected to be gradually launched starting in 2027.

In contrast, Zhipu has struggled to raise over HKD 70 billion this year through placements and convertible bonds, which, converted to USD, is less than one-seventh of Anthropic's single round of ammunition.

Only the super giants can afford the money.

The secondary market's clearing has always been merciless.

Zhipu went public on the Hong Kong Stock Exchange on January 8, with an issue price of HKD 116.2. The spring's computing power shortage did not stop the bulls' enthusiasm; on June 22, its stock price surged to HKD 2980, with a market value briefly touching the HKD 1.33 trillion ceiling.

Then came a long decline.

With the lifting of the lock-up period in July, two rounds of placements drove the issue price down from HKD 1588 to HKD 714. On July 16, the next day after Kimi K3's debut, Zhipu's stock plummeted by 28.48%. On September 22, with the joint launch of Opus 5.5 and GPT-6, Zhipu's decline exceeded 10% again.

On September 25, Zhipu's stock dipped to HKD 610.5, with a total market value shrinking to around HKD 300 billion, nearly an 80% retracement from its historical peak.

MiniMax's trajectory was equally tragic. From a peak of HKD 1330, it plunged 17.98% on the day of the lock-up release, closing at HKD 216 in mid-July. Nearly 60% of its revenue in the first half of the year relied on overseas customers, and overseas is precisely the battlefield where the American giants' price-cutting storm first struck.

Valuation metrics are also tightening. According to media reports, Jefferies, in its September research report, lowered the valuation multiple for Zhipu's cloud business corresponding to the 2026 forecast annual revenue from 50 times to 30 times, a decrease of 40%.

Chinese models once seized market share from giants with extreme cost-effectiveness, but now this path has also been traversed across the ocean. Developers come for the cheap prices and will naturally leave for even cheaper options.

The Ghost of $0.1

On August 25, 2006, Jeff Barr typed an article introducing the beta version of Amazon EC2 on a hot beach in Cabo San Lucas, Mexico.

At the end of the technical description, he wrote down a number destined to enter commercial history: developers could rent a virtual computing instance for $0.1 per hour.

The subsequent story is well-known. AWS actively lowered prices over a hundred times in the following decade, with S3 storage rates halving by 51% in 2014 alone.

Cloud computing has never achieved dominance by selling equivalent computing power at increasingly higher prices. Scale, self-developed chips, and extreme operational efficiency have squeezed every marginal profit, which has been continuously fed back to the end market, thereby attracting more developers. The continuously lowering prices of AWS ultimately grew into a cash cow with nearly $130 billion in annual revenue and an operating profit margin of 35.4%.

Price wars are the final encirclement tactic of scale monopolists.

Today, those who can push model prices down to one-tenth and immediately fill the shelves of the three major public clouds globally are not just selling tokens but also digesting tokens. Anthropic is slashing API prices while using its vast inference computing power to feed Claude Code, Cowork, and its own end-to-end Agent products.

Chinese large model teams have not yet laid down their weapons. In September, DeepSeek once again lowered the fees for the Flash series, Xiaomi pushed MiMo-Flash to the limit of $1 per million tokens, and Zhipu also stubbornly priced GLM-5.3-Flash in the low-profit zone. But the rules of the game on the table have already been rewritten.

Low prices themselves are no longer an exclusive moat card.

On the scatter plot updated in real-time by Artificial Analysis, not a single Chinese company's name is absent. GLM, DeepSeek, Kimi, Qianwen, MiniMax are all still present.

However, they can no longer draw that line of execution.

Original link

-- Price

--
--
--

This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.

You may also like

iconiconiconiconiconiconiconiconicon
Customer Support:@weikecs
Business Cooperation:@weikecs
Quant Trading & MM:bd@weex.com
VIP Program:support@weex.com