
Kimi K3 Major Release: Is the China-US Large Model Competition Landscape Shifting Dramatically?
TechFlow Selected TechFlow Selected

Kimi K3 Major Release: Is the China-US Large Model Competition Landscape Shifting Dramatically?
The ultimate outcome depends on who can establish a closed loop of capabilities, revenue, data, and computing power. Currently, the US has a higher win rate, but China is changing the rules of the game.
TL;DR
If you were to bet on a game of Texas Hold'em on the sidelines, absolutely do not bet rashly in the early stages of the game.
- Starting hand, the US currently holds the advantage. Frontier models, high-end chips, dollar capital, cloud platforms, and global software entry points constitute the US's neatest hand. If the cards were revealed today, the US has a higher win rate.
- The board updates too fast; advantage is not a winning position. The leading window for large models is shrinking from a stable generation gap to a fluctuating time lag. Kimi K3's local overtaking in some long-context code evaluations shows that China's speed in reading, analyzing, and recombining cards has significantly accelerated.
- Compute determines chip depth; algorithms relabel chip face value. The US controls high-end chips, HBM, interconnects, CUDA, and super-large clusters, possessing more opportunities to bet simultaneously and iterate through trial and error; China, unable to access the same volume of high-face-value chips, seeks to increase the purchasing power of each chip.
- The US holds its hole cards close; China expands the table. OpenAI and Anthropic aim to turn capability leadership into repeatable "model rent"; Chinese vendors exchange low prices, open weights, and local deployment for distribution rights, competing for "ecosystem rent" formed by cloud services, industrial deployment, and development standards.
- Token price is just the bet amount; unit task cost corresponds to the pot. Cheapness is far from a showdown. The real contest is how much it ultimately costs to complete the same task. The US guards key tasks with the highest cost of failure, while China attempts to qualify a larger volume of ordinary tasks for the pot.
- Capital purchases both the right to trial and error and starts the countdown on returns. US private AI investment scale significantly leads, allowing simultaneous bets on more routes; China's capital is more dispersed, relying more on major tech firms, industrial capital, and policy forces. Both sides are burning not just money, but time borrowed from the future.
- Many scenarios do not equal a made hand. China lacks no business hand histories with real wins and losses, but only when results from transactions, fulfillment, quality inspection, and equipment operation can flow back into training will industrial scale translate into model advantage.
- Finally, it depends on who can control drawdowns and dodge table-flipping variables. Employment shocks, fraud, privacy incidents, Agent loss of control, capital bubbles, escalated blockades, and even the early arrival of AGI could all reprice today's most valuable chips.
Prologue: Players Seated, Texas Deal
Lights hang low from the ceiling, the green felt table absorbs the cold light from四周,leaving only chips, poker cards, and two pairs of eyes unwilling to blink first. The air mixes the psychedelic scent of coffee, cold smoke, server heat dissipation, and newly printed banknotes. It is the smell of someone preparing to mortgage the future.
The dealer shuffles with head bowed, fingers clean, movements gentle. Time is always like this. It deals the cards for you, but does not take responsibility for you.
Many people have sat at the AI table. Europe came carrying regulatory documents, Japan and Korea bit down on their positions in the chip supply chain, and the Middle East piled energy and oil capital onto the table stack by stack. But as the blinds rose, those truly able to keep calling round after round with frontier models, super-large clusters, and global entry points are mainly left to the US and China.
The stacks of chips before the US are piled like the Manhattan skyline. OpenAI, Anthropic, Google, and Meta are shown openly, unstoppable; Nvidia, cloud platforms, dollar capital, and global software entry points press underneath, solid and substantial. China's chips are not as neat, but are thickening rapidly: DeepSeek, Qwen, Kimi, Zhipu, MiniMax, StepFun, and the internet platforms, open-source communities, domestic chips, and vast application markets behind them have become thick enough that opponents can no longer bet leisurely according to old odds.
The dealer speaks not, flipping the cards, on which are赫然 written in aqueous gloss: Large Models.
ChatGPT was indeed shocking when it appeared, but initially entering user life, it was perhaps just a smarter Siri. People had it write poems, craft stories, explain quantum mechanics, and screenshot its serious nonsense to forward everywhere. Only a few years later, that chat box has unknowingly drilled into search, programming, office work, customer service, and research processes.
It has truly begun to take over work, or rather, productivity.
I. Frontier Model Capability: First Look at the Board, Then See if It Can Become a Hand
Whether playing Texas Hold'em or other card games, there is a simple truth: Being dealt a big card does not equal winning. Because the distance from a big card to a made hand is still a long and hard road.
Holding an A and K is certainly imposing, but if the flop has no connection, they are temporarily just an Ace-high. Conversely, two inconspicuous small cards, once connected with the community cards, might instead sneak through to form a straight.
Large models are exactly this starting hand.
1. First Look at the Board: US Has Big Cards, China Has Sat in the Front Row
Different Benchmarks measure different capabilities and cannot be directly mixed into a total score. But laying out several core leaderboards as of July 17, 2026, one can still see the general shape of the current table.
Big cards are increasing, but no single card can sweep all.
If looking only at the strongest models, the US still holds the larger starting hand. Anthropic, OpenAI, Google, xAI, and Meta constitute a whole row of frontier model fortresses, capable of taking the lead in turns across general reasoning, Coding, multimodal, and Agent. The US advantage is not some company accidentally rushing to the top of the list, but the overall thickness of frontier model supply.
But the Kimi K3 released today significantly narrows the distance between China and the largest board face. Third-party Artificial Analysis scores K3 at 57 points, ranking 4th on the comprehensive intelligence list, only two or three points behind Claude Fable 5 at about 60 points and GPT-5.6 Sol at about 59 points. Chinese models thereby begin to 贴近 the global comprehensive capability ceiling.
Coding is the card K3 played most beautifully. In public group tests, K3's Program Bench is 77.8 points, slightly higher than GPT-5.6 Sol's 77.6 points; SWE Marathon is 42.0 points, higher than Sol's 39.0 points; Terminal-Bench 2.1 is 88.3 points, closely following Sol's 88.8 points. In the Frontend Code Arena where users blindly select works, K3 tops the list with 1679 points and ranks first in six of seven frontend subfields.
The board face has therefore changed. The US still controls the ceiling for comprehensive capabilities, general experience, and systematic Agent capabilities, but China has already been able to win a few hands on the high-value community card of Coding.
The competition between Chinese and US frontier models is shifting from comprehensive catch-up to local wins and losses.
2. Game Pace: Generation Gap Is Becoming Time Lag
Benchmarks are still important, but they are gradually becoming less so.
The key is that they can increasingly only tell us what the freeze-frame of the table looks like at this second, but it is difficult to tell us who will still smile at the top of the list a few weeks later.
When model iteration was slower, one lead often meant a generation gap of half a year or even longer. At that time, a high score on the leaderboard was not just a beautiful report card, but also represented a wide enough technical moat. In March 2026, Stanford's "2026 AI Index" showed that US head models had squeezed into a narrow interval of less than 25 Arena Elo points, Qwen and DeepSeek also entered the frontier region; since 2025, Chinese and US models have swapped positions multiple times, and by March 2026, the public performance gap between top models on both sides was about 2.7%.
So, what benchmarks are losing is not the ability to measure, but the ability to predict the endgame.
"Chinese and US large models are only three months apart" is a statement appearing precisely against this background. This actually does not mean China is fixed three months behind the US in all directions, but that the competition between the two sides is changing from a relatively stable "generation gap" to a constantly opening and closing "time lag." Moonshot AI released the code-specialized model K2.7 Code on June 12, to launching the flagship K3 on July 17, only 35 days apart. Code, math, multimodal, Agent, and real product experience each follow their own clock, some separated by months, sometimes only weeks.
Accelerating the catch-up further is distillation. Student models do not need to see the teacher model's parameters, only need to learn extensively from its answers, code solutions, and tool usage methods, to possibly master part of the judgment path faster along the road the other party has already walked. Distillation itself is a common technology; the controversy lies only in whether competitors massively call another party's commercial model without permission.
In June 2026, Anthropic wrote to US Senators, accusing operators related to Alibaba and the Qwen team of interacting with Claude about 28.8 million times through nearly 25,000 fake accounts, attempting to extract its Agent reasoning, software engineering, and long-range task capabilities. Related claims come from Anthropic and should still be distinguished from independent investigation conclusions. But it at least reveals one thing: US frontier model companies no longer view only parameters, chips, and training code as strategic assets; even the answers generated by models are beginning to be viewed as potential outlets for leaking capabilities.
The US is still playing new cards more frequently, but China's speed in reading, analyzing, and recombining cards is already far faster than in the past. The top of the list has become a short-term action right, no longer naturally equaling long-term winning position.
II. Recalculating Win Rate: How Compute, Algorithms, Data, and Talent Change the Board
The tree wants to be quiet but the wind does not stop; the cards want to be quiet but the heart does not stop.
A player clearly holding better cards might lose because chips are too shallow, information is insufficient, or they simply did not understand the opponent's range; another player starts slightly weaker, but as long as they can see a few more rounds at a lower cost, they might扳 back the win rate bit by bit.
Compute, algorithms, data, and talent together determine precisely this portion of unrealized win rate. They determine how many routes an AI company can attempt, how expensive one trial and error is, whether experience can be retained, and whether they can sit at the table again after the next round of model updates.
1. Compute: US Controls Highest Face Value, China Fights for Economic Usability
Compared to the base model starting hand, compute is the US's true gold chip. If the cards are bad, you can wait for the next community card; if chips run out, it is hard to continue sitting at the table.
The US controls the casting rights for high-end AI compute: Nvidia defines accelerators, interconnects, and software stacks, TSMC undertakes advanced manufacturing, Japanese and Korean enterprises supply HBM, and cloud vendors organize tens of thousands of chips into training clusters. Chips, network, storage, power supply, and CUDA bite together; frontier competition is long past comparing single-card scores, but rather seeing whether an "AI Factory" can twist tens of thousands of chips into a whole.
China has already crossed the threshold of "whether there are domestic AI chips." Huawei CloudMatrix organizes Ascend, Kunpeng, network, and software into a unified system, used for DeepSeek model training and inference. What truly remains to be crossed is moving from "able to run" to "economically practical": chips must be supplied sufficiently, connected stably, and 调度 able; model migration cannot pay unacceptable engineering costs; running a ten-thousand-card cluster for a month cannot let most time be eaten by communication, failures, and compatibility issues.
Effective compute is more honest than chip count. Theoretical compute must undergo layers of loss from memory bandwidth, node communication, software operators, failures, and utilization rates. Ten thousand cards written on paper looks imposing, but the ability truly continuously used for training may be far lower than simple multiplication. The US advantage is that chips, HBM, interconnects, and software have been jointly optimized around the same training demand; China's difficulty is that every time one link is 补 up, the bottleneck may immediately shift to the next link.
Once China cannot cross the threshold of "economic practicality," the US can continuously map chip advantage to model win rate; once crossed, export controls will still increase costs, but will find it difficult to make China leave the table.
2. Algorithms and Engineering: DeepSeek Relabels Chip Face Value
In Texas Hold'em, of course, the more chips the better. But if the opponent can wipe two zeros off the 10,000 face value chips in your hand, the purchasing power on the field will be recalculated.
China once played such a stunning card, DeepSeek.
The DeepSeek-V3 technical report provides a rare detailed account: the final complete training used about 2.788 million H800 GPU hours. This number does not calculate early exploration, failed experiments, hardware procurement, and infrastructure, so "only spent about 6 million USD to train a frontier model" omits more than half the bill. But it still shocked the industry because it tore open an old odds ratio many viewed as common sense: there is no fixed exchange rate between model capability and compute investment.
V3 adopts MoE architecture, activating only part of parameters per token; multi-head latent attention compresses cache, FP8 training reduces computation and communication costs, load balancing reduces expert idleness. R1 pushes efficiency further into the post-training stage: on tasks with verifiable results like math and code, the model iterates through trial and error via reinforcement learning, turning part of expensive manual reasoning demonstrations into automatically adjudicable reward signals. That is to say, DeepSeek did not凭空 obtain more chips; it made chips work more efficiently.
The US is still better at exploring new routes. Transformer, Scaling Law, RLHF, and various Agent frameworks were mostly first pushed to the frontier by US research institutions and enterprises; deeper capital pools also allow laboratories to bet on multiple unproven directions simultaneously. China was previously stronger at reproduction, compression, and engineering optimization. After DeepSeek, this boundary began to loosen: Chinese teams are no longer just making others' routes cheaper, but also beginning to propose methods sufficient to rewrite industry cost expectations.
However, algorithm dividends will not permanently replace hardware. Papers will diffuse, leading models will also absorb the same techniques; efficiency improvements will also stimulate more calls, letting saved compute be quickly eaten by new demand. Players have learned to control bets more precisely, but blinds are also continuing to rise.
Therefore, the strategic value of engineering efficiency for China is more about being able to increase experiment counts with limited compute, shorten verification cycles, and buy time for domestic hardware maturity.
3. Data: Data with Results Is Appreciating, and the US Currently Leads
Professional players reviewing hands naturally do not just record whether they got an A or K. They must save complete actions: who bet first, how raises happened on the turn, how community cards changed, what the opponent finally showed, and where their own judgment went wrong.
Data for models is exactly so.
Early large models mainly learned language, knowledge, and code from the public internet; the US gained first-mover advantage relying on English web pages, GitHub, paper publications, and global digital platforms. As high-quality public corpora are repeatedly used, what is truly scarce is no longer just text the model hasn't read, but data with clear results capable of judging task success or failure.
The value of a customer service dialogue lies not only in what the agent said, but also in whether the problem was solved; the value of a piece of program code lies not only in whether it looks reasonable, but also in whether it can pass tests after modification. Data that cannot point to results is mostly just noise.
Early 2026, Alibaba trained a custom version Qwen-Coder for programming Agent Qoder, introducing real software tasks, product environments, and engineering rewards into training. Alibaba disclosed that after iteration, code online retention rate increased by 3.85%, tool anomaly rate decreased by 61.5%, and token consumption decreased by 14.5%. Numbers come from the vendor and await external verification, but it pierces the most expensive part of professional data: not that text, but a verifiable result behind the text.
The US possesses broader feedback entry points. ChatGPT, Claude, Gemini, GitHub, and office suites connect global consumers, developers, and enterprises; model companies can observe where users modify answers, which code programmers retain, and where Agents fail.
China's opportunity lies in denser business processes. ByteDance's short videos and ads, Meituan's food delivery and distribution, Alibaba's e-commerce and fulfillment, new energy vehicle intelligent driving, factory quality inspection and equipment operation, every link carries real wins and losses.
What China truly lacks is high-quality complete hand histories. Massive industrial data is still cut across different enterprises, departments, and local systems. Only when tasks can be standardized, results verified, and data flow under compliant conditions will business traces become model assets. Many scenarios are just many raw materials; whether a short chain of "model execution, reality feedback, result flowback" can be formed determines whether the next round of training eats nutrition, or a warehouse of uncountable old cards.
4. Talent: China Is Talent Upstream, US Is Talent Amplifier
In the first few rounds of the game, competition between China and the US always looks like individual tech companies going down to face off. On the US side, out come OpenAI, Anthropic, Google, Meta; on the China side, out come DeepSeek, Zhipu, Moonshot AI, Alibaba, ByteDance.
Tech companies release models, hoard compute, fight for users, and also represent their respective countries continuously "going to war" at the table. But companies are ultimately just carriers of great power competition in the commercial world. What truly determines whether a company can read what cards, dare to bet on which route, and whether it can extend a lead once gained, is still the talent behind the organization.
Talent can be researchers proposing new architectures, engineers building training systems, or product teams pushing models to market. If the scope is expanded, it can also include universities, laboratories, and competition systems that continuously cultivate and transport new people for the entire large model system.
Large model success is often told as a flash of inspiration from a certain scientist; this is a fallacy with a tendency to create gods. Like professional Texas Hold'em although played by one person sitting at the table, behind top players are coaches, Solvers, hand databases, physical management, and long-term capital management. Truly stable advantage and winning position come not from some divine stroke, but from a whole set of continuous review and correction systems.
AI companies are also so. Top researchers can make decisions on whether a route is worth betting on, but behind them hundreds of system engineers will invest into computation loss, handle failures, and turn one accidental success into training and products that can be repeatedly replicated.
Great power talent competition is essentially comparing who can provide a table more worth sitting down at. Top researchers will choose places with higher salaries, deeper compute, stronger peers, more important problems, and more allowance for failure. The US has long attracted global experts relying precisely on this "table selection right."
The most paradoxical place about this table is: China is not without talent. Quite the opposite, China is already one of the most important upstream sources of global top AI talent. The problem is that after many talents complete early training, they ultimately connect to the US amplifier.
US large model advantage is not the result of a closed system being self-sufficient. It is more like a huge global talent centrifuge: sucking in the smartest batch of young people from China, India, Europe, and around the world into the US PhD system, then sending them into top laboratories at Stanford, MIT, Berkeley, CMU, finally transporting them into companies like OpenAI, Anthropic, Google DeepMind, Meta, Microsoft, NVIDIA.
Universities still provide methods, papers, and talent, but the entities truly able to amplify methods to ten-thousand-card clusters, global products, and commercial revenue are increasingly concentrating in enterprises. Stanford AI Index statistics also show that publishing entities for important frontier models have become highly industrialized. Universities are responsible for inventing hand types, enterprises then control the truly large-stake tables.
China's hand is not completely the same. China's advantage lies not in necessarily having more少数 stars, but in a larger engineering talent pool, shorter responsibility chains, and task density closer to industrial scenes.
The value of this engineering force only truly appears after large models enter complex systems. Frontier models need a few top scientists, but model landing needs hundreds of system engineers, data engineers, inference optimization engineers, chip adaptation engineers, and product teams. The US is good at pushing a few top talents to capability boundaries; China is better at spreading large numbers of engineers into industrial scenes. The former suits impacting general capability ceilings; the latter suits rapid transformation, deployment, and delivery.
This is also the change DeepSeek, Moonshot AI, and other companies bring to the Chinese large model narrative. They prove Chinese companies can participate in the game with another organizational method: tighter teams, shorter feedback chains, stronger engineering compression capabilities, and space letting young people enter core tasks faster.
Perhaps the true win or loss of the Chinese and US large model game is not now, but depends on when the next generation of top students depart from Beijing, Shanghai, Guangzhou, etc., which location they will view as their home ground. Will they fly to Silicon Valley, connecting to the US super amplifier; or stay in China, at a new table under expansion, connecting domestic models, chips, systems, and industrial scenarios into a closed loop.
III. Playing Style: US Occupies High Ground, China Surrounds Cities from Rural Areas
When model lead compresses from generation gap to time lag, the question is no longer just "who plays a big card first," but whether this card can be transformed into long-term advantage.
1. US Holds Hole Cards to Prove Pricing Power
OpenAI, Anthropic, and Google keep strongest models in subscription products, APIs, and cloud platforms. Developers can call them, but cannot download full weights, freely modify, or bypass platform deployment.
Closed source is first a pricing strategy.
Whenever OpenAI and Anthropic enter the public market, they must face Silicon Valley and Wall Street pricing logic. Capital markets will not pay long-term for one benchmark topping, caring more about revenue, retention, gross margin, and bargaining power. Only by controlling the entry to the most advanced models can they charge for every call, every subscription user, and every enterprise contract, turning temporary lead into assets that can be repeatedly sold.
These two companies bear the heavy asset costs of chips, data centers, and energy, but hope to obtain software platform valuations. If weights are fully open, cloud vendors and application companies can quickly host, package, and lower prices; base models could become homogeneous raw materials. Trainers bear the heaviest risk, but profits flow to compute and entry points.
So, what the US holds is not just technical secrets, but model rent. API, Agent frameworks, enterprise permissions, data connections, and security audits all add locking layers around the model. The deeper customers connect, the higher migration costs, and temporary lead is more likely to pass through the next leaderboard shuffle.
2. China Reveals Some Paths to Exchange for Distribution Rights
Today, Xi Jinping proposed at WAIC (World Artificial Intelligence Conference) adhering to openness and win-win, encouraging open source, cooperation and sharing, and listing helping Global South countries strengthen capability building and bridge the digital intelligence gap as an important direction for global AI governance. The Conference Chair Statement released the same day further proposed that open-source ecosystems should be built responsibly, improving accessibility of AI technology and services under the premise of respecting enterprise autonomous choice and intellectual property protection.
Most Chinese model companies do not have global entry points like ChatGPT, Google, and Workspace. Hence locking large models inside their own platforms may not exchange for high profits like the Silicon Valley model.
Therefore, open weights are first a distribution means.
DeepSeek, Qwen, GLM compete for developers with downloadable weights, low-price APIs, interface compatibility, and local deployment. The Kimi K3 released today pushes this route to the extreme of "generosity": total parameters 2.8 trillion, native visual support, highest context reaching 1 million tokens. Moonshot AI calls it the first open 3T-level model and promises to release full weights before July 27.
China allows others to take this hand of cards from their own hands, to observe and explore themselves, letting more countries possess another set of technical options not fully dependent on US APIs. This is the logic of "surrounding cities from rural areas": not first fighting for the most expensive, most closed central market, but using open weights, low-price APIs, and hardware adaptation to enter scenarios with larger quantities, tighter budgets, and data unwilling to leave domains.
For many developing countries, what is most important is not necessarily models winning two more points on leaderboards. They care more about whether models support local languages, whether data must leave the country, whether prices are affordable, whether deployment is possible on domestic servers, and whether suppliers will suddenly cut off services due to geopolitical changes.
Xi Jinping announced in his speech that in the next 5 years China will provide 5,000 AI special research training quotas for developing countries, build international AI application cooperation centers for ASEAN, Arab League, African Union, CELAC, SCO, and BRICS, and promote intelligent weather warning solutions landing in 30 countries. The World Artificial Intelligence Cooperation Organization was also declared established in Shanghai.
Thereby, China is forming a three-layer mutually cooperating distribution system.
The first layer is models: using open weights and low-price APIs to lower usage thresholds.
The second layer is infrastructure: using cloud computing, domestic chips, local deployment, and industry solutions to undertake actual demands.
The third layer is institutional networks: through training, cooperation centers, and international organizations, turning one model download into longer-term technical relationships.
Of course, openness can exchange for global attention and adoption rates, but cannot expect massive cloud business to automatically recover costs for them. They must then find cash flow from official APIs, enterprise services, private deployment, and upper-layer products.
China must within a predictable time dimension turn open weights into development standards, turn development standards into cloud and application revenue, and turn international adoption into long-term ecosystems. Otherwise, so-called distribution rights are just temporary download volumes, not commercial closed loops capable of feeding the next generation of models.
3. US Collects Model Rent, China Fights for Ecosystem Rent
The US wants to collect model rent: relying on highest capability and platform locking, letting calls, subscriptions, and enterprise contracts continuously pass through their own charging entry points. China is more inclined to fight for ecosystem rent: lowering the model itself's price, using openness to exchange for developers, cloud workloads, enterprise deployment, and technical standards, then recovering value from one of those layers.
Both routes may fall into their own traps. Closed-source platforms most fear capability gaps shrinking; high walls without obviously better capability support will turn from moats into toll booths; open routes easily win applause but lose cash, overseas cloud and application companies take revenue, original vendors bear the most expensive training costs.
The Chinese route must complete one transformation: from open weights to development standards, from development standards to cloud and deployment revenue, then from revenue and real tasks obtain compute and data needed for next-generation models. The US must also prove its lead is sufficient to deepen into user workflows, becoming that layer of intelligence not easily replaced.
The US is proving: my cards are scarce enough, so pay for every look.
China is betting: as long as enough people use the same hand of cards, diffusion itself will also generate bargaining power.
IV. Judging Pot Odds: Cheap Is Not Favor, It Is Prerequisite for Scale
Pot odds answer a cold question: to fight for chips on the table, how much more price must be paid? No matter how beautiful the cards, if every call is ridiculously expensive, long-term will burn out principal.
1. API Price List Is Just the First Bill
As of July 17, 2026, public API prices for several representative models are as follows, unit is per million tokens. (Prices adjust anytime, also do not represent enterprise long-term contract prices. Kimi K3 input price calculated by cache miss.)
This table still shows domestic model price advantages, but can no longer be summarized as "domestic models uniformly dozens of times cheaper." DeepSeek continues pressing low prices to the extreme, Kimi K3 attempts to price upward relying on frontier code capabilities and 1 million token context. Chinese model tactics are diverging from single price wars into two routes of low-price volume and high-end premium.
Things are not this simple, Tokens are not this generous.
Enterprise complete bills at least need to include tokens, tool calls, failure retries, manual checks, and redundancy reserved for latency and stability. Price lists only list the first item; once errors occur, enterprises need to pay for all subsequent correction links.
Truly fair comparison is not "how much for one million tokens," but completing the same task, how much money is finally deducted from the account.
However, even so, low price is still a very fierce card. Customer service, summarization, document processing, and batch code review may call millions of times daily; single price differences are rapidly amplified by scale. Cheap means startup teams dare to trial and error, SMEs can enter the market, Agents also have opportunities to walk from demo tables into daily work. China is attempting to let more ordinary tasks qualify for the pot.
2. Capital Double-Edged Sword: Financing Scale Purchases Trial Rights, Also Starts Return Countdown
If further putting four frontier laboratories together, capital structure differences become more intuitive.
Financing magnitude differences sufficiently explain that US top laboratories can move chip, data center, and talent budgets for future several years to today at once.
What capital supports is actually large model companies' trial rights. As long as money is deep enough, ten research routes can run simultaneously, letting nine fail after which the tenth continues living, also can lock data centers and power contracts before revenue forms.
But all fate's gifts are priced in the dark. The more money, the louder the return clock ticks. Huge valuations, credit, and infrastructure commitments are all demanding debts from future cash flows. OpenAI and Anthropic must turn capability lead into subscriptions, APIs, and enterprise contracts, proving they bear heavy industry costs but possess software platform profit structures. Capital gave them deeper chips, also quietly wrote profit dates under the table.
China's capital structure is relatively more dispersed. Internet giants use cloud and consumer business as base, startups accept industrial capital and local even central policy support, public compute facilities bear part of basic investment. This can let key routes not immediately stop due to short-term profit insufficiency, also may create duplicate construction, low utilization clusters, and projects responsible only to subsidies.
US risk is hot money is enough, easily capitalizing demands still to be verified in advance; China risk is hot money not concentrated enough, frontier research may be forced to turn to short-term delivery when most needing long-term investment. The former must prove high-price models sufficient to cover huge investments; the latter must prove low price and openness are not permanent subsidies, but can exchange for cloud revenue, enterprise deployment, and industrial efficiency.
In this round of power, what both sides burn is far from just money; the most precious is time borrowed from the future.
3. Power Limits Strategy Space: China Possesses Infrastructure Depth
Training is concentrated burst load; search, office, and Agent inference are day-and-night non-stop loads. When models are called billions of times, electricity prices, grid connection, heat dissipation, and chip utilization all enter every task's cost.
China energy infrastructure advantage is first scale.
As "Infrastructure Maniac," China possesses larger power systems, still continuously building wind power, photovoltaic, energy storage, nuclear power, and cross-regional transmission networks. When data center construction needed by large models brings new large-scale loads, China obviously possesses larger supply depth and stronger engineering expansion capabilities.
US problem is capital and chips are already waiting outside the door, but grids cannot complete expansion via one software update. US Department of Energy research shows high-voltage transmission projects from development, approval to completion average about 10 years, typical interval 5 to 17 years; power projects from application grid connection to operation average time has also extended from about 2 years in 2008 to about 5 years in 2023.
Just in June 2026, US Federal Energy Regulatory Commission further required six regional grid operators to explain or reform large load grid connection rules. Federal system decentralization characteristics determine regional transmission line approvals are slow, transformer and other equipment delivery cycles are long, different states, power generation enterprises, grid operators, and data centers must also repeatedly negotiate who bears costs.
In comparison, China Mobile Ningxia Zhongwei Data Center Park fully put into use in 2026, cumulative invested IT power reaches 332 MW, intelligent compute scale exceeds 100 EFLOPS, green power ratio stably above 80%, comprehensive electricity price about 0.36 Yuan/kWh. This project's significance lies not only in cheap electricity prices, but that power sources, grids, energy storage, and data centers can be constructed synchronously, turning energy resources directly into compute supply.
Therefore, at this energy infrastructure layer, China's advantage is actually clearer than advantages on Benchmarks. International Energy Agency expects global data center electricity use will increase from about 485 TWh in 2025 to about 950 TWh in 2030. US and China will contribute nearly 80% of the increment. As competition shifts from model training to billions of continuous inferences, China's energy infrastructure advantage will become more important: training can wait for one cluster schedule, inference services need stable, low-price power uninterruptedly year-round.
However, energy advantages ultimately still need conversion through chip efficiency, software optimization, and cluster utilization rates. Cheap electricity consumed by low-efficiency chips and idle machine rooms cannot automatically become low-cost intelligence. What truly needs comparison is how much electricity, chip depreciation, and manual operations are needed to complete one million real tasks.
Who can first achieve lower unit task costs has more ability to drag this competition into long-term war of attrition. On this point, China's whole-nation system may hold more advantage.
4. China Needs to Leapfrog Is a Cost Closed Loop
Putting API, capital, and energy together, win or loss is easily lightly attributed to "who has more money." US indeed can use high investments to impact model ceilings, then recover costs via global subscriptions, cloud, and enterprise software; China attempts to use engineering optimization, low-price models, and infrastructure to lower call thresholds, then find revenue and data from large-scale use.
Chinese route risk lies in every link pressing prices, but no place leaving sufficient profit. Low-price APIs if only bring call volumes, open weights if only bring download volumes, local data centers if only have construction scale, three added together may still be a huge loss. US route risk is opposite: every link can collect high prices, but costs are high enough that significant capability gaps must be continuously maintained.
Current pot odds still favor the US because it can sell technical lead to global most expensive customers. But once model capabilities gradually approach, unit task cost weight will rapidly become heavy, China's low price, electricity, and engineering efficiency will also show edges.
Premise is these advantages ultimately converge into the same cash flow, not leaving three separate good-looking but lonely account books.
V. Collecting the Pot: Who Can Spread Models into Real Work
Board lead will not automatically push chips to front. Only when users continuously use, enterprises willing to pay, tasks leave verifiable results, does technical advantage count as cashed in. LLM competition true pot has only four things: revenue, entry points, feedback data, and workflows reorganized by models.
1. US First Occupies Entry Points, China Closer to Transactions
ChatGPT first changed habits of people seeking answers, subsequently entering writing, research, code, and enterprise workspaces; Google embeds Gemini into search, Workspace, and cloud; Anthropic relies on Claude and Claude Code entering knowledge work and software development. Behind US model companies originally were browsers, office suites, code repositories, and global cloud platforms.
Entry points are advantages harder to catch up than leaderboards. Once enterprises connect identity, permissions, data, and procurement, changing models is no longer just changing a name, but needing to dismantle part of workflows again. OpenAI published Signals data shows user daily message count increases about 50% half a year after registration, attempted task types double, users mainly using non-English already account for over half of active users.
Chinese platforms are farther from global office entry points, but closer to specific transactions. Alibaba connects Qwen into Taobao and Tens of billions of item catalogs; users can compare, order, check logistics, and handle after-sales in dialogue. Models connect not just web pages, but merchants, orders, and fulfillment systems. Whether users placed orders, whether recommendations were accepted, whether logistics and after-sales completed, every step has results.
US general models first occupy user entry points, then connect external services; Chinese super platforms can from day one put models into transaction closed loops. The former possesses stronger global distribution; the latter sticks to denser behavior feedback.
2. Coding Is First High-Value Test Field
Coding earliest formed stable payments because results are easy to verify: whether completion was accepted, whether projects can compile, whether tests pass, whether Bugs are fixed, which large model is useful is clear at a glance.
In Coding aspect, US originally held the neatest hand: GitHub, Microsoft, OpenAI, Anthropic, and large amounts of development tools control global code entry points, head models also long led complex repository tasks. Kimi K3 let the model capability column appear new gaps. According to Moonshot AI disclosed evaluations, it exceeds GPT-5.6 Sol and Claude Fable 5 in Program Bench and SWE Marathon, in Terminal-Bench 2.1 only lags GPT-5.6 Sol by 0.5 points, and demonstrated long-range execution capabilities in partial GPU kernel optimization, compiler development, and chip design tasks. But models winning several evaluations does not equal already taking development environments. Kimi Code and open weights are 补 distribution; China still needs to catch up on users, toolchains, and feedback closed loops accumulated by GitHub, Codex, and Claude Code.
Code has already exposed future market shape: strongest models undertake complex tasks, cheap models responsible for completion, testing, and batch review; enterprises then switch between closed-source cloud and local models according to difficulty and sensitivity. Ultimately taking away value may not be some model, but possibly that system knowing when to use which model.
Who controls development environments can correct models faster. Whether code is retained or withdrawn, where tests fail, how developers modify, all will become hand histories for next round training.
3. To C Agent: Trump Card for Sinking User Mindshare
If frontier models compare board faces, then To C Agents fight for who can turn model capabilities into subconscious actions for thousands upon thousands of users.
When a person wants to search information, organize files, make presentations, plan travel, or solve work problems, who do they open first?
US currently most representative are OpenAI and Anthropic two cards.
ChatGPT has already formed the strongest AI native brand mindshare. OpenAI also splits this card into two routes: ChatGPT Work (CodeX) responsible for research, documents, tables, presentations, and other complete deliveries. Users need not leave ChatGPT, can walk from raising questions all the way to getting finished products.
Anthropic occupies mindshare relatively narrower, but deeper, and more valuable, even once exceeding OpenAI in disclosed annualized revenue. Claude Code has become priority choice for many programmers handling complex projects; Claude Cowork pushes same working methods to researchers, analysts, and other knowledge workers, letting Claude directly handle local files, desktop applications, and multi-step tasks.
US is building a very clear product cognition: encountering a complex work, need not first think which several software to open, first hand it to Agent.
And China actually originally did not walk this road.
Initially, Chinese enterprises always attempted pressing Agents into scenarios enterprises were already good at.
Alibaba connects Qwen into Taobao and Tmall's 4 billion items. Users can search, compare, order, query logistics, and handle after-sales in dialogue. Behind Qwen is Alibaba's already operating over twenty years, commodities, merchants, payment, logistics, and after-sales this whole transaction chain.
ByteDance bets chips on content production. Seedance 2.0 allows ordinary users to control complete audio-visual creation with text, images, audio, and video, fully accessing content distribution network composed of Doubao, Douyin, Jianying, and creator ecosystems.
Meituan's "Xiao Tuan" upgraded based on LongCat has deeply intervened into life scenarios like eating, drinking, having fun, travel, and medical consultation. During 2026 "May Day" holiday, Meituan announced "Xiao Tuan" served over 100 million person-times. Compared to general Agents, what Meituan holds is not just answers, but also merchants, reviews, locations, inventory, fulfillment, and transaction results.
Thus, Chinese and US To C Agents form two different expansion methods—US extends outward from AI native entry points; China surrounds inward from existing scenarios.
Chinese path risk lies in mindshare fragmentation; opening Qwen when shopping, opening Doubao when creating, using WorkBuddy when working; this itself is a kind of chaos. Scenarios are deep, entry points remain scattered. China possesses many good-positioned tables, but has not yet appeared a super entry point like ChatGPT capable of gathering all user mindshare under one name.
And now, Chinese major firms have all realized this point, beginning to align with US already verified general Agent forms, transforming chat boxes into workbenches.
Tencent WorkBuddy already can from one instruction complete data organization, data analysis, document production, and content creation; Kimi's general Agent, Kimi Claw, and Xiaomi MiMo Claw are also attempting to take over files, tools, and long-range tasks. What they do is becoming more and more similar to ChatGPT Work and Claude Cowork: users only propose goals, Agents responsible for breaking down steps, calling tools, finally delivering finished products.
Because one good card can only win one hand, but seizing user default entry points can "cheat-style" turn every subsequent application scenario into training experience, thereby "cheat-style" see all hands.
VI. Control Drawdown: Institutional and Social Pressure Capacity
Texas Hold'em is never a game of only win no loss. Even strongest players will get bad cards, will be overtaken by small probabilities when judging correctly, also will watch a huge pot pushed to the opposite side.
What truly distinguishes professional players from ordinary players is not the former never losing, but they can control drawdowns, avoid emotional loss of control, not letting previous hand loss destroy next hand judgment.
AI competition is similarly so. Models improve productivity, may also compress employment; lower content costs, will also amplify fraud and deepfakes; after entering enterprise systems, it can improve efficiency, may also spread one error to payments, code, and key data.
1. Regulation Determines Pool Entry Range
US closer to wide-range playing style. Enterprises first release products, then markets, courts, state-level legislation, and regulatory agencies draw boundaries. It gives enterprises more trial opportunities, also means part of copyright, privacy, and product harm will first be borne by society, then corrected afterwards.
China more emphasizes pre-defining boundaries. Filing, content governance, platform responsibilities, and industry pilots can reduce obvious risks, also facilitate forming unified standards; but if boundaries are blurred, enterprises and locals to avoid responsibility may layer additions, turning stop-loss into premature folding.
US bears higher volatility, exchanging for more trials; China lowers loss of control probability, exchanging for more stable advancement. Truly high-minded systems are not forever loose, nor forever tight, but know when to expand range, when to take chips back.
2. Employment Impact Determines How Much Drawdown Society Can Bear
LLMs first affect work with higher proportions of language and information processing: customer service, translation, administration, content, basic programming, legal retrieval, and partial analysis. International Labour Organization estimates about one-quarter of global jobs exist some degree of generative AI exposure, but employment at highest exposure level accounts for about 3.3%. This is more like tasks being reorganized, not all professions disappearing neatly.
Danger is not some job disappearing, but change speed exceeding speed society rearranges jobs, income, and security. Enterprises can introduce AI in several months, laborers may need several years to relearn skills; economies long-term still growing, part of people already close to bankruptcy.
US labor markets more flexible, enterprises dare reorganize jobs, but costs fall more on individuals and families. China can through vocational education, industrial policies, and large organization coordination training, but faces larger employment populations and higher stability requirements.
Excellent risk management is not guaranteeing never drawdown, but ensuring one drawdown will not make people lose ability to re-enter the game.
3. Most Dangerous Is Not Losing Hand, But Tilt
Tilt is players losing rationality after encountering continuous bad luck, beginning to chase losses, expand bets, abandoning originally effective strategies. Society may also repeatedly Tilt between technical fanaticism and regulatory panic.
If AI continuously brings fraud, false content, employment anxiety, and privacy incidents, public will turn from curiosity to distrust. Enterprises may still believe next-generation models can solve everything, society may due to several serious accidents demand comprehensive restrictions. Former easily forms bubbles; latter may abandon long-term gains due to short-term losses.
Mature social acceptance is not requiring everyone remain optimistic, but letting public know technical boundaries, possessing right to know, right to exit, and right to appeal. Risk governance also cannot only stare at what models "said," but must ask what they "have right to do": when Agents execute transfers, modify code, or call key systems, minimum permissions, secondary confirmations, log audits, and accident responsibilities must enter processes.
True pressure capacity is after failures occur still maintaining judgment, controlling drawdowns, correcting strategies, not letting one part's losses turn whole technological progress into bad debt no one willing to continue bearing.
VII. Table-Flipping Variables: Possible External Shocks Making Gamble Invalid
All comparisons above imply one premise: game will continue according to current rules. Models gradually improve, compute remains scarce, capital continuously invests, supply chains maintain operation, society also willing to bear shocks.
Reality may not be so.
Technical revolution turning points are often not raises, calls, or folds at the table, but some variable suddenly changing chip value, rewriting rules, even directly flipping the table.
- Whether AGI arrives early. If some system can stably complete long-term scientific research, programming, and engineering design, and substantively participate in next-generation model R&D, several months lead may roll into uncatchable fault lines. Short-term, this favors US possessing frontier laboratories, top chips, and cloud platforms; but whether capabilities leak, replicate, and distill will determine whether advantage can be monopolized.
- Whether compute bottlenecks are broken. If new architectures, low-precision computation, or specialized chips reduce computation needed for equivalent capabilities by an order of magnitude, part of US super-large cluster advantages will be repriced; if long-chain reasoning and Agents continue swallowing more tokens, chips, energy, cloud, and capital will further concentrate.
- Whether blockades escalate, supply chains interrupt. If US continues expanding restrictions to cloud compute, HBM, equipment maintenance, model access, and talent cooperation, China short-term costs will still significantly rise. But blockades more persistent, China less will take US supply as reliable foundation. More extreme geopolitical conflicts will simultaneously strike US design alliances and China manufacturing markets; competition will retreat from who innovates faster to who has thicker inventory, more alternative capacity.
- Whether AI bubble bursts. If Agent saved costs insufficient to cover model bills, capital markets will reprice according to cash flow and profits. Bubble bursting will not make LLMs disappear, but will clear participants relying on valuations, subsidies, and low-utilization facilities from the table.
- Whether AI unemployment triggers social backlash. Society need not wait for large-scale unemployment to react. As long as entry-level positions decrease, career entry narrows, productivity gains mainly flow to platforms and capital, automated approvals, employment protection, and new distribution policies may arrive early.
- Whether open models change power structures. If open weights continuously press gaps to "most tasks sufficient," power will shift from few laboratories to cloud platforms, application companies, and countries possessing local data. But openness not naturally belongs to China; truly depends on who can precipitate download volumes into toolchains, standards, revenue, and next-generation training resources.
- Whether major safety accidents cause regulatory emergency brakes. Once high-permission Agents cause serious financial, medical, network, or infrastructure accidents, model licensing, liability insurance, and deployment approvals may all tighten rapidly. Safety may also become industrial tool, re-dividing market boundaries foreign models, data, and compute can enter.
These variables will not line up to come. Blockades may accelerate domestic chips; open models may lower costs also amplify abuse; bubble bursting may delay AGI, may also drive resources toward fewer companies. Employment shocks and safety accidents may also wake both countries' regulators on the same night.
When history truly turns, movement often comes from outside the board face in progress.
It will let most valuable chips suddenly depreciate, also give originally lagging side, opening a road no one calculated.
Conclusion
Vision returns to that green felt table at the beginning, machine room top lights still coldly shining.
US still holds larger starting hand. It controls frontier models, high-end chips, capital, and global entry points; if revealing cards immediately today, win rate still higher. China is however changing calling methods: using algorithms to relabel chip face values, using openness to expand tables, using industrial scenarios to find hand histories with real wins and losses, then attempting to turn cost advantages into a self-reinforcing loop.
This competition will continue consuming electricity, capital, talent, and social patience, until one side truly establishes closed loops between capabilities, revenue, data, and compute, or until both sides discover, price of winning this game has become high enough not to look like victory.
Dealer does not urge, only places hand on card stack. Outside door occasionally comes a bit of wind sound, very light very light, like a piece of bad news not yet written into financial reports.
People at the table have not risen.
Join TechFlow official community to stay tuned
Telegram:https://t.me/TechFlowDaily
X (Twitter):https://x.com/TechFlowPost
X (Twitter) EN:https://x.com/BlockFlow_News













