China Is Closing the A.I. Gap by Learning From America's Best Models

U.S. laboratories say Chinese competitors are extracting capabilities from proprietary systems at industrial scale. The dispute reveals a deeper problem for the frontier-model business: intelligence may be easier to imitate than its creators expected.

Category: Strategy. Written by Jaime Garcia, Founder, SnowRock. Published . 19 min read.

In short

The leading American artificial-intelligence companies are confronting an uncomfortable possibility. The capabilities they have spent billions of dollars developing may be transferable faster, more cheaply and more systematically than their business models assumed.

Anthropic has accused Alibaba and other Chinese technology companies of creating large networks of unauthorized accounts to extract outputs from its most advanced models. According to the company, those outputs were then used to train competing systems through a process known as model distillation. The allegation is significant because distillation can allow a smaller model to absorb portions of the behavior of a more powerful one without reproducing the full cost of its original development. The frontier laboratory pays for the chips, researchers and energy to produce the teacher. The competitor studies the teacher responses and builds a cheaper student.

American companies argue that this process is being conducted systematically by Chinese laboratories and that it is helping narrow the technological distance between the two countries. The concern has intensified as Chinese models approach the performance of leading American systems. Z.ai GLM-5.2, for example, is reported to compete closely with top U.S. models in several advanced domains, including cybersecurity, and DeepSeek previously showed that Chinese developers could produce highly capable systems at costs far below those of the largest American laboratories. Some researchers estimate that China now trails the United States in frontier artificial intelligence by only several months.

Distillation alone does not explain that progress. It may nevertheless be accelerating it. The resulting dispute is not simply about whether Chinese companies violated commercial terms. It concerns whether the frontier-model advantage can remain proprietary long enough to support the enormous investment being made in it.

What model distillation actually does

Model distillation is not a new or inherently malicious technique. Researchers have used it for more than a decade to make systems smaller, faster and less expensive. A large, capable model acts as a teacher and a smaller model acts as a student. The teacher generates answers, probability distributions, classifications or reasoning patterns across a large set of examples, and the student is trained to reproduce those outputs. The student does not need to understand the teacher internal architecture or possess its original training data. It learns from observable behavior, which can produce a smaller system that keeps much of the larger model practical usefulness while requiring fewer computing resources.

The technique was initially developed as an efficiency measure. A company might train an enormous model in a data center and then distill its capabilities into a smaller version suitable for a smartphone, vehicle or low-cost server. In that context the developer is distilling its own technology. The conflict begins when one company uses another company proprietary model as the teacher.

Behavior can be copied without copying the model

A proprietary model is not normally distributed as downloadable software. Customers interact with it through a website or an application programming interface. They submit a prompt and receive a response. That interface protects the underlying code, model weights and training data. It does not prevent outsiders from observing behavior.

A competitor can submit millions of carefully designed prompts, collect the responses and use those examples to train another system. The goal is not to recreate the original model exactly. It is to reproduce enough of its useful behavior. A model can be asked to solve mathematical problems, write software, analyze documents, identify security vulnerabilities or respond to complex professional scenarios, and each answer becomes a training example for the competing system. At sufficient scale, these examples transfer meaningful capability. The process resembles interviewing an expert repeatedly and converting the answers into a training manual. The competitor does not receive the expert brain. It receives a structured record of how the expert behaves.

Why frontier laboratories object

OpenAI, Anthropic and other proprietary-model providers prohibit this form of distillation under their service agreements, and their position is understandable. The companies have spent extraordinary sums developing advanced systems and do not want competitors using paid access to reproduce those capabilities and sell substitutes. From their perspective, large-scale unauthorized querying is not ordinary product use. It is industrial extraction.

Anthropic alleges that Chinese companies created tens of thousands of accounts to disguise coordinated activity and evade usage controls, and that the resulting conversations were designed to collect enough high-quality output to train rival systems. It has characterized these operations as deliberate attempts to harvest American technological capability. If the allegations are accurate, the scale distinguishes the conduct from an ordinary developer experimenting with a public chatbot. The objective would not be to use the service. It would be to manufacture a competitor.

The practice is not limited to China

The moral clarity of the complaint is weakened by the fact that American companies use distillation too. Model developers routinely study one another systems, comparing outputs, testing capabilities and using stronger models to improve weaker ones. Elon Musk has publicly acknowledged that xAI distilled technology from OpenAI and described the practice as common across the industry. Open-source developers use distillation extensively, and when the teacher model is released under terms permitting copying and modification, this is generally treated as legitimate and beneficial.

The dispute is therefore not about the technique itself. It is about authorization, scale and ownership. American companies support distillation when they control the source model or when the system is openly licensed, and object when a competitor extracts behavior from a closed model without permission. That distinction is commercially coherent. It is legally and geopolitically difficult to enforce.

It is not yet clear whether unauthorized distillation violates federal law. A company may argue that the process misappropriates trade secrets, since the behavior of a proprietary model can contain valuable knowledge unavailable through ordinary public sources, and if a competitor obtains it through deception, account manipulation or violation of access restrictions, the provider may claim that protected information was acquired improperly. The difficulty is defining the secret. The model internal weights and training data are clearly proprietary, but the outputs delivered to customers are intentionally exposed. A user who submits a prompt is authorized to receive the answer.

The legal question is whether collecting those answers at scale and using them to train another system turns permitted access into unlawful appropriation. Copyright law offers no simple answer, because distillation generally copies functionality or behavior rather than reproducing a specific text verbatim, and copyright traditionally protects expression, not methods or general capabilities. Contract law may provide a clearer claim, since a company that violates terms prohibiting automated collection or competitive training could be liable for breach, but enforcement against a foreign company operating through thousands of accounts may have little practical effect. U.S. courts have not established a comprehensive rule for when behavioral imitation of an AI system becomes illegal. The technology is advancing faster than the doctrine.

Why China creates greater alarm

The United States does not treat Chinese distillation as an ordinary commercial dispute. Artificial intelligence is increasingly viewed as a strategic technology with implications for military systems, surveillance, intelligence, cyber operations, scientific research and economic power. A Chinese company that narrows the model gap is therefore not merely gaining market share. It may be strengthening the technological position of a geopolitical rival.

American laboratories argue that China progress would be materially slower without access to outputs from U.S. systems, but that claim is difficult to prove. Chinese companies possess substantial engineering talent, domestic research institutions, enormous data resources and strong state support, and are capable of developing advanced models independently. Distillation may let them avoid some experimentation, identify productive training methods or transfer specific capabilities more quickly. It is likely an accelerant rather than a complete explanation. The distinction matters for policy: if distillation is the principal reason China is catching up, restricting it could preserve American leadership, but if China progress reflects a broader and increasingly self-sustaining research ecosystem, a crackdown may provide only temporary friction.

DeepSeek changed the strategic assumption

DeepSeek emergence altered the industry view of Chinese competition. The company demonstrated that strong model performance did not necessarily require the same scale of expenditure associated with leading U.S. laboratories. Its systems appeared to use computing resources more efficiently and were released under more open terms. OpenAI subsequently alleged that DeepSeek had distilled capabilities from its models, but even if true, the broader lesson was hard to dismiss: a Chinese laboratory had converted available techniques, engineering talent and potentially external model outputs into a system competitive enough to disrupt assumptions throughout Silicon Valley.

The market had previously treated American leadership as structurally secure because U.S. companies controlled the best chips, the largest computing clusters and the most heavily funded laboratories. DeepSeek suggested that efficiency, imitation and rapid diffusion could offset part of that advantage. The frontier might be expensive to reach first. It might be much cheaper to follow.

GLM-5.2 raises the stakes

The release of GLM-5.2 strengthens the concern that China model ecosystem is maturing. The system reportedly approaches leading American models across several demanding categories and performs strongly in cybersecurity, which carries unusual geopolitical weight. Advanced models can inspect code, identify vulnerabilities, generate software and reason through complex technical systems, functions that support both network defense and offensive operations. A country with broadly capable cyber models may improve its ability to discover flaws in foreign infrastructure, automate reconnaissance or accelerate software exploitation, and may also strengthen domestic defense.

American officials and technology executives therefore view the model race partly through a national-security lens. The concern is not only whether Chinese companies can build competitive commercial products. It is whether they can convert those systems into strategic capability.

How industrial distillation works

Large-scale distillation can be hard to distinguish from ordinary use when it is distributed across many accounts. A single account generating millions of repetitive queries would be easy to identify and suspend, but a network of accounts can obscure the pattern. Operators may vary prompts, access the system from different locations and imitate normal customer behavior, using intermediaries, resellers or automated account creation to reduce the visibility of the campaign.

The prompts themselves can be carefully designed. A competitor does not need random answers. It needs examples that expose reasoning, domain knowledge and behavior across specific categories, targeting advanced mathematics, software engineering, cybersecurity, legal reasoning, scientific analysis, tool use, instruction following and model-safety behavior. The collected outputs can then be filtered, ranked and incorporated into the training of the student model. The teacher may never reveal its internal structure. Its responses still provide a map of what strong performance looks like.

Detection is probabilistic

Frontier laboratories monitor usage patterns to identify suspected distillation, examining account creation, query volume, repeated prompt structures, geographic signals and similarities among groups of users. Thousands of newly created accounts submitting related technical prompts and exporting responses at unusual speed can strongly suggest coordinated collection, but the evidence is rarely perfect. A legitimate company conducting model evaluation may produce similar traffic, a research institution may run large benchmark suites, and a customer may build a high-volume commercial product through the same interface.

Providers must decide when suspicious activity becomes clear enough to justify suspension. If controls are too weak, systematic extraction continues. If controls are too aggressive, legitimate customers are blocked. This creates an adversarial balance: the distiller tries to resemble an ordinary user while the provider tries to identify intent from behavior, and as both sides improve their automation, the contest becomes increasingly difficult.

Why prevention is so hard

A model provider cannot fully prevent users from learning from its outputs without restricting the product itself, because every useful interaction reveals something about the system. Rate limits can reduce large-scale collection, identity verification can make account networks harder to create, behavioral monitoring can detect coordinated activity, and watermarking or response signatures may help identify outputs used elsewhere. None of these measures is complete. A determined competitor can spread requests over time, use human operators, purchase access through intermediaries or combine outputs from several models.

The more widely available the system becomes, the harder the behavior is to contain. The problem resembles other forms of digital copying: once information can be accessed, controlling how it is studied and reused becomes increasingly expensive. Frontier laboratories can create friction. They cannot restore absolute scarcity.

Cooperation among U.S. laboratories

Anthropic, OpenAI and Google have begun sharing information about suspected distillation campaigns, which may include account patterns, technical indicators and evasion methods. The logic resembles collaboration among banks combating fraud or technology companies identifying cyber threats: a network blocked by one provider may migrate to another, and shared intelligence can make that movement harder.

The arrangement also raises competitive and governance questions. The laboratories are commercial rivals with access to detailed information about users and developers, and joint enforcement could create a private security regime in which a small group of companies decides who is permitted to access advanced artificial intelligence. Collaboration may be necessary to address coordinated abuse, but it should still be governed by clear standards, privacy protections and methods for legitimate users to challenge incorrect restrictions. The solution to unauthorized extraction should not become arbitrary exclusion.

The appeal to Congress

Anthropic has asked lawmakers to create stronger mechanisms for cooperation between the federal government and frontier-model companies, seeking greater support for identifying and disrupting distillation operations linked to China. Possible measures include formal threat-information sharing, new penalties for evading model-access controls, restrictions on intermediaries facilitating prohibited access, support for identity verification, and expanded government authority over entities associated with foreign intelligence or military programs.

Legislation could improve enforcement inside the United States, but its reach would diminish rapidly outside the country. A Chinese company with no meaningful U.S. assets may have little reason to comply with an American judgment, and domestic rules may mainly affect cloud providers, payment processors, account resellers and other intermediaries connecting foreign users to U.S. models. That can raise the cost of distillation. It is unlikely to eliminate it.

Export controls as an indirect weapon

American laboratories also support restrictions on China access to advanced semiconductors. Distillation reduces the need to recreate every stage of frontier research, but the student model still requires substantial computing power to train and operate. The most advanced processors are designed by American companies and manufactured through a supply chain heavily influenced by U.S. export policy, so restricting access to those chips can slow Chinese model development regardless of where the training data originated. This makes semiconductor controls a more powerful instrument than litigation over model outputs.

The strategy has limits. China is investing heavily in domestic chips and alternative computing systems, restricted hardware can be obtained through third countries, resale networks or cloud access, and export controls may encourage faster Chinese self-sufficiency. The policy can preserve an advantage. It may simultaneously strengthen the incentive to escape American technological dependence.

Distillation challenges the frontier-lab business model

The strategic danger extends beyond China. If proprietary model behavior can be captured and reproduced quickly, frontier laboratories may struggle to maintain pricing power. Their business depends on an enormous upfront investment: they must train the strongest system, operate it at scale and charge enough to recover the cost. A competitor using distillation may avoid part of the original research expense and offer a slightly weaker model at a substantially lower price, and for many customers that trade is attractive. The laboratory at the frontier bears the economics of invention. The follower benefits from the economics of imitation.

pays for the chips, researchers, energy and experimentationTeacher
learns from observable outputs at a fraction of the costStudent
the lead that separates them, and shrinkingMonths
The economics of teacher and student. The frontier lab bears the economics of invention. The follower benefits from the economics of imitation. Source: SnowRock analysis, illustrative..

This does not make the original innovation worthless. The teacher remains necessary before the student can learn. It does reduce the duration of the advantage. A frontier lead lasting several months may not justify monopoly-like valuations if competitors can repeatedly close the gap, and the largest laboratories may find themselves funding a global research process whose benefits diffuse throughout the industry.

Model outputs as a strategic externality

Every time a frontier model is made publicly available, it produces value beyond the customer immediate use. Its responses can educate developers, reveal effective reasoning patterns and create data for subsequent systems. The provider receives revenue from the interaction. The industry receives information. This is a strategic externality: the laboratory cannot fully capture the value it creates because portions of that value leak through usage, and the more capable the model becomes, the more valuable its outputs become as training material.

This creates a paradox. Commercial distribution is necessary to generate revenue, adoption and feedback, yet the same distribution exposes the model to imitation. The provider must choose between limiting access and weakening growth or expanding access and accelerating diffusion. There is no stable solution.

Open models complicate the complaint

The distinction between proprietary and open models is central to the dispute. Open-weight developers often intend their systems to be studied, modified and distilled, aiming for broad adoption rather than control over every downstream use. This accelerates industry progress, letting researchers compare methods, create specialized versions and deploy models where a closed commercial service would be unsuitable. American frontier laboratories benefit from this ecosystem too, hiring researchers familiar with open work, incorporating public discoveries and building on decades of academic research. Their proprietary systems are not created in isolation. They emerge from a global knowledge base.

This does not give competitors the right to violate access terms. It does make the rhetoric of pure technological ownership harder to sustain. Modern artificial intelligence is built through constant borrowing, publication, imitation and recombination, and the legal and political fight concerns where legitimate learning ends and misappropriation begins.

Is distillation really the main threat?

Some researchers argue the importance of distillation is overstated. A collection of strong answers can help a model imitate visible behavior, but it cannot automatically reproduce the teacher full internal capabilities. The student still needs suitable architecture, training expertise, computing resources, evaluation systems and original data. Distillation may work well for particular tasks while failing to transfer deeper generalization, so a model could learn to answer familiar benchmark questions without acquiring the broader ability to perform in novel environments.

This means a top Chinese system cannot be explained simply as a copied American model. Its development requires substantial independent capability. Focusing too heavily on distillation may let U.S. companies attribute competitive pressure to unfair conduct rather than acknowledging that Chinese laboratories are becoming genuinely strong, and that misdiagnosis could produce ineffective policy. A competitor cannot be stopped through access controls when it has developed the capacity to innovate independently.

The agent era may reduce distillation importance

The next generation of artificial intelligence is increasingly focused on agents. An agent does more than answer a prompt. It operates software, searches for information, uses tools, manages long assignments and responds to changing conditions, capabilities that are harder to reproduce through simple input-output imitation. A model may produce excellent answers in isolation while failing to navigate a workplace, maintain context across hours or recover from unexpected errors. Training agents requires interactive environments, long sequences of feedback and experience with real or simulated tools, and a competitor cannot capture the full process merely by collecting final responses from a public interface.

This may reduce the value of traditional distillation. It will not eliminate imitation. Developers can still observe agent behavior, reproduce tasks and use stronger models to generate trajectories for weaker systems. But the training target becomes more complex, and the next competitive frontier may depend less on extracting answers and more on constructing superior environments in which models can learn to act.

Cybersecurity makes the problem harder

Cybersecurity is especially vulnerable to model imitation because much of the work can be expressed digitally. A model can inspect code, explain vulnerabilities and generate technical recommendations entirely through text and software, so high-quality responses become valuable training examples. At the same time, security providers cannot simply close access, because businesses and governments need advanced systems to identify vulnerabilities before attackers exploit them, and restricting the strongest models could weaken defenders. The same capabilities that create national-security concern are valuable precisely because they can improve national security.

This dual-use structure makes broad prohibitions difficult. A policy aimed at preventing Chinese military access may also block legitimate researchers, multinational companies and allied governments. The issue is not whether access should be controlled. It is whether controls can distinguish risk accurately enough to preserve defensive value.

China structural advantages

China possesses several advantages that distillation restrictions cannot remove. It has a large domestic technology industry, substantial scientific talent and a government willing to coordinate investment around strategic objectives. Chinese companies serve enormous consumer and enterprise markets, generating data and use cases for model development, and the country has strong capabilities in manufacturing, telecommunications, surveillance systems and industrial deployment.

These advantages matter because the competition will not be decided only by benchmark performance. It will be decided by diffusion. A slightly weaker model embedded throughout factories, logistics networks, public administration and military systems may create more strategic value than a superior model confined to commercial applications. The United States may lead at the frontier while China competes through scale of implementation. Distillation can help close technical gaps. Deployment can convert those gains into national power.

The risk of a defensive policy

American companies have a legitimate interest in protecting proprietary systems, and the government has a legitimate interest in preserving strategic leadership. The policy response can nevertheless become too defensive. An industry focused on preventing imitation may invest less attention in creating the next advantage, and restrictions can slow competitors but rarely create durable leadership by themselves.

The United States retains substantial strengths in research universities, capital markets, semiconductor design, cloud computing and entrepreneurial talent. The most effective strategy may be to extend those advantages rather than relying primarily on legal or technical barriers, which means investing in computing infrastructure, scientific research, workforce development and rapid commercial adoption. A country cannot preserve leadership indefinitely by preventing others from learning. It must continue learning faster.

Distillation as evidence of a maturing market

The prevalence of distillation may indicate that artificial intelligence is becoming a more conventional technology market. Early industries are often dominated by breakthrough inventions and large performance differences, while mature industries develop standard methods, competing suppliers and increasingly interchangeable products. Knowledge spreads, production becomes more efficient, margins decline and customers gain alternatives. Distillation accelerates this movement by shortening the period during which one model possesses a unique capability.

The result may be economically beneficial for users, because cheaper, smaller models can make artificial intelligence accessible to more companies and countries. It is less attractive for the laboratories that financed the frontier, whose innovation becomes a source of industry-wide progress rather than permanently defensible property.

The central strategic question

The dispute over Chinese copycats ultimately raises a larger question: can machine intelligence remain proprietary? The internal model weights can be protected, the computing infrastructure can be secured, and the customer interface can be governed by contract. But useful intelligence reveals itself through behavior. When a system explains, reasons and acts, it teaches observers something about how those capabilities can be reproduced, and the stronger the system becomes, the more valuable the lesson. Useful intelligence reveals itself through behavior. The stronger the system becomes, the more valuable the lesson.

American companies are asking government to help preserve the distance between teacher and student. That may slow the transfer. It is unlikely to stop it. The long-term advantage will not belong simply to the company or country that builds the strongest model first. It will belong to the one that can improve continuously, deploy effectively and create economic value faster than its capabilities can be copied. Distillation is not merely a method for training artificial intelligence. It is evidence that the frontier itself may be difficult to own.

The SnowRock take

Set the geopolitics aside and one instruction remains for an operator. If a nation-state can compress a frontier lead to a few months by studying a model behavior, so can your vendor competitors, which means the supremacy of whichever provider you buy from today is a wasting asset. Do not anchor a five-year plan to it.

In practice that means the same discipline we advise everywhere. Buy the model you need for the task in front of you, not the one with the best headline. Keep the freedom to swap providers without re-architecting your business. And put your durable investment where it cannot be distilled out of you, in your proprietary data, your workflows and the judgment to know which few tasks are worth automating at all. The frontier may be impossible to own. What happens inside your own operations is not. Rent the model. Own the workflow.

That is the whole lesson, read from the buyer side. The labs will keep spending to stay briefly ahead, and their leads will keep eroding. None of it should decide whether your business works. Build so that it does not.