Experts say exploiting Anthropic’s Fable isn’t how Kimi K3 got so good

The already tense landscape of US-China technological competition has been further inflamed by recent allegations from White House science advisor Michael Kratsios. Kratsios publicly accused Moonshot, a prominent Chinese artificial intelligence company, of illicitly copying Anthropic’s advanced Fable large language model (LLM) to develop its Kimi K3, currently recognized as the largest available open-weight LLM. Compounding the gravity of the claim, Kratsios asserted that Moonshot achieved this feat using semiconductor chips, specifically Grace Blackwell 300s, that are explicitly prohibited from export to China under stringent U.S. national security regulations.

Kratsios articulated the administration’s strong stance on the matter, stating via social media, "Large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology and undermining American research is unacceptable." This assertion comes amidst ongoing discussions within the U.S. government regarding potential bans on Chinese open-weight AI models, a prospect that has already generated considerable unease and debate across the global AI sector. Moonshot has yet to issue a public response regarding its training methodologies, and Kratsios refrained from providing further granular details concerning the provenance of his allegations.

The White House advisor’s statements find an echo in earlier remarks made by Treasury Secretary Scott Bessent, who previously noted the detection of "watermarks of our U.S. large language models on many of the Chinese models, and that that’s unacceptable." While the precise nature and composition of these "watermarks" remain undisclosed, and the Treasury Department has not elaborated on the issue, such claims point to sophisticated methods of identifying intellectual property infringement in the rapidly evolving domain of AI.

The Technical Debate: Distillation and its Efficacy

At the heart of the intellectual property dispute lies the concept of "distillation"—a process wherein a smaller, "student" LLM learns from a larger, more capable "teacher" LLM by systematically querying it to understand its internal workings and replicate its capabilities. This can involve asking the teacher model to articulate its chain-of-thought, or more commonly, using its prompts and responses to train the student model through a technique known as supervised fine-tuning (SFT). SFT is often where a model "picks up its manners," acquiring the stylistic and behavioral characteristics of its teacher.

However, experts in the field express significant skepticism that simple distillation alone could account for the advanced capabilities and rapid development of Moonshot’s Kimi K3, especially given the tight timeline. Braden Hancock, a distinguished researcher at the Laude Institute and co-founder of Snorkel AI, voiced these doubts to TechCrunch. "I don’t think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock explained. He highlighted the temporal impossibility, noting, "There’s just not even frankly time, right? Fable’s only been publicly available since July 1st. You can’t distill that much data, train a model, and release it in two weeks." The sheer volume of data required for effective distillation and the computational overhead for training a model of Kimi K3’s scale typically demand a much longer development cycle.

Nathan Lambert, an AI researcher at the Allen Institute for AI, further elaborated on this technical limitation in a recent podcast. Lambert posits that "distillation has becoming less and less impactful over time as the Chinese models get closer to the frontier and the training regime shifts to [reinforcement learning]." He argued that if basic distillation were sufficient, competitors would easily replicate the performance of models like Kimi K3 or GLM, a phenomenon not observed from supervised fine-tuning alone.

The transition to more complex and resource-intensive training methodologies, such as reinforcement learning (RL) from human feedback (RLHF) or even AI feedback (RLAIF), becomes crucial for achieving frontier-level performance. To distill Fable-like capabilities through these advanced techniques would necessitate an intricate process where a larger model’s "agent" grades the responses of a smaller model, iteratively adjusting its parameters based on the feedback. Such advanced reinforcement learning runs can involve tens of millions of agents, requiring monumental computational infrastructure. Attempting to execute such extensive RL processes via a frontier lab’s API would be prohibitively expensive and inherently slow, potentially offering little to no performance uplift compared to indigenous development. The consensus among many experts is that while distillation can play a role in initial model development or fine-tuning, it is unlikely to be the sole or primary driver behind creating a leading-edge LLM from scratch in such a short timeframe. This technical nuance underscores the complexity of proving IP theft in the AI domain, where independent innovation can sometimes mirror the outputs of existing models.

A History of Accusations and Industry Practices

This is not the first instance where Anthropic has publicly accused Chinese companies of systematically distilling its models. Earlier in the year, Anthropic directly named Moonshot, DeepSeek, and MiniMax, alleging that it had discovered millions of exchanges between its models and users identified at these companies through IP addresses and other metadata. These queries were described as "distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use." Despite these previous allegations, Anthropic has not yet responded to queries specifically concerning the alleged distillation of its Fable model.

It is also critical to acknowledge that the practice of distillation, or the use of synthetic data derived from querying other models, is a widespread and often blurry area within the AI industry, not confined solely to Chinese firms. Elon Musk, for instance, testified earlier this year that his company, SpaceXAI, utilized OpenAI models to develop its own LLM, Grok, and openly stated that such practices were common across the industry. The precise demarcation between legally permissible "reverse engineering" or the creation of synthetic datasets and outright intellectual property theft remains a contentious and evolving legal and ethical challenge. This ambiguity complicates the enforcement of IP rights in a field characterized by rapid iteration and knowledge sharing.

The Illicit Chip Trade: A National Security Concern

Beyond the intellectual property dispute, Kratsios’s allegations delve into a critical national security concern: Moonshot’s alleged acquisition and use of advanced Nvidia chips, specifically Grace Blackwell 300s, and access to GB300-equipped servers in Thailand. These chips are subject to stringent U.S. export controls, explicitly banned from shipment to China, as part of a broader strategy by the U.S. Department of Commerce to curb China’s access to cutting-edge semiconductor technology that could be leveraged for military modernization or to gain a decisive advantage in critical AI applications.

The Grace Blackwell 300s represent the pinnacle of current AI accelerator technology, offering unparalleled computational power essential for training and deploying large, sophisticated LLMs. Their restricted availability is intended to slow China’s progress in developing advanced AI capabilities. However, as Sam Bresnick, a research fellow at Georgetown’s Center for Security and Emerging Technology, points out, a black market for these prohibited chips undeniably exists. This illicit trade facilitates the circumvention of export controls, enabling entities in China to access technology deemed critical by the U.S. In a tangible example of this challenge, the founder of Supermicro, a U.S. server manufacturer, was indicted in May for allegedly smuggling advanced chips into China, highlighting the real-world complexities and vulnerabilities in the export control regime.

The U.S. government has been actively pursuing measures to strengthen these controls. In 2024, President Joe Biden’s Department of Commerce proposed federal "know-your-customer" (KYC) rules for data centers globally, aiming to create a reporting mechanism for companies conducting large training runs on state-of-the-art hardware. Bresnick is a strong proponent of such regulations, stating, "If you are letting a company conduct huge training runs on your state-of-the-art hardware, there needs to be a reporting mechanism for who that company is and what they’re doing." However, under the subsequent administration, progress on these proposed KYC rules appears to have stalled. Despite this, existing regulations mandate that exporters shipping advanced chips abroad are supposed to ensure that these technologies are used only for approved purposes, placing a significant burden of due diligence on American companies.

Broader Implications for the US-China AI Rivalry

The allegations against Moonshot underscore the multifaceted nature of the US-China AI rivalry, which encompasses not only technological innovation but also intellectual property protection, national security, and global economic leadership. The potential ban on Chinese open-weight models, a topic currently under discussion, reflects the deep-seated concern within the U.S. government that such models could pose risks, whether through data exploitation, biased outputs, or their potential use in applications antithetical to U.S. interests.

However, many experts caution against underestimating the indigenous technical capabilities of Chinese AI researchers. Braden Hancock emphasized this point, noting, "[I]n general, Americans are understating the technical expertise of these Chinese teams." He highlighted that "one of the founders of Moonshot was a CMU PhD student. These are legitimate researchers and engineers doing solid work." Hancock contended that while a complete halt in American model development might slow China’s progress, it would not stop it entirely, stating, "They’re not just riding coattails here." This perspective is crucial for a balanced understanding of the global AI landscape, acknowledging China’s significant investment in AI research, development, and talent cultivation over the past decade.

The implications of these developments are far-reaching. For intellectual property, the case highlights the urgent need for clearer international norms and enforcement mechanisms in the age of AI. The technical challenges of proving "copying" in LLMs, coupled with the ease of information dissemination and the blurred lines of synthetic data generation, present formidable obstacles to traditional IP frameworks. For national security, the alleged circumvention of export controls on advanced chips represents a direct challenge to U.S. efforts to maintain its technological edge and prevent the proliferation of dual-use technologies. The incident also serves as a stark reminder of the complexities of global supply chains and the difficulty of enforcing controls in a deeply interconnected world.

Ultimately, this unfolding dispute is a microcosm of the broader geopolitical competition for dominance in artificial intelligence. It will likely prompt further tightening of export controls, increased scrutiny of international data center operations, and intensified efforts to protect proprietary AI technologies. The outcome of these allegations, and the policy responses they trigger, will undoubtedly shape the future trajectory of AI development, international tech collaboration, and the delicate balance of power between the world’s leading technological nations.

Related Posts

I tried out OpenAI’s new AI keypad — which will be fun for some coders and slightly mystifying to everyone else

This debut marks a significant strategic pivot for the leading AI research and deployment company, traditionally known for its groundbreaking software and language models. Developed in collaboration with specialty keyboard…

Why Cognition bought Poke: AI personality is becoming a competitive advantage

The burgeoning landscape of artificial intelligence witnessed a significant strategic maneuver with the acquisition of The Interaction Company of California, the innovator behind the popular AI assistant Poke, by Cognition,…

You Missed

Japan’s Luxury Sector Shines as Jewellery Sales Soar 19% Amidst Inflationary Pressures and Yen Depreciation

Japan’s Luxury Sector Shines as Jewellery Sales Soar 19% Amidst Inflationary Pressures and Yen Depreciation

The APOE2 Gene Variant Offers Enhanced Neuronal Protection Against DNA Damage and Cellular Senescence, Unlocking New Avenues for Alzheimer’s Research

The APOE2 Gene Variant Offers Enhanced Neuronal Protection Against DNA Damage and Cellular Senescence, Unlocking New Avenues for Alzheimer’s Research

The Hidden Environmental Cost of the Puffer Jacket: Unpacking the Footprint of a Cold-Weather Staple

The Hidden Environmental Cost of the Puffer Jacket: Unpacking the Footprint of a Cold-Weather Staple

The Evolution of Modern Storage: A Comprehensive Guide to High-End Sideboards and Credenzas in Interior Design

The Evolution of Modern Storage: A Comprehensive Guide to High-End Sideboards and Credenzas in Interior Design

Volker Türk Becomes First UN Human Rights Chief to Secure Two Full Terms Amidst Significant International Division

Volker Türk Becomes First UN Human Rights Chief to Secure Two Full Terms Amidst Significant International Division

Ralph W. Hemecker, Acclaimed Television Director and Showrunner, Dies at 65

Ralph W. Hemecker, Acclaimed Television Director and Showrunner, Dies at 65