Key Takeaways
- DeepSeek’s V4 model is a large, open‑weight system (1.6 trillion parameters) that is priced for mass deployment but remains hampered by compute shortages.
- While V4 trails the top U.S. frontier models (GPT‑5.4, Gemini 3.1‑Pro) in raw performance, it offers a one‑million‑token context window via a novel hybrid‑attention architecture at far lower inference cost than its predecessor V3.
- U.S. officials allege that V4 was trained on smuggled Nvidia Blackwell chips and that DeepSeek has used industrial‑scale distillation attacks to extract capabilities from American models, saving hundreds of millions—or even billions—of R&D dollars.
- Chinese government subsidies, especially through ties to Huawei, likely further depress DeepSeek’s pricing, though current chip shortages make those prices moot for large‑scale deployment.
- The real contest is not just performance but adoption: cheap, open models diffuse quickly in the Global South, where Hugging Face downloads already favor Chinese systems over U.S. counterparts.
- To preserve its lead, the United States should combine defensive measures (threat‑intelligence sharing, industry collaboration) with offensive steps—sanctions, Entity List designations, and multilateral condemnation of distillation as industrial espionage.
Overview of DeepSeek V4 Release
When DeepSeek unveiled its V4 model, the immediate headlines focused on whether it had narrowed the gap with American frontier systems. As Michael C. Horowitz notes, “U.S. models remain ahead—but that’s the wrong question.” The release arrived a day after the White House formally accused Chinese actors of running industrial‑scale campaigns to extract capabilities from U.S. frontier models, underscoring the timing’s geopolitical relevance.
Performance Gap Relative to U.S. Frontier Models
On raw performance, V4 is impressive but still trails the American frontier. DeepSeek claims that “V4‑Pro beats every other open‑weight model, but says it falls short of OpenAI’s GPT‑5.4 and Google’s Gemini 3.1‑Pro.” This admission places V4 roughly seven months behind the leading U.S. systems, a gap corroborated by independent expert estimates cited in the Council on Foreign Relations analysis.
Open‑Source Scale, Cost Advantage, and Deployment Challenges
V4 is large in scale—the Pro version boasts 1.6 trillion parameters—and is priced for mass deployment at least four times cheaper than comparable American competitors. Yet DeepSeek itself admits that “it currently cannot serve its v4 pro model to most customers because it lacks the chips to do so.” The resulting compute shortage renders the low price moot for large‑scale customers, limiting the model’s real‑world impact despite its attractive cost profile.
Algorithmic Innovations in V4
Despite the hardware constraints, V4 does contain real engineering achievements. Notably, “a hybrid attention architecture that enables a one‑million‑token context window at a fraction of V3’s inference compute cost.” This advance shows that Chinese labs are keeping pace with algorithmic efficiency gains occurring in closed‑source U.S. labs, even when they struggle to procure the latest silicon.
Accusations of Illicit Compute and IP Theft
U.S. government officials have asserted that “V4 was still trained on smuggled Nvidia Blackwell chips, which are banned in China.” The V4 report, however, remains silent on the exact chips used, a omission that fuels suspicion. Separately, Anthropic and OpenAI have alleged that DeepSeek engaged in “industrial‑scale” distillation attacks on its Claude models, “creating over twenty‑four thousand fake accounts and conducting more than sixteen million interactions to extract capabilities and improve its own systems.” Such tactics, if proven, represent a massive transfer of U.S. intellectual property that helps offset DeepSeek’s R&D expenses.
Effect of U.S. Controls, Subsidies, and Supply‑Chain Constraints
Analysts argue that DeepSeek’s low prices are likely enabled by Chinese government subsidies, particularly given the firm’s direct integration with Huawei, and by the savings from distillation attacks that “save it hundreds of millions or even billions of dollars in R&D costs.” Meanwhile, the broader U.S. strategy to constrain China’s access to AI compute has yielded a provisional seven‑month lead, but V4 demonstrates that China continues to exploit loopholes to obtain capable, if not directly competitive, models. Closing those loopholes—restricting access to U.S. AI chips, models, and chipmaking tools—could stretch the American advantage from months to years.
The Adoption Frontier: Why Diffusion Matters More Than Raw Benchmarks
Horowitz reframes the competition: “Understanding what V4 reveals about the AI competition requires asking a different one: who is winning the most important U.S.-China competition, the adoption race?” He argues that success in converting AI into global power depends less on having the absolute best model and more on deploying “good‑enough solutions that can be deployed quickly and at scale.” Second‑best models that are cheap and open enjoy enormous competitive value because they diffuse easily, especially in markets where buyers are not choosing between GPT‑5 and Claude Sonnet 4.6 but between accessible tools with different value propositions.
Global South Uptake and Hugging Face Evidence
Empirical support for the adoption thesis appears in Hugging Face download statistics: “Chinese AI models already have more downloads on Hugging Face, the open-source AI platform, than those from the United States.” This trend suggests that countries in the Global South are gravitating toward Chinese‑offered tools, valuing affordability and openness over marginal performance gains. The diffusion of V4‑class models could therefore translate into broader economic and strategic influence for Beijing, even if the models never match the peak performance of U.S. frontier systems.
U.S. Policy Responses to Counter Distillation and Chip Smuggling
To safeguard its lead, the United States should blend defensive and offensive measures. Defensive steps include sharing threat intelligence with U.S. firms and removing impediments to industry collaboration. Offensively, policymakers could: explore sanctions against firms engaged in distillation attacks; cut off their access to U.S. financial markets and dollar transactions; add those entities to the Department of Commerce’s Entity List, warning cloud providers, chip vendors, and equipment suppliers of regulatory exposure; and multilateralize pressure by building consensus that distillation constitutes industrial espionage—already signaled by a State Department directive instructing diplomatic staff to raise the issue abroad.
Looking Ahead: Sustaining the U.S. Lead
Jessica Brandt cautions that DeepSeek V4 “is undoubtedly a capable model, though it appears to be an incremental advance over other Chinese offerings rather than a competitive challenge to the U.S. frontier.” Nonetheless, the model underscores that China will continue to exploit any available pathways to narrow the gap. As Horowitz warns, “The United States is still ahead. It cannot afford to assume that advantage is self-sustaining.” By tightening export controls, confronting illicit distillation, and fostering rapid, scalable deployment of its own AI innovations, the United States can aim to extend its lead from a fleeting seven‑month edge to a durable, multi‑year advantage in the global AI race.
https://www.cfr.org/articles/deepseek-v4-signals-a-new-phase-in-the-u-s-china-ai-rivalry

