On April 24, 2026, DeepSeek launched the V4 preview: the Pro model a 1.6-trillion-parameter mixture-of-experts with 49 billion active per pass, one million tokens of context. The world’s most capable open model family went live on Huawei’s Ascend chips and CANN software, not CUDA. Alibaba Cloud and Tencent Cloud served it within hours. Huawei’s own chips had trained part of V4-Flash.
Jensen Huang had warned on the Special Competitive Studies Project’s Memos to the President podcast: “The day that DeepSeek comes out on Huawei first, that is a horrible outcome for our nation.”
That day arrived on April 24, 2026.
The controls didn’t fail because they were poorly designed. They failed because they worked exactly as designed, and the mechanism they created produced the opposite of the intended result. Washington wanted to contain China’s frontier AI. What it built instead was a captive market with guaranteed demand, revenue, bug reports, production experience, and ecosystem maturity. Every restriction functioned as a forced subsidy.
Nvidia’s moat was never the chips. Chips are physical objects: you can count them, license them, and stop them at a border. The moat was the software wrapped around the silicon. CUDA: the compiler, the libraries, the profilers, the two decades of accumulated engineering knowledge, and the fact that virtually every AI framework, every research paper, and every pre-trained model assumes CUDA underneath. The silicon was the drawbridge. The software was the castle.
Here’s the part Washington got backwards: you can stop chips at a border, but you can’t stop software the same way. Software isn’t a thing that ships. It’s a habit, a set of skills, a web of dependencies that lives in heads and repos and CI pipelines. Deny the chips and you don’t dissolve the demand for compute. You redirect it. The controls were aimed at the drawbridge. The castle was never in range.
I want to trace how that happened. Not from the headlines down, but from the engineering floor up. Because the CUDA moat didn’t break where everyone thought it would. It didn’t break at the ISA or the compiler layer. It broke from the abstraction layers down. That matters for how you think about moats, policy, and what happens when a billion-dollar ecosystem is walled off from half its market.
What Washington Thought It Was Doing#
The export control campaign against China’s AI compute didn’t arrive all at once. It rolled out in escalating waves, each one tighter than the last, and each one justified by the same logic: deny frontier chips, slow frontier AI, and preserve America’s lead.
October 7, 2022: the Bureau of Industry and Security imposed licensing requirements on advanced computing chips bound for the PRC. A100s and H100s were captured by performance thresholds, not named individually. The idea was surgical. Block the chips that power large-scale model training, and you block the models themselves.
October 17, 2023: Washington closed the loophole. Nvidia had been shipping H800 and A800 variants to China, designed to sit just below the original thresholds. The new rules added a performance-density metric that caught those chips, too.
January 2024: Reuters reported that Nvidia planned mass production of the H20 for the second quarter. The chip was engineered to comply with the 2023 rules. Below the license threshold, no approval was needed. A compliant chip for a captive market.
April 9, 2025: Commerce reversed course. The H20 now required licenses over supercomputer-diversion risk, and Nvidia disclosed a charge of up to $5.5 billion. August brought another reversal: Commerce began issuing H20 licenses again.
January 13, 2026: H200 sales reopened on a case-by-case basis, with third-party testing, adequate US supply, no military use, and Chinese purchases capped at 50 percent of US-customer volume. David Sacks, the White House AI czar, explained the logic: allowing advanced US chips “discourages” Chinese rivals from “redoubling efforts.” He was right about the incentive structure. He was wrong about what scarcity actually does.
March 17, 2026: Beijing approved H200 sales. Nvidia said it had licenses and orders for many Chinese customers. As of June, zero H200 units had shipped. The US required China-only use. Beijing told Chinese firms to restrict Nvidia chips to overseas deployments. A stalemate dressed up as a deal.
April 2026: the MATCH Act proposed chokepoint-tool controls targeting CXMT, Hua Hong, Huawei, SMIC, and YMTC. Commerce also sent “is-informed” letters to Hua Hong fabs, halting certain tool and material shipments.
The logic was consistent across every phase. Deny compute, deny progress, and maintain the gap. But the mechanism had a flaw that no one in Washington seemed to model: a denied market doesn’t disappear. It reorients.
Restriction as Subsidy#
Here’s how the backfire actually worked.
When you cut a company off from its preferred supply, you don’t eliminate demand. You guarantee it. Every Chinese cloud provider, AI lab, and enterprise that needed accelerator compute suddenly had only one viable path: domestic silicon. That guaranteed demand became revenue. Revenue funded R&D. R&D produced better chips. Better chips attracted more users. More users filed more bug reports. Bug reports drove ecosystem maturity.
It’s a flywheel. And every export control spun it faster.

The R2 failure arc tells this story with unusual clarity. In August 2025, DeepSeek tried to train its next model on Huawei’s Ascend. It failed. Not partially, and not with degraded performance. The system couldn’t complete a single successful training run. DeepSeek reverted to Nvidia’s H20 and delayed the model. That failure should’ve been the end of the Ascend story.
It was the beginning.
Failure on that scale produces data: bug reports, error logs, and specificity. It’s the kind of specificity a software team needs to fix a stack built for a different world. DeepSeek went to work with Huawei and Cambricon to rewrite pieces of the codebase. By April 2026, V4 was running on Ascend. Huawei’s own chips had contributed to V4-Flash training, and the entire Ascend SuperNode line supported V4 from day one.
The Technical Spine: Where the Moat Actually Broke#
Everyone assumed the CUDA moat lived at the ISA layer: the instruction set, the compiler, and the bare-metal interface between software and silicon. If you were going to break CUDA, the thinking went, you’d have to replace PTX, SASS, and nvcc. Nobody could do that.
The moat broke somewhere else entirely. It broke from the abstraction layers down.

CANN, Huawei’s Compute Architecture for Neural Networks, was fully open-sourced in August 2025 at the Ascend Computing Industry Summit. The scope was precise: the compiler and virtual ISA interface were open, the rest of the stack was open, the license was permissive, and forks and ports were allowed from day one. This is the critical difference. CUDA’s license bans translation layers that target its API. CANN’s license invites them. Eric Xu made it clear: this was a community play, not just a product play.
vLLM-Ascend tracks upstream vLLM release for release. Version 0.23.0rc1 in July 2026 added Ascend 950 support, context parallelism, and KV-cache offload. DeepSeek-V4 support landed in v0.20.2 in June, and Mooncake connector support followed in v0.22.1rc1. torch_npu tracks PyTorch on CANN. Triton-Ascend and TileLang-Ascend give kernel authors CUDA-like dialects to work in.
The moat didn’t break because someone cloned CUDA. It broke because someone built a parallel stack that connected the same abstractions and let you do the same work in a different dialect.
A field study published on arXiv in July 2026 makes this concrete. Researchers ran DeepSeek-V4-Flash workloads on a 16-device Ascend 910 cluster. They documented 12 source patches across eight limitation categories. Real friction. Working workloads. The patches tell you exactly where the ecosystem still hurts, and the working workloads tell you exactly where it doesn’t matter anymore.
The SLAI team published its own work: full-parameter post-training of the DeepSeek-V4 family on an Ascend SuperPOD, hitting 34.22 percent MFU. That’s not a toy benchmark. That’s production-scale fine-tuning on domestic silicon.
What the Market Shows#
The numbers are harder to argue with than the technical details.
IDC’s 2025 accelerator survey for China came in at roughly 4 million units. Nvidia held 2.2 million, about 55 percent. AMD had 160,000, roughly 4 percent. Chinese vendors took 1.65 million, or 41 percent. Huawei led with 812,000, followed by T-Head at 265,000, and Kunlunxin and Cambricon at 116,000 each.
Morgan Stanley’s Charlie Chan estimated that Huawei held 62 percent of China’s AI-accelerator segment as of May 2026, with Cambricon at 14 percent, and Baidu and Alibaba at 5 percent each. Morgan Stanley projected a $67 billion China AI-GPU total addressable market by 2030, with 86 percent of supply coming from Chinese vendors. Bernstein put Nvidia’s share at roughly 8 percent and falling. A Bloomberg survey, summarized by The Next Web, found that Chinese AI companies were raising domestic chip allocations from 30 to 46 percent of their budgets over 12 months.
Huang described the result bluntly on the SCSP podcast: “In China, we have now dropped to zero,” and the export controls had “already largely backfired.” In a separate CNBC interview on May 21, 2026, he said Nvidia had “largely conceded” China to Huawei because “we’ve evacuated that market.”
The hardware roadmap tells its own story. The 950PR is the only domestic chip with FP8 support. It sits between the H100 and H200 in performance and, by Huawei’s own claims, delivers roughly 2.8 times the H20’s performance (unverifiable: Hopper lacks native FP4). Mass production started in April 2026, targeting about 750,000 units, with full-scale shipments expected in the second half. Prices are up 20 percent on demand. The 950DT comes in the fourth quarter of 2026, and the Ascend 960 in the fourth quarter of 2027. CloudMatrix 384 is in production at Wuhu. CXMT is developing its own HBM.
This isn’t a story about catching up. It’s a story about building a parallel ecosystem that doesn’t need to catch up anymore.
What the Voices Outside Washington Say#
CSIS published a paper in March 2026 with a finding that should’ve grabbed headlines and didn’t: the controls’ “principal effect” was accelerating indigenous adoption. They “inadvertently catalyze” semiconductor autonomy. An earlier CSIS brief from April 2025 found that China had “doubled down,” and that controls “cannot impede long-run innovation and are at best a short-term palliative.”
Brookings’s John Villasenor wrote in January 2025 that DeepSeek showed the controls “may actually be accelerating” Chinese AI progress. Scarcity, he noted, fostered more efficient training innovation. Omdia’s Su Lian Jye called V4’s Ascend launch “real, tangible progress” toward infrastructure self-sufficiency.
These aren’t contrarian takes. They’re mainstream analyses from institutions that generally support the strategic rationale for export controls. The gap isn’t in the diagnosis. It’s in the response.
What the Controls Actually Did Slow#
I don’t want to write this as a clean victory lap. That would be dishonest, and the data doesn’t support it.
The controls did things. They just didn’t do the thing Washington wanted.
DeepSeek’s V4 Pro costs up to 12 times more to run than Flash because of constraints in high-end compute capacity. That’s a real constraint, reported by Reuters in April 2026, and it limits how widely the Pro tier can be offered. SMIC faced yields that Reuters reported near 20 percent for the 910C in late 2024, versus the 70 percent threshold for commercial viability. An unnamed senior US official alleged in February 2026 that V4 was trained on smuggled Blackwell chips. No evidence was provided, so it should be treated as the allegation it is.
DeepSeek-V3 was trained on 2,048 H800s, the compliant-chip path that the controls couldn’t block. That compliance window matters. China’s open-model leadership started on hardware that was legal to buy.
The controls shaped the trajectory. They didn’t stop it. There’s a difference.
The Human Cost of a Shifted Ecosystem#
Here’s the part that doesn’t fit in a policy brief.
Engineering teams in China didn’t all choose to migrate to Ascend. They were pushed. The researchers who wrote those 12 patches in the field study didn’t set out to document the growing pains of a new stack. They set out to train models. The patches were the cost of doing business in a market reoriented without their consent.
DeepSeek engineers watched R2 training runs fail on Ascend, went back to Nvidia chips, and then returned to the drawing board to rewrite pieces of the codebase for Huawei and Cambricon. They did the work because the alternative was to wait for chips that might never arrive.
And if you’ve spent 15 years building CUDA expertise, you know the compiler quirks, memory-layout tricks, and kernel-optimization patterns almost by feel. Now that skill set has a market-specific expiration date. Not because CUDA is dying, but because a parallel ecosystem is maturing fast enough that a global company can no longer assume CUDA is the only path.
That’s not a geopolitical abstraction. That’s your relationship to your craft changing under your feet. And it’s happening to people on both sides of the Pacific.
Two Worlds#
The strategic threat isn’t that CUDA gets replaced. It’s that it doesn’t have to be.
What’s emerging isn’t a successor ecosystem. It’s a parallel one: self-sustaining, well-funded, and growing. And the cost of living in both worlds is going up for everyone. Global companies now need two stacks, two codebases, and two teams. Developers need two skill sets. Open-source projects need two CI pipelines.
Write once, run anywhere is dead.
The question isn’t whether one ecosystem wins. The question is what it costs you to exist in both, and whether the policy that created this bifurcation was worth the price.
I think the answer depends on whether you measure success in chips denied or ecosystems created. By the first metric, the controls worked. By the second, they didn’t just fail. They produced the opposite of what they intended. A captive market became a self-sustaining one. A containment strategy became a subsidy. A moat became a mirror.
And when DeepSeek shipped on Huawei first, the mirror showed us what was always going to happen. You can’t wall off half a market and expect the wall to hold.
The Copycat Premise#
Underneath all of this sits a premise that most of Washington never examined: China copies, and America invents. Cheap knockoffs, the thinking goes, and that’s the ceiling of what they can do.
The evidence in this article is the counterargument. A 1.6-trillion-parameter model that leads every open model in the world, built to run on a domestic chip and a domestic software stack that barely existed four years ago. A field study that documented twelve patches and eight limitation categories, and the engineers who closed most of those gaps within months. That isn’t copying. That’s world-class data science, world-class systems engineering, and world-class software engineering, done by people who live and work in China.
They’re the kind of people we would have hired in America. The same policy instinct that denied China chips also denied China’s best minds entry to the United States: visas refused, work blocked, citizens deported. Secretary of State Marco Rubio announced in May 2025 that the US would “aggressively” revoke visas of Chinese students, a move Reuters noted could break “a crucial pipeline of talent for U.S. technology companies.”
That pipeline didn’t break. It went home. We didn’t just create the demand for a domestic ecosystem. We exported the talent that would build it, and Chinese academics are heading back in growing numbers. Now those minds are concentrated in one place, and the state holding them has a vested interest in global AI dominance.
Smugness isn’t a character flaw here. It’s a strategic decision. It told Washington it didn’t need to look, didn’t need to learn, didn’t need to compete with a country it had already dismissed.
Playing poker against China is a bad idea. Not because they cheat. Because they’re holding the cards we threw away, and they’ve been playing with them for years while we were sure we’d already won.
Sources: DeepSeek V4 preview release, DeepSeek V4 launch, V4 on Huawei, Chinese firms’ Ascend order scramble, H200 export reopening, Chinese H200 approval, IDC market data, DeepSeek R2 delay, SMIC yields, Nvidia SEC filing, H20 license resumption, BIS October 2022 controls, CSET October 2023 explainer, H20 production plan, MATCH Act, Hua Hong shipment halt, Blackwell allegation, V4 withheld from US chipmakers, CNBC interview with Jensen Huang, SCSP interview with Jensen Huang, CSIS March 2026 analysis, CSIS April 2025 analysis, Brookings analysis, Ascend field study, SLAI T-Rex, DeepSeek-V3 report, vLLM-Ascend releases, torch_npu, The Next Web overview, Reuters: US to revoke Chinese student visas, AP: Chinese student deported at Houston airport, CNA: Chinese academics returning home, and Tom’s Hardware market and roadmap report.
Related: AI’s Architect Problem: Why We’re Building on Borrowed Land
