By Djellal Djouad
Three signals landed on the global tape in 72 hours. Each one looked discrete. Read together, they confirm the first vector of our China AI Disruption Thesis — and they have direct, dated implications for US hyperscaler equities and credit.
May 26, 2026 — DeepSeek made permanent a 75% price cut on its flagship V4-Pro model: $0.435 per million input tokens, $0.87 per million output tokens. (Reuters)
May 27, 2026 — Xiaomi cut MiMo-V2.5 API prices by up to 99% globally: cached input at $0.0028 per million tokens, output at $0.28 per million. (Caixin)
May 28, 2026 — Shanghai Futures Exchange (SHFE) is designing futures contracts for AI tokens — the smallest unit of information processed by AI models, per Reuters reporting. (Yahoo/Reuters)
The market is filing each of these under "China noise". That is the wrong filing cabinet. These three events are the same event seen from three angles: AI tokens are being financialized as a commodity, in China, on the same exchange that prices copper. And the price discovery curve they imply is incompatible with the back half of 2026 capex calls coming out of US hyperscalers.
Why these three signals are the same signal
Our China AI Disruption Thesis lays out five convergent shocks. Vector 1 is token commoditization — the structural collapse of unit economics on AI inference. The standard objection from the sell-side has been: "sure, prices fall, but Jevons takes over — cheaper tokens mean more tokens consumed, and revenue holds." The events of this week dismantle that defense in three distinct ways.
Signal 1 — SHFE: tokens get a futures market
SHFE is not Hong Kong. It is not a crypto exchange. SHFE is the venue that prices Chinese copper, zinc, gold, silver, and crude. Putting AI tokens next to copper on the same exchange is a sovereign statement: AI tokens are infrastructure, not software.
Xiao Feng (HashKey CEO), quoted in the Reuters report:
"Tokens function as the 'digital fuel' or 'raw material' that powers AI models."
This is the linguistic register of a commodity, not a SaaS product. Reuters also reports:
"China's daily token usage increased 1,000-fold since early 2024, reaching over 140 trillion by March 2026."
No commodity gets a futures contract without that kind of consumption density. The SHFE filing is the financial plumbing being built around an already-commoditized underlying. Note the geopolitical split: US (CME / ICE) is preparing GPU compute futures — hedging the input cost (chip rental). China (SHFE) is preparing AI token futures — hedging the output (AI service pricing). Same supply chain, opposite ends. Both jurisdictions are financializing AI simultaneously, and they have decided the relevant unit is on different sides of the cost curve.
Signal 2 — DeepSeek makes -75% permanent
DeepSeek's V4-Pro discount was originally a promotional cap. As of May 31, 2026, it becomes the floor. New official pricing per the DeepSeek API docs: $0.435 per million uncached input tokens, $0.87 per million output tokens, with cache hit reduced to one-tenth of launch price as of April 26, 2026.
Compare to the US frontier:
OpenAI GPT-5.5 — doubled output pricing to $30 per million tokens at launch.
Anthropic Claude Opus 4.7 — shipped with an updated tokenizer that inflates actual cost by up to 35%.
DeepSeek V4-Pro — output at $0.87 per million tokens (now permanent).
The gap between Chinese and US frontier on output is now approximately 30-35x before cache discounts (Decrypt). This is not a temporary subsidy. It is consistent with a mixture-of-experts architecture, Chinese-sourced training data, Chinese inference hardware, and Chinese inference deployment. The unit economics support the price.
Signal 3 — Xiaomi MiMo at -99%: the commodity floor reveals itself
Then Xiaomi entered. On May 27, MiMo-V2.5 prices were cut up to 99% globally. The new structure (CTOL Digital):
Input (cache hit): $0.0028 per million tokens
Input (cache miss): $0.14 per million tokens
Output: $0.28 per million tokens
MiMo-V2.5-Pro cache hit: ¥0.025 per million tokens
Critically, Fuli Luo, head of the Xiaomi MiMo team, made the structural point explicit:
"Operating at these newly reduced API prices, our production inference engine is running at near full capacity, and we can still essentially break even."
That is the sentence that should reframe every US AI-infra long thesis in the back half of 2026. Break-even at $0.28 per million output tokens — at full capacity — is a statement about the floor of the price curve, not the ceiling. The arithmetic only works if the underlying capex (Huawei silicon, Chinese fab capacity, Chinese energy mix) sits at a fraction of what NVIDIA/TSMC/PJM-grid US deployments cost. That capex disparity is Vector 2 and Vector 3 of our thesis — which we'll connect below.
The Jevons trap — and the lag the sell-side is ignoring
The most common counter-thesis right now goes like this: "tokens get cheaper → enterprises use 10x more → token revenue stays flat or grows." This is the Jevons paradox applied to AI. It is also exactly the argument Rich Privorotsky (Head of European One Delta trading, Goldman Sachs) raised in his desk note this week — and not in a bullish framing.
Privorotsky's actual question was not "will Jevons hold", it was: is there a meaningful lag where cheaper tokens cannibalize more expensive inferences before new use cases emerge?
That distinction matters. The Jevons argument assumes the elasticity is instantaneous. The market structure says it is not. Three points:
Cannibalization runs first. Most current enterprise AI workloads are interchangeable with cheaper open-weight inference. A 95-99% price cut on tokens against an indistinguishable workload produces immediate revenue compression for the incumbent. New use cases (agentic workflows, complex reasoning) take quarters to deploy.
Diminishing token returns are now documented at the enterprise. Jellyfish analyzed 12,000 developers across 200 companies in Q1 2026: the top 20% of token spend produced 23 merged PRs/quarter, versus 11 PRs for the bottom 20%. 2.1x token consumption for 2.1x throughput at the high end — but the cost-per-PR ramped sharply. Linear in tokens, linear in cost, no efficiency curve in between. (Tom's Hardware)
Enterprise procurement is starting to push back. Andrew Macdonald, COO of Uber, on the Rapid Response podcast last week:
"If you're not actually able to draw a direct line to how many useful features and functionality you're shipping to your users, that trade becomes harder to justify."
Uber burned its entire 2026 AI tools budget in four months. Macdonald coined the term "tokenmaxxing" to describe employees inflating usage numbers without proportional output. That is the same dynamic Amazon, Meta, and Microsoft are now grappling with internally (Tom's Hardware). When CFOs start asking outcome-linked questions instead of renewing seat counts, the entire "agentic AI eats 1000x more tokens" demand-side bull case meets balance-sheet reality.
The bottom line on Jevons: yes, total token consumption probably grows. But if it grows at $0.28 instead of $30, the revenue line per inference compresses by 90%+ — and the elasticity has to be greater than 10x just to keep aggregate revenue flat. That is a heroic assumption on a 12-18 month horizon, especially when CFOs are visibly cutting.
Vector 2 connects: Chinese hardware now supports the price floor
Xiaomi's "break even at full capacity" is only credible if the underlying compute stack is cheap enough. This is where Vector 2 — Chinese hardware parity — provides the substrate.
Huawei Atlas 800 / Ascend 910C — 600,000 units in 2026
Per Digitimes / Tom's Hardware reporting:
600,000 Ascend 910C units targeted for 2026 — double 2025 output.
Including other Ascend models, 1.6 million dies expected across China's AI sector this year.
Atlas 800 32B all-in-one delivers approximately 60-70% of NVIDIA H100 inference at ~30% of system cost — a cost-per-performance ratio ~2.0-2.3x superior to NVIDIA on inference.
Primary deployment customers: Alibaba, Tencent, DeepSeek — the same companies pushing API prices toward zero.
Atlas 950 launch in Q4 2026: 8,192 Ascend chips per node, 500,000+ processors per system — direct NVL144 / Blackwell competitor.
Kirin X90: not just servers
Geekbench 6 benchmark on Huawei MateBook Fold:
Kirin X90 multi-core: 11,640 points at 30W.
Apple M2 multi-core: 10,064 points.
Intel i7-13700H: 12,203 points.
Built on SMIC 7nm (N+2) — still a process generation behind TSMC, but functionally at parity for the workloads the buyer actually runs.
Translation: a Chinese laptop chip beats Apple M2 multi-core. The narrative "China is years behind" stops being true at the price point the median enterprise customer cares about. Combine this with the open-weight delivery: Qwen 3.5-8B now runs at 20-30 tokens/sec on a 4-year-old MacBook M2 Air with 16GB (InsiderLLM benchmarks). What used to require a data center in 2024 now runs locally for free. That collapses cloud inference revenue on entire classes of workloads — and it is the demand-side mirror of the SHFE supply-side filing.
CXMT: the DRAM transmission mechanism
ChangXin Memory Technologies (CXMT) — Chinese DRAM:
Global DRAM market share Q4 2025: 7.67% (up from 3.97% Q2 2025) — almost a doubling in 6 months.
IPO cleared review for ~6.5 trillion KRW (~$5B) — fresh capital for HBM3 + DDR5-8000 mass production.
LPDDR6 launch H2 2026 — potentially first globally, ahead of Samsung and Micron.
HBM3 samples already shipped to Chinese AI chip designers (Caixin).
Morgan Stanley's memory supercycle call assumes Chinese DRAM stays at ~5% market share through 2028. Q4 2025 is already at 7.67%. The framework is wrong on its starting condition.
Vectors 3-5: how this lands on US AI-infra equities
Now the connection to US tape. The China-side commoditization (Vectors 1-2) only matters for US equities to the extent the US balance sheets are built on the opposite assumption. They are. Three concrete pressure points.
Vector 3 — PJM $329.17/MW-day is the binding constraint, not the headline number
From the PJM 2026/2027 Base Residual Auction Report:
$329.17/MW-day clearing price — hit the FERC-approved cap.
Without the cap, simulations suggest $500+/MW-day (RTO Insider).
Data centers: 63% of price increase — $9.3 billion in costs to ratepayers.
PJM missed reserve margin target by 0.2pp (309 MW short) — load growing faster than generation.
The number to internalize: $329.17 today is the floor, not the ceiling, of US AI compute power input pricing. The grid cannot connect the queue in time. Sell-side targets that extrapolate hyperscaler capex through 2028 assume this constraint resolves. The physics does not allow it.
Vector 4 — Hyperscaler bond wall: 2026 tracking $230-240B
From BofA / UBS credit strategists:
$28 billion/year average hyperscaler IG bond issuance, 2020-2024.
2025: $121 billion issued by Amazon, Alphabet, Meta, Microsoft, Oracle combined.
2026: $230-240 billion projected — 8x the pre-AI baseline.
2026 hyperscaler capex: $602B projected (MUFG).
Individual deals from the recent vintage: Oracle $18B (Sept), Meta $30B (Oct, largest-ever non-M&A IG deal), Alphabet $17.5B (Nov), Amazon $15B (Nov). Maturities cluster at 8-10 years on the AI-infra issuance.
This is the duration mismatch in plain English: 8-10 year IG bonds funding 18-24 month revenue streams that are commoditizing at -75% to -99% per cycle. When token revenue compresses in line with Xiaomi/DeepSeek price discovery, the operating cash flow servicing those bonds compresses with it. The bonds don't disappear. The credit spreads adjust.
Vector 5 — Trade catalyst calendar
Two hard-dated catalysts anchor the window into Q1 2027:
November 10, 2026 — Expiration of the US-China tariff truce signed in November 2025.
November 27, 2026 — Expiration of China's suspension of gallium, germanium, and antimony export restrictions to the US.
The market is partially pricing the first and largely ignoring the second. Both are likely to re-escalate. They will land on top of a US AI infra complex already pressured by Vectors 1-4. This is the Q1 2027 window the book centers on.
What this means for the trade
We do not call levels or set price targets. We frame which exposures asymmetric to the data, and we publish dated catalysts. Five transmission mechanisms make this thesis tradeable:
Dispersion long, index variance short — single-stock AI-infra dispersion is the cleanest expression of the thesis. The narrative is concentrated in 5-8 names; the consensus assumes they move as a block; the catalyst calendar disagrees. See: Dispersion Trading Through the China AI Thesis.
Hyperscaler IG credit spreads widening — duration mismatch from Section 4. Credit prints before equity.
NVIDIA mid-range margin compression — Huawei Atlas 800 cost-parity is a quarterly headwind, not a future story.
Memory equities re-rating — CXMT 7.67% > 5% breaks Morgan Stanley's supercycle duration call.
Long Chinese AI hardware names selectively — Atlas 800 / CXMT customers have a supply-side moat in inference compute that the West underestimates.
The other side — engaging with the Privorotsky bull case
A research note worth its name presents the strongest version of the opposite trade. Rich Privorotsky, Head of European One Delta trading at Goldman Sachs, has written exactly that in his recent desk note, raising the right questions. We engage with them here.
What Privorotsky has right
Five points where the bull case is genuinely strong:
Token rationalization becomes the narrative itself. Privorotsky: "La rationalisation des dépenses en jetons pourrait devenir tout aussi importante que le discours sur la croissance de l'IA lui-même, en particulier lorsque l'objectif de « 90% de la production pour 10% du coût » devient de plus en plus viable grâce aux alternatives open source." Correct. Open-weight inference is now a real procurement option for the median enterprise workload.
Near-term mechanical momentum can persist. Month-end liquidity dynamics, steady retail flows, and the fact that institutional investors are more skeptical than they should be — together, paradoxically, support the tape. We agree this can extend the trend into late Q2.
The most consensual contrarian view is also a warning. Privorotsky: "L'une des opinions les plus répandues est peut-être tout simplement la conviction que le prix du pétrole ne peut que baisser davantage." Translated: when everyone agrees a trade is over, the trade often isn't. Worth holding in mind.
Long-term, AI does change the world. "À plus long terme, le développement de l'IA pourrait tout à fait se révéler exact et […] tout comme Internet, il pourrait fondamentalement changer le monde." We agree. Our thesis is not that AI fails. It is that current valuations do not survive the path to that future.
Human ingenuity resolves shortages — historically. "L'ingéniosité humaine résout les problèmes… les pénuries de mémoire sont résolues, les pénuries d'énergie finissent par attirer les investissements, les contraintes s'atténuent." True over decades. The relevant question is the timing inside the catalyst window.
Where the framing breaks
Three points where the bull case quietly skips the trajectory between near-term momentum and long-term resolution — which is precisely the window our thesis prices.
The 2000 parallel is a trap, not a comfort. Privorotsky writes: "Cela s'est finalement avéré vrai également en 2000." The Internet did structurally change the world. But the Nasdaq Composite fell 78% from March 2000 to October 2002 while that was being proven true. Cisco — the index AI-infra incumbent of its era, with a working business and real customers — lost 89% peak-to-trough and did not recover the 2000 high for fifteen years. The long-term destination was correct. The equity tape lived a different story between.
Human ingenuity is not instant ingenuity. Power grids take 4-7 years to build out. The PJM 2026/2027 BRA reports the auction missed its reserve margin by 309 MW already — and it cleared at the $329/MW-day FERC cap. Memory capacity scales on 2-3 year fab build-out cycles; CXMT is already deploying (global DRAM share doubled to 7.67% in six months). Both shortages get resolved. Neither resolves inside the next twelve months. Q1 2027 is inside that window.
Jevons elasticity has to do more than reverse the price cut — it has to do it on a calendar. If tokens fall 90-99%, aggregate revenue holds only if consumption elasticity is >10x and arrives this fiscal year. The empirical data points the other way: Jellyfish's Q1 2026 study of 12,000 developers found token consumption scales linearly with output (top 20% spend ≈ 2.1x output of bottom 20% — no efficiency curve). Cheaper tokens may yet drive 10x demand in 2027-2028. They are not driving it on the 2026 P&L.
The window the bull case implicitly leaves open
Privorotsky's two anchors — near-term mechanical momentum and long-term human ingenuity — are compatible with our Q1 2027 thesis once you place them on a calendar:
Near-term mechanical = pre-November 2026. Holds until the tariff truce expiration (10 Nov) and the critical-minerals suspension expiration (27 Nov) re-introduce the trade catalyst.
Long-term human ingenuity = post-Q1 2028. Grid, fab, and energy capacity arrive on multi-year construction horizons.
The gap between is roughly Q4 2026 through Q1 2027. That gap is what the book frames. It is where the catalyst calendar binds and the relief construction has not yet delivered.
Said differently: we agree with the bull case on the destination. We differ on the path. The desk's job is to position for the path, not to argue with the destination.
Verdict
The events of May 26-28, 2026 — DeepSeek -75% permanent, Xiaomi -99%, SHFE token futures filing — are not three separate Chinese narratives. They are three sides of the same prism. Tokens have been commoditized. The financial plumbing is being built to price them like copper. The US balance sheets — hyperscaler capex, IG bond issuance, PJM grid prices — were not designed for that pricing regime. Vector 1 of our thesis is now confirmed by market structure, not just by data points.
Vectors 2 through 5 are setting up into Q1 2027. The window the book identified is now operational, and the catalyst calendar is hard-dated.
Read the full thesis
The complete framework — five convergent shocks, dated catalysts through Q1 2027, the precise transmission mechanism into US equity volatility surface — is laid out in the research book and the free explainer:
→ The China AI Disruption Thesis — full explainer (free, no paywall, no email gate)
→ Book hub on crossvol.com — chapter summaries, datasets, footnoted references
→ The 4-lens framework (the desk method) — how we read AI-infra equity dispersion through dealer gamma, vega, risk reversal and term structure
→ The China AI Disruption Thesis — Kindle on Amazon (B0H11WH3R9, ~320 pages, footnoted)
CrossVol Research publishes non-consensus research on derivatives, volatility, dealer positioning, and macro flow — across equities, FX, futures, futures options, rates, credit, and commodities. Eleven languages. Cross-asset by default. We do not call price targets. We frame which exposures are asymmetric to the data, and we publish dated catalysts.
Primary sources for this article:
— CrossVol Research
CrossVol Team & Djellal Djouad
Further reading


