$30 billion. Fifty percent. Nine months. In February 2026, Nvidia invested $30B directly into OpenAI. In June, OpenAI unveiled a chip its design partner says runs inference for about 50% less than a standard AI GPU. Same two companies. This is the paradox at the heart of the new OpenAI inference chip, codenamed Jalapeño — and it is not the decoupling the headlines promised. The accurate read is narrower and more interesting: OpenAI is buying partial control of its inference unit costs while keeping training firmly on Nvidia.
Key Takeaways
- The split is the story: training stays on Nvidia GPUs, only inference moves in-house — and a $30B Nvidia investment entangles the two further.
- “50% cheaper” is a vendor figure from Broadcom’s CEO, not an OpenAI benchmark; OpenAI’s only official claim is “performance per watt substantially better.”
- Whoever wins the inference-silicon war, the HBM is Korean — Samsung and SK Hynix are projected at ~88% of ASIC-bound HBM.
This is the next stop in a bottleneck that keeps moving. In The AI Memory Tax, the constraint sat in HBM wafer allocation. Now it has shifted one notch further — to who controls the silicon that runs inference. The governing question is no longer “who has the most GPUs.” It is “who controls the cost of intelligence.”
FIG. 01 — WHOSE CHIP RUNS WHAT
Training vs inference: who owns each layer
Training (stays on Nvidia)
Inference (Jalapeño ASIC)
Nvidia GPU (general-purpose)
Custom inference ASIC
Flexible, high precision
Fixed-function, specialized
Still in flux — ASIC unfit
One known workload
One-time training cost
Billed 24/7 for product life
Locked to Nvidia + $30B stake
Partial cost control, hedged
SOURCE: OpenAI, TechCrunch, TheNextWeb, Tom's Hardware
Why this OpenAI inference chip is a hedge, not a decoupling
Start with the cleanest fact: Jalapeño is inference-only. Training — where the model architecture is still in flux and ASICs are a poor fit — stays on Nvidia GPUs. OpenAI keeps multi-year commitments with Nvidia, AMD’s MI450, Cerebras, and Oracle. The February $30B Nvidia stake plus a Vera Rubin training-and-inference capacity deal means the two companies are more entangled, not less.
So “decoupling” overstates it. TechCrunch, TheNextWeb, and others land on the same phrase: diversification at the edges, not a divorce. OpenAI is trying to escape being a captive customer, not to replace Nvidia.
Nvidia is defending the same ground
Nvidia is not standing still on inference. Its Rubin CPX — a GPU purpose-built for massive-context inference — is slated to ramp in Q3–Q4 2026, defending the exact market Jalapeño targets. Nvidia’s CUDA software moat remains intact. New Street analysts model Nvidia’s 90%+ inference share falling to roughly 20–30% by 2028 — pressure, not displacement. “Nvidia’s pricing power on notice” is the precise framing, not “Nvidia replaced.”

The inference economics behind the OpenAI inference chip
Here is why a few percentage points matter existentially. In 2026, inference overtook training to become the dominant line item in AI compute spending — roughly 55–80%+ of the total. Training is a one-time cost; inference is billed 24/7 for the life of a product. The inference-chip market is around $20.5B in 2026.
For OpenAI specifically, inference cost is climbing from $8.4B in 2025 to an estimated $14.1B in 2026, against an adjusted gross margin around 33% (down from ~40% in 2025). Profitability is not expected before 2030. When inference is the dominant cost and margins are thin, controlling unit price is controlling survival.
Where the savings actually come from
That is what Jalapeño aims at: Nvidia’s ~75% margin, subtracted from OpenAI’s own cost base. One caveat the headlines flatten — Broadcom’s gross margin (78.6%, segment basis) is higher than Nvidia’s (~73.5%), so “cheaper because Broadcom takes less” is wrong. The real sources are ASIC efficiency plus removing Nvidia’s markup. This is the same unit-economics logic behind why model pricing collapsed — see the GLM versus GPT-5.5 reliability gap on how token cost reshapes the business case.
FIG. 02 — THE BOTTLENECK KEEPS MOVING
From memory allocation to inference silicon
PREDECESSOR
HBM wafer allocation bottleneck
The AI Memory Tax: the constraint sat in who could get HBM wafer capacity.
2026
Inference becomes the dominant cost
Inference overtakes training at 55-80%+ of AI compute spend, billed continuously.
JUN 2026
Inference silicon goes in-house
OpenAI and Broadcom unveil Jalapeño to pull Nvidia's margin out of OpenAI's cost base.
STILL EXTERNAL
HBM, TSMC and CoWoS stay outside
Memory, 3nm manufacturing and sold-out CoWoS packaging remain third-party dependencies.
SOURCE: Spheron, Astute Group, OpenAI, TheByteDive analysis
Separating the claim from what is verified
The single most important discipline here: do not let one number become “Nvidia killer.” The “~50% cheaper” figure comes from a single source — Broadcom CEO Hock Tan, in a Bloomberg interview, citing early lab tests. OpenAI’s official statement claims only “performance per watt substantially better than current state-of-the-art.” No percentage. No published TFLOPS, memory, power, node, or MLPerf benchmark. The detailed technical report is promised “in the coming months.”
The nine-month tape-out is real but hard to reproduce: typical high-performance ASIC cycles run 18–36 months. Jalapeño’s speed came from OpenAI using its own models to accelerate design, Broadcom reusing mature IP, and specializing for a single known workload. The ASIC over-specialization risk is real too — if transformer inference patterns shift sharply, a fixed-function chip can strand.
FIG. 03 — WHY UNIT COST IS EXISTENTIAL
The inference economics in three figures
$14.1B
OpenAI estimated 2026 inference cost from $8.4B
~33%
OpenAI adjusted gross margin
55-80%+
inference share of AI compute spend
~75%
Nvidia margin Jalapeño targets
SOURCE: Sacra, Spheron, hashrateindex
On timing: the first deployment in late 2026 is a small-scale prototype (per Hock Tan to CNBC), with volume in 2027–28 and 10 gigawatts targeted by 2029. “Already deployed at scale” is wrong. Reported node (TSMC 3nm) and HBM stack count (~8, estimated from photos) are press-and-photo attributions, not OpenAI specs.
The Korea landing: whoever designs the chip, the memory is Korean
For semiconductor watchers, the OpenAI inference chip story has a clean punchline: the inference-silicon war is a tailwind for Korea regardless of who wins the design race. Jalapeño’s HBM stacks are supplied by Samsung and SK Hynix. Samsung’s HBM4 reportedly exceeded Broadcom’s performance bar and carries an OpenAI commitment for up to 800 million Gb of 12-high HBM4 in H2 2026. JPMorgan projects the two Korean firms at a combined ~88% share of ASIC-bound HBM, with ASIC-driven HBM demand up 82% year over year (33% of the 2026 HBM market).
The honest limit on Korea’s position
But the design layer is not Korea’s. ASIC co-design is a US duopoly — Broadcom plus Marvell hold 80%+ of hyperscaler custom silicon, the “tollgate” that collects margin no matter whose chip wins. Manufacturing is TSMC’s, and CoWoS packaging is sold out for 2026 (Nvidia pre-booked ~60%) — a structural cap on Jalapeño’s ramp. Korea’s own answer to the design gap is in progress: Rebellions and FuriosaAI are running the same inference-ASIC playbook.
The timing hook for investors
SK Hynix lists on Nasdaq via ADR on July 10, raising about $29.4B — effectively selling this inference-silicon demand to US investors. The tracking metric for Korean chip watchers: ASIC-bound HBM share and HBM4 design-win lead, not the chip-design scoreboard.

Bottom Line
The OpenAI inference chip is not OpenAI leaving Nvidia. It is OpenAI buying a measure of control over its inference unit costs — trading one dependency (Nvidia) for four (Broadcom, TSMC, Celestica, Korean memory) to claw back bargaining power. The capability is a record; the relationship is more tangled than ever.
Career Takeaway. Inference unit cost is becoming the decisive variable in AI business margins. When adopting AI internally, track two things together: the token-price trend (down ~1,000x in three years) and vendor lock-in (GPU versus ASIC). The question worth asking is not “which model is best,” but “who controls the cost of running it.”
Frequently Asked Questions (FAQ)
Q. Is the OpenAI inference chip really decoupling from Nvidia? A. No. Jalapeño is inference-only, while training stays on Nvidia GPUs, and a $30B Nvidia investment ties the two companies closer together. The accurate description is diversification at the edges, not a divorce.
Q. Is the “50% cheaper” claim verified? A. Not independently. The figure comes from Broadcom CEO Hock Tan in a Bloomberg interview citing early lab tests. OpenAI’s only official claim is “performance per watt substantially better,” with no benchmark, node, or power data published yet.
Q. When does Jalapeño actually ship at scale? A. The first deployment in late 2026 is a small-scale prototype, with volume in 2027–28 and 10 gigawatts targeted by 2029. Claims of large-scale deployment today are inaccurate.
Q. How does this affect Korean chipmakers? A. It is a tailwind regardless of who wins the design race. Samsung and SK Hynix are projected at ~88% of ASIC-bound HBM, so whichever inference chip leads, the memory is Korean. The design layer, however, stays with the US Broadcom–Marvell duopoly.
Q. Why does inference cost matter so much? A. Inference is now 55–80%+ of AI compute spending and is billed continuously, unlike one-time training. For OpenAI, inference cost rises to about $14.1B in 2026 against a ~33% margin, so controlling unit price is existential.
FIG. 04 — CLAIM VS VERIFIED VS INTERPRETATION
Separating the headline from the fact
| Claim | Verified | Interpretation |
|---|---|---|
| Decoupling from Nvidia | Inference-only; training stays on Nvidia + $30B stake | Diversification at the edges, not a divorce |
| ~50% cheaper inference | Single source — Broadcom CEO, Bloomberg, early lab | Vendor figure; OpenAI official = 'performance per watt' |
| Fully independent design | Depends on Broadcom, TSMC, Celestica, Korean HBM | Swaps one dependency for four — bargaining power |
SOURCE: OpenAI, Bloomberg via TechTimes, TheNextWeb
References
- OpenAI and Broadcom unveil LLM-optimized inference chip — OpenAI
- OpenAI and Broadcom announce strategic collaboration (10GW) — OpenAI
- OpenAI and Broadcom Unveil LLM-Optimized Intelligence Processor — Broadcom
- Broadcom and OpenAI unveil custom-built Jalapeño inference processor — Tom’s Hardware
- OpenAI unveils first custom AI inference chip Jalapeño with Broadcom — VentureBeat
- OpenAI unveils its first custom chip, built by Broadcom — TechCrunch
- Jalapeño chip: a way out from Nvidia (diversification not divorce) — TheNextWeb
- OpenAI’s first custom AI chip targets 50% cheaper inference (Hock Tan/Bloomberg) — TechTimes
- AI Inference Cost Economics in 2026 — Spheron
- Broadcom’s Custom AI Empire (6th customer) — kavout
- AI Chip Design Partner Duopoly: Broadcom & Marvell — hashrateindex
- Samsung, SK Hynix ride OpenAI/HBM supply boon — KED Global
- Nvidia secures 60% of CoWoS capacity — Astute Group
- NVIDIA unveils Rubin CPX inference GPU — NVIDIA
- OpenAI revenue, valuation & funding — Sacra
Disclaimer. This analysis is for informational purposes only and is not investment advice. Figures attributed to vendors (including the “~50% cheaper” claim) are unverified company statements, not independently benchmarked results. Do your own research before making any decisions.
