Skip to content

OpenAI Unveils In-House AI Chip, Will ChatGPT's Cost Decrease?

Jun 25, 17:42
OpenAI Unveils In-House AI Chip, Will ChatGPT's Cost Decrease?
TL;DR
· OpenAI and Broadcom announced the first self-developed AI inference chip, Jalapeño, aiming for initial deployment by the end of 2026.
· This chip is designed for LLM inference and is not a training GPU replacement. Early performance claims are still based on company testing standards.
· OpenAI has taken the first step to reduce reliance on NVIDIA, but details regarding cost reduction, production scale, and real-world workload performance have not been disclosed.


On June 24, OpenAI and Broadcom unveiled Jalapeño, a self-designed AI accelerator tailored for large language model inference, marking OpenAI's first chip referred to as an "Intelligence Processor." According to the joint announcement, Jalapeño is still in the testing and sampling phase, with the goal of starting initial deployment by the end of 2026 and expanding over the following years. For OpenAI, the focus is not on immediately replacing all NVIDIA GPUs but gradually shifting the growing inference requests in models like ChatGPT, Codex, API, and future agents toward a more suitable software and hardware combination.


Starting with Inference, Targeting the Daily Query Cost


The demand for AI chips is roughly divided into training and inference. Training determines the model's capability threshold, while inference determines whether, after model training, it can answer a massive number of user requests at an acceptable cost.


Jalapeño targets the latter. OpenAI stated that this chip will be used for large language model inference workloads, serving ChatGPT, Codex, API, and future agent products. For the average user, it will not directly change the chat interface but may impact the cost, speed, and scalability of processing requests in the background.


This is also the common direction of AI companies in self-developed chip trend. Google has TPU, and Amazon and Meta are also advancing custom accelerators. They do not necessarily aim to fully replace NVIDIA but to move the most stable and largest-scale internal workloads to chips that are more suitable for their own models and software stack.


OpenAI has long relied on external GPU supply, especially NVIDIA's high-end accelerators. As model invocation increases, relying solely on purchasing general-purpose GPUs brings a dual pressure: cost and supply constraints from external supply chains, and general chips may not necessarily achieve the highest efficiency in specific models and service modes. The direct change with Jalapeño is to shift OpenAI from solely buying computing power to participating in defining that computing power.


Broadcom is responsible for silicon implementation and connectivity technology, with the semiconductor manufacturing partner undisclosed


Jalapeño was designed from scratch by OpenAI, with Broadcom providing the silicon implementation, network, and connectivity technology, and Celestica involved in the board, rack, and system-level implementation. The official announcement did not disclose the wafer manufacturing partner, so it cannot be simply stated that Broadcom is responsible for manufacturing.


Broadcom has accumulated deep expertise in custom ASICs and data center networking, and in recent years has also become a supplier to several AI giants behind their self-developed chips. For OpenAI, choosing to collaborate with Broadcom seems more like entrusting the architectural requirements of the in-house model company to a mature semiconductor and system supply chain to turn them into deployable products.


The performance aspect still needs to leave room for boundaries. Both companies stated that Jalapeño aims to combine the throughput capabilities of today's leading AI accelerators and achieve a significant improvement in performance/watt compared to current state-of-the-art solutions. The official statement also emphasized that engineering samples have been running machine learning workloads in the lab at the target frequency and power consumption, with final performance still being measured, and a more detailed technical report will be released in the coming months.


Third-party benchmarks have not been publicly disclosed yet, nor have key metrics such as exact throughput, latency, power consumption, and the cost reduction per inference been revealed. Some market interpretations may directly compare it to NVIDIA's Blackwell or Google's TPU, but until real deployment data is available, cost reduction and performance leadership can only be considered based on the company's early testing.


9 Months to Tapeout, AI-Assisted Chipmaking Becomes Another Clue


Another detail highlighted by the official announcement about Jalapeño is the rapid time from initial design to tapeout, which took only 9 months. This time frame does not equate to the full production cycle but rather refers to the speed at which the chip design entered the manufacturing preparation stage.


In traditional chip development, the process from architecture design, validation, EDA flow to tapeout typically requires a long cycle. OpenAI stated that this process was accelerated by its models. The company is trying to prove that AI can not only write code and generate content but also enter the hardware engineering flow to help shorten the chip design and validation time.


If this approach can be reused in subsequent products, Jalapeño will not just be a single-generation chip but the starting point of OpenAI's multi-generation custom computing platform. OpenAI and Broadcom previously announced a custom 10GW-class AI accelerator collaboration in October 2025, scheduled to start deployment in the second half of 2026 and be completed by the end of 2029. Jalapeño is the first public product sample of this collaboration.


However, fast design speed does not equate to fast deployment speed. Whether Jalapeño can enter OpenAI's core infrastructure will depend on subsequent mass production, packaging, high-bandwidth memory supply, server system integration, and data center scheduling. Any bottleneck in any of these stages could affect the initial deployment pace by the end of the year.


Reducing Reliance on NVIDIA is a Direction, Not a Completed Replacement


Jalapeño is most easily interpreted by the market as the "OpenAI Challenge to NVIDIA." More accurately, OpenAI is beginning to look for custom paths outside of NVIDIA for some of its inference workloads.


NVIDIA's strength lies not only in the performance of a single GPU, but also in the CUDA ecosystem, training and inference toolchains, system-level interconnect, developer base, and supply scale. Even if OpenAI successfully deploys Jalapeño, it will not immediately break free from NVIDIA. Especially in cutting-edge large model training, general R&D, and diverse workloads, high-end GPUs remain a key infrastructure.


The inference stage is more suitable for custom chips to enter. The request patterns are relatively stable, the model structure is more controllable, and the service targets are clearer. If OpenAI can migrate high-frequency, standardized inference tasks from ChatGPT, Codex, or the API to Jalapeño, it may achieve practical benefits in terms of cost and supply elasticity.


For Broadcom, AI giants developing chips in-house does not mean bypassing chip companies, but rather bringing more custom projects. However, the profit potential of such projects still depends on component costs, delivery scales, and system complexity, especially the price and availability of key supplies such as high-bandwidth memory.


What Jalapeño really needs to answer next is not whether it can match Blackwell or TPU at launch, but how large-scale it can go live by the end of the year, what real workloads it can run, and what level it can bring down the cost of inference. For OpenAI, this "Mexican Pepper" is just the first taste to break free from the GPU bottleneck. Whether it can become the main course will depend on the deployment data.



Recommended

Eight-Year Investment U-Turn: Why Did Ethereum Suddenly Abandon Poseidon?

Aug 16, 10:00
Eight-Year Investment U-Turn: Why Did Ethereum Suddenly Abandon Poseidon?

The Wall Street Journal: How is AI Trading Stealing the Limelight from Cryptocurrency?

Aug 15, 14:00
The Wall Street Journal: How is AI Trading Stealing the Limelight from Cryptocurrency?

Tencent Still Has a Dream

Aug 15, 11:27
Tencent Still Has a Dream

To Catch North Korean Hackers, They Set Up a Fake Project

Aug 15, 10:00
To Catch North Korean Hackers, They Set Up a Fake Project

From Litigation to Settlement: Positive Signal Released by HTX's Negotiation with FCA

Aug 14, 19:32
From Litigation to Settlement: Positive Signal Released by HTX's Negotiation with FCA

11,742 Shipping Addresses Exposed Alongside Trezor Orders

Aug 14, 19:01
11,742 Shipping Addresses Exposed Alongside Trezor Orders