Big Tech

OpenAI Says Its AI-Designed Chip Gives Faster Answers

A smiling humanoid AI races forward with blurred legs, holding a large bag labelled “OPENAI CHIPS” and a handful of crisps.
An editorial illustration of an AI racing ahead after grabbing a bag of “OpenAI Chips”, reflecting OpenAI’s claim that its AI-designed chip can deliver faster answers.

The first measured results for an OpenAI chip show faster AI responses and more work per watt than the commercial systems used in the company’s comparisons, according to figures published on 25 August 2026. The custom inference processor, called Jalapeño, delivered between 1.7 and 3.6 times lower end-to-end latency across three large public models in the reported tests. These are OpenAI’s results, produced with a public SemiAnalysis benchmark, not an independent verdict that Jalapeño is faster for every model or workload.

Jalapeño is designed for inference, the stage when a trained model processes a prompt and generates an answer. That is different from training, where vast clusters teach a model from data. For users, inference hardware affects the pause before an answer begins, the speed at which tokens appear and the number of simultaneous requests a service can handle without slowing down.

How the OpenAI Chip Was Tested

OpenAI tested Jalapeño with GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T. The workloads ran through InferenceX, a public benchmarking system developed by semiconductor research firm SemiAnalysis. Its methodology measures the trade-off between total throughput and responsiveness rather than selecting one convenient point from a chip’s maximum specification.

Across the three models, OpenAI reports 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems. For highly interactive workloads, where each user needs rapid token delivery, the company says performance was 2.1 to 4.1 times higher.

The Kimi K2.5 result offers one concrete example. OpenAI says Jalapeño achieved about 1.5 times the peak performance per watt and 3.4 times lower end-to-end latency than the relevant commercial comparison. That combination matters because chip benchmarks normally involve a trade-off: serving more total requests can make each individual user’s response slower.

SemiAnalysis says its InferenceX benchmark is open and reproducible, with recipes, data and workflow provenance available for inspection. However, its analysts also said they ran the Jalapeño benchmark in a laboratory with OpenAI engineers. The workload is public, but the hardware is not yet generally available for outside laboratories to buy and retest. The fairest description is independently defined methodology with company-assisted execution.

Why Lower Latency Can Matter More Than a Bigger Score

A faster benchmark is useful only if it changes the service people experience. End-to-end latency includes more than raw arithmetic. It reflects how long requests wait, how quickly the system processes the prompt and how rapidly it generates the completion. Lower latency can make voice interfaces feel more natural, coding agents finish steps sooner and complex multi-call workflows spend less time waiting between actions.

Throughput matters at the operator level. More tokens per unit of power can let a data centre serve additional demand without a proportional increase in electricity. That is commercially important as LiveAIWire’s analysis of Big Tech’s $1.09 trillion in future data-centre leases shows how aggressively the industry is committing to compute capacity. Efficiency does not erase total energy demand if usage continues growing, but it can change the cost of each answer.

OpenAI rates Jalapeño at 700 watts and says the tested workloads sustained power at or below 550 watts. A chip’s power figure is not the energy cost of an entire service. Memory, networking, cooling, servers and idle capacity all contribute. The published result therefore supports a performance-per-watt claim about the tested systems, not a complete calculation of the environmental or financial cost of running ChatGPT.

AI Helped Design Parts of the Chip

The unusual part of Jalapeño’s development is that OpenAI says its own models assisted the engineering process. The company moved from initial design to manufacturing tape-out in nine months, using AI to explore implementations, support verification loops and optimise arithmetic circuits. Tape-out means the design files were completed for fabrication. It does not mean a production deployment was finished nine months after work began.

For selected GPT-OSS compute blocks, OpenAI says machine-generated implementations ran 1.5 to 1.8 times faster than versions produced by human experts. That is evidence about particular circuit blocks, not proof that AI autonomously designed the whole processor or outperformed engineers across every component. People set objectives, reviewed outputs, integrated the architecture and remained responsible for verification.

Broadcom handled silicon implementation and contributes networking and connectivity technology, while Celestica provides board, rack and systems expertise. The partners unveiled Jalapeño in June, when final performance measurements were still pending. The new publication supplies those first detailed figures.

Jalapeño Is Also a Strategic Move Away From One Supplier

OpenAI says it will continue using Nvidia accelerators. A custom chip nevertheless gives the company another source of capacity and more control over how hardware, models, kernels and serving software fit together. LiveAIWire’s reporting on the proposed Nvidia and OpenAI data-centre financing shows how closely the two companies remain connected even as OpenAI builds more of its own stack.

The move also sits inside a geopolitical competition over advanced chips, manufacturing and access. As LiveAIWire’s analysis of the US-China AI chip divide explains, compute supply has become a strategic constraint rather than an ordinary hardware purchase. Owning a design does not remove dependence on fabrication, packaging, memory, networking or power, but it can reduce reliance on one accelerator roadmap.

What the Results Do Not Establish Yet

OpenAI has not published a purchase price, full cost-of-ownership comparison or production reliability record for Jalapeño. The benchmark covers three public models and selected operating points. Results may differ with proprietary models, longer contexts, different precision, changing software or data-centre-scale deployment. A system that leads at one latency target may not lead at another.

Production qualification is still under way. OpenAI says initial deployment in its own compute infrastructure is planned by the end of 2026, with later generations already in development. Until production systems run sustained customer traffic, the figures remain promising engineering results rather than proof of fleet-wide performance.

There is also no direct evidence yet that ordinary ChatGPT users are receiving answers from Jalapeño. The company is describing what the processor can do in tests and where it plans to deploy it. Any claim that the chip has already made the public service faster would go beyond the announcement.

The First Results Put Pressure on the Inference Market

If the measured advantages hold at scale, Jalapeño gives OpenAI a credible way to reduce latency and increase the amount of inference served from limited power. It also demonstrates a feedback loop the company wants to repeat: models help engineers design hardware, and that hardware serves later models more efficiently.

The headline result is therefore real but bounded. OpenAI has shown strong numbers on a public benchmark, and SemiAnalysis participated in the laboratory work. External replication, production economics and long-term reliability are still missing. Jalapeño may give faster answers, but the next decisive test will not be another company chart. It will be whether the chip delivers the same advantage inside a live service used by millions of people.

About the Author

Stuart Kerr is Technology Correspondent at LiveAIWire, covering artificial intelligence, cybersecurity and the social impact of emerging technology. LiveAIWire is an independent, human-led technology publication using AI-assisted research, editorial production and original AI-assisted editorial illustrations under his direction.