OpenAI says Jalapeño delivers up to 1.9x more AI work per watt
Written by Rebecca Uffindell 1 minute ago

OpenAI has published the first performance results for Jalapeño, its first custom inference chip, reporting between 1.5 and 1.9 times more AI work per watt at peak throughput and between 1.7 and 3.6 times lower end-to-end latency than the comparison systems it tested.
The company tested Jalapeño across three public models — GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T — using InferenceX, a public benchmark from SemiAnalysis. OpenAI said the results showed a stronger combination of throughput, power efficiency and latency across the tested operating range.
For data centre infrastructure, the power figures are particularly notable. Jalapeño has a published chip power rating of 700W, although OpenAI said measured sustained power remained at or below 550W on the workloads tested.
OpenAI Puts Performance Per Watt in Focus
OpenAI said it evaluated Jalapeño at a matched user experience, measuring how much useful AI work each system could complete per unit of power while still meeting latency requirements.
Across the three public models, Jalapeño delivered between 1.5 and 1.9 times more work per watt at peak throughput compared with the systems against which it was tested.
The strongest result came with GPT-OSS 120B, where OpenAI reported approximately 1.9 times higher peak mixed tokens per second per kilowatt. On DeepSeek R1, the improvement was around 1.7 times, while Kimi K2.5 recorded approximately 1.5 times higher peak performance per watt.
The figures are OpenAI’s own test results, but they give a clearer picture of how the company is approaching inference efficiency.
Lower Latency Came With the Efficiency Gains
The results were not limited to power efficiency. OpenAI reported between 1.7 and 3.6 times lower end-to-end latency across the three public models.
DeepSeek R1 produced the largest reported latency improvement, with Jalapeño recording 1.65 seconds compared with 5.99 seconds for the comparison system. For Kimi K2.5, OpenAI reported 1.56 seconds against 5.31 seconds, while GPT-OSS 120B recorded 1.03 seconds compared with 1.80 seconds.
That combination is significant because infrastructure optimised for throughput often involves trade-offs in response speed.
OpenAI said Jalapeño was designed to deliver higher throughput and lower latency within the same architecture, particularly for interactive and agentic workloads where delays can accumulate across multiple sequential steps.
The Chip Was Designed as Part of a Wider System
OpenAI said the results came from designing the chip, memory, networking, software, and rack-scale system together around language-model inference.
Different stages of inference place different demands on infrastructure. Processing the initial prompt, known as prefill, is compute-intensive, while generating the response token by token is more constrained by memory bandwidth.
Moving data between chips and other resources can introduce additional delays. OpenAI said Jalapeño was designed to reduce that movement and keep model state local where possible, while activating the appropriate combination of compute, memory, and networking for different stages of inference.
This makes the chip part of a broader system-level approach to optimisation rather than a standalone accelerator test.
Jalapeño is the First Step in a Broader Roadmap
OpenAI plans to begin deploying Jalapeño within its own compute infrastructure by the end of 2026.
The company described the chip as the first generation of a multigenerational platform, with Gen 2 already deep in development and Gen 3 taking shape. Production qualification, software development, and further performance validation are continuing ahead of deployment.
That does not mean OpenAI is moving away from third-party accelerators.
The company said meeting AI demand would require compute from multiple sources and that it would continue to deploy accelerators from NVIDIA and other partners for both training and inference.
Jalapeño therefore adds custom silicon to OpenAI’s infrastructure strategy rather than replacing its existing accelerator suppliers.
Power Efficiency is Becoming Part of the Scaling Equation
The first Jalapeño results show how OpenAI is thinking about the infrastructure demands created by more capable models and agentic workloads.
Instead of measuring progress only through accelerator performance, the company is placing greater emphasis on the amount of useful AI work produced from the power and hardware available.
OpenAI said producing more useful work from the same resources could allow it to serve greater demand while lowering the cost of delivering successful results.
For data centres, that makes the reported efficiency gains relevant. AI infrastructure growth remains closely tied to the availability of power, and increasing the amount of inference work delivered per watt gives operators another route to expanding capacity alongside building more physical infrastructure.
Jalapeño’s first results remain vendor-reported benchmarks ahead of deployment. But with OpenAI preparing to introduce the chip into its infrastructure this year and already developing subsequent generations, custom silicon and system-level power efficiency are becoming a more explicit part of its strategy for scaling AI inference.
Written by Rebecca Uffindell 1 minute ago
Tags:
AI infrastructure Jalapeño nvidia OpenAI power efficiencyMost Viewed News
Wed 12 Aug 2026
South Yorkshire Launches Women in Tech Taskforce
