OpenAI Reveals Test Results for Custom 'Jalapeño' AI Chip Developed with Broadcom and CelesticaOpenAI has officially disclosed performance benchmarks for Jalapeño, its custom artificial intelligence chip co-designed in partnership with Broadcom and Celestica. Initially unveiled in June without granular technical specifications, new data showcases significant operational efficiency gains in handling large-scale inference workloads.
According to Richard Ho, Vice President of Hardware at OpenAI, the defining architectural advantage of Jalapeño is its ability to break the traditional engineering trade-off between throughput (processing volume per unit) and latency (response speed). Typically, AI chips must sacrifice execution speed to maximize processing capacity, or vice versa.
To evaluate real-world performance, OpenAI utilized the InferenceX benchmark suite across three open-source AI models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Benchmark results demonstrated:
Throughput-per-Watt Efficiency: Jalapeño delivered 1.5x to 1.9x higher throughput per watt compared to reference hardware architectures.
Latency Reduction: Response speeds improved dramatically, achieving 1.7x to 3.6x lower latency while maintaining peak processing volume.
OpenAI plans to deploy Jalapeño chips in limited quantities by the end of this year, with broader enterprise scale-out scheduled throughout 2027. However, management emphasized that Jalapeño is not intended to replace OpenAI’s existing GPU infrastructure, noting that current compute arrangements remain highly effective for primary training and large-scale operations.
Why is balancing latency and throughput so critical to modern AI? During real-time inference, high batch sizes increase system throughput but create queuing delays that degrade the user experience. By optimizing bandwidth, memory access, and processing pipelines specifically for inference algorithms, Jalapeño enables OpenAI to handle thousands of parallel requests concurrently without compromising response speed.
Developing custom application-specific integrated circuits (ASICs) in collaboration with Broadcom allows OpenAI to decentralize its hardware. While flagship GPUs like NVIDIA's still dominate the market for training large-scale models, application-specific inference chips like Jalapeño reduce operating costs for continuous background tasks, API calls, and light agent-based reasoning.
Measuring performance in terms of throughput per watt reflects a significant shift in hyperscale computing. As data centers worldwide face power grid constraints and soaring electricity costs, hardware performance is a key determinant of profitability. Achieving nearly double the throughput per watt enables cloud operations to generate significantly more tokens per megawatt without scaling data center physical space.
Source: OpenAI
OpenAI Reveals Test Results for Custom 'Jalapeño' AI Chip Developed with Broadcom and CelesticaOpenAI has officially disclosed performance benchmarks for Jalapeño, its custom artificial intelligence chip co-designed in partnership with Broadcom and Celestica. Initially unveiled in June without granular technical specifications, new data showcases significant operational efficiency gains in handling large-scale inference workloads.
According to Richard Ho, Vice President of Hardware at OpenAI, the defining architectural advantage of Jalapeño is its ability to break the traditional engineering trade-off between throughput (processing volume per unit) and latency (response speed). Typically, AI chips must sacrifice execution speed to maximize processing capacity, or vice versa.
To evaluate real-world performance, OpenAI utilized the InferenceX benchmark suite across three open-source AI models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Benchmark results demonstrated:
Throughput-per-Watt Efficiency: Jalapeño delivered 1.5x to 1.9x higher throughput per watt compared to reference hardware architectures.
Latency Reduction: Response speeds improved dramatically, achieving 1.7x to 3.6x lower latency while maintaining peak processing volume.
OpenAI plans to deploy Jalapeño chips in limited quantities by the end of this year, with broader enterprise scale-out scheduled throughout 2027. However, management emphasized that Jalapeño is not intended to replace OpenAI’s existing GPU infrastructure, noting that current compute arrangements remain highly effective for primary training and large-scale operations.
Why is balancing latency and throughput so critical to modern AI? During real-time inference, high batch sizes increase system throughput but create queuing delays that degrade the user experience. By optimizing bandwidth, memory access, and processing pipelines specifically for inference algorithms, Jalapeño enables OpenAI to handle thousands of parallel requests concurrently without compromising response speed.
Developing custom application-specific integrated circuits (ASICs) in collaboration with Broadcom allows OpenAI to decentralize its hardware. While flagship GPUs like NVIDIA's still dominate the market for training large-scale models, application-specific inference chips like Jalapeño reduce operating costs for continuous background tasks, API calls, and light agent-based reasoning.
Measuring performance in terms of throughput per watt reflects a significant shift in hyperscale computing. As data centers worldwide face power grid constraints and soaring electricity costs, hardware performance is a key determinant of profitability. Achieving nearly double the throughput per watt enables cloud operations to generate significantly more tokens per megawatt without scaling data center physical space.
Source: OpenAI
Comments
Post a Comment