📡 Breaking news
0/0
Analyzing latest trends...
AI Text-to-Speech.

OpenAI Unveils Jalapeño AI Chip Benchmarks High Throughput with Ultra-Low Latency.

OpenAI Unveils Jalapeño AI Chip Benchmarks High Throughput with Ultra-Low Latency.
OpenAI Reveals Test Results for Custom 'Jalapeño' AI Chip Developed with Broadcom and Celestica

OpenAI has officially disclosed performance benchmarks for Jalapeño, its custom artificial intelligence chip co-designed in partnership with Broadcom and Celestica. Initially unveiled in June without granular technical specifications, new data showcases significant operational efficiency gains in handling large-scale inference workloads.

According to Richard Ho, Vice President of Hardware at OpenAI, the defining architectural advantage of Jalapeño is its ability to break the traditional engineering trade-off between throughput (processing volume per unit) and latency (response speed). Typically, AI chips must sacrifice execution speed to maximize processing capacity, or vice versa.

To evaluate real-world performance, OpenAI utilized the InferenceX benchmark suite across three open-source AI models: GPT-OSS 120B, DeepSeek R1, and Kimi K2.5 1T. Benchmark results demonstrated:

  • Throughput-per-Watt Efficiency: Jalapeño delivered 1.5x to 1.9x higher throughput per watt compared to reference hardware architectures.

  • Latency Reduction: Response speeds improved dramatically, achieving 1.7x to 3.6x lower latency while maintaining peak processing volume.

OpenAI plans to deploy Jalapeño chips in limited quantities by the end of this year, with broader enterprise scale-out scheduled throughout 2027. However, management emphasized that Jalapeño is not intended to replace OpenAI’s existing GPU infrastructure, noting that current compute arrangements remain highly effective for primary training and large-scale operations.

Why is balancing latency and throughput so critical to modern AI? During real-time inference, high batch sizes increase system throughput but create queuing delays that degrade the user experience. By optimizing bandwidth, memory access, and processing pipelines specifically for inference algorithms, Jalapeño enables OpenAI to handle thousands of parallel requests concurrently without compromising response speed.

Developing custom application-specific integrated circuits (ASICs) in collaboration with Broadcom allows OpenAI to decentralize its hardware. While flagship GPUs like NVIDIA's still dominate the market for training large-scale models, application-specific inference chips like Jalapeño reduce operating costs for continuous background tasks, API calls, and light agent-based reasoning.

Measuring performance in terms of throughput per watt reflects a significant shift in hyperscale computing. As data centers worldwide face power grid constraints and soaring electricity costs, hardware performance is a key determinant of profitability. Achieving nearly double the throughput per watt enables cloud operations to generate significantly more tokens per megawatt without scaling data center physical space.

 

Source: OpenAI 

💬 AI Content Assistant

Ask me anything about this article. No data is stored for your question.

Comments

Popular posts from this blog

Databricks Secures $5B Funding at $190B Valuation as Annual Revenue Reaches $7B.

Singapore Polytechnic and Ministry of Manpower Launch CASTLE Lab to Protect 50 SMEs.

OpenAI Acquires InstantDB Team to Build Real-Time Infrastructure for AI Agents.

Meta Emerges as One of Azure Biggest AI Clients Alongside OpenAI.

India Orders Google to Takedown 57 Firebase Accounts Tied to $2.4B Banking Fraud.

NVIDIA to Raise AI Server Prices by 15%+ as High-Bandwidth Memory Costs Surge.

Double the Commits, Double the Pressure Inside GitHub 7-Hour Outage and Azure Pivot.