Moonshot AI Optimizes Kimi K2.6 on AMD MI355X Chips, Achieving 3.2x Faster TTFT via Custom 'UMBP' Storage SystemFollowing the global attention captured by its flagship Kimi K3 model which demonstrated reasoning capabilities outperforming Anthropic’s Claude 3 Opus across several benchmarks Moonshot AI has detailed its deep hardware-level optimization work on the Kimi K2.6 model running on AMD Instinct MI355X accelerators. By fully exploiting AMD's architectural features, the engineering team achieved dramatic improvements in API throughput and latency reduction.
At the core of Moonshot’s technical breakthrough is a custom-engineered storage architecture named Unified Memory & Bandwidth Pool (UMBP). Built specifically for Key-Value (KV) cache management during prompt processing, UMBP addresses a critical bottleneck in traditional LLM serving systems, which often lack optimizations tailored for agentic, multi-step workflows.
UMBP dynamically calculates available system bandwidth, applies intelligent cache eviction policies optimized for agent interactions, and leverages hardware-level capabilities on AMD chips such as GPU-Direct Storage (GDS) and GPU-initiated NVMe protocols to maximize raw hardware data transfer throughput.
By integrating UMBP directly into the SGLang inference engine, Moonshot AI achieved remarkable performance gains:
Time to First Token (TTFT): Reduced by 3.2x, drastically lowering response latency for end users.
Token Generation Rate: Increased by 7.7%.
First-Call Cache Hit Rate: Improved by 9.5% for users returning to long-paused or cold sessions.
Moonshot AI confirmed that its engineering pipeline will next focus on applying these UMBP optimizations to Kimi K3, while expanding inference engine support to include vLLM and AMD's proprietary ATOM execution framework.
The Importance of KV Cache for Modern AI Agents: AI no longer engages in short question-and-answer conversations. It must manage massively long context windows and multi-turn agent loops. Every time a new command is sent, repeatedly processing the same context causes server latency. Moonshot AI's creation of UMBP to directly pull KV cache from NVMe to the GPU via GPU-Direct Storage significantly reduces the Time to First Message (TTFT) wait time.
The choice of SGLang (Structured Generation Language), an inferential engine highly effective at managing RadixTree Cache and Complex Prompt Workflows, demonstrates a trend among leading AI vendors moving away from off-the-shelf systems to developing their own custom memory layers. Future expansion to vLLM will allow this system to better support large clusters.
The growth of the AMD ROCm ecosystem and the MI350 Series chips, which previously led the market to believe that high-level model optimization was only possible on NVIDIA chips (CUDA Ecosystem), demonstrates that the performance of Moonshot AI on the MI355X is crucial evidence that AMD chips possess the bandwidth and hardware architecture capable of efficiently supporting top-tier LLMs, provided the software is properly optimized.
Source: AMD
Moonshot AI Optimizes Kimi K2.6 on AMD MI355X Chips, Achieving 3.2x Faster TTFT via Custom 'UMBP' Storage SystemFollowing the global attention captured by its flagship Kimi K3 model which demonstrated reasoning capabilities outperforming Anthropic’s Claude 3 Opus across several benchmarks Moonshot AI has detailed its deep hardware-level optimization work on the Kimi K2.6 model running on AMD Instinct MI355X accelerators. By fully exploiting AMD's architectural features, the engineering team achieved dramatic improvements in API throughput and latency reduction.
At the core of Moonshot’s technical breakthrough is a custom-engineered storage architecture named Unified Memory & Bandwidth Pool (UMBP). Built specifically for Key-Value (KV) cache management during prompt processing, UMBP addresses a critical bottleneck in traditional LLM serving systems, which often lack optimizations tailored for agentic, multi-step workflows.
UMBP dynamically calculates available system bandwidth, applies intelligent cache eviction policies optimized for agent interactions, and leverages hardware-level capabilities on AMD chips such as GPU-Direct Storage (GDS) and GPU-initiated NVMe protocols to maximize raw hardware data transfer throughput.
By integrating UMBP directly into the SGLang inference engine, Moonshot AI achieved remarkable performance gains:
Time to First Token (TTFT): Reduced by 3.2x, drastically lowering response latency for end users.
Token Generation Rate: Increased by 7.7%.
First-Call Cache Hit Rate: Improved by 9.5% for users returning to long-paused or cold sessions.
Moonshot AI confirmed that its engineering pipeline will next focus on applying these UMBP optimizations to Kimi K3, while expanding inference engine support to include vLLM and AMD's proprietary ATOM execution framework.
The Importance of KV Cache for Modern AI Agents: AI no longer engages in short question-and-answer conversations. It must manage massively long context windows and multi-turn agent loops. Every time a new command is sent, repeatedly processing the same context causes server latency. Moonshot AI's creation of UMBP to directly pull KV cache from NVMe to the GPU via GPU-Direct Storage significantly reduces the Time to First Message (TTFT) wait time.
The choice of SGLang (Structured Generation Language), an inferential engine highly effective at managing RadixTree Cache and Complex Prompt Workflows, demonstrates a trend among leading AI vendors moving away from off-the-shelf systems to developing their own custom memory layers. Future expansion to vLLM will allow this system to better support large clusters.
The growth of the AMD ROCm ecosystem and the MI350 Series chips, which previously led the market to believe that high-level model optimization was only possible on NVIDIA chips (CUDA Ecosystem), demonstrates that the performance of Moonshot AI on the MI355X is crucial evidence that AMD chips possess the bandwidth and hardware architecture capable of efficiently supporting top-tier LLMs, provided the software is properly optimized.
Source: AMD
Comments
Post a Comment