Indonesian Political, Business & Finance News

Huawei Launches OceanStor M900 Context Memory Storage to Accelerate AI Inference in Hyperscale Data Centres

| Source: ANTARA_ID Translated from Indonesian | Technology
Huawei Launches OceanStor M900 Context Memory Storage to Accelerate AI Inference in Hyperscale Data Centres
Image: ANTARA_ID

Shanghai, (ANTARA/PRNewswire) - Huawei has officially introduced the OceanStor M500 Context Memory Storage at the HUAWEI CONNECT 2026 event. David Wang, Deputy Chairman of the Board and Rotating Chairman of Huawei, presented the product, which is designed to meet the AI inference requirements of hyperscale data centres. The OceanStor M900 provides shared memory space for SuperPoD with capacities up to the petabyte scale and performance reaching terabytes per second. The introduction of this solution marks a shift in AI infrastructure from a compute-centric approach towards close collaboration between computing, networking, and data storage, thereby optimising SuperPoD computing power.

By 2026, AI is rapidly transitioning from mere technological breakthroughs to large-scale implementation phases. AI applications are also evolving from chatbots into AI agents capable of completing complex tasks independently. The use of AI agents is becoming increasingly prevalent across various critical sectors, marking the beginning of the era of agentic AI.

As large-scale AI models expand to reach 10 trillion parameters, SuperPoD is becoming the preferred choice for AI infrastructure. Current major AI models support context windows of over one million tokens, while multi-turn inference and complex tasks are becoming more common. Simultaneously, the volume of KV cache data generated during the inference process continues to grow. This condition makes on-chip memory and DRAM increasingly limited in terms of both capacity and cost efficiency. Consequently, the industry is recognising the need for tiered storage systems that coordinate on-scale memory, DRAM, and SSD to form a shared memory space with massive capacity.

Huawei presents the OceanStor M900 Context Memory Storage to overcome memory capacity limitations during inference with very long contexts and multi-turn interactions. The OceanStor M900 utilises the UnifiedBus network to build a tiered global KV cache at the petabyte scale through a single-hop connection. This approach optimises the potential computing power of SuperPoD and accelerates AI inference in hyperscale data centres. The OceanStor M900 Context Memory Storage offers three key capabilities:

Breaking Capacity Limits to Support Large-Scale AI with Massive Memory Capacity

With the support of high-speed UnifiedBus interconnection networks, KV cache can be aggregated and shared globally through tiered data storage. The SuperPoD KV cache capacity is expanded from on-chip memory and DRAM to SSDs, allowing a single cluster to provide a capacity of 64 PB. The KV cache capacity available on each NPU also increases from gigabytes to terabytes. With this capacity, more context can be stored, shared, and reused, significantly increasing the KV cache hit ratio.

Enhancing Inference Performance to Optimise Computing Power

The OceanStor M900 is the industry’s first architecture to integrate CPU, network controller units, and NAND controller units. The product provides native KV semantics so that SuperPoD NPUs connect directly to SSDs via a single hop. This approach eliminates the need for protocol conversion and CPU forwarding. As a result, access latency drops from milliseconds to 60 microseconds, a 90% reduction. A single cluster can generate an aggregate access bandwidth of up to 40 TB/s, which is 1.5 times higher than similar solutions. In typical AI programming scenarios, this architecture doubles the token throughput of the inference cluster and halves the time required to generate the first token. Consequently, computing power can be utilised more productively.

Reducing Token Costs to Drive More Economical Large-Scale AI Adoption

The OceanStor M900 utilises the industry’s first KV-based adaptive storage technology. This technology predicts the lifecycle of the KV cache based on data values and intelligently distributes it across various storage media. The technology supports up to 24 drive writes per day (DWPD), extending SSD endurance by up to 16 times and maintaining stability for three years. By reducing the need for media replacement and O&M costs, this solution lowers the long-term costs of large-scale AI inference infrastructure while accelerating AI adoption.

AI is increasingly being used in core production systems across various industries. Furthermore, AI infrastructure is evolving from compute-centric models towards closer collaboration between computing, networking, and data storage. Context memory storage will become a vital component that continues to enhance the capacity and efficiency of KV cache access in hyperscale inference systems.

View JSON | Print