{
    "success": true,
    "data": {
        "id": 1989215,
        "msgid": "huawei-launches-oceanstor-m900-context-memory-storage-to-accelerate-ai-inference-in-hyperscale-data-centres-1789786508",
        "date": "2026-09-19 09:29:25",
        "title": "Huawei Launches OceanStor M900 Context Memory Storage to Accelerate AI Inference in Hyperscale Data Centres",
        "author": "",
        "source": "ANTARA_ID",
        "tags": "",
        "topic": "Technology",
        "summary": "Huawei has unveiled the OceanStor M900 Context Memory Storage at HUAWEI CONNECT 2026, designed to optimise AI inference for hyperscale data centres. The new solution addresses memory capacity limitations in large-scale AI models by providing a massive shared memory space for SuperPoD architectures.",
        "content": "<p>Shanghai, (ANTARA\/PRNewswire) - Huawei has officially introduced the\nOceanStor M500 Context Memory Storage at the HUAWEI CONNECT 2026 event.\nDavid Wang, Deputy Chairman of the Board and Rotating Chairman of\nHuawei, presented the product, which is designed to meet the AI\ninference requirements of hyperscale data centres. The OceanStor M900\nprovides shared memory space for SuperPoD with capacities up to the\npetabyte scale and performance reaching terabytes per second. The\nintroduction of this solution marks a shift in AI infrastructure from a\ncompute-centric approach towards close collaboration between computing,\nnetworking, and data storage, thereby optimising SuperPoD computing\npower.<\/p>\n<p>By 2026, AI is rapidly transitioning from mere technological\nbreakthroughs to large-scale implementation phases. AI applications are\nalso evolving from chatbots into AI agents capable of completing complex\ntasks independently. The use of AI agents is becoming increasingly\nprevalent across various critical sectors, marking the beginning of the\nera of agentic AI.<\/p>\n<p>As large-scale AI models expand to reach 10 trillion parameters,\nSuperPoD is becoming the preferred choice for AI infrastructure. Current\nmajor AI models support context windows of over one million tokens,\nwhile multi-turn inference and complex tasks are becoming more common.\nSimultaneously, the volume of KV cache data generated during the\ninference process continues to grow. This condition makes on-chip memory\nand DRAM increasingly limited in terms of both capacity and cost\nefficiency. Consequently, the industry is recognising the need for\ntiered storage systems that coordinate on-scale memory, DRAM, and SSD to\nform a shared memory space with massive capacity.<\/p>\n<p>Huawei presents the OceanStor M900 Context Memory Storage to overcome\nmemory capacity limitations during inference with very long contexts and\nmulti-turn interactions. The OceanStor M900 utilises the UnifiedBus\nnetwork to build a tiered global KV cache at the petabyte scale through\na single-hop connection. This approach optimises the potential computing\npower of SuperPoD and accelerates AI inference in hyperscale data\ncentres. The OceanStor M900 Context Memory Storage offers three key\ncapabilities:<\/p>\n<p>Breaking Capacity Limits to Support Large-Scale AI with Massive\nMemory Capacity<\/p>\n<p>With the support of high-speed UnifiedBus interconnection networks,\nKV cache can be aggregated and shared globally through tiered data\nstorage. The SuperPoD KV cache capacity is expanded from on-chip memory\nand DRAM to SSDs, allowing a single cluster to provide a capacity of 64\nPB. The KV cache capacity available on each NPU also increases from\ngigabytes to terabytes. With this capacity, more context can be stored,\nshared, and reused, significantly increasing the KV cache hit ratio.<\/p>\n<p>Enhancing Inference Performance to Optimise Computing Power<\/p>\n<p>The OceanStor M900 is the industry\u2019s first architecture to integrate\nCPU, network controller units, and NAND controller units. The product\nprovides native KV semantics so that SuperPoD NPUs connect directly to\nSSDs via a single hop. This approach eliminates the need for protocol\nconversion and CPU forwarding. As a result, access latency drops from\nmilliseconds to 60 microseconds, a 90% reduction. A single cluster can\ngenerate an aggregate access bandwidth of up to 40 TB\/s, which is 1.5\ntimes higher than similar solutions. In typical AI programming\nscenarios, this architecture doubles the token throughput of the\ninference cluster and halves the time required to generate the first\ntoken. Consequently, computing power can be utilised more\nproductively.<\/p>\n<p>Reducing Token Costs to Drive More Economical Large-Scale AI\nAdoption<\/p>\n<p>The OceanStor M900 utilises the industry\u2019s first KV-based adaptive\nstorage technology. This technology predicts the lifecycle of the KV\ncache based on data values and intelligently distributes it across\nvarious storage media. The technology supports up to 24 drive writes per\nday (DWPD), extending SSD endurance by up to 16 times and maintaining\nstability for three years. By reducing the need for media replacement\nand O&amp;M costs, this solution lowers the long-term costs of\nlarge-scale AI inference infrastructure while accelerating AI\nadoption.<\/p>\n<p>AI is increasingly being used in core production systems across\nvarious industries. Furthermore, AI infrastructure is evolving from\ncompute-centric models towards closer collaboration between computing,\nnetworking, and data storage. Context memory storage will become a vital\ncomponent that continues to enhance the capacity and efficiency of KV\ncache access in hyperscale inference systems.<\/p>",
        "url": "https:\/\/jawawa.id\/newsitem\/huawei-launches-oceanstor-m900-context-memory-storage-to-accelerate-ai-inference-in-hyperscale-data-centres-1789786508",
        "image": ""
    },
    "sponsor": "Okusi Associates",
    "sponsor_url": "https:\/\/okusiassociates.com"
}