To put an end to the "long context disaster", LUISUAN releases a new paradigm for AI storage! CES live coverage: Jensen Huang has just set the tone, we have a reciprocal plan

When AI learns to "forget", this is no longer a philosophical issue, but the reality of expensive computing power hitting a storage bottleneck.

At CES 2026 in Las Vegas, NVIDIA founder Huang Jenxun pointedly pointed out: "Context is the new bottleneck, and storage must be re-architected." As large model parameters approach trillions and context length reaches one million Tokens, the massive key-value cache (KV Cache) generated during the inference process is overwhelming expensive GPU memory.

In order to meet the challenge, NVIDIA has launched a new "Inference Context Memory Storage Platform" (Inference Context Memory Storage Platform), the core of which is a four-layer storage architecture built around "HBM→DRAM→Rack SSD→Network SSD", designed to provide a home for massive KV Cache (Key Value Cache).

On the other side of the world, LUISUAN Technology released a solution based on the same layering concept at almost the same time. This is not a coincidence, but the resonance of top engineering minds when facing the same technical problem. The two parties are highly consistent in the top-level design, and in terms of the implementation of the most critical performance core - the third layer (L3), the LUISUAN Solution proposes a more extreme path: through pure hardware protocol offloading, the performance potential of network storage is released to unprecedented heights.

01. New bottleneck: “Capacity disaster” of context memory

The dilemma described by Jen-Hsun Huang in his CES keynote speech is being played out in every line of inference code.

A doctor inputs N examination reports of a patient into the AI ​​model in batches, hoping that it will give a comprehensive judgment. But after the fourth report was input, because the GPU memory was full, the AI ​​had to "forget" the earliest check results, and the final diagnostic conclusion was completely different - this is "long context catastrophic forgetting".

The root cause is that each new word (Token) generated by large model reasoning relies on the memory of all previous contexts (KV Cache). When processing a novel or a complex technical document, the size of the KV Cache can reach tens of GB or even TB, far exceeding the endurance limit of a single or multiple GPU memory. If it cannot provide a "place" for all contexts, the model's "memory" and judgment will decline sharply.

Traditional stand-alone storage solutions cannot solve this systemic problem. To this end, pioneers in the field of AI have unanimously adopted the same response strategy: hierarchical storage, transformed into a pool.

To put an end to the

02. Source of consensus: Why is the four-level storage architecture inevitable?

The "long context catastrophic forgetting" scenario revealed by Jen-Hsun Huang in his speech intuitively reflects the limitations of the existing architecture: when processing ultra-long documents or multiple rounds of conversations, the GPU memory cannot accommodate the entire context memory and is forced to discard early information, resulting in a decrease in the quality of reasoning.

The root cause is that the KV Cache volume generated by large model inference is moving from the GB level to the TB level, far exceeding the carrying limit of expensive graphics memory. To break this bottleneck, a hierarchical storage system must be built that takes into account both capacity and performance.

To put an end to the

L1: GPU HBM (video memory): stores the most active and hottest KV Cache segments.

L2: Node DRAM (memory): serves as the next-level extension buffer of HBM.

L3: Rack-level SSD (Rack SSD): Forms a scalable, shared "context memory" resource pool.

L4: Network/Object Storage: carries massive amounts of cold data and model weights.

This solution uses BlueField-4 DPU as the core storage processor, achieves high-speed, lossless RDMA access through Spectrum-X Ethernet, and is deeply integrated with Dynamo inference scheduling software to achieve up to 5 times higher throughput and energy efficiency.

To put an end to the
To put an end to the

03. Concepts on the same frequency: LUISUAN’s four-level architecture blueprint

Based on the exact same layering concept, LUISUAN Technology describes a four-level architecture with equivalent structures, especially in the core L3 layer, proposing a more focused and extreme hardware optimization idea.

To put an end to the

04. Optimizing the core: When network storage gets “local-level” performance

How to provide a storage device at the other end of the network with an access experience comparable to a local NVMe SSD? This is the ultimate challenge of L3 layer design.

NVIDIA's path is to manage and optimize this data pool through a highly intelligent DPU (data processing unit), which works in depth with its own GPU and scheduler to build a powerful "storage brain."

The LUISUAN GP plan has chosen a more pure and low-level "hardened" path. Its core is a dedicated device that does not have a general-purpose CPU and uses ASIC/FPGA to implement full protocol offloading. Its design goal is single and powerful: to completely eliminate the software protocol stack overhead in network storage access.

To put an end to the

The key to its technology lies in:

Protocol offloading: Completely offload complex NVMe-oF network storage protocol processing from the host CPU to the dedicated chip of the LUISUAN GP device for solidification and execution.

The path is extremely simple: Data uses RDMA (Remote Direct Memory Access) technology to establish a direct channel between the GPU memory and the network SSD to achieve end-to-end zero-copy transmission.

Latency removal: This hardware-level optimization reduces the extra delay of remote access from milliseconds to microseconds, making "network SSD like local SSD" a reality.

It is this focus on overcoming basic bottlenecks that enables LUISUAN GP to maximize the performance of standard commercial SSDs in a more transparent and efficient manner, providing a flexible, independently scalable top-level cache resource pool for AI clusters.

05. Different paths to the same destination: providing multiple choices for the future of AI

Mr. Huang’s insights point out the key direction for the development of AI infrastructure. The LUISUAN Technology solution proves that in this clear direction, by deeply cultivating the underlying hardware offloading technology, it can provide the industry with an excellent choice with ultimate performance and open architecture.

This is not only technological progress, but also ecological enrichment. It means that no matter what technology stack customers are in, they can find solutions that match their needs to cope with the coming AI era with true "long memory". Great ideas lead the way, and diverse realizations jointly pave the way to the future.

LUISUAN Technology Co.,Ltd. ← Return to news list