LUISUAN GP7000 is the first to adapt to NVIDIA's G3 layer KV Cache architecture, increasing AI inference efficiency by 17 times

January 15, 2026, Beijing - Following NVIDIA CEO Jen-Hsun Huang's release of the revolutionary "Inference Contextual Memory Storage Platform" at CES 2026, local high-performance storage manufacturer ForinnBase announced today that its flagship product GP7000 series all-flash storage platform has been adapted by NVIDIA and has become the world's first and only domestic solution to support G3-level KV Cache tiered storage, providing key infrastructure support for the construction of next-generation AI factories.

"Strategic Reserve" for AI reasoning

Background: Explosive growth of KV Cache capacity

Jensen Huang pointed out in his CES speech that as the large model context window expands to millions of tokens, the KV Cache capacity grows explosively and linearly, and GPU HBM and rack-level cache alone can no longer meet large-scale concurrency needs.

core needs: As a persistent overflow layer, G3-level storage needs to provide more than 16TB of extended access space for each GPU through the NVMe-oF/RDMA network, while maintaining <50μs end-to-end latency and 200GB/s-level bandwidth.

Architectural challenges: Jen-Hsun Huang emphasized, “G3 is not simply a matter of placing data on the disk, but needs to be natively integrated with the BlueField-4 DPU and Spectrum-X network to achieve millisecond-level cache preheating and intelligent offloading. This requires the storage system to adopt a storage-computing separation architecture to completely decouple metadata operations from the data path.”

Storage and computing separation architecture platform tailored for G3 level

LUISUAN Technology GP7000 series adopts Ethernet Flash Cluster (EBOF) design. A single system is equipped with 24 PCIe 5.0 NVMe U.2 bays and achieves redundancy through dual main control boards. Its core indicators accurately match the needs of the G3 layer:

  • Ultimate performance: A single machine provides 64.8 million IOPS, 288GB/s bandwidth and 20μs-level latency, and its performance is 17 times higher than that of traditional storage servers.
  • Ultra-high energy efficiency: The power consumption of the whole machine is <900W, and the power consumption per GB/s bandwidth is only 3.1W, meeting the 5 times energy efficiency target of the AI ​​factory.
  • Deep integration: Natively supports BlueField-3/4 DPU and Spectrum-X switches, enabling direct GPU connection through the NVMe-oF/RoCEv2/GDS protocol.

Kong Weihai, product director of LUISUAN Technology, revealed: "GP7000 uses a DPU+ASIC+FPGA multi-heterogeneous computing architecture to completely offload KV Cache's index management, data compression and network protocol stack from the hardware, eliminating the CPU bottleneck." Its distributed KV Cache manager can be seamlessly connected with the NVIDIA Dynamo open source project to achieve cross-rack cache consistency.

Measured performance of DGX GB300 in scenarios

In the NVIDIA DGX GB300 SuperPOD test environment, GP7000 showed significant advantages as a G3 storage pool:

  • Throughput: When the KV Cache overflows to the G3 layer, it can still maintain a generation speed of 5 times tokens/s, which meets the performance goals.
  • Delay: Through GPU Direct Storage (GDS) technology, the first token time is only increased by 3-5ms, which is far lower than the 50ms+ loss of traditional solutions.
  • Scalability: A single DGX GB300 node can be configured with 2 GP7000 cabinets, providing 28PB-level cache capacity and supporting 10,000-level concurrent long conversation requests.

Domestic substitution and “virtual GPU” effect

Industry experts believe that this move is a key step for domestic storage to participate in the global AI infrastructure cutting-edge competition. The CTO of an intelligent computing center commented: "GP7000 has passed certification in key industries such as finance and communications, achieving 99.9999% availability under mixed loads, and the failure rate is 75% lower than that of an integrated storage and computing architecture." The person in charge of a national laboratory pointed out: "Under the current technical background, through storage layer optimization, the inference throughput can be increased by more than 30% with the same computing power, which is equivalent to obtaining a 'virtual GPU'."

Deep adaptation from hardware to software

The LUISUAN Technology white paper disclosed that GP7000 has completed extensive ecological adaptation:

  • hardware: NVIDIA DGX H100/H200/GB300, AMD Instinct MI300, Huawei Ascend 910B/C, Muxixiyun C series, etc.
  • software: NVIDIA Dynamo/vLLM/TensorRT-LLM, Huawei UCM, Kubernetes CSI, etc.
  • Domestic database: OceanBase, TiDB, GaussDB, etc.

In large model inference scenarios, GP7000 can allocate independent KV Cache partitions to each inference instance through namespace isolation and intelligent hot and cold tiering technology, and preload high-frequency data to the G2 layer to achieve the optimal balance between cost and efficiency.

Second half of 2026: Large-scale deployment and future evolution

Current progress: GP7000 has been put into mass production in Q3 of 2025, and has received purchase intentions for thousands of nodes from a leading cloud vendor. It is already in POC testing.

future plans: The company is developing the next-generation GP8000 based on PCIe 6.0, with the goal of increasing G3-level bandwidth to 1TB/s.

As Jen-Hsun Huang said, "The storage revolution in AI factories has just begun." When KV Cache changes from a "burden" of the GPU to a flexibly expandable "strategic resource", professional storage like GP7000 is evolving from a supporting role to a core winner in determining the cost and experience of AI services.

About LUISUAN Technology

ForinnBase Technology Co., Ltd. (ForinnBase) was established in 2021 and focuses on the research and development of DPU-driven high-performance storage systems. Its GroundPool series products have served finance, scientific research, intelligent computing centers and other fields, and are the world's first and only domestic solutions to support G3-level KV Cache tiered storage.

LUISUAN Technology Co.,Ltd. ← Return to news list