Solidigm and LUISUAN Technology jointly released the technical white paper "Storage Expansion Solution for SOHO AI Inference"

Content summary:As large language models continue to evolve toward long context, multi-round sessions, RAG, and Agent applications, KV Cache has gradually become an important factor affecting the capacity, cost, and scalability of AI inference systems. For NVIDIA DGX Spark, which is targeted at individual developers, small teams, and enterprise edge deployment scenarios, local storage capacity limitations make long-term preservation and large-scale expansion of KV Cache challenging.

This article proposes a storage expansion solution based on LUISUAN Technology storage nodes and Solidigm NVMe SSD. Remote NVMe storage resources are connected to DGX Spark through the RDMA network to provide scalable storage capacity for KV Cache Offload. This solution adopts an architecture design that separates computing and storage to achieve independent expansion and centralized management of KV Cache capacity while maintaining local reasoning capabilities.

NVIDIA KV Cache multi-level storage architecture

In this layered cache architecture, HBM provides the highest performance, DRAM provides larger capacity, local SSD is used for KV Cache Offload, and network storage further expands the cache space. Test results show that the RDMA storage pool built with LUISUAN Technology storage nodes can provide larger KV Cache capacity while maintaining end-to-end inference performance comparable to local NVMe SSD - the remote JBOF solution has no obvious performance loss in TTFT and system throughput indicators, and performs better than the local storage configuration in some scenarios.

Solution architecture

Storage expansion solution system architecture diagram

The solution is based on an architectural design that separates computing and storage. By introducing the LUISUAN Technology GP Spark 3000 storage node and Solidigm D7-PS1010 NVMe SSD in addition to DGX Spark, a shared storage pool based on the RDMA network is built to provide KV Cache with a capacity basis for independent expansion and centralized management.

LUISUAN Technology GP Spark storage

LUISUAN Technology GP Spark 3000 switching storage platform is a switching storage device for AI cluster scenarios, supporting up to 8 DGX Sparks sharing a unified NVMe-oF storage resource pool. The hardware is equipped with NVIDIA BlueField-3 DPU, equipped with 4 100GbE or 8 25GbE switching ports, 8 E1.S PCIe 5.0 SSDs and 1GbE BMC management interfaces. A single machine can achieve 2.66 million IOPS and 50GB/s throughput, with end-to-end latency less than 20 microseconds and power consumption controlled within 200W. Its core value lies in breaking the "cache island" problem in traditional multi-machine clusters - any cache written by the Prefill node can be directly read by any Decode node with a delay of no more than 20 microseconds. The P99 delay is reduced by 78%, the Cache hit rate is increased by 39%, and the overall TCO is reduced by more than 30%.

Solidigm D7-PS1010 PCIe Gen5 NVMe SSD

This solution uses Solidigm D7-PS1010 enterprise-class NVMe SSD as the underlying storage medium for KV Cache Offload. D7-PS1010 is based on the PCIe Gen5 x4 interface and TLC NAND design. The E1.S specification has a capacity of up to 7.68 TB, a sequential read bandwidth of up to 14.5 GB/s, and a random read of up to 3.3 million IOPS. It provides 1 DWPD enterprise-level endurance and 2.5 Million Hours MTBF, meeting the data integrity and system stability requirements of long-term AI inference platforms.

Remote storage performance test

The test uses fio to verify the storage pool and simulates the actual storage form of KV Cache Chunk with 2 MB Block Size: first sequentially writes 400 GB data to simulate the KV Cache persistence process, and then randomly reads the same data set to simulate the cache recovery and loading process. All tests access remote JBOF storage resources through the RDMA network.

KV Cache persistence performance

Test results show that the remote storage achieved a sustained write bandwidth of 9.94 GiB/s (10.4 GB/s), and it only took about 41 seconds to complete the entire 400 GB data write.

Table 1 KV Cache persistence performance test results

KV Cache loading performance

In the read test, the remote JBOF storage was able to provide a read bandwidth of 12.3 GiB/s (13.2 GB/s), maintaining stable data throughput in high concurrency scenarios. Tests have shown that the RDMA-based storage architecture can effectively reduce the performance impact caused by remote access, allowing remote NVMe SSDs to have access capabilities close to those of local storage devices.

Table 2 KV Cache loading performance test results

KV Cache Offload inference performance verification

Based on the vLLM and LMCache environments (Qwen3-1.7B model), KV Cache Tester was used to conduct verification in two configurations: local SSD and remote JBOF, focusing on comparing TTFT (Time to First Token) and system throughput.

Single prompt test

Single Prompt test instructions

Tests cover five context lengths from 2,000 to 32,000 tokens. The results show that the first request TTFT difference is only +1.1% to +2.3%; in cache hit requests, the TTFT of the remote JBOF is no higher than the local SSD, and the difference ranges from −0.8% to −13.1%.

Single Prompt test TTFT comparison
Table 3 Comparison of TTFT between two storage configurations

Cache Hit Rate Test

The context length is fixed at 32,000 tokens and the cache hit rate is adjusted in steps of 10%. Under the condition that the model, context length and other inference parameters are consistent, there is no obvious difference in the performance indicators of local SSD and remote RDMA JBOF.

Local SSD test results

Local SSD cache hit rate test results

RDMA JBOF test results

RDMA remote JBOF cache hit rate test results

Summary and Outlook

This article verifies the feasibility of remote storage based on RDMA network in the AI ​​KV Cache Offload scenario: compared with local high-performance SSD, the performance overhead introduced by remote storage is smaller, and the overall end-to-end inference performance is basically at the same level. As a storage expansion solution for edge and SOHO AI scenarios, this architecture provides a capacity basis for independent expansion and centralized management of KV Cache while maintaining the compact deployment form and local reasoning capabilities of DGX Spark.

In the future, as long contexts, multi-round sessions, RAGs, and multi-agent collaboration workloads continue to grow, shared remote storage can further support cross-node prefix cache, session cache, and agent memory reuse, providing a new implementation path for small teams and edge deployments to build more scalable AI inference infrastructure.

Reference

1. Solidigm D7-PS1010 E1.S Product Brief

2. KV Cache Tester (stress testing tool)

LUISUAN Technology Co.,Ltd. ← Return to news list