This AI SSD tech makes 8 RTX 5090s perform like 46 GPUs in inference

https://cdn.mos.cms.futurecdn.net/XvSQGhRtTGEfJ9TvnASgdZ-1920-80.png
  • GenStorAIGE AI90 shifts AI memory beyond traditional GPU HBM limitations using SSDs
  • PT200Z SSD supports constant cache updates during demanding inference workloads efficiently
  • Eight RTX 5090 GPUs gain dramatically larger effective inference memory capacity

GenStorAIGE has introduced its AI90 inference acceleration platform at WAIC 2026, taking a storage-centric approach to expanding effective AI memory capacity.

Rather than depending solely on GPU high-bandwidth memory, the platform incorporates PCIe Gen5 solid-state drives directly into the memory hierarchy itself.

This allows portions of the Key-Value Cache used by large language models to sit outside GPU memory entirely.

A three-tier memory architecture built around SSD offloading

AI90 combines HBM, system DRAM, and SSD into a unified three-tier memory structure for handling inference workloads.

By transparently offloading KV Cache data onto SSDs, the platform reduces pressure on GPU memory while supporting significantly larger workloads and longer context windows.

According to GenStorAIGE, this architecture cuts first-token...

Copyright of this story solely belongs to techradar.com. To see the full text click HERE