Weka's new storage platform caches 100% of a model's pre-calculated tokens, so it never has to redo the work

https://images.ctfassets.net/jdtwqhzvc2n1/UUXwemP15RHK5kBVdXsy3/ca6758f1de3542c0d26cb65269c5c5c0/cheap-flash-storage-smk1.jpg?w=800&q=75

GPU memory is the most expensive resource in production AI, and it's also the one running out fastest.

Long context windows and multi-turn conversations force AI models to repeatedly recompute information they've already processed, consuming GPU memory and compute that could otherwise serve additional users or generate new responses.

Instead of treating GPU memory as the limiting resource, why not extend it with much cheaper storage technologies?

Weka, for one, believes that cheap flash storage can close that gap. The company's NeuralMesh 6 software platform, launching alongside its first self-designed hardware line, Wekapod 3, extends what Weka calls Augmented Memory Grid, an approach that aggregates NAND flash to behave like GPU memory at a fraction of the cost.

This is an active and increasingly crowded category. Dell, NetApp, Pure Storage and VAST have all repositioned toward AI infrastructure over the past two years and Weka is one of several...

Copyright of this story solely belongs to venturebeat.com. To see the full text click HERE

Read more