d-Matrix Claims 20x Bandwidth Density Over NVIDIA Rubin In New Chip

https://hothardware.com/contentimages/NewsItem/71463/content/16x9_2133x1200_highres-d-matrix-pavehawk-photo.jpg

In AI inference, you have two phases of the workload: prefill and decode. In most cases, decode is overwhelmingly the more time-consuming portion of the workload because it's strictly memory-bandwidth bound. You can have all the compute in the world, and that's great for prefill, but decode doesn't care. There have been various strategies to attack this problem, including startups like Groq (whose tech was recently purchased by NVIDIA) and d-Matrix offering chips with massive SRAM to deliver enormous memory bandwidth. d-Matrix now has a new product it's showing off, and it attacks the bandwidth problem directly from a different angle: below.

You know how AMD's second-generation 3D V-Cache, used on the Ryzen 9000 processors, places the SRAM cache die underneath the CPU coresto improve thermals? That's basically what's going on in d-Matrix's "Raptor" accelerator chip, except the Silicon Valley startup is using up to four layers...

Copyright of this story solely belongs to hothardware.com. To see the full text click HERE