'A disk in a planet-scale computer': Meta has so many expensive GPUs that it's buying SSDs to kill idle…
- Meta rebuilt storage systems after slow data repeatedly stalled expensive AI GPUs
- SSD caching dramatically reduced AI dataset loading times from hours to minutes
- Meta replaced complex metadata lookups with a faster unified storage architecture
Meta says storage systems have failed to keep pace with AI computing power, creating delays that leave costly GPUs waiting instead of processing workloads efficiently.
According to the company's engineers, storage bottlenecks remain a major cause of GPU stalls, increasing operating costs while slowing research progress and extending development timelines.
To address those delays, Meta redesigned its storage architecture, arguing that faster movement of data can unlock greater value from expensive AI hardware investments.
Meta's engineers explained that the company operates hundreds of exabyte-scale storage clusters supporting Facebook, Instagram, Reality Labs, Meta AI, advertising systems, databases, and internal data warehouses.
Those services rely on a foundational storage layer called Tectonic, which manages object storage, file...
Copyright of this story solely belongs to techradar.com. To see the full text click HERE