Analysis: AI Data Growth Outpacing Compute, Storage Emerging as New Bottleneck in AI Infrastructure
August 15th
According to a recent analysis by Western Digital, as the scale of artificial intelligence applications rapidly expands, AI data center construction is transitioning from a mere GPU computing power competition to a data storage capacity competition, where storage planning has become a core element of AI infrastructure.
The article quotes IDC's prediction that by 2030, the global annual data creation will reach 718ZB. Data generated by AI systems does not disappear when the computing tasks end. Training data, model checkpoints, embedding vectors, inference logs, prompts, outputs, and evaluation data will all continue to accumulate.
Western Digital stated that many current AI infrastructure plans excessively focus on GPU utilization but overlook the data precipitation throughout the AI lifecycle. Data generated during the training and inference processes will become crucial assets for model iteration, quality assessment, and compliance audits. Storage costs will directly impact the long-term operational efficiency of AI systems.
As data scales enter the PB or even EB level, a single storage architecture is insufficient to meet the demands. Enterprises need to adopt a tiered storage strategy, using high-performance flash storage for training and real-time inference, and high-capacity HDD and object storage for long-term data retention, historical records, and low-access scenarios.
The analysis suggests that in the future, a key metric for AI infrastructure competition will not just be the number of GPUs but rather the cost of data storage per PB, energy consumption, recovery efficiency, and data lifecycle management capabilities. If enterprises continue to view storage as a post-computing ancillary process, they may face issues such as uncontrollable data costs and decreased model iteration efficiency.