Utilizing Tech - The Podcast Series about New and Emerging Technologies · Tech Field Day - Part of The Futurum Group

07x03: Benchmarking AI Data Infrastructure with MLCommons

·31 min·3 clips
ML Commons measures if storage can keep $40,000 NVIDIA GPUs from starving for data during AI training.
This episode focuses on benchmarking storage performance for AI training workloads, featuring host Stephen Foskett and co-host Ace Striker from Solidigm. Their guest is Curtis Anderson, co-chair of the Storage Working Group at ML Commons, who explains the group's practical approach to storage benchmarking. ML Commons develops benchmarks that measure how well a storage subsystem keeps AI training accelerators like GPUs utilized, rather than reporting traditional metrics like IOPS. The benchmark currently focuses on the training phase of the AI pipeline, simulating the data workload a storage system would face during model training. It measures accelerator utilization as its core metric, essentially testing whether storage can keep expensive GPUs from idling due to data starvation. The test supports various AI model types, including image recognition, recommender systems, and large language models, each imposing a different workload. Practitioners can use the results to determine how many GPUs a given storage solution can support without bottlenecking their training jobs. The benchmark uses software to emulate accelerators, allowing vendors and researchers to test without a physical rack of costly GPUs. Curtis Anderson notes the initial focus is on NVIDIA GPUs due to market dominance, but support for other accelerators like Graphcore or Tenstorrent is planned. A future goal is to incorporate data preparation into the benchmark, a phase Meta's research indicated can consume 50% of an AI project's total compute energy. The working group aims for two benchmark releases per year, with a version 1.0 announcement expected around mid-May, followed by a submission window and peer review. A key insight is that the benchmark answers a practitioner's direct question: will this storage purchase keep my specific GPU cluster busy? It bridges the terminology gap between storage engineers focused on IOPS and AI practitioners who think in terms of samples per second and GPU counts. The emulation approach democratizes testing, enabling participation from academia, open-source projects, and vendors without vast hardware budgets. Surprisingly, inference workloads are currently a lower priority for storage benchmarking, as they are less storage-intensive than large-scale training. Data preparation is acknowledged as highly complex and unique per application, making it a challenging but critical next frontier for standardization. The benchmark results are intended for practical procurement and sizing decisions, not just for generating marketing "angel numbers." The tone is educational and conversational, with hosts asking clarifying questions to unpack technical concepts for a broad audience. The style is interview-based and industry-focused, emphasizing real-world applicability over theoretical performance. This episode is ideal for IT infrastructure architects, data center managers, and AI practitioners responsible for selecting or sizing storage for AI projects. Listeners seeking an introductory overview of AI hardware benchmarking or debates on industry trends might find the technical deep-dive less applicable.
Listen to the show on