机器学习训练在真实处理内存系统上的实验评估 (An Experimental Evaluation of Machine Learning Training on a Real Processing-in-Memory System)

Training machine learning (ML) algorithms is a computationally intensive process, which is frequently memory-bound due to repeatedly accessing large training datasets. As a result, processor-centric systems (e.g., CPU, GPU) suffer from costly data movement between memory units and processing units, which consumes large amounts of energy and execution cycles. Memory-centric computing systems, i.e., with processing-in-memory (PIM) capabilities, can alleviate this data movement bottleneck. Our goal is to understand the potential of modern general-purpose PIM architectures to accelerate ML training. To do so, we (1) implement several representative classic ML algorithms (namely, linear regression, logistic regression, decision tree, K-Means clustering) on a real-world general-purpose PIM architecture, (2) rigorously evaluate and characterize them in terms of accuracy, performance and scaling, and (3) compare to their counterpart implementations on CPU and GPU. Our evaluation on a real memory-centric computing system with more than 2500 PIM cores shows that general-purpose PIM architectures can greatly accelerate memory-bound ML workloads, when the necessary operations and datatypes are natively supported by PIM hardware. For example, our PIM implementation of decision tree is $27\times$ faster than a state-of-the-art CPU version on an 8-core Intel Xeon, and $1.34\times$ faster than a state-of-the-art GPU version on an NVIDIA A100. Our K-Means clustering on PIM is $2.8\times$ and $3.2\times$ than state-of-the-art CPU and GPU versions, respectively. To our knowledge, our work is the first one to evaluate ML training on a real-world PIM architecture. We conclude with key observations, takeaways, and recommendations that can inspire users of ML workloads, programmers of PIM architectures, and hardware designers & architects of future memory-centric computing systems.

翻译：训练机器学习（ML）算法是一个计算密集型的过程，它由于反复访问大型训练数据而经常受到内存限制。因此，面向处理器的系统（例如CPU、GPU）在内存单元和处理单元之间遭受昂贵的数据移动，这会消耗大量的能量和执行周期。具有处理内存（PIM）功能的内存中心计算系统可以缓解这种数据移动瓶颈。我们的目标是理解现代通用PIM架构加速ML训练的潜力。为此，我们（1）在真实的通用PIM架构上实现几种典型的ML算法（即线性回归、逻辑回归、决策树、K-Means聚类），（2）从准确性、性能和可扩展性等角度严格评估和表征它们，(3) 并与它们在CPU和GPU上的对应实现进行比较。我们在具有2500多个PIM核心的真实内存中心计算系统上进行的评估表明，当所需操作和数据类型被PIM硬件本地支持时，通用PIM架构可以大大加速内存受限的ML工作负载。例如，我们对决策树的PIM实现比基于8核Intel Xeon的最先进CPU版本快27倍，比基于NVIDIA A100的最先进GPU版本快1.34倍。我们的PIM K-Means聚类比最先进CPU和GPU版本分别快2.8倍和3.2倍。据我们所知，我们的工作是第一个在真实世界的PIM架构上评估ML训练的工作。我们总结了关键观察结果、要点和建议，可以启发ML工作负载的用户、PIM架构的程序员以及未来内存中心计算系统的硬件设计师和架构师。