波斯:多达100万个三亿个参数的混合系统增强深学习建议 (Persia: A Hybrid System Scaling Deep Learning Based Recommenders up to 100 Trillion Parameters)

Xiangru Lian,Binhang Yuan,Xuefeng Zhu,Yulong Wang,Yongjun He,Honghuan Wu,Lei Sun,Haodong Lyu,Chengjun Liu,Xing Dong,Yiqiao Liao,Mingnan Luo,Congfei Zhang,Jingru Xie,Haonan Li,Lei Chen,Renjie Huang,Jianying Lin,Chengchun Shu,Xuezhong Qiu,Zhishan Liu,Dongying Kong,Lei Yuan,Hai Yu,Sen Yang,Ce Zhang,Ji Liu

Deep learning based models have dominated the current landscape of production recommender systems. Furthermore, recent years have witnessed an exponential growth of the model scale--from Google's 2016 model with 1 billion parameters to the latest Facebook's model with 12 trillion parameters. Significant quality boost has come with each jump of the model capacity, which makes us believe the era of 100 trillion parameters is around the corner. However, the training of such models is challenging even within industrial scale data centers. This difficulty is inherited from the staggering heterogeneity of the training computation--the model's embedding layer could include more than 99.99% of the total model size, which is extremely memory-intensive; while the rest neural network is increasingly computation-intensive. To support the training of such huge models, an efficient distributed training system is in urgent need. In this paper, we resolve this challenge by careful co-design of both the optimization algorithm and the distributed system architecture. Specifically, in order to ensure both the training efficiency and the training accuracy, we design a novel hybrid training algorithm, where the embedding layer and the dense neural network are handled by different synchronization mechanisms; then we build a system called Persia (short for parallel recommendation training system with hybrid acceleration) to support this hybrid training algorithm. Both theoretical demonstration and empirical study up to 100 trillion parameters have conducted to justified the system design and implementation of Persia. We make Persia publicly available (at https://github.com/PersiaML/Persia) so that anyone would be able to easily train a recommender model at the scale of 100 trillion parameters.

翻译：深层次的学习模型主导了当前生产建议系统。此外,近年来,从谷歌2016年模型(含10亿参数)到脸书最新模型(含12万亿美元参数)的模型比例大幅增长。模型能力的每次跳跃都带来质素的显著提升,这使我们相信100万亿参数的时代即将到来。然而,即使是在工业规模的数据中心内部,这类模型的培训也具有挑战性。这种困难来自于培训计算-模型的简单嵌入层的惊人异质性,其中可能包括总模型规模的99.99%以上,这是非常记忆密集型的;而休息神经网络则越来越需要计算密集。为了支持这种巨大的模型的培训,一个高效分布式的培训系统是迫切需要的。在本文件中,我们通过谨慎地共同设计优化算法和分布式系统架构来解决这一挑战。具体地说,为了确保培训模式的效率和培训准确性,我们设计了一个新型的混合培训算法,在这个系统中,嵌入层层和密度稠密的神经网络由不同的同步机制处理; 然后,我们构建一个名为双级的模型设计系统,用来进行快速的模型设计。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/

【PAISS 2021 教程】概率散度与生成式模型，92页ppt

专知会员服务

34+阅读 · 2021年11月30日

深度学习优化算法，73页ppt，Optimization Algorithms on Deep Learning

专知会员服务

135+阅读 · 2021年6月16日

神经常微分方程教程，50页ppt，A brief tutorial on Neural ODEs

专知会员服务

74+阅读 · 2020年8月2日

使用TensorFlow建立深度学习模型，563页pdf，Deep Learning Pipeline Building a Deep Learning Model with TensorFlow

专知会员服务

149+阅读 · 2020年1月2日