了解深点击区速率预测模型的超称现象 (Towards Understanding the Overfitting Phenomenon of Deep Click-Through Rate Prediction Models)

Deep learning techniques have been applied widely in industrial recommendation systems. However, far less attention has been paid to the overfitting problem of models in recommendation systems, which, on the contrary, is recognized as a critical issue for deep neural networks. In the context of Click-Through Rate (CTR) prediction, we observe an interesting one-epoch overfitting problem: the model performance exhibits a dramatic degradation at the beginning of the second epoch. Such a phenomenon has been witnessed widely in real-world applications of CTR models. Thereby, the best performance is usually achieved by training with only one epoch. To understand the underlying factors behind the one-epoch phenomenon, we conduct extensive experiments on the production data set collected from the display advertising system of Alibaba. The results show that the model structure, the optimization algorithm with a fast convergence rate, and the feature sparsity are closely related to the one-epoch phenomenon. We also provide a likely hypothesis for explaining such a phenomenon and conduct a set of proof-of-concept experiments. We hope this work can shed light on future research on training more epochs for better performance.

翻译：在工业建议系统中广泛采用了深层次的学习技术,然而,对建议系统中模型的过分适应问题的关注却少得多,相反,这个问题被确认为深神经网络的关键问题。在点击率预测中,我们观察到一个令人感兴趣的一个时代的过度适应问题:模型性能在第二个时代之初显示急剧退化。在CTR模型的实际应用中,这种现象被广泛看到。因此,通常通过只进行一个时代的培训才能取得最佳的绩效。为了了解单一时代现象背后的根本因素,我们对从Alibaba的显示广告系统收集的成套生产数据进行了广泛的实验。结果显示,模型结构、具有快速趋同率的优化算法和特征偏执与单一时代现象密切相关。我们还为解释这种现象和进行一套概念验证实验提供了一种可能的假设。我们希望这项工作能够为今后对如何培训更深入地进行更好的表现的研究提供启发。

相关内容

过拟合

关注 8

过拟合，在AI领域多指机器学习得到模型太过复杂，导致在训练集上表现很好，然而在测试集上却不尽人意。过拟合（over-fitting）也称为过学习，它的直观表现是算法在训练集上表现好，但在测试集上表现不好，泛化性能差。过拟合是在模型参数拟合过程中由于训练数据包含抽样误差，在训练时复杂的模型将抽样误差也进行了拟合导致的。

高效可扩展图神经网络的研究进展，Recent Advances in Efficient and Scalable Graph Neural Networks

专知会员服务

78+阅读 · 2022年3月15日

NLP必读经典文献100篇

专知会员服务

124+阅读 · 2020年9月8日

Linux导论，Introduction to Linux，96页ppt

专知会员服务

81+阅读 · 2020年7月26日

图像分类技巧集，17页ppt《Bag of Tricks for Image Classification》

专知会员服务

96+阅读 · 2020年3月12日