高层面强有力的角基转移学习 (Robust angle-based transfer learning in high dimensions)

Transfer learning aims to improve the performance of a target model by leveraging data from related source populations, which is known to be especially helpful in cases with insufficient target data. In this paper, we study the problem of how to train a high-dimensional ridge regression model using limited target data and existing regression models trained in heterogeneous source populations. We consider a practical setting where only the parameter estimates of the fitted source models are accessible, instead of the individual-level source data. Under the setting with only one source model, we propose a novel flexible angle-based transfer learning (angleTL) method, which leverages the concordance between the source and the target model parameters. We show that angleTL unifies several benchmark methods by construction, including the target-only model trained using target data alone, the source model fitted on source data, and distance-based transfer learning method that incorporates the source parameter estimates and the target data under a distance-based similarity constraint. We also provide algorithms to effectively incorporate multiple source models accounting for the fact that some source models may be more helpful than others. Our high-dimensional asymptotic analysis provides interpretations and insights regarding when a source model can be helpful to the target model, and demonstrates the superiority of angleTL over other benchmark methods. We perform extensive simulation studies to validate our theoretical conclusions and show the feasibility of applying angleTL to transfer existing genetic risk prediction models across multiple biobanks.

翻译：转让学习旨在通过利用相关源群的数据改进目标模型的性能,据了解,这种数据在目标数据不足的情况下特别有用。在本文件中,我们研究了如何利用有限的目标数据和在不同源群中培训的现有回归模型来培训高维脊回归模型的问题。我们考虑一个实际的设置,即只有匹配源群模型的参数估计数,而不是个人源数据,才能获得适合的来源模型的参数估计数,而不是个人源数据。在仅使用一个源模型的设置下,我们提议一种创新的灵活角度转移学习(tragle TL)方法,利用源和目标模型参数的一致性。我们表明,角度TL通过构建,将若干基准方法统一起来,包括仅使用目标数据培训的目标型模型、源数据配置的源模型以及远程转移学习方法,其中只包括源参数估计数和目标数据,而不是个人源数据。我们还提供算法,以便有效地纳入多种源模型,因为一些源模型可能比其他来源模型更有帮助。我们的高维度分析提供了解释和洞察一些基准方法,即当一种源模型能够帮助进行广泛的实验性研究时,我们进行跨生物实验室的模型的推算。

相关内容

MoDELS

关注 43

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/