不可忽略的失踪的深生成型样元混合模型模型 (Deep Generative Pattern-Set Mixture Models for Nonignorable Missingness)

We propose a variational autoencoder architecture to model both ignorable and nonignorable missing data using pattern-set mixtures as proposed by Little (1993). Our model explicitly learns to cluster the missing data into missingness pattern sets based on the observed data and missingness masks. Underpinning our approach is the assumption that the data distribution under missingness is probabilistically semi-supervised by samples from the observed data distribution. Our setup trades off the characteristics of ignorable and nonignorable missingness and can thus be applied to data of both types. We evaluate our method on a wide range of data sets with different types of missingness and achieve state-of-the-art imputation performance. Our model outperforms many common imputation algorithms, especially when the amount of missing data is high and the missingness mechanism is nonignorable.

翻译：我们提出一个变式自动编码结构,用小不列颠(1993年)建议的模式设定混合物来模拟可忽略和不可忽略的数据缺失。我们的模型明确学会根据观察到的数据和缺失面罩将缺失的数据分组成缺失模式。我们的方法所依据的假设是,在缺失情况下的数据分布在概率上半由观察到的数据分布样本监督。我们的设置交换了可忽略和不可忽略的缺失的特征,因此可以适用于这两种类型的数据。我们评估了各种数据集中不同类型缺失的方法,并取得了最先进的估算性能。我们的模型比许多常见估算算法要强得多,特别是当缺失的数据数量高,而且缺失机制不亮的时候。

相关内容

MoDELS

关注 44

ACM/IEEE第23届模型驱动工程语言和系统国际会议，是模型驱动软件和系统工程的首要会议系列，由ACM-SIGSOFT和IEEE-TCSE支持组织。自1998年以来，模型涵盖了建模的各个方面，从语言和方法到工具和应用程序。模特的参加者来自不同的背景，包括研究人员、学者、工程师和工业专业人士。MODELS 2019是一个论坛，参与者可以围绕建模和模型驱动的软件和系统交流前沿研究成果和创新实践经验。今年的版本将为建模社区提供进一步推进建模基础的机会，并在网络物理系统、嵌入式系统、社会技术系统、云计算、大数据、机器学习、安全、开源等新兴领域提出建模的创新应用以及可持续性。官网链接：http://www.modelsconference.org/