通过子抽样对要素分析和主要构件分析的构件数量进行信任度间数 (Confidence Intervals for the Number of Components in Factor Analysis and Principal Components Analysis via Subsampling)

Factor analysis (FA) and principal component analysis (PCA) are popular statistical methods for summarizing and explaining the variability in multivariate datasets. By default, FA and PCA assume the number of components or factors to be known \emph{a priori}. However, in practice the users first estimate the number of factors or components and then perform FA and PCA analyses using the point estimate. Therefore, in practice the users ignore any uncertainty in the point estimate of the number of factors or components. For datasets where the uncertainty in the point estimate is not ignorable, it is prudent to perform FA and PCA analyses for the range of positive integer values in the confidence intervals for the number of factors or components. We address this problem by proposing a subsampling-based data-intensive approach for estimating confidence intervals for the number of components in FA and PCA. We study the coverage probability of the proposed confidence intervals and provide non-asymptotic theoretical guarantees concerning the accuracy of the confidence intervals. As a byproduct, we derive the first-order \emph{Edgeworth expansion} for spiked eigenvalues of the sample covariance matrix when the data matrix is generated under a factor model. We also demonstrate the usefulness of our approach through numerical simulations and by applying our approach for estimating confidence intervals for the number of factors of the genotyping dataset of the Human Genome Diversity Project.

翻译：系数分析(FA)和主要组成部分分析(PCA)是用来总结和解释多变量数据集变异性的流行统计方法。默认情况下,FA和CPA将假定已知的元素或因素数量或因素数量,但在实践中,用户首先估计系数或要素数量,然后使用点估计进行FA和CPA分析。因此,用户实际上忽略了因素或组成部分数量点估计值的任何不确定性。对于点估计值的不确定性不可忽略的数据集,谨慎的做法是在因素或组成部分数量的信任间隔中进行FA和CPA分析正整数值的范围。我们通过提出基于次级抽样的数据密集方法来解决这一问题,以估计FA和CPA组成部分数量的信任间隔。我们研究拟议的信任间隔的涵盖范围,并提供有关信任间隔的准确性非约束性理论保证。作为副产品,我们得出第一级 emph {Edgeworth 扩展} 进行FAA和CPA分析是谨慎的。当我们通过模拟数据基数的模型模型显示人类基因组方法的精度值时,我们通过模拟模型模型显示我们的数据基数的数值,我们用模拟模型显示人类基因组方法的可靠性系数组的模型显示我们的数据基数。

相关内容

PCA

关注 3

在统计中，主成分分析（PCA）是一种通过最大化每个维度的方差来将较高维度空间中的数据投影到较低维度空间中的方法。给定二维，三维或更高维空间中的点集合，可以将“最佳拟合”线定义为最小化从点到线的平均平方距离的线。可以从垂直于第一条直线的方向类似地选择下一条最佳拟合线。重复此过程会产生一个正交的基础，其中数据的不同单个维度是不相关的。这些基向量称为主成分。

不可错过！《机器学习100讲》课程，UBC Mark Schmidt讲授

专知会员服务

76+阅读 · 2022年6月28日

ICLR 2022杰出论文公布：7篇论文获得，清华朱军课题组摘得

专知会员服务

60+阅读 · 2022年4月22日

不可错过！UIUC最新《统计强化学习》课程！

专知会员服务

54+阅读 · 2020年9月7日