Testing of the Effect of Missing Data Estimation and Distribution in Morphometric Multivariate Data Analyses

Testing of the Effect of Missing Data Estimation and Distribution in Morphometric Multivariate Data Analyses
复制标题

DOI:
10.1093/sysbio/sys047
复制
发表时间:
2012-12-01
期刊:
影响因子:
6.5
通讯作者:
Jackson, Donald A.
Jackson, Donald A.
中科院分区:
生物学1区
文献类型:
--
作者:
Brown, Caleb Marshall;Arbour, Jessica H.;Jackson, Donald A.

文献摘要

被引文献

相似文献

缺失数据是生物学数据集中不可避免的问题,而缺失数据删除和估计技术在形态学数据集上的性能却知之甚少。在这里,一种新的方法是用来衡量引入误差的多个技术的代表性样品。对大样本的鳄鱼头骨进行了测量和主成分分析。将23个不同比例的缺失数据引入数据集中,估计,分析,并与原始结果进行比较,使用Procrustes叠加。以前的工作调查缺失数据的影响随机输入缺失值,一种非生物现象。在这里,丢失的数据被引入到数据集使用三种方法:纯粹随机,作为相应的测量之间的欧几里得距离的函数(模拟解剖区域),并作为每个分类群所占据的样本的一部分的函数(模拟不平等的缺失数据在罕见的分类群)。高尔的距离被认为是最好的非估计方法,和贝叶斯PCA的最好的估计方法。样本量小的类群和那些形态上最不同的标本有最高的估计误差。缺失数据的分布对几乎所有方法和比例的估计误差都有显著影响。分类学上有偏差的缺失数据往往表现出与随机相似的趋势,但错误率更高。解剖偏见的缺失数据表现出更大的偏离随机比分类的偏见,和幅度依赖于估计方法。
Missing data are an unavoidable problem in biological data sets and the performance of missing data deletion and estimation techniques in morphometric data sets is poorly understood. Here, a novel method is used to measure the introduced error of multiple techniques on a representative sample. A large sample of extant crocodilian skulls was measured and analyzed with principal component analysis (PCA). Twenty-three different proportions of missing data were introduced into the data set, estimated, analyzed, and compared with the original result using Procrustes superimposition. Previous work investigating the effects of missing data input missing values randomly, a non-biological phenomenon. Here, missing data were introduced into the data set using three methodologies: purely at random, as a function of the Euclidean distance between respective measurements (simulating anatomical regions), and as a function of the portion of the sample occupied by each taxon (simulating unequal missing data in rare taxa). Gower's distance was found to be the best performing non-estimation method, and Bayesian PCA the best performing estimation method. Specimens of the taxa with small sample sizes and those most morphologically disparate had the highest estimation error. Distribution of missing data had a significant effect on the estimation error for almost all methods and proportions. Taxonomically biased missing data tended to show similar trends to random, but with higher error rates. Anatomically biased missing data showed a much greater deviation from random than the taxonomic bias, and with magnitudes dependent on the estimation method.