assignPOP: An R package for population assignment using genetic, non-genetic, or integrated data in a machine-learning framework
assignPOP: An R package for population assignment using genetic, non-genetic, or integrated data in a machine-learning framework
复制标题
DOI:
10.1111/2041-210x.12897
复制
发表时间:
2018-02-01
影响因子:
6.6
通讯作者:
Ludsin, Stuart A.
中科院分区:
文献类型:
--
作者:
Chen, Kuan-Yu;Marschall, Elizabeth A.;Ludsin, Stuart A.
1. The use of biomarkers (e.g., genetic, microchemical and morphometric characteristics) to discriminate among and assign individuals to a population can benefit species conservation and management by facilitating our ability to understand population structure and demography.2. Tools that can evaluate the reliability of large genomic datasets for population discrimination and assignment, as well as allow their integration with non-genetic markers for the same purpose, are lacking. Our R package, assignPOP, provides both functions in a supervised machine-learning framework.3. assignPOP uses Monte-Carlo and K-fold cross-validation procedures, as well as principal component analysis, to estimate assignment accuracy and membership probabilities, using training (i.e., baseline source population) and test (i.e., validation) datasets that are independent. A user then can build a specified predictive model based on the relative sizes of these datasets and classification functions, including linear discriminant analysis, support vector machine, naive Bayes, decision tree and random forest.4. assignPOP can benefit any researcher who seeks to use genetic or non-genetic data to infer population structure and membership of individuals. assignPOP is a freely available R package under the GPL license, and can be downloaded from CRAN or at . A comprehensive tutorial can also be found at https://github.com/alexkychen/assignPOP. A comprehensive tutorial can also be found at https://alexkychen.github.io/assignPOP/.