BREEDING AND GENETICS SYMPOSIUM: Really big data: Processing and analysis of very large data sets

BREEDING AND GENETICS SYMPOSIUM: Really big data: Processing and analysis of very large data sets
复制标题

DOI:
10.2527/jas.2011-4584
复制
发表时间:
2012-03-01
影响因子:
3.3
通讯作者:
Coffey, M.
Coffey, M.
中科院分区:
农林科学2区
文献类型:
--
作者:
Cole, J. B.;Newman, S.;Coffey, M.

文献摘要

被引文献

相似文献

现代动物育种数据集非常庞大,而且还在不断扩大,部分原因是最近出现了高密度 SNP 阵列和廉价的测序技术。用于高效数据仓储和分析的高性能计算方法正在开发中。使用共享集群时,财务和安全考虑非常重要。需要良好的软件工程实践,并且最好尽可能使用现有的解决方案。尽管全序列数据需要更大的存储容量,但基因型的存储要求并不高。遗传评估的中间和结果文件的存储要求要高得多,特别是当必须存储多次运行以进行研究和验证研究时。对于低遗传力的性状,基因组选择的准确性得到了最大的提高,并且人们对新的健康和管理性状越来越感兴趣。收集足够的表型以进行准确的评估可能需要许多年的时间,并且需要针对老年公牛的高可靠性证据来估计标记效应。应用于大型数据集的数据挖掘算法可能有助于识别数据中意想不到的关系,而改进的可视化工具将提供见解。使用大数据进行基因组选择需要大量的计算能力,特别是当对大部分群体进行基因分型时。理论改进使得大型分子关系矩阵的求逆成为可能,允许求解大型方程组,并产生了方差分量估计的快速算法。最近的工作表明,将 BLUP 与基因组关系 (G) 矩阵相结合的单步方法与传统 BLUP 具有类似的计算要求,限制因素是许多基因型的 G 的构建和反转。为 14,000 人创建 G 的简单算法需要运行近 24 小时,但自定义库和并行计算将其缩短至 15 m。大数据集也给遗传评估的实施带来了挑战,必须以不破坏从传统评估到基因组评估的过渡的方式克服这些挑战。处理时间很重要,尤其是在开发用于农场决策的实时系统时。这些系统的最终价值是缩短研究结果的时间,提高基因组评估的准确性,并加快遗传改良的速度。
Modern animal breeding data sets are large and getting larger, due in part to recent availability of high-density SNP arrays and cheap sequencing technology. High-performance computing methods for efficient data warehousing and analysis are under development. Financial and security considerations are important when using shared clusters. Sound software engineering practices are needed, and it is better to use existing solutions when possible. Storage requirements for genotypes are modest, although full-sequence data will require greater storage capacity. Storage requirements for intermediate and results files for genetic evaluations are much greater, particularly when multiple runs must be stored for research and validation studies. The greatest gains in accuracy from genomic selection have been realized for traits of low heritability, and there is increasing interest in new health and management traits. The collection of sufficient phenotypes to produce accurate evaluations may take many years, and high-reliability proofs for older bulls are needed to estimate marker effects. Data mining algorithms applied to large data sets may help identify unexpected relationships in the data, and improved visualization tools will provide insights. Genomic selection using large data requires a lot of computing power, particularly when large fractions of the population are genotyped. Theoretical improvements have made possible the inversion of large numerator relationship matrices, permitted the solving of large systems of equations, and produced fast algorithms for variance component estimation. Recent work shows that single-step approaches combining BLUP with a genomic relationship (G) matrix have similar computational requirements to traditional BLUP, and the limiting factor is the construction and inversion of G for many genotypes. A naive algorithm for creating G for 14,000 individuals required almost 24 h to run, but custom libraries and parallel computing reduced that to 15 m. Large data sets also create challenges for the delivery of genetic evaluations that must be overcome in a way that does not disrupt the transition from conventional to genomic evaluations. Processing time is important, especially as real-time systems for on-farm decisions are developed. The ultimate value of these systems is to decrease time-to-results in research, increase accuracy in genomic evaluations, and accelerate rates of genetic improvement.