课题基金 / 基金详情

Next generation imputation for huge data sets

Next generation imputation for huge data sets
大数据集的下一代插补
批准号:
BB/L020726/1
负责人:
John Hickey
金额:
$59.29万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --

项目摘要

项目成果

John Hickey的其他基金

相似基金

相关文献

中文摘要
翻译
从基因组测序中获得的知识在提高家畜育种中遗传变化的方向和速度以及动物科学中的生物发现方面具有巨大的潜力。然而,需要对大量的个体进行测序才能释放这种潜力,而目前家畜测序的成本是每个个体数百或数千英镑。这将仍然是常规使用这些数据的障碍,直到单位成本达到数十英镑的数量级。在降低成本的同时保持结果数据质量的一个有希望的方法是使用称为低覆盖率下一代测序(LcNGS)的技术。有了LcNGS,大量的个体可以以每个个体的低成本对他们的序列进行采样,但每个个体的序列都将有大量的缺失信息。准确性是通过使用一种称为补偿的过程来推断缺失数据来恢复的。在家畜中,家畜种群的系谱结构使这一过程变得更加有效。使用来自芯片的单核苷酸多态(SNP)数据进行推算已在家畜中成功应用。然而,由于几个原因,这些方法对于从LcNGS数据进行推算并不是最优的。(I)单核苷酸多态性芯片的基因分型非常准确,只是偶尔会因为技术问题而丢失数据点。相比之下,LcNGS数据对特定基因座上的真实基因型的确定性要低得多,并且缺失的数据随机分布在整个基因组中。(2)与序列数据相比,SNP-CHIP基因类型只覆盖基因组中存在的遗传变异的一小部分,因此输入序列数据的计算技术需要更有效地用于实际用途。(3)LcNGS产生的数据范围正在迅速演变,需要下一代推算算法非常灵活。该算法结合了启发式和概率两种方法,从一个新的角度解决了这些问题。启发式算法使用继承的基本原理,因此速度快、精度高。它们非常适合动物育种,因为它们使用谱系来推断来自大家庭的密切相关个体的丰富性,基因组的很大一部分在一对个体之间共享。然而,如果这样的数据在整个或部分基因组中缺乏或不可靠,启发式方法可能会失败。概率算法主要使用隐马尔可夫模型来在统计上模拟遗传,与启发式算法相比,它在计算上要求更高、速度更慢、精度更低。它们主要是为了应用于血统结构不适合于利用启发式算法的能力的人类群体而开发的,例如小兄弟姐妹关系。由于这两种方法在信息恢复和计算效率方面具有互补的优势,因此所提出的算法将通过结合这两种方法而获得协同效应。因此,总体目标是开发一个通用的推算系统,该系统能够在数以百万计的动物的数量级的数据集中进行推算,并且能够处理从LcNGS中可能出现的各种数据类型。将采用新的启发式方法来开发数据,这些数据可以与概率方法相结合,并组合成一种新的混合算法。将开发高效的数据处理和存储框架,并开发一个用户界面,以确保该算法在计算上高效、易于使用,并便于用户使用。该算法将使用一系列真实和模拟的数据集以及历史和真实的SNP芯片数据进行基准测试,以确保它仍然向后兼容当前或以前的技术。该算法的可用性将使育种者能够以较低的单位成本积累数百万只动物的序列数据,进而促进更高的选择准确性和育种目标的创新。
英文摘要
Knowledge gained from genome sequencing has great potential for increasing the direction and rate of genetic change in livestock breeding, and biological discovery in animal science. However huge numbers of individuals will need to be sequenced to unlock this potential, and the current cost of sequencing for livestock is several hundreds or thousands of pounds per individual. This will remain a barrier for using this data routinely until the unit cost is of the order of tens of pounds. One promising approach to reducing costs whilst maintaining the quality of the resulting data is to use technology called next-generation sequencing with low coverage (lcNGS). With lcNGS, large numbers of individuals can have their sequences sampled at low cost per individual, but each individual sequence will have substantial missing information. Accuracy is restored by inferring missing data using a process known as imputation. In livestock this process is made more efficient by pedigree structures in livestock populations.Imputation using single nucleotide polymorphism (SNP) data from chips has been successfully applied in livestock. However, these methods are not optimal for the imputation from lcNGS data for several reasons. (i) SNP-chip genotypes are highly accurate and data points are missing only occasionally due to technical issues. In contrast, lcNGS data has much less certainty over the true genotype at a particular locus, and the missing data is randomly spread over the whole genome. (ii) SNP-chip genotypes cover only a small fraction of the genetic variation present in the genome in comparison to sequence data, so the computational techniques for imputing sequence data need to be much more efficient for practical use. (iii) The range of the data produced by lcNGS is rapidly evolving, requiring next-generation imputation algorithms to be very flexible. The imputation algorithm proposed will address these issues from a novel direction by combining two approaches: heuristic and probabilistic. Heuristic algorithms use basic principles of inheritance and so are fast, and accurate. They are well-suited to animal breeding since they use pedigree to make inferences from the abundance of closely-related individuals from large families, with large portions of the genome shared between pairs of individuals. However, heuristic methods can fail if such data is lacking or is unreliable across all or parts of the genome. Probabilistic algorithms primarily use Hidden Markov Models to mimic inheritance statistically and are computationally more demanding, slower, and inherently less accurate than heuristic algorithms. They have been developed primarily for application to human populations in which the pedigree structures, for example small sibships, are not well-suited to exploiting the power of heuristic algorithms. The proposed algorithm will obtain synergy from combining the two approaches as they have complementary strengths in the recovery of information and computational efficiency.The overall objective is therefore to develop a generic imputation system that is capable of imputing in data sets of the order of millions of animals, can cope with the wide variety of data types that may appear from lcNGS. New heuristic approaches will be adopted to develop data that can be integrated with probabilistic approaches and combined into a novel hybrid algorithm. Efficient data handling and storage frameworks, and a user interface will be developed to ensure the algorithm is computationally efficient, easy-to-use, and readily available to users. The algorithm will be benchmarked using a range of real and simulated data sets and historical, real SNP-chip data to ensure it remains backwards compatible to current or previous technology. The availability of the algorithm will enable breeders to accumulate sequence data on millions of animals at low unit cost, and in turn prompt greater accuracy of selection and innovation in breeding goals.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
MOESM1 of Potential of gene drives with genome editing to increase genetic gain in livestock breeding programs
MOESM1 的基因驱动与基因组编辑增加牲畜育种计划遗传增益的潜力
DOI: 10.6084/m9.figshare.c.3666634_d1
发表时间: 2017
期刊:
影响因子: --
作者: [Gonen S]
通讯作者: Gonen S
DOI: 10.1186/s12711-017-0300-y
发表时间: 2017-03-03
期刊: Genetics, selection, evolution : GSE
影响因子: --
作者: [Antolín R, Nettelblad C, Gorjanc G, Money D, Hickey JM]
通讯作者: Hickey JM
MOESM5 of A hybrid method for the imputation of genomic data in livestock populations
MOESM5 家畜种群基因组数据插补的混合方法
DOI: 10.6084/m9.figshare.c.3708046_d5
发表时间: 2017
期刊:
影响因子: --
作者: [AntolA­N R]
通讯作者: AntolA­N R
A family-based phasing algorithm for sequence data
基于家族的序列数据定相算法
DOI: 10.1101/504480
发表时间: 2018
期刊:
影响因子: --
作者: [Battagin M]
通讯作者: Battagin M
共 9 条
    A general method for the imputation of genomic data in crop species
    • 批准号:
      BB/R002061/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $40.3万
    • 财政年份:
      2017
    • 负责人:
      John Hickey
    • 依托单位:
    Analysis of quantitative genetic traits in a huge data set
    • 批准号:
      BB/N006178/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $83.83万
    • 财政年份:
      2016
    • 负责人:
      John Hickey
    • 依托单位:
    15AGRITECHCAT3 Precision Breeding: Broilers from Sequence to Consequence
    • 批准号:
      BB/N004728/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $157.66万
    • 财政年份:
      2015
    • 负责人:
      John Hickey
    • 依托单位:
    Developing next generation genetic improvement tools from next generation sequencing
    • 批准号:
      BB/M009254/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $43.54万
    • 财政年份:
      2015
    • 负责人:
      John Hickey
    • 依托单位:
    国内基金
    海外基金
    细胞周期蛋白依赖性激酶Cdk1介导卵母细胞第一极体重吸收致三倍体发生的调控机制研究
    • 批准号:
      82371660
    • 项目类别:
      面上项目
    • 资助金额:
      49.00万元
    • 批准年份:
      2023
    • 负责人:
      魏喆
    • 依托单位:
    Next Generation Majorana Nanowire Hybrids
    二次谐波非线性光学显微成像用于前列腺癌的诊断及药物疗效初探
    • 批准号:
      30470495
    • 项目类别:
      面上项目
    • 资助金额:
      20.0万元
    • 批准年份:
      2004
    • 负责人:
      邓小元
    • 依托单位: