NIRG: FARSPhase: a Flexible, widely Applicable, Robust, and Scalable phasing algorithm for human genetics
NIRG: FARSPhase: a Flexible, widely Applicable, Robust, and Scalable phasing algorithm for human genetics
批准号:
MR/M000370/1
负责人:
John Hickey
金额:
$48.26万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2015
资助国家:
英国
项目状态:
已结题
起止时间:
2015 至 --
中文摘要
在计算遗传学中,定相是对二倍体基因型的潜在单倍体结构的建模。这对许多遗传学研究很重要,因为遗传实际上发生在单倍体水平,即使我们只能用目前的主流技术直接观察二倍体基因型。在许多应用中,单倍型比单独的基因型提供更丰富和更有用的信息。单倍型阶段的应用包括理解遗传变异和疾病的相互作用,使血统模型能够用于遗传力分析,基因关联研究和基因组预测,未分型遗传变异的插补,为测序区分个体的优先级,识别基因型,检测基因型错误,推断人类人口统计学历史,推断重组点,检测复发突变和选择特征,人类遗传学数据集,将可能在未来分阶段进行,可以分为:(i)巨大的人口名义上无关的个人(例如500,000人,英国生物库);(ii)此类人群的较小子集(例如,在个别研究中收集的数据);(三)大(例如50,000人)或小型(例如1 000人)从相互隔离的人群中收集的数据集,这些人群之间具有高度的相关性(例如:Orcades -奥克尼,deCODE -冰岛,维京海盗-瑞典);(iv)有和没有系谱信息的数据集;(v)结合了联合收割机几个特征的数据集(例如苏格兰一代);以及(vi)具有不同类型的基因组信息的数据集(例如,单核苷酸多态性、低或高覆盖度序列、短或较长的序列读段等)。人类遗传学数据有许多定相方法,这些方法可以大致分为两组:(i)启发式方法(例如长距离定相(LRP));和(ii)概率方法(例如隐马尔可夫模型(HMM))。分阶段是计算密集型的,不同数据集的大小和特征使它们或多或少适合于特定的方法。与HMM相比,LRP在计算上是快速的,但仅适用于个体共享相对较近的祖先(例如在10代内)并且因此共享相对较长的单倍型(例如5至10 cM长度)的情况。孤立的群体(例如在Orcades,Orkney)非常适合LRP,但具有数十万名义上不相关个体的巨大群体也可能适合(例如UK Biobank)。目前的HMM应用于如此庞大的人口是计算上棘手的。然而,HMM比LRP更适合于这样的群体的子集,因为HMM仅要求个体由于共享非常远的亲属(例如50至100代前)而共享短的单倍型(例如<lcM)JRP和HMM方法在许多方面是互补的。一个模型长单倍型,其他短单倍型。HMM方法更灵活,可以更好地模拟数据中的不确定性。LRP方法在计算上更有效,并且在它们适合的场景中也更准确。LRP方法也更适合于并入谱系信息。一个组合算法可以利用这种互补性。本提案的目标是开发FARSPhase:一个灵活的,广泛适用的,鲁棒的,可扩展的,人类遗传学的定相算法,结合了LRP,其他遗传学和HMM方法的最佳功能到一个单一的框架。除了满足小数据集的分阶段需求外,如果成功的话,这项研究将使巨大的数据集分阶段,从而为更强大的分析提供可能性。所开发的算法将被合并到一个用户友好的软件包中,该软件包使用软件工程中的最佳实践构建,其性能将在广泛的模拟和真实的数据集中进行测试,这些数据集反映了人类遗传学未来可能的分阶段情景。
英文摘要
In computational genetics, phasing is the modelling of the underlying haploid structure of diploid genotypes. It is important for many genetic studies because inheritance actually takes place at the haploid level, even though we can only directly observe diploid genotypes with current mainstream technologies. In many applications haplotypes provide richer and more useful information than genotypes alone. Applications of haplotype phase include understanding the interplay of genetic variation and disease, enabling identity-by-descent models for use in heritability analysis, gene association studies and genomic prediction, imputation of un-typed genetic variation, prioritizing individuals for sequencing, calling genotypes, detecting genotype error, inferring human demographic history, inferring points of recombination, detecting recurrent mutation and signatures of selection, and modelling cis-regulation of gene expression.Human genetics data sets that will likely be phased in the future can be categorised into: (i) huge populations of nominally unrelated individuals (e.g. 500,000 individuals, UK Biobank); (ii) smaller subsets of such populations (e.g. data collected in individual studies); (iii) large (e.g. 50,000 individuals) or small (e.g. 1,000 individuals) data sets collected from isolated populations with high degrees of relatedness within them (e.g. Orcades - Orkney, deCODE - Iceland, VIKING - Sweden); (iv) data sets with and without pedigree information; (v) data sets that combine several of these features (e.g. Generation Scotland); and (vi) data sets with different types of genomic information (e.g. single nucleotide polymorphisms, low- or high-coverage sequence, short or longer sequence reads, etc.).There are many phasing methods for human genetics data and these can be broadly classified into two groups: (i) heuristic methods (e.g. Long-Range Phasing (LRP)); and (ii) probabilistic methods (e.g. Hidden Markov Models (HMM)). Phasing is computationally intensive and the size and features of different data sets make them more or less suited to particular methods. LRP is computationally fast in comparison to HMM, but is only applicable to situations where individuals share relatively recent ancestry (e.g. within 10 generations) and thus share relatively long haplotypes (e.g. 5 to 10 cM length). Isolated populations (e.g. as in Orcades, Orkney) are ideally suited to LRP but huge populations with hundreds of thousands of nominally unrelated individuals may also be suitable (e.g. UK Biobank). Application of current HMM to such huge populations is computationally intractable. However, HMM are more suited to subsets of such populations than LRP because HMM only require that individuals share short haplotypes (e.g. <1 cM) due to sharing very distant relatives (e.g. 50 to 100 generations ago).LRP and HMM methods are complementary in many ways. One models long haplotypes, the other short haplotypes. HMM methods are more flexible and can better model uncertainty in the data. LRP methods are computationally much more efficient and are also more accurate in scenarios to which they are suited. LRP methods are also more amenable to incorporation of pedigree information. A combined algorithm could exploit this complementarity.The objective of this proposal is to develop FARSPhase: a Flexible, widely Applicable, Robust, and Scalable, phasing algorithm for human genetics that combines the best features of LRP, other heuristics, and HMM methods into a single framework. As well as meeting the phasing needs for small data sets, if successful, this research will enable huge data sets be phased and thereby opening the possibility of more powerful analysis. The developed algorithm will be combined into a user friendly software package built using best practices in software engineering and its performance will be tested in a wide range of simulated and real data sets that reflect the likely future phasing scenarios for human genetics.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1186/s12711-017-0300-y
发表时间:
2017-03-03
期刊:
Genetics, selection, evolution : GSE
影响因子:
--
作者:
[Antolín R, Nettelblad C, Gorjanc G, Money D, Hickey JM]
通讯作者:
Hickey JM
MOESM3 of A hybrid method for the imputation of genomic data in livestock populations
用于家畜种群基因组数据插补的混合方法的 MOESM3
DOI:
10.6084/m9.figshare.c.3708046_d3
发表时间:
2017
期刊:
影响因子:
--
作者:
[AntolAN R]
通讯作者:
AntolAN R
MOESM8 of A hybrid method for the imputation of genomic data in livestock populations
MOESM8 家畜种群基因组数据插补的混合方法
DOI:
10.6084/m9.figshare.c.3708046_d8
发表时间:
2017
期刊:
影响因子:
--
作者:
[AntolAN R]
通讯作者:
AntolAN R
A family-based phasing algorithm for sequence data
基于家族的序列数据定相算法
DOI:
10.1101/504480
发表时间:
2018
期刊:
影响因子:
--
作者:
[Battagin M]
通讯作者:
Battagin M
DOI:
10.1186/s12711-016-0221-1
发表时间:
2016-06-22
期刊:
Genetics, selection, evolution : GSE
影响因子:
--
作者:
[Battagin M, Gorjanc G, Faux AM, Johnston SE, Hickey JM]
通讯作者:
Hickey JM
共 9 条
A general method for the imputation of genomic data in crop species
-
批准号:BB/R002061/1
-
项目类别:Research Grant
-
资助金额:$40.3万
-
财政年份:2017
-
负责人:John Hickey
-
依托单位:
Analysis of quantitative genetic traits in a huge data set
-
批准号:BB/N006178/1
-
项目类别:Research Grant
-
资助金额:$83.83万
-
财政年份:2016
-
负责人:John Hickey
-
依托单位:
15AGRITECHCAT3 Precision Breeding: Broilers from Sequence to Consequence
-
批准号:BB/N004728/1
-
项目类别:Research Grant
-
资助金额:$157.66万
-
财政年份:2015
-
负责人:John Hickey
-
依托单位:
Developing next generation genetic improvement tools from next generation sequencing
-
批准号:BB/M009254/1
-
项目类别:Research Grant
-
资助金额:$43.54万
-
财政年份:2015
-
负责人:John Hickey
-
依托单位:
15AGRITECHCAT3 Innovative NextGen pig breeding using DNA sequence data
-
批准号:BB/N004736/1
-
项目类别:Research Grant
-
资助金额:$149.1万
-
财政年份:2015
-
负责人:John Hickey
-
依托单位:
Next generation imputation for huge data sets
-
批准号:BB/L020726/1
-
项目类别:Research Grant
-
资助金额:$59.29万
-
财政年份:2014
-
负责人:John Hickey
-
依托单位: