课题基金 / 基金详情

Inferring ancestry and relatedness of human genomes using ancient DNA samples' and falls within the EPSRC Artificial Intelligence and Healthcare Techn

Inferring ancestry and relatedness of human genomes using ancient DNA samples' and falls within the EPSRC Artificial Intelligence and Healthcare Techn
使用古代 DNA 样本推断人类基因组的祖先和相关性,属于 EPSRC 人工智能和医疗保健技术的范围
批准号:
2420820
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2020
资助国家:
英国
项目状态:
未结题
起止时间:
2020 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
在每一个基因组位置,两个个体通过宗谱关系联系在一起,从而产生共同的祖先。从个体到这个祖先的时间距离被称为时间到最近的共同祖先(TMRCA)。这可以通过用树表示他们的家谱关系来推广到一组个体。沿着基因组移动,树的拓扑结构可以随着减数分裂期间基因组的重组而改变。因此,一组样本的进化史可以用一个被称为祖先重组图(ARG)的图来简洁地表示,该图由跨越不同基因组块的单个树组成。从现代DNA样本的高质量测序数据中重建ARG有多种计算密集型的方法。由于环境条件和污染,古样品质量下降;通常,它们只能在非常低的覆盖率下进行测序,这使得将它们纳入一组现代样本的ARG的任务具有挑战性。在现代和古代DNA样本之间重建精确的联合ARG具有多种潜在的应用,我们的目标是探索。它可以用于祖先推断,通过恢复现代样本从各种古代祖先群体继承的祖先比例,使我们能够重建历史事件,如人口迁移。我们可以进一步利用ARG拓扑结构来检测自然选择,通过定位基因组中与某些个体或古代群体异常共享的区域。寻找处于正选择或负选择下的区域,特别是具有已知生物功能的区域,在医疗保健相关应用中可能特别有用,例如,这些区域已被用于确定制药环境中的药物靶标。最后,变异的表型影响可以通过测试来自某些群体的祖先是否与某些表型更密切相关来评估。该项目的第一个目标是建立一个相关性推断算法,该算法可以推断现代和古代样本之间的树的拓扑结构和tmrca,并使用它来重建联合ARG,数据来自英国生物银行和其他来源。跨对样本共享的远程染色体区域为该分析提供了信息,但很难在低覆盖率的古代DNA中检测到,因此我们的算法将需要隐式或显式地模拟单倍型共享,尽管缺乏相位信息,或者存在嘈杂的计算相位。对于这个算法,我们利用了深度学习(DL),它在过去十年中改变了许多科学领域。群体遗传学传统上专注于开发复杂的参数模型,并没有明显受益于DL的进步。测序数据具有空间结构,因此来自多个样本的序列可以堆叠形成图像并使用计算机视觉方法(例如卷积神经网络)进行分析,并根据样本顺序无关的事实进行调整(即需要可交换网络)。由于我们已经获得了现代样本的ARG,我们的目标是探索使用基于注意力和图形的方法来提取ARG信息,这将有助于推断现代和古代样本之间的祖先。总的来说,我们期望作出两项主要贡献。首先是算法开发,允许使用深度学习重建现代和古代DNA样本的联合系谱树,解决后者的质量问题。第二种是联合ARG推理,使用真实的英国生物银行和古代数据。然后,我们的目标是分析该ARG,以回答与自然选择和祖先来自某些古代群体的表型影响有关的问题。
英文摘要
At every genomic position, two individuals are connected through genealogical relationships that lead to a common ancestor. The chronological distance from the individuals to this ancestor is termed time to the most recent common ancestor (TMRCA). This can be generalized to a set of individuals by representing their genealogical relationships by a tree. Moving along the genome, the topology of the trees can change as the genome is broken up by recombination during meiosis. Hence, the evolutionary history of a set of samples can be compactly represented by a graph, called the ancestral recombination graph (ARG), comprised by the individual trees spanning different chunks of the genome. There are multiple computationally intensive methods to reconstruct the ARG from high-quality sequencing data of modern DNA samples. Ancient samples are of degraded quality due to environmental conditions and contamination; usually they can only be sequenced at very low coverage making the task of incorporating them into the ARG of a set of modern samples challenging. Reconstructing an accurate joint ARG between modern and ancient DNA samples has multiple potential applications that we aim to explore. It can be used for ancestry inference, by recovering the ancestry proportion a modern sample inherits from various ancient ancestral groups, enabling us to reconstruct historical events such as population migrations. We can further exploit ARG topology to detect natural selection, by locating regions of the genome that are unusually shared from certain individuals or ancient groups . Finding regions under positive or negative selection, particularly with known biological functionality, can be especially useful in healthcare-related applications and such regions have, for example, been leveraged to determine drug targets in pharmaceutical settings. Finally, the phenotypic impact of variants can be evaluated by testing whether ancestry from certain groups is more closely related to certain phenotypes.The project's first goal is to build a relatedness inference algorithm that can infer tree topology and TMRCAs between modern and ancient samples and use it to reconstruct a joint ARG with data from the UK BioBank and other sources. Long-range chromosomal regions that are shared across pairs of samples are informative for this analysis but hard to detect in low coverage ancient DNA, so our algorithm will need to to implicitly or explicitly model haplotype sharing despite the lack of phasing information, or in the presence of noisy computational phasing. For this algorithm, we leverage Deep Learning (DL), which has transformed many scientific fields in the past decade. Population genetics has traditionally focused on developing complex parametric models and has not yet significantly benefited from DL advances. Sequencing data has a spatial structure, so sequences from multiple samples can be stacked to form an image and analysed using computer vision approaches (e.g. Convolutional Neural Networks),adjusting for the fact that sample order is irrelevant (i.e. require exchangeable networks). As we already have access to an ARG for modern samples, our aim is to explore the use of attention- and graph-based methods to extract ARG information that will help infer ancestry between modern and ancient samples.Overall, we expect to make two main contributions. The first is algorithmic development that will allow the use of DL for reconstructing joint genealogical trees for modern and ancient DNA samples, tackling quality issues for the latter. The second is joint ARG inference using real UK Biobank and ancient data. We then aim to analyse this ARG to answer questions relating to natural selection and phenotypic impact of having ancestry from certain ancient groups.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金