课题基金 / 基金详情

A Software Framework for Exploring 1,000 Genomes of African Descent

A Software Framework for Exploring 1,000 Genomes of African Descent
用于探索 1,000 个非洲人后裔基因组的​​软件框架
批准号:
9301024
负责人:
Kathleen C Barnes
金额:
$44.74万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2015
资助国家:
美国
项目状态:
已结题
起止时间:
2015-07-01 至 2019-06-30

项目摘要

项目成果

Kathleen C Barnes的其他基金

相似基金

相关文献

中文摘要
翻译
 描述(由申请人提供):我们建议创建新的软件和分析方法,旨在使探索一个独特的数据集成为可能,即美洲非洲裔人口哮喘联盟(CAAPA)测序的1,004个基因组。这个数据集的大小超过130太字节,目前无法用基于比对的工具进行探索,研究人员只能使用包含单核苷酸变体的小得多的文件。我们建议的软件将使这个数据集和其他类似的数据集可用于实时搜索,这是任何这种规模的基因组数据库都不可能实现的功能。自20世纪90年代初以来,科学家们利用DNA序列数据库研究了一系列问题,包括新基因发现、突变检测、更大结构变异的研究以及进化过程。使用BLAST和类似程序搜索所有已知基因和基因组的能力一直被认为是有能力的,世界各地的序列搜索引擎都提供这种能力。然而,CAAPA数据集的巨大规模使得使用当前工具搜索数据本身是不可能的。人们不能寻找特定的突变,不能提取和重新分析任何特定基因或调控区域的数据,也不能寻找结构变异。更新、更快的下一代序列比对程序,如最初由我们团队开发的Bowtie,允许更快地将NGS读数与基因组进行比对,但即使是这些程序也无法实时搜索CAAPA规模的数据。需要设计和构建不同的体系结构来容纳这些非常大的数据集。CAAPA勘探系统将使用高效的数据库、非常快的存储和快速搜索算法的组合来实现我们的目标。该项目旨在实现几个目标,从而极大地提高CAAPA的价值。首先,这些数据将提供给一个非常大的研究人员社区,他们不仅可以使用它来研究CAAPA人群中哮喘和过敏的遗传学,还可以将这些受试者与其他群体进行比较。数据目前驻留在硬盘上,只有一小部分项目PI可以使用,这种情况限制了其价值。第二,创建一个与DBGaP一致的认证系统,我们将创建一个其他项目可以使用的数据共享模型,这将消除从人类受试者共享基因组数据的一些技术障碍。第三,作为建立数据库的一部分,我们将使用新发布的人类基因组构建(Hg20)重新调用所有SNP,创建一组一致的变体,我们也将通过项目数据库免费共享这些变体。第四,我们将确定所有细菌污染物,包括在样本采集时已知患有血液感染的一组受试者中的细菌。第五,我们将确定CAAPA人群特有的结构变异,然后我们可以探索这些变异与哮喘风险的任何关联。
英文摘要
 DESCRIPTION (provided by applicant): We propose to create new software and analysis methods designed to make possible the exploration of a unique dataset, the 1,004 genomes sequenced by the Consortium on Asthma among African-Ancestry Populations in the Americas (CAAPA). The size of this dataset, over 130 Terabytes, currently prevents it from being explored with alignment-based tools, and researchers instead are limited to using the much smaller files containing single-nucleotide variants. Our proposed software will make this dataset and others like it available for real- time searching, a capability that is not yet possible for any genomic database of this size. Since the early 1990s, scientists have used DNA sequence databases to study a wide range of problems, including novel gene discovery, mutation detection, the investigation of larger structural variants, and evolutionary processes. The ability to search all known genes and genomes using BLAST and similar programs has long been assumed, and sequence search engines throughout the world provide this ability. However, the vast size of the CAAPA dataset makes it impossible to search the data itself using current tools. One cannot look for specific mutations, extract and re-analyze data for any particular gene or regulatory region, or look for structural variants. Newer, fast next-generation sequence alignment programs such as Bowtie, originally developed in our group, allow far faster alignment of NGS reads to the genome, but even these programs cannot search data on the scale of CAAPA in real time. Different architectures need to be designed and built to accommodate these very large datasets. The CAAPA exploration system (CESYS) will use a combination of a highly efficient database, very fast storage, and fast search algorithms to achieve our goals. This project aims to accomplish several goals that will dramatically enhance the value of CAAPA. First, the data will be made available to a very large community of researchers, who can use it not only to study the genetics of asthma and allergy in the CAAPA populations, but also to compare these subjects to other groups. The data currently resides on hard drives and is available only to a small number of the project's PIs, a situation that limits its value. Second, b creating an authentication system consistent with dbGaP, we will create a data sharing model that other projects can use and that will remove some of the technical barriers to sharing genome data from human subjects. Third, as part of building the database, we will re-call all the SNPs using the newly released human genome build (hg20), creating a consistent set of variants that we will also share freely through the project database. Fourth, we will identify all bacterial contaminants, including those in a subset of subjects known to have bloodstream infections at the time of sample collection. Fifth, we will identify structural variants unique to he CAAPA population, which we can then explore for any association with the risk of asthma.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
PRIDE Academy: Impact of Ancestry and Gender to omics of lung diseases
  • 批准号:
    10077882
  • 项目类别:
  • 资助金额:
    $46.98万
  • 财政年份:
    2019
  • 负责人:
    Kathleen C Barnes
  • 依托单位:
PRIDE Academy: Impact of Ancestry and Gender to omics of lung diseases
  • 批准号:
    10378108
  • 项目类别:
  • 资助金额:
    $46.98万
  • 财政年份:
    2019
  • 负责人:
    Kathleen C Barnes
  • 依托单位:
Multi-omic studies of asthma severity in an African ancestry population
  • 批准号:
    10094181
  • 项目类别:
  • 资助金额:
    $68.95万
  • 财政年份:
    2018
  • 负责人:
    Kathleen C Barnes
  • 依托单位:
Multi-omic studies of asthma severity in an African ancestry population
  • 批准号:
    10331294
  • 项目类别:
  • 资助金额:
    $66.08万
  • 财政年份:
    2018
  • 负责人:
    Kathleen C Barnes
  • 依托单位:
海外基金