课题基金 / 基金详情

CAREER: BCSP: Methods for analyzing sequencing data from repetitive genomes

CAREER: BCSP: Methods for analyzing sequencing data from repetitive genomes
职业:BCSP:分析重复基因组测序数据的方法
批准号:
1349906
负责人:
Benjamin Langmead
金额:
$53.59万
依托单位:
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-05-15 至 2019-04-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
来自DNA测序仪的数据日益推动了我们对生物系统如何工作的理解。在过去的几年里,测序技术有了很大的进步,但测序仪产生的数据集很笨拙,很难解释。当被研究的基因组包含许多重复的DNA片段时尤其如此,就像大多数哺乳动物和植物的情况一样。该项目的目标是开发改进的计算和统计方法来分析DNA测序数据,为研究具有重复基因组的生物体的科学家提供更快、更准确和更具解释性的结果。这些方法将作为开放源码软件工具向研究界免费提供。一个成功的项目将导致这些工具在生物研究界得到广泛采用。重复序列与细胞调控过程有关,并与人类疾病有关。综合教育计划还寻求通过向计算机科学专业的学生传授在大数据基因组学时代制作有用的基因组学软件所需的一整套技能,来改进分析测序数据的软件。PI将开发一个计算生物学的本科辅修课程,一个涵盖分析大型测序数据集的方法的研究生课程,以及一个竞争性的课程项目。成功的努力将导致更多训练有素的计算机科学家加入计算生物学和基因组的研究,并为其做出贡献。植物、哺乳动物和其他高等真核生物的基因组包含许多重复的DNA序列。例如,80%的玉米基因组被重复的DNA片段所覆盖。与此同时,计算工具通常将DNA建模为字符串。这有好处;它允许这些工具借用为其他字符串(如书籍和网页)开发的方法,并将其应用于DNA。但对于重复的基因组,串提取无法捕捉到通过进化相互关联的重复DNA序列的普遍存在。PI提出了一个广泛的研究议程,基于这样的想法,即分析来自重复基因组的测序数据需要特殊的、重复感知的计算方法。第一个项目探索了将序列读数与重复家族进行比对的准确和高效的方法。PI提出了利用对齐问题之间的相似性来产生比当前方法更准确的解决方案的方法。第二个项目探索用于预测由读对齐器报告的对齐是正确的概率的新方法,即,对齐器正确地识别了读取器的起始点。下游分析工具使用此数量来权衡它们对从比对中得出的证据的置信度。但准确估计这一数量是困难的,目前还没有广泛适用的方法可用。PI提出了一种串联模拟方法,由此模拟真实数据集的属性的模拟器可以提供训练样本,从而使我们能够准确地预测真实数据的这些概率。这些方法解决了日常常见基因组分析中的主要缺陷,这些缺陷因重复的DNA而变得缓慢和不准确。PI还将进行一套综合的课程建设和推广工作。这些课程的目的是让更多的学生更早地注意到计算生物学,并为研究生和高年级本科生提供强大的计算生物学课程。具体地说,PI将在约翰·霍普金斯大学开发和实施计算生物学本科辅修课程。其次,PI将设计一门新的研究生水平的课程,涵盖分析非常大的序列数据集合的现代方法。最后,PI将开发一个名为大序列数据五项的竞争性项目,测试学生在并行计算机系统上设计可扩展、可用的基因组分析工具的能力。
英文摘要
Our understanding of how biological systems work is increasingly fueled by data from DNA sequencers. Sequencing has improved dramatically over the past several years, but the datasets produced by sequencers are unwieldy and difficult to interpret. This is especially true when the genome being studied contains many repeated stretches of DNA, as is the case for most mammals and plants. The goal of this project is to develop improved computational and statistical methods for analyzing DNA sequencing data, providing faster, more accurate, and more interpretable results to scientists studying organisms with repetitive genomes. These methods will be implemented as open source software tools made freely available to the research community. A successful project will result in these tools being widely adopted in the biological research community. Repetitive sequences are implicated in cellular regulation processes and associated with human disease. The integrated education plan also seeks to improve software for analyzing sequencing data by teaching computer science students the complete set of skills needed to make usable genomics software in the era of big data genomics. The PI will develop an undergraduate minor in computational biology, a graduate class covering methods for analyzing large sequencing datasets, and a competitive class project. A successful effort will result in more trained computer scientists joining and contributing to the study of computational biology and genomics.The genomes of plants, mammals and other higher eukaryotes contain many repeated DNA sequences. 80% of the maize genome, for example, is covered by repetitive stretches of DNA. At the same time, computational tools typically model DNA as a string. This has advantages; it allows these tools to borrow methods developed for other strings, such as books and web pages, and apply them to DNA. But for repetitive genomes, the string abstraction fails to capture the prevalence of repeated DNA sequences related to each other through evolution. The PI proposes a broad research agenda based on the idea that analyzing sequencing data derived from repetitive genomes requires special, repeat-aware computational methods. The first project explores accurate and efficient methods for aligning sequence reads to repeat families. The PI proposes methods that exploit similarities between alignment problems to yield solutions that are more accurate than current approaches. The second project explores novel methods for predicting the probability that an alignment reported by a read aligner is correct, i.e., that the aligner correctly identified the read's point of origin. Downstream analysis tools use this quantity to weigh their confidence in evidence derived from the alignment. But estimating this quantity accurately is difficult, and there are no widely applicable approaches available now. The PI proposes a tandem simulation approach, whereby a simulator mimicking properties of a real dataset can provide training examples that in turn allows us to accurately predict these probabilities for real data. These methods address major deficiencies in everyday common genomics analyses, which are made slower and less accurate by repetitive DNA.The PI will also conduct an integrated set of curriculum building and outreach efforts. These have the goal of bringing computational biology to the attention of more students earlier in their training, and to provide graduate and upper undergraduate students with a strong computational biology curriculum. Specifically, the PI will develop and implement an undergraduate minor in computational biology at Johns Hopkins University. Second, the PI will design a new graduate-level course covering contemporary methods for analyzing very large collections of sequence data. Finally, the PI will develop a competitive project called the Big Sequence Data Pentathlon that tests students' ability to design scalable, usable genomics analysis tools on a parallel computer system.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
内蒙古自治区神经型布氏杆菌病临床特点、BCSP31基因扩增测序及流行病学调查
  • 批准号:
    82160248
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    34万元
  • 批准年份:
    2021
  • 负责人:
    赵世刚
  • 依托单位: