课题基金 / 基金详情

项目摘要

项目成果

Xiaohui Xie的其他基金

相似基金

相关文献

中文摘要
翻译
描述(申请人提供):DNA测序已成为许多生物学和医学领域不可或缺的工具。最近在下一代测序(NGS)方面的技术突破使快速和廉价地对数十亿个碱基进行测序成为可能。已经建立了一些基于NGS的工具,包括芯片序列、RNA序列、甲基序列和外显子/全基因组测序,从而为研究疾病、基因组和表观基因组提供了一种全新的方法。基于NGS的方法的广泛使用需要更好和更有效的工具来分析和解释NGS高通量数据。虽然已经开发了一些计算工具,但它们在绘制和研究基因组重复、重复和其他所谓的不可映射区域中的基因组特征方面还不够。在这个项目中,将开发计算算法和软件,将NGS的基因组可及性扩展到这些以前未被研究的区域。这些算法将从一种新的方法开始,将原始读数从NGS映射到参考基因组,然后使用机器学习方法来解决模糊映射的读数,并将集成到芯片序列的综合分析流水线中。更具体地说,研究的三个目标是开发:(1)读映射的数据结构和高效算法,以快速识别所有映射位置。与现有方法不同,本研究的重点是快速识别每个阅读的所有候选位置,而不是一个或几个位置。(2)用于读数分析的机器学习算法,以解决用于芯片序列分析和遗传变异检测的模糊映射读数。这项工作将开发概率模型,通过汇集整个阅读集合的信息来解决模糊映射的阅读。(3)一个综合性的芯片-序列分析流水线,系统地研究位于基因组不可映射区的基因组特征。这些算法将使用公开可用的数据和来自已建立的湿实验室合作者的数据进行测试和改进。除了发现位于重复、重复或其他以前无法获得的区域内的新基因组特征外,这项工作还将为NGS社区提供(A)更快、更准确地绘制短序列阅读图谱的工具,(B)扩大NGS基因组可获得性的一般方法,以及(C)用于NGS数据分析的通用、模块化、开放源码的算法工具箱,(D)在所有公开可用的CHIP-SEQ数据集中全面分析重复区域的蛋白质-DNA相互作用。这项工作是计算机科学家和网络实验室生物学家之间的密切合作,他们正在开发NGS分析来研究生物医学问题。特别是,我们将与Sanford-Burnham医学研究所的Timothy Osborne合作研究参与胆固醇和脂肪酸代谢的调节剂,与加州大学欧文分校的Kyoko Yokomori合作研究粘附素、Nipbl及其在Cornelia de Lange综合征中的作用,以及与加州大学欧文分校的Ken Cho合作研究FoxH1和SchNurri在发育和生长控制中的作用。 与公共卫生相关:DNA测序已成为基础生物医学研究以及发现新疗法和帮助生物医学研究人员了解疾病机制的不可或缺的工具。下一代测序能够以相对较低的成本快速产生数十亿个碱基,对于如何高效和准确地分析大量的序列数据提出了巨大的计算挑战。这项研究的目标是开发开源软件,以提高下一代测序分析工具的效率和准确性,从而让生物医学研究人员充分利用下一代测序来研究生物学和疾病。
英文摘要
DESCRIPTION (provided by applicant): DNA sequencing has become an indispensable tool in many areas of biology and medicine. Recent techno- logical breakthroughs in next-generation sequencing (NGS) have made it possible to sequence billions of bases quickly and cheaply. A number of NGS-based tools have been created, including ChIP-seq, RNA-seq, Methyl- seq and exon/whole-genome sequencing, enabling a fundamentally new way of studying diseases, genomes and epigenomes. The widespread use of NGS-based methods calls for better and more efficient tools for the analysis and interpretation of the NGS high-throughput data. Although a number of computational tools have been devel- oped, they are insufficient in mapping and studying genome features located within repeat, duplicated and other so-called unmappable regions of genomes. In this project, computational algorithms and software that expand genomic accessibility of NGS to these previously understudied regions will be developed. The algorithms will begin with a new way of mapping raw reads from NGS to the reference genome, followed by a machine learning method to resolve ambiguously mapped reads, and will be integrated into a comprehen- sive analysis pipeline for ChIP-seq. More specifically, the three aims of the research are to develop: (1) Data structures and efficient algorithms for read mapping to rapidly identify all mapping locations. Unlike existing methods, the focus of this research is to rapidly identify all candidate locations of each read, instead of one or only a few locations. (2) Machine learning algorithms for read analysis to resolve ambiguously mapped reads for both ChIP-seq analysis and genetic variation detection. This work will develop probabilistic models to resolve ambiguously mapped reads by pooling information from the entire collection of reads. (3) A comprehensive ChIP- seq analysis pipeline to systematically study genomic features located within unmappable regions of genomes. These algorithms will be tested and refined using both publicly available data and data from established wet-lab collaborators. In addition to discovering new genomic features located within repeat, duplicated or other previously unac- cessible regions, this work will provide the NGS community with (a) a faster and more accurate tool for mapping short sequence reads, (b) a general methodology for expanding genomic accessibility of NGS, and (c) a versatile, modular, open-source toolbox of algorithms for NGS data analysis, (d) a comprehensive analysis of protein-DNA interactions in repeat regions in all publicly available ChIP-seq datasets. This work is a close collaboration between computer scientists and web-lab biologists who are developing NGS assays to study biomedical problems. In particular, we will collaborate with Timothy Osborne of Sanford- Burnham Medical Research Institute to study regulators involved in cholesterol and fatty acid metabolism, with Kyoko Yokomori of UC Irvine to study Cohesin, Nipbl and their roles in Cornelia de Lange syndrome, and Ken Cho of UC Irvine to study the roles of FoxH1 and Schnurri in development and growth control. PUBLIC HEALTH RELEVANCE: DNA-sequencing has become an indispensable tool for basic biomedical research as well as for discovering new treatments and helping biomedical researchers understand disease mechanisms. Next-generation sequencing, which enables rapid generation of billions of bases at relatively low cost, poses a significant computational challenge on how to analyze the large amount of sequence data efficiently and accurately. The goal of this research is to develop open-source software to improve both the efficiency and accuracy of the next-generation sequencing analysis tools, and thereby allowing biomedical researchers to take full advantage of next-generation sequencing to study biology and disease.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Machine learning methods to increase genomic accessibility by next-gen sequencing
  • 批准号:
    8683213
  • 项目类别:
  • 资助金额:
    $22.13万
  • 财政年份:
    2012
  • 负责人:
    Xiaohui Xie
  • 依托单位:
Machine learning methods to increase genomic accessibility by next-gen sequencing
  • 批准号:
    8518436
  • 项目类别:
  • 资助金额:
    $22.06万
  • 财政年份:
    2012
  • 负责人:
    Xiaohui Xie
  • 依托单位:
海外基金