Machine learning methods to increase genomic accessibility by next-gen sequencing
Machine learning methods to increase genomic accessibility by next-gen sequencing
批准号:
8683213
负责人:
Xiaohui Xie
金额:
$22.13万
依托单位国家:
美国
项目类别:
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-08-01 至 2016-06-30
关键词:
AlgorithmsAnusAreaBindingBiologicalBiological AssayBiologyBiomedical ResearchBruck-de Lange syndromeChIP-seqCholesterolChromatinCollaborationsCollectionCommunitiesComputational algorithmComputer softwareComputersDNA SequenceDNA-Protein InteractionDataData AnalysesData SetDetectionDiseaseExonsFacioscapulohumeralFoundationsGenerationsGenetic VariationGenomeGenomicsGoalsGrowth and Development functionInternetLocationMachine LearningMapsMedical ResearchMedicineMethodologyMethodsMuscular DystrophiesProceduresPublishingReadingResearchResearch InstituteResearch PersonnelRoleScientistSequence AnalysisSoftware EngineeringSpeedStatistical ModelsStructureTestingUncertaintyWorkbasecohesincomputerized toolscostepigenomefatty acid metabolismfunctional genomicsgenome sequencinggenome-wideimprovedinsertion/deletion mutationnext generation sequencingnovelopen sourcepublic health relevancetooltranscription factortranscriptome sequencingxenopus development
中文摘要
点击翻译按钮获取中文摘要
英文摘要
DESCRIPTION (provided by applicant): DNA sequencing has become an indispensable tool in many areas of biology and medicine. Recent techno- logical breakthroughs in next-generation sequencing (NGS) have made it possible to sequence billions of bases quickly and cheaply. A number of NGS-based tools have been created, including ChIP-seq, RNA-seq, Methyl- seq and exon/whole-genome sequencing, enabling a fundamentally new way of studying diseases, genomes and epigenomes. The widespread use of NGS-based methods calls for better and more efficient tools for the analysis and interpretation of the NGS high-throughput data. Although a number of computational tools have been devel- oped, they are insufficient in mapping and studying genome features located within repeat, duplicated and other so-called unmappable regions of genomes. In this project, computational algorithms and software that expand genomic accessibility of NGS to these previously understudied regions will be developed. The algorithms will begin with a new way of mapping raw reads from NGS to the reference genome, followed by a machine learning method to resolve ambiguously mapped reads, and will be integrated into a comprehen- sive analysis pipeline for ChIP-seq. More specifically, the three aims of the research are to develop: (1) Data structures and efficient algorithms for read mapping to rapidly identify all mapping locations. Unlike existing methods, the focus of this research is to rapidly identify all candidate locations of each read, instead of one or only a few locations. (2) Machine learning algorithms for read analysis to resolve ambiguously mapped reads for both ChIP-seq analysis and genetic variation detection. This work will develop probabilistic models to resolve ambiguously mapped reads by pooling information from the entire collection of reads. (3) A comprehensive ChIP- seq analysis pipeline to systematically study genomic features located within unmappable regions of genomes. These algorithms will be tested and refined using both publicly available data and data from established wet-lab collaborators. In addition to discovering new genomic features located within repeat, duplicated or other previously unac- cessible regions, this work will provide the NGS community with (a) a faster and more accurate tool for mapping short sequence reads, (b) a general methodology for expanding genomic accessibility of NGS, and (c) a versatile, modular, open-source toolbox of algorithms for NGS data analysis, (d) a comprehensive analysis of protein-DNA interactions in repeat regions in all publicly available ChIP-seq datasets. This work is a close collaboration between computer scientists and web-lab biologists who are developing NGS assays to study biomedical problems. In particular, we will collaborate with Timothy Osborne of Sanford- Burnham Medical Research Institute to study regulators involved in cholesterol and fatty acid metabolism, with Kyoko Yokomori of UC Irvine to study Cohesin, Nipbl and their roles in Cornelia de Lange syndrome, and Ken Cho of UC Irvine to study the roles of FoxH1 and Schnurri in development and growth control.
期刊论文(4)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1186/1471-2105-15-42
发表时间:
2014-02-05
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Kim J, Li C, Xie X]
通讯作者:
Xie X
DOI:
10.1186/1471-2164-16-s2-s1
发表时间:
2015
期刊:
BMC genomics
影响因子:
4.4
作者:
[Li Y, Xie X]
通讯作者:
Xie X
DOI:
10.1186/1471-2105-14-s5-s11
发表时间:
2013
期刊:
BMC bioinformatics
影响因子:
3
作者:
[Li Y, Xie X]
通讯作者:
Xie X
Machine learning methods to increase genomic accessibility by next-gen sequencing
-
批准号:8350385
-
项目类别:
-
资助金额:$22.0万
-
财政年份:2012
-
负责人:Xiaohui Xie
-
依托单位:
Machine learning methods to increase genomic accessibility by next-gen sequencing
-
批准号:8518436
-
项目类别:
-
资助金额:$22.06万
-
财政年份:2012
-
负责人:Xiaohui Xie
-
依托单位:
海外基金