课题基金 / 基金详情

Genome analysis based on the integration of DNA sequence and shape

Genome analysis based on the integration of DNA sequence and shape
基于DNA序列和形状整合的基因组分析
批准号:
8795204
负责人:
Remo Rohs
金额:
$30.47万
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-02-01 至 2018-01-31

项目摘要

项目成果

Remo Rohs的其他基金

相似基金

相关文献

中文摘要
翻译
描述(由申请人提供):目前的基因组分析技术主要基于由字母A、C、G和t组成的一维DNA序列。然而,蛋白质将DNA识别为三维(3D)物体。DNA形状在单核苷酸分辨率上的细微差别在转录因子(tf)的结合特异性中起着至关重要的作用,包括那些参与胚胎发育和人类癌症的转录因子。该项目涉及开发一系列基因组分析工具,通过整合来自DNA序列和DNA 3D结构(或“DNA形状”)的信息。这些新工具的基础是在基因组尺度上预测局部DNA形状的多种特征的高通量(HT)方法。数据将通过web服务器接口以UCSC基因组浏览器跟踪格式提供给社区。这些工具将使用户能够分析任何数量或长度的DNA序列的形状,包括整个基因组和DNA甲基化的影响。HT形状预测将基于x射线晶体学、核磁共振光谱和羟基自由基裂解数据进行验证。这些预测将与ENCODE项目ORChID相结合,该项目通过羟基自由基裂解实验推断DNA小槽的几何形状。HT方法将用于研究同源tf如何在体内选择不同的靶点,尽管它们共享核心结合基序或在体外具有相似的结合特性。为了研究这个问题,我们将研究侧翼序列对TF结合位点(TFBSs)多个结构特征的影响。本研究的最初重点将是同位结构域和基本螺旋-环-螺旋(bHLH) tf。其他蛋白家族随后将被纳入并用于构建一个全面的TFBS数据库,该数据库提供来自JASPAR和其他motif数据库的结合motif的形状特征。单核苷酸多态性(snp)的结构效应也将进行分析。一些snp与有害的功能有关,而另一些则没有明显的影响。HT形状预测方法将用于基于DNA形状预测非编码区snp的功能。我们将把snp对DNA结构的定量影响与表达数量性状位点(quantitative trait loci, eqtl)和全基因组关联研究(genome-wide association study, GWAS)信号联系起来,开发一种预测snp功能影响的工具。HT形状预测方法将用于设计具有不同AT/GC含量但形状相似的DNA序列。序列和形状对结合的相对贡献将通过多元线性回归(MLR)和支持向量回归(SVR)等分析模型进行检验。对于序列和形状集成被证明是有利的系统,将基于扩展字母表开发新的motif查找工具,该字母表将序列与信息结构特征相结合,通过机器学习和特征选择方法进行选择。序列+形状的motif将通过motif扫描进行测试,与仅序列的motif相比,并集成到MEME套件中。这种序列形状整合的目标是提高在基因组中发现体内TFBSs的准确性。
英文摘要
DESCRIPTION (provided by applicant): Current techniques for genome analysis are mainly based on the one-dimensional DNA sequence, comprised of the letters A, C, G, and T. However, proteins recognize DNA as a three-dimensional (3D) object. Nuances in DNA shape at single nucleotide resolution play a crucial role in the binding specificity of transcription facors (TFs), including those involved in embryonic development and human cancer. This project involves the development of a battery of tools for genome analysis, through the integration of information derived from the DNA sequence and the 3D structure of DNA, or "DNA shape". The basis for these novel tools is a high- throughput (HT) method for the prediction of multiple features of local DNA shape at the genomic scale. Data will be made available to the community in the UCSC Genome Browser track format through a web server interface. These tools will enable users to analyze the shape of any number or length of DNA sequences, including whole genomes and the effect of DNA methylation. HT shape predictions will be validated based on X-ray crystallography, NMR spectroscopy, and hydroxyl radical cleavage data. Predictions will be combined with ORChID, an ENCODE project that infers DNA minor groove geometry from hydroxyl radical cleavage experiments. The HT method will be used to study how paralogous TFs select different target sites in vivo despite sharing core-binding motifs or having similar binding properties in vitro. To study this question, we will investigate the effect of flanking sequences on multiple structural features of TF binding sites (TFBSs). The initial focus of this study will be homeodomains and basic helix-loop-helix (bHLH) TFs. Other protein families will later be included and used to construct a comprehensive TFBS database that provides shape features for binding motifs derived from JASPAR and other motif databases. Structural effects of single nucleotide polymorphisms (SNPs) will also be analyzed. Some SNPs are associated with deleterious functions, whereas others have no apparent effect. The HT shape prediction method will be used to predict the function of SNPs in non-coding regions based on DNA shape. We will correlate quantitative effects of SNPs on DNA structure with expression quantitative trait loci (eQTLs) and genome-wide association study (GWAS) signals, to develop a predictive tool for the functional effect of SNPs. The HT shape prediction approach will be used to design DNA sequences with different AT/GC contents but similar shapes. The relative contributions of sequence and shape to binding will be tested with analytic models including multiple linear regression (MLR) and support vector regression (SVR). For systems in which the integration of sequence and shape proves advantageous, novel motif finding tools will be developed based on an extended alphabet that combines sequence with informative structural features, selected by machine learning and feature selection approaches. Sequence+shape motifs will be tested by motif scanning, compared to sequence-only motifs, and integrated into the MEME Suite. The goal of this sequence-shape integration is to increase the accuracy of finding in vivo TFBSs in the genome.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Quantitative Modeling of Transcription Factor-DNA Binding
Quantitative Modeling of Transcription Factor-DNA Binding
Quantitative Modeling of Transcription Factor-DNA Binding
Quantitative Modeling of Transcription Factor-DNA Binding
海外基金