High dimensional statistical data integration for studying regulatory variation
High dimensional statistical data integration for studying regulatory variation
批准号:
9344668
负责人:
Sunduz Keles
金额:
$32.5万
依托单位国家:
美国
项目类别:
财政年份:
2007
资助国家:
美国
项目状态:
已结题
起止时间:
2007-04-26 至 2020-06-30
关键词:
AddressBindingBinding SitesBioconductorBiologic CharacteristicCellsChIP-seqChromatinCollectionCommunitiesComputer AnalysisComputer softwareDNADNA MethylationDNA-Protein InteractionDataData SetData SourcesDerivation procedureDevelopmentDiagnosisDiseaseElementsGalaxyGenerationsGeneticGenetic TranscriptionGenomeGenomicsGenotypeHistonesHumanIndividualInternationalInvestigationJointsKnowledgeLettersLocationMapsMessenger RNAMethodologyMethodsPhenotypeProtein IsoformsRNARNA analysisRNA-Binding ProteinsRNA-Protein InteractionRegulationRepetitive SequenceResearchResearch PersonnelResourcesSamplingSourceStatistical Data InterpretationStatistical MethodsStatistical ModelsTechnologyTissuesTrainingUntranslated RNAValidationVariantbasecell typecrosslinking and immunoprecipitation sequencingdata integrationepigenomeexperienceexperimental studygenetic variantgenome wide association studygenome-widegenomic datagenomic profileshigh dimensionalityhigh throughput technologyhistone modificationhuman diseaseimprovedinnovationnext generation sequencingnovelprotein profilingreference genomesimulationtooltraittranscription factorwhole genome
中文摘要
项目摘要
下一代测序(NGS)技术彻底改变了遗传学和基因组学领域。
基因组学,允许快速和廉价的测序数十亿个碱基。虽然
每种数据类型的基本分析工具都是丰富的统计方法,
可以整合不同的数据源,以解决关键的、具有挑战性的问题,
缺乏我们建议为关键的、广泛使用的应用开发综合方法
迫切需要可靠的统计整合工具。我们方法的核心是
将多种适当的数据类型与新颖的统计方法有效地结合起来。
首先,尽管迄今为止,大量的蛋白质-DNA相互作用和组蛋白
修改是映射的,系统的方法,允许用户查询这些数据,
缺乏可验证的假设。第二,与生成
(epi)基因组图谱,全基因组关联研究(GWAS)已经成功
在识别疾病和性状相关的遗传变异(GV)。然而,我们的能力,
确定因果变异,并阐明基因型影响的机制
表型受到重大障碍的阻碍。第三,尽管效用的含义是,
映射到参考基因组上的多个位置(多读段)已经很好地
建立了一些NGS应用程序,如RNA-seq和ChIP-seq,
用于询问RNA结合蛋白的新兴数据类型CLIP-seq的方法
依赖于仅使用唯一映射到参考基因组的读段(uni-reads),
不可靠的推论我们计划通过开发(1)快速
用于多个ChIP-seq的联合分析的可扩展的综合统计方法
数据集,使个人数据水平的推断和联合效应的识别;
(2)一个统计分析框架,用于将GWAS结果与日益增长的
(3)一个整合的多读段
通过CLIP-seq研究RNA-蛋白质相互作用的映射框架
实验这些项目将通过以下方法的结合来完成:
开发、模拟、计算分析和实验验证。方法
将使用ENCODE和REMC的数据集进行开发和评估
as novel新datasets数据集from collaborators合作者.该项目产生的统计资源将
以公开可用的软件传播。总的来说,这些目标将大大
提高研究人员可用的全基因组数据类型的实用性。
英文摘要
Project Summary
Next generation sequencing (NGS) technologies revolutionized the fields of genetics and
genomics by allowing rapid and inexpensive sequencing of billions of bases. Although
basic analysis tools for each individual data type are abundant, statistical methods that
can integrate different sources of data for addressing key, challenging questions are
lacking. We propose to develop integrative methods for critical, widely used, applications
urgently requiring reliable statistical integration tools. At the core of our methods is
effective integration of multiple appropriate data types with novel statistical methods.
First, although, to date, large numbers of protein-DNA interactions and histone
modifications are mapped, systematic methods that allow users to query these data and
generate testable hypotheses are lacking. Second, in parallel to generation of
(epi)genomic profiles, genome-wide association studies (GWAS) have been successful
at identifying disease and trait-associated genetic variants (GVs). However, our ability to
identify causal variants and elucidate the mechanisms by which genotypes influence
phenotypes is hampered by significant obstacles. Third, although the utility of reads that
map to multiple locations on the reference genome (multi-reads) has been well
established for some NGS applications such as RNA-seq and ChIP-seq, all the analysis
methods for the emerging data type CLIP-seq that interrogates RNA binding proteins
rely on using only reads that map uniquely to reference genome (uni-reads) leading to
unreliable inference. We plan to address these critical challenges by developing (1) Fast
and scalable integrative statistical methods for joint analysis of multiple ChIP-seq
datasets to enable both individual data level inference and identification of joint effects;
(2) A statistical analysis framework for integrating GWAS results with the increasing
numbers of genome-wide maps of functional annotations; (3) An integrative multi-read
mapping framework for studying RNA-protein interactions through CLIP-seq
experiments. The projects will be accomplished through a combination of methodological
development, simulation, computational analysis, and experimental validation. Methods
will be developed and evaluated using datasets from the ENCODE and REMC as well
as novel datasets from collaborators. Statistical resources generated from the project will
be disseminated in publicly available software. Collectively, these aims will significantly
improve the utility of genome-wide data types that are available to researchers.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Statistical methods for co-expression network analysis of population-scale scRNA-seq data
-
批准号:10740240
-
项目类别:
-
资助金额:$40.76万
-
财政年份:2023
-
负责人:Sunduz Keles
-
依托单位:
Functionally relevant mapping of human GWAS SNPs on model organisms
-
批准号:10056966
-
项目类别:
-
资助金额:$40.05万
-
财政年份:2020
-
负责人:Sunduz Keles
-
依托单位:
Statistical Power Calculations for ChIP-seq experiments
-
批准号:8284083
-
项目类别:
-
资助金额:$18.41万
-
财政年份:2012
-
负责人:Sunduz Keles
-
依托单位:
High dimensional statistical data modeling and integration for studying regulatory variation
-
批准号:10413927
-
项目类别:
-
资助金额:$37.88万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Analysis Methods and Software for ChIP-seq Data
-
批准号:8785690
-
项目类别:
-
资助金额:$29.8万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Methods for the Analysis of ChlP-chip Data
-
批准号:7253510
-
项目类别:
-
资助金额:$28.24万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Analysis Methods and Software for ChIP-seq Data
-
批准号:8605900
-
项目类别:
-
资助金额:$29.95万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Analysis Methods and Software for ChIP-seq Data
-
批准号:8370723
-
项目类别:
-
资助金额:$29.52万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Methods for the Analysis of ChlP-chip Data
-
批准号:7799293
-
项目类别:
-
资助金额:$28.19万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
High dimensional statistical data modeling and integration for studying regulatory variation
-
批准号:10610872
-
项目类别:
-
资助金额:$37.88万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Methods for the Analysis of ChlP-chip Data
-
批准号:7413330
-
项目类别:
-
资助金额:$28.47万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
Statistical Methods for the Analysis of ChlP-chip Data
-
批准号:7616521
-
项目类别:
-
资助金额:$28.47万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
High dimensional statistical data modeling and integration for studying regulatory variation
-
批准号:10213308
-
项目类别:
-
资助金额:$36.46万
-
财政年份:2007
-
负责人:Sunduz Keles
-
依托单位:
国内基金
海外基金
登录
查看更多内容
帽结合蛋白(cap binding protein)调控乙烯信号转导的分子机制
-
批准号:32170319
-
项目类别:面上项目
-
资助金额:58.00万元
-
批准年份:2021
-
负责人:董春海
-
依托单位:
帽结合蛋白(cap binding protein)调控乙烯信号转导的分子机制
-
批准号:--
-
项目类别:--
-
资助金额:58万元
-
批准年份:2021
-
负责人:董春海
-
依托单位:
ID1 (Inhibitor of DNA binding 1) 在口蹄疫病毒感染中作用机制的研究
-
批准号:31672538
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2016
-
负责人:孙跃峰
-
依托单位:
番茄EIN3-binding F-box蛋白2超表达诱导单性结实和果实成熟异常的机制研究
-
批准号:31372080
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2013
-
负责人:杨迎伍
-
依托单位:
P53 binding protein 1 调控乳腺癌进展转移及化疗敏感性的机制研究
-
批准号:81172529
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2011
-
负责人:杨其峰
-
依托单位:
DBP(Vitamin D Binding Protein)在多发性硬化中的作用和相关机制的蛋白质组学研究
-
批准号:81070952
-
项目类别:面上项目
-
资助金额:35.0万元
-
批准年份:2010
-
负责人:刘师莲
-
依托单位:
研究EB1(End-Binding protein 1)的癌基因特性及作用机制
-
批准号:30672361
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2006
-
负责人:徐宁志
-
依托单位: