Analysis of genomics datasets at a massive scale
Analysis of genomics datasets at a massive scale
批准号:
9978590
负责人:
Kasper Daniel Hansen
金额:
$46.32万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2022-07-31
关键词:
AddressAllelesArchivesBioconductorBiologyCatalogsCloud ComputingCommunitiesComputer softwareComputing MethodologiesDataDiseaseEventExcisionExonsGene ExpressionGene Expression ProfileGene Expression ProfilingGenesGenetic TranscriptionGenetic VariationGenomeGenotype-Tissue Expression ProjectGoalsHumanHuman BiologyImpact evaluationLeadLicensingLightLinear ModelsMapsMeasuresMetadataMethodsModelingMorphologic artifactsNucleotidesPathogenicityPatternProceduresProcessQuality ControlRNARNA SplicingRegulationResearchResolutionResourcesSample SizeSamplingSpliced GenesStatistical MethodsSystematic BiasThe Cancer Genome AtlasTissuesTranscriptUpdateVariantVisualizationWorkbasecell typedata resourcegenome-widegenomic datahuman RNA sequencinghuman diseaseinsightopen sourcereconstructionstatisticstranscriptometranscriptome sequencing
中文摘要
点击翻译按钮获取中文摘要
英文摘要
Understanding both normal and pathogenic patterns of human gene expression can help shed light on the biology
of human disease. Thousands of studies have now been undertaken measuring gene expression in different
tissues and diseases. By aggregating and analyzing all available human RNA-sequencing data using a high
powered computational and statistical framework, we will provide a transformative resource for characterizing
human gene expression patterns including rare transcriptional events, cellular networks, and genetic variation.
In Aim 1 we propose to uniformly process all publicly available human transcriptome sequencing data and collect it
into a publicly available resource called the Transcriptome Aggregation Resource (TAR); at least 150,000 samples
will be processed using cloud computing. This resource will contain single-base resolution maps of expression,
de novo mapped exon-exon splice junctions and allele specific expression across a set of common variations.
We will supplement the expression data with cleaned and predicted metadata. In Aim 2 we will develop statistical
and computational methods necessary to fully realize the potential of this resource. Specifically we will remove
unwanted variation at scale and develop mixture models to summarize the large data resource at the gene,
junction and single base levels. In Aim 3 we will analyze this resource to address fundamental questions in
expression biology, include a systematic study of expression outliers and allele specific expression at the gene,
junction and single base resolution. We will infer well-powered co-expression networks over both expressed
genes and splicing patterns.
This work will contribute significantly to our understanding of gene expression by analyzing genomics data at a
massive scale.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Data Science Tools to Increase Insight in Genomics Data
-
批准号:10623491
-
项目类别:
-
资助金额:$46.23万
-
财政年份:2023
-
负责人:Kasper Daniel Hansen
-
依托单位:
Analysis of genomics datasets at a massive scale
-
批准号:10223343
-
项目类别:
-
资助金额:$46.17万
-
财政年份:2017
-
负责人:Kasper Daniel Hansen
-
依托单位:
海外基金