Deep tensor genomic imputation
Deep tensor genomic imputation
批准号:
10557916
负责人:
William Stafford Noble
金额:
$38.38万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
未结题
起止时间:
2021-02-01 至 2025-01-31
关键词:
3-DimensionalArchitectureAvocadoAwarenessBindingBiochemicalBiological AssayCell LineCellsCellular AssayChromatinChromatin Interaction Analysis by Paired-End Tag SequencingCollectionComplexComputer softwareCouplesDNADNA MethylationDNA SequenceDNA sequencingDataData SetDiseaseEpigenetic ProcessEvaluationFutureGene ExpressionGenetic TranscriptionGenetic VariationGenomeGenomicsGenotype-Tissue Expression ProjectGoalsHealthHi-CHigh-Throughput Nucleotide SequencingHumanIndividualInternetInvestigationJointsLearningMachine LearningMeasurementMeasuresMethodologyMethodsMethylationModelingMolecularPatternPositioning AttributeProcessPropertyRegulatory ElementResolutionResourcesSamplingScientistSystemTechniquesTechnologyTissuesTrainingUnited States National Institutes of HealthUntranslated RNAValidationVariantWorkbiological systemscell typecostdata standardsdeep neural networkexperimental studygenetic manipulationgenome-widegenomic datagenomic locushistone modificationimprovedin silicoinventionlarge datasetsmodel organismnext generationopen sourcepredictive modelingsyntaxtranscription factorweb portal
中文摘要
项目摘要/摘要
高通量测序分析使科学家能够测量转录因子等生化特性
结合、组蛋白莫迪fi阳离子和基因在几乎任何细胞系或原生组织中的表达(“生物样本”)。
不幸的是,在每个生物样本中测量所有可能的生化特性是不可行的,这既是因为
样品供应有限,而且成本过高。我们之前已经开发了一种状态-
一种称为Avocado的最先进的推算方法,它可以fill在这样的数据集中的洞。牛油果情侣张量
使用深度神经网络的因式分解。该方法可扩展到大型数据集,并提供更准确的
与ChromImpute或PREDICTD等竞争方法不同。我们已经应用了牛油果
系统地提供给NIH ENCODE数据集,并通过ENCODE网络公开提供这些估算
波塔尔。
在这里,我们建议在四个重要方面扩展牛油果。首先,我们将扩展Avocado以处理单细胞
数据集,从而有效地将每个单细胞实验转变为测量多个
每个单元格的并行属性。其次,我们将扩展Avocado以使用Hi-C等数据,这些数据衡量
DNA的三维特性。扩展包括转换Avocado的3D张量(生物样本分析
基因组位置)到具有两个基因组位置轴的4D张量。这一延期将适用于各种各样的人。
数据类型,包括各种类型的Hi-C数据、Sprite、GAM、Chia-PET和Plac-Seq。第三,我们将
增强牛油果使用可识别变异的基因组序列以实现调控基因的高分辨率归属
ProfiLes.最后,我们将利用输入的数据来推断顺式调控序列注释和分子
调控的非编码变体在最全面的细胞环境集合之一中的影响。
这个项目生产的所有软件都将是开源的,所有输入的数据和潜在的
因子分解将通过与NIH 4D核基因组相关的门户网站公开提供
对联合体进行编码,为这些数据集的用户提供宝贵的公共资源。
英文摘要
Project Summary/Abstract
High-throughput sequencing assays allow scientists to measure biochemical properties like transcription factor
binding, histone modifications, and gene expression in nearly any cell line or primary tissue (“biosample”).
Unfortunately, measuring all possible biochemical properties in every biosample is infeasible, both because of
limited sample availability and because the cost would be prohibitive. We have previously developed a state-of-
the-art imputation method, called Avocado, that can fill in the holes in such data sets. Avocado couples tensor
factorization with a deep neural network. The method is scalable to large data sets and provides more accurate
imputations than competing methods such as ChromImpute or PREDICTD. We have already applied Avocado
systematically to the NIH ENCODE data set and made the imputations publicly available via the ENCODE web
por tal.
Here, we propose to extend Avocado in four important ways. First, we will extend Avocado to handle single-cell
data sets, thereby effectively turning each single-cell experiment into an in silico co-assay that measures multiple
properties of each cell in parallel. Second, we will extend Avocado to work with data such as Hi-C, which measures
three-dimensional properties of DNA. The extension involves converting Avocado's 3D tensor (biosample assay
genomic position) to a 4D tensor with two genomic position axes. This extension will apply to a wide variety
of data types, including various types of Hi-C data, SPRITE, GAM, ChIA-PET and PLAC-seq. Third, we will
enhance Avocado to use variant aware genomic sequence to enable high-resolution imputation of regulatory
profiles. Finally, we will leverage the imputed data to infer cis-regulatory sequence annotations and the molecular
impact of regulatory non-coding variants in one of the most comprehensive collections of cellular contexts.
All of the software produced by this project will be open source, and all of the imputed data and latent
factorizations will be made publicly available via the web portals associated with the NIH 4D Nucleome and
ENCODE Consortia, providing a valuable public resource for users of these data sets.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Deep tensor genomic imputation
-
批准号:10096947
-
项目类别:
-
资助金额:$39.86万
-
财政年份:2021
-
负责人:William Stafford Noble
-
依托单位:
Optimization and joint modeling for peptide detection by tandem mass spectrometry
-
批准号:9214942
-
项目类别:
-
资助金额:$33.23万
-
财政年份:2017
-
负责人:William Stafford Noble
-
依托单位:
Project 2: UW-CNOF Data Analysis and Modeling
-
批准号:9021413
-
项目类别:
-
资助金额:$63.28万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9983850
-
项目类别:
-
资助金额:$27.7万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9353379
-
项目类别:
-
资助金额:$229.07万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
University of Washington Center for Nuclear Organization and Function
-
批准号:9916567
-
项目类别:
-
资助金额:$8.44万
-
财政年份:2015
-
负责人:William Stafford Noble
-
依托单位:
Machine learning methods to impute and annotate epigenomic maps
-
批准号:8814095
-
项目类别:
-
资助金额:$28.51万
-
财政年份:2014
-
负责人:William Stafford Noble
-
依托单位:
Machine learning methods to impute and annotate epigenomic maps
-
批准号:8925082
-
项目类别:
-
资助金额:$28.29万
-
财政年份:2014
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8642168
-
项目类别:
-
资助金额:$20.78万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8840551
-
项目类别:
-
资助金额:$21.9万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
BIGDATA: DA: Interpreting massive genomic data sets via summarization
-
批准号:8599826
-
项目类别:
-
资助金额:$21.48万
-
财政年份:2013
-
负责人:William Stafford Noble
-
依托单位:
The MEME suite of motif-based sequence analysis tools
-
批准号:8324604
-
项目类别:
-
资助金额:$32.59万
-
财政年份:2009
-
负责人:William Stafford Noble
-
依托单位:
The MEME suite of motif-based sequence analysis tools
-
批准号:8129528
-
项目类别:
-
资助金额:$32.59万
-
财政年份:2009
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:8288063
-
项目类别:
-
资助金额:$62.17万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:8038072
-
项目类别:
-
资助金额:$62.76万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7194479
-
项目类别:
-
资助金额:$62.39万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7797540
-
项目类别:
-
资助金额:$59.36万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7581004
-
项目类别:
-
资助金额:$60.77万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:8470188
-
项目类别:
-
资助金额:$60.22万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
Machine learning analysis of tandem mass spectra
-
批准号:7365198
-
项目类别:
-
资助金额:$60.25万
-
财政年份:2007
-
负责人:William Stafford Noble
-
依托单位:
海外基金