A Modeling Framework for Multi-View Data, with Applications to the Pioneer 100 Study and Protein Interaction Networks
A Modeling Framework for Multi-View Data, with Applications to the Pioneer 100 Study and Protein Interaction Networks
批准号:
9752596
负责人:
Jacob Bien
金额:
$32.37万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-08-01 至 2021-06-30
关键词:
AddressAdoptionAgreementAlgorithmsBiologyBiomedical ResearchClinical DataCommunitiesComplexComputer softwareConflict (Psychology)DataData SetDetectionDevelopmentDimensionsDiseaseFoundationsFutureGene ExpressionGeneticGenomicsGoalsHealthHumanIndividualMeasurementMeasuresMedical GeneticsMeta-AnalysisMethodologyMethodsModelingParticipantPatientsPrincipal Component AnalysisProteinsProteomicsRecordsResearch PersonnelResourcesSet proteinStatistical Data InterpretationStatistical MethodsTechnologyTestingTimeTrustValidationVariantgenomic dataimprovedinsightmetabolomicsnovel strategiesopen sourceunsupervised learning
中文摘要
生物医学研究的新进展使收集多种数据“观点”成为可能--例如,
遗传、代谢和临床数据--针对单个患者。这样的多视角数据有望提供更深层次的
对患者健康和疾病的洞察,如果只有一个数据视图是可能的。但是,在
为了实现这一承诺,需要新的统计方法。
这项建议涉及开发用于分析多视角数据的统计方法。这些方法可以
用于回答以下基本问题:数据视图是否包含有关
或者每个数据视图包含一组不同的信息?这个问题的答案将提供
洞察数据视图,以及洞察观察结果。如果两个数据视图包含冗余信息
关于观测,那么这两个数据视图是相互关联的。此外,如果每个数据视图都告诉
同样的“故事”的观察,那么我们可以相当确信这个故事是真的。fi。
研究人员将开发一个统一的fi框架来对多视图数据进行建模,然后将应用于
一系列设置。在目标1中,这一框架将适用于多视角多变量数据(例如,单一集合
具有临床和遗传测量的患者),以确定单个聚集性是否可以
充分描述所有数据视图中的患者,或者患者是否单独聚集在每个数据中
查看。在目标2中,该框架将应用于多视点网络数据(例如,一组单一蛋白质,两者都有
测量的二元和共复相互作用),以便确定节点是否属于
跨数据视图的社区,或每个数据视图中的一组单独的社区。在目标3中,框架
将应用于多视图多变量数据,以确定观测值是否可以嵌入
跨所有数据视图的单个潜在空间,或者它们是否属于每个数据视图中的单独潜在空间。
在目标1-3中,开发的方法将应用于先锋100的研究和蛋白质相互作用组。在……里面
目标4(A),将利用多个数据视图的可用性来开发调整参数的方法
无监督学习中的选择。在目标4(B)中,将对在目标2中识别的蛋白质群落进行验证
试验性的。AIM 5将开发高质量的开源软件。
本提案中开发的方法将用于确定是否从多个数据视图进行fi绑定
是相同或不同的。这些方法在多视角数据集上的应用,包括先锋100研究
以及蛋白质相互作用组,将提高我们对人类健康和疾病的理解,以及
生物学。
英文摘要
New advances in biomedical research have made it possible to collect multiple data “views” — for example,
genetic, metabolomic, and clinical data — for a single patient. Such multi-view data promises to offer deeper
insights into a patient's health and disease than would be possible if just one data view were available. However, in
order to achieve this promise, new statistical methods are needed.
This proposal involves developing statistical methods for the analysis of multi-view data. These methods can
be used to answer the following fundamental question: do the data views contain redundant information about the
observations, or does each data view contain a different set of information? The answer to this question will provide
insight into the data views, as well as insight into the observations. If two data views contain redundant information
about the observations, then those two data views are related to each other. Furthermore, if each data view tells the
same “story” about the observations, then we can be quite confident that the story is true.
The investigators will develop a unified framework for modeling multi-view data, which will then be applied in
a number of settings. In Aim 1, this framework will be applied to multi-view multivariate data (e.g. a single set
of patients, with both clinical and genetic measurements), in order to determine whether a single clustering can
adequately describe the patients across all data views, or whether the patients cluster separately in each data
view. In Aim 2, the framework will be applied to multi-view network data (e.g. a single set of proteins, with both
binary and co-complex interactions measured), in order to determine whether the nodes belong to a single set of
communities across the data views, or a separate set of communities in each data view. In Aim 3, the framework
will be applied to multi-view multivariate data in order to determine whether the observations can be embedded in
a single latent space across all data views, or whether they belong to a separate latent space in each data view.
In Aims 1–3, the methods developed will be applied to the Pioneer 100 study, and to the protein interactome. In
Aim 4(a), the availability of multiple data views will be used in order to develop a method for tuning parameter
selection in unsupervised learning. In Aim 4(b), protein communities that were identified in Aim 2 will be validated
experimentally. High-quality open source software will be developed in Aim 5.
The methods developed in this proposal will be used to determine whether the findings from multiple data views
are the same or different. The application of these methods to multi-view data sets, including the Pioneer 100 study
and the protein interactome, will improve our understanding of human health and disease, as well as fundamental
biology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A Modeling Framework for Multi-View Data, with Applications to the Pioneer 100 Study and Protein Interaction Networks
-
批准号:9361170
-
项目类别:
-
资助金额:$34.0万
-
财政年份:2017
-
负责人:Jacob Bien
-
依托单位:
海外基金