Theory and Methods for Tree-Informed High-Dimensional Compositional Data Analysis
Theory and Methods for Tree-Informed High-Dimensional Compositional Data Analysis
批准号:
2113458
负责人:
Shulei Wang
金额:
$15.0万
依托单位国家:
美国
项目类别:
Continuing Grant
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-08-01 至 2024-07-31
中文摘要
组成数据,即某些整体的部分的定量测量,受到约束,需要其分析不同于标准的无约束多元统计分析。高维成分数据自然出现在广泛的现代科学应用中,包括人类微生物组研究、营养科学、基因组学研究和地球化学。在这些科学应用中,由树结构表示的层次关系通常可用于组合数据的不同组件。由于这些数据的组成性质和树状结构,这些数据对以数据驱动的方式获得可靠且具有科学意义的见解构成了独特的挑战。目前分析这类树木组成数据的努力主要是为个人应用而设计的;需要统一框架的新方法和新理论。在这种需求的推动下,该项目旨在开发新的统计理论、方法和计算工具,以实现更强大、更有效的分析。该项目将为致力于统计学与其他科学领域交叉研究的学生提供跨学科研究机会。该项目还将开发用户友好的开源软件,实现新的统计方法,使广泛的科学界受益。本项目旨在研究统计分析应如何可靠而有效地考虑数据的组成性质和树状结构。通过统一的框架,该项目将开发新颖的原则性方法,并深入了解树状结构在树状成分数据分析中的作用。具体而言,本研究将研究树信息成分数据分析的三个基本主题:1)树信息成分数据的独立性和条件独立性测试;2)树信息成分数据的差分成分测试;3)树信息成分数据的度量学习。该项目产生的方法和理论将为不同科学领域的树木成分数据分析提供更可靠和强大的实用工具,最终有助于促进科学和健康方面的知识。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Compositional data, that is, quantitative measurements of the parts of some whole, is subject to constraints that necessitate its analysis be distinct from that of standard unconstrained multivariate statistical analysis. High-dimensional compositional data naturally arises in a wide range of modern scientific applications, including human microbiome studies, nutritional science, genomics studies, and geochemistry. In these scientific applications, a hierarchical relationship represented by a tree structure is often available for the compositional data’s different components. Because of compositional nature and tree structure, these data pose a unique challenge to gaining reliable and scientifically meaningful insights in a data-driven way. Current efforts on analyzing such tree-informed compositional data are primarily designed for individual applications; there is need for new methodology and theory in a unified framework. Motivated by this need, this project aims to develop novel statistical theories, methodologies, and computational tools for more robust and efficient analysis. The project will provide interdisciplinary research opportunities for students who aim to work on the intersection between statistics and other scientific areas. The project will also develop user-friendly open-source software implementing the new statistical methods to benefit a broad scientific community. This project aims to study how statistical analysis should take data's compositional nature and tree structure into account reliably and efficiently. Through a unified framework, the project will develop novel and principled methodologies and provide a deep understanding of the tree structure's role in tree-informed compositional data analysis. Specifically, this research will study three fundamental topics in tree-informed compositional data analysis: 1) independence and conditional independence test for tree-informed compositional data, 2) testing for differential components in tree-informed compositional data, and 3) metric learning for tree-informed compositional data. The resulting methods and theories from the project will lead to more robust and powerful practical tools for tree-informed compositional data analysis in different scientific fields, ultimately helping advance knowledge in science and health.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
Robust differential abundance test in compositional data
成分数据中稳健的差异丰度测试
DOI:
10.1093/biomet/asac029
发表时间:
2022
期刊:
Biometrika
影响因子:
2.7
作者:
[Wang, Shulei]
通讯作者:
Wang, Shulei
国内基金
海外基金
Computational Methods for Analyzing Toponome Data
-
批准号:60601030
-
项目类别:青年科学基金项目
-
资助金额:17.0万元
-
批准年份:2006
-
负责人:Axel Mosig
-
依托单位: