课题基金 / 基金详情

Geometry Aware Exploratory Data Analysis and Inference Methods for Complex Data

Geometry Aware Exploratory Data Analysis and Inference Methods for Complex Data
复杂数据的几何感知探索性数据分析和推理方法
批准号:
2311034
负责人:
Paromita Dubey
金额:
$27.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2026-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
复杂的大数据经常出现在科学领域,并已成为当代数据科学的标准内容。众所周知,很难分析存在于度量空间中的数据,缺乏基本的向量空间运算,如加法和标量乘法,并且数据元素之间没有排序。这些数据以直方图、网络、图像、系统发育树等样本的形式出现,并出现在许多领域,如健康监测、神经科学、商业和经济研究、气候和环境研究、进化遗传学、社会科学和人口学。当观察到的复杂数据是动态的时,例如,当数据是时变的或在其他连续的域上观察到时,挑战被放大。该项目将通过创建一个理论上合理和用户友好的实用工具包,克服这些挑战,完成几项重要的数据分析任务,从而推动现代数据分析的前沿。新方法仅植根于数据元素之间的成对距离,并在设计上自由调整,将立即迎合科学家和工程师使用不同数据表示的需求,例如在纵向功能磁共振研究、在线检测病毒系统发育中的突变、了解微生物多样性组成、监测电子健康分析中的每日血糖分布、时变基因调控网络、了解社会进化趋势等,为从业者提供一系列现成的工具,以便在进入下游建模任务之前对复杂数据进行探索性分析。该奖项还将支持研究生的培训,并为本科生提供研究机会。基于距离的无模型方法推动了开发面向复杂非欧几里德数据的统计方法的成功,而对环境数据空间或数据分布的要求最低。这项研究旨在扩大对象数据分析的方法论武器库,为常见的数据分析工作开发新的严格证明合理的算法,并建立推理程序,这些程序位于统计学的核心,构成大多数科学家试图用数据回答的基础。为了解决物体数据中缺乏矢量空间结构和数据元素之间缺乏排序的关键挑战,新的发展将基于深度剖面和传输等级的概念,深度剖面是由数据定律指示的距离分布,传输等级是使用深度剖面之间的最佳传输映射构建的对象数据的中心向外排序方案。具体的子项目将侧重于基于排名的目标数据分组和分类、离群值检测和以模式为中心的数据分析程序。具有严格理论的推理框架将被设计用于新颖的双样本测试、独立性测试、变点检测和定位,所有这些都将基于距离并且易于实现。最后,新的工具将扩大到包括对时变对象数据的探索性分析和降维,无论是在观测时间密集的情况下,还是在更具挑战性的情况下,当只观测到时间上的稀疏测量时都是如此。理论和方法的发展将涉及来自经验过程和U过程理论、M估计和功能数据分析的工具。高效和可扩展的软件实现以及用于吸引对象数据的可视化代码将免费提供给实践者。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Complex big data appear routinely in the sciences and has become standard fare in contemporary data science. It is known to be difficult to analyze data that live in metric spaces, lacking fundamental vector space operations like addition and scalar multiplication and with no ordering between the data elements. Such data show up in the form of samples of histograms, networks, images, phylogenetic trees, and so on, and in multitudes of fields such as health monitoring, neuroscience, business and economics research, climate and environmental studies, evolutionary genetics, social sciences, and demography. Challenges are magnified when the observed complex data are dynamic, for example, when the data are time-varying or observed on other continuous domains. This project will push the frontiers in the state of the art of modern data analysis by creating a theoretically sound and user-friendly practical toolkit that will overcome these challenges for several important data analysis tasks. The new methods, being rooted only in pairwise distances between the data elements and tuning free by design, will immediately cater to the needs of scientists and engineers working with diverse representations of data, for example in longitudinal fMRI studies, online detection of the mutations in the virus phylogeny, understanding microbial diversity compositions, monitoring daily blood glucose distributions in electronic health analytics, time-varying gene-regulatory networks, understanding trends in social evolution and many more, offering practitioners a bundle of off-the-shelf tools to carry out exploratory analysis on the complex data before moving on to the downstream modeling tasks. The award will also support graduate students' training and offer research opportunities to undergraduates.Model-free distance-based approaches drive the success of developing statistical methods oriented to complex non-Euclidean data with minimal requirements on the ambient data space or the data distribution. This research aims to expand the arsenal of methodology in object data analysis by developing new rigorously justified algorithms for common data analysis jobs and building inference procedures that lie at the heart of statistics and constitute the basis of what most scientists attempt to answer with data. To address the key challenge of the lack of a vector space structure in object data and the absence of ordering among the data elements, the new developments will be based on the concepts of depth profiles, which are the distributions of distances as dictated by the law of the data, and the transport ranks, that are center-outward ordering schemes for object data constructed using optimal transport maps between the depth profiles. Specific sub-projects will focus on rank-based object data clustering and classification, outlier detection, and mode-centric data analysis procedures. Inferential frameworks with rigorous theory will be designed for novel two-sample tests, independence tests, change point detection, and localization, all of which will be distance-based and easily implementable. Finally, the new tools will be broadened to include exploratory analysis and dimension reduction for time-varying object data, both when the observations are dense in time and the more challenging case when only sparse measurements in time are observed irregularly. Theory and methodology development will involve tools from the empirical process and U-process theory, M-estimation, and functional data analysis. Efficient and scalable software implementations together with codes for appealing visualizations, which are extremely challenging for object data, will be made freely available for practitioners.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金