课题基金 / 基金详情

Geometry Aware Exploratory Data Analysis and Inference Methods for Complex Data

Geometry Aware Exploratory Data Analysis and Inference Methods for Complex Data
复杂数据的几何感知探索性数据分析和推理方法
批准号:
2311034
负责人:
Paromita Dubey
金额:
$27.5万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2026-06-30

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
复杂的大数据在科学中经常出现,并已成为当代数据科学的标准内容。众所周知,由于缺乏基本的向量空间运算,如加法和标量乘法,并且数据元素之间没有排序,因此很难分析度量空间中的数据。这些数据以直方图、网络、图像、系统发育树等样本的形式出现,并出现在健康监测、神经科学、商业和经济研究、气候和环境研究、进化遗传学、社会科学和人口统计学等众多领域。当观察到的复杂数据是动态的,例如,当数据时变或在其他连续域上观察时,挑战就会放大。该项目将通过创建一个理论上合理且用户友好的实用工具包来推动现代数据分析技术的前沿,该工具包将克服几个重要数据分析任务的这些挑战。新方法仅基于数据元素之间的成对距离,并通过设计自由调整,将立即满足科学家和工程师处理不同数据表示的需求,例如纵向功能磁共振成像研究、病毒系统发育突变的在线检测、了解微生物多样性组成、监测电子健康分析中的每日血糖分布、时变基因调控网络、了解社会进化的趋势以及更多,为从业者提供了一套现成的工具,以便在进行下游建模任务之前对复杂数据进行探索性分析。该奖项还将支持研究生的培训,并为本科生提供研究机会。基于无模型距离的方法推动了开发面向复杂非欧几里得数据的统计方法的成功,这些方法对环境数据空间或数据分布的要求最小。本研究旨在通过为常见的数据分析工作开发新的严格合理的算法,并建立位于统计学核心的推理程序,从而扩展对象数据分析的方法论库,并构成大多数科学家试图用数据回答的基础。为了解决对象数据中缺乏矢量空间结构和数据元素之间缺乏排序的关键挑战,新的发展将基于深度剖面的概念,深度剖面是由数据规律决定的距离分布,以及传输等级,这是使用深度剖面之间的最佳传输图构建的对象数据的中心向外排序方案。具体子项目将侧重于基于等级的对象数据聚类和分类、离群值检测和以模式为中心的数据分析程序。将为新的双样本测试、独立性测试、变化点检测和定位设计具有严格理论的推理框架,所有这些都将基于距离且易于实现。最后,新工具将被扩展到包括探索性分析和时变目标数据的降维,无论是在观测时间密集的情况下,还是在只有稀疏的时间测量不规则观察的更具挑战性的情况下。理论和方法论的发展将涉及经验过程和u过程理论、m估计和功能数据分析的工具。高效和可伸缩的软件实现,以及用于吸引人的可视化的代码(这对对象数据来说是极具挑战性的),将免费提供给从业者。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Complex big data appear routinely in the sciences and has become standard fare in contemporary data science. It is known to be difficult to analyze data that live in metric spaces, lacking fundamental vector space operations like addition and scalar multiplication and with no ordering between the data elements. Such data show up in the form of samples of histograms, networks, images, phylogenetic trees, and so on, and in multitudes of fields such as health monitoring, neuroscience, business and economics research, climate and environmental studies, evolutionary genetics, social sciences, and demography. Challenges are magnified when the observed complex data are dynamic, for example, when the data are time-varying or observed on other continuous domains. This project will push the frontiers in the state of the art of modern data analysis by creating a theoretically sound and user-friendly practical toolkit that will overcome these challenges for several important data analysis tasks. The new methods, being rooted only in pairwise distances between the data elements and tuning free by design, will immediately cater to the needs of scientists and engineers working with diverse representations of data, for example in longitudinal fMRI studies, online detection of the mutations in the virus phylogeny, understanding microbial diversity compositions, monitoring daily blood glucose distributions in electronic health analytics, time-varying gene-regulatory networks, understanding trends in social evolution and many more, offering practitioners a bundle of off-the-shelf tools to carry out exploratory analysis on the complex data before moving on to the downstream modeling tasks. The award will also support graduate students' training and offer research opportunities to undergraduates.Model-free distance-based approaches drive the success of developing statistical methods oriented to complex non-Euclidean data with minimal requirements on the ambient data space or the data distribution. This research aims to expand the arsenal of methodology in object data analysis by developing new rigorously justified algorithms for common data analysis jobs and building inference procedures that lie at the heart of statistics and constitute the basis of what most scientists attempt to answer with data. To address the key challenge of the lack of a vector space structure in object data and the absence of ordering among the data elements, the new developments will be based on the concepts of depth profiles, which are the distributions of distances as dictated by the law of the data, and the transport ranks, that are center-outward ordering schemes for object data constructed using optimal transport maps between the depth profiles. Specific sub-projects will focus on rank-based object data clustering and classification, outlier detection, and mode-centric data analysis procedures. Inferential frameworks with rigorous theory will be designed for novel two-sample tests, independence tests, change point detection, and localization, all of which will be distance-based and easily implementable. Finally, the new tools will be broadened to include exploratory analysis and dimension reduction for time-varying object data, both when the observations are dense in time and the more challenging case when only sparse measurements in time are observed irregularly. Theory and methodology development will involve tools from the empirical process and U-process theory, M-estimation, and functional data analysis. Efficient and scalable software implementations together with codes for appealing visualizations, which are extremely challenging for object data, will be made freely available for practitioners.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金