Manifold Coordinates with Physical Meaning
Manifold Coordinates with Physical Meaning
批准号:
2015272
负责人:
Marina Meila
金额:
$15.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-08-15 至 2023-07-31
中文摘要
为高维现象寻找低维有意义的描述符一直是科学发现的动力之一。例如,Cavalli-Sforza引入的“基因图谱”将人类基因组的变异绘制成二维地理位置,绘制出史前迁徙的图表。在这个例子中,科学家们的直觉引导了基因组在空间维度上的映射。这个项目将开发一个通用的统计框架来扩展和自动化这个过程。科学家提供了一个具有科学意义的描述符列表,可以用来在低维度上展开数据。这个列表被称为“字典”。字典在专家知识(用科学领域的概念表示)和学习算法使用的数据的较低层次表示(另一方面)之间进行调解。这个项目将促进科学家和机器之间的知识转移。统计学习算法将取代手动检查单个描述符与数据变化的相关性的任务;该算法将立即对整个字典执行此任务。输出是来自字典的一小组描述符,它们一起捕获了数据中的大部分变化;我们称之为可解释嵌入坐标(IEC)。与抽象的主成分或主方向不同,这些坐标总是有意义和可解释的,因为它们是从科学家提供的描述符字典中选择出来的。在这个方案中,假设数据位于一个光滑的低维流形上或附近;字典由流形上的光滑函数组成。新方法在字典中可解释的、有意义的函数之间寻找流形中的坐标。可解释嵌入坐标(IEC)将被表述为一个非参数、非线性稀疏泛函回归问题。主要思想是将该问题转化为函数梯度空间中的线性稀疏回归。这使得人们可以将发达的稀疏恢复方法应用于IEC,而不会牺牲问题的原始非线性。将给出恢复的统计和几何保证。这些新方法将被整合到大数据无监督学习平台megaman中,由美拉集团分发和维护。在华盛顿大学科学研究所的支持下,PI Meila将在一个关于无监督学习的在线主动训练实验室中传播这些想法和方法。该项目是Meila当前研究项目“无监督学习的无监督验证”的一部分,该项目旨在设计基于数学的方法来解释、验证和验证机器学习算法的输出,以用于科学数据和科学发现。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Finding low dimensional meaningful descriptors for high-dimensional phenomena has been one of the motors of scientific discovery. For instance, the "genetic maps" introduced by Cavalli-Sforza map the variation of the human genomes into 2-dimensional geographic locations, charting prehistorical migrations. In this example, the scientists' intuition guided the mapping of the genomes on the spatial dimensions. This project will develop a general statistical framework to expand and automate this process. A scientist provides a list of descriptors with scientific meaning, that could be used to unfold the data in low dimensions. This list is called a "dictionary". The dictionary mediates between expert knowledge, expressed with the concepts of the scientific domain, on one hand, and the lower level representations of the data used by learning algorithms, on the other. This project will facilitate the tranfer of knowledge between scientist and machine. A statistical learning algorithm will replace the task of manually checking individual descriptors for correlation with the data variation; the algorithm will perform this task on the whole dictionary at once. The output is a small set of descriptors from the dictionary, which together capture most of the variation in the data; we call them Interpretable Embedding Coordinates (IEC). Unlike Principal Components or Principal Directions, which are abstract, these coordinates are always meaningful and interpretable, because they are selected from the dictionary of descriptors supplied by the scientist.In this project it is assumed that the data lie on or near a smooth low-dimensional manifold; the dictionary consists of smooth functions on the manifold. The new method finds coordinates in the manifold among the interpretable, meaningful functions in the dictionary. Interpretable Embedding Coordinates (IEC) will be formulated as a non-parametric, non-linear sparse functional regression problem. The main idea is to tranform this problem into a linear sparse regression in the space of function gradients. This allows one to apply the well-developed aresenal of sparse recovery methods to IEC, without sacrificing the original non-linearity of the problem. Statistical and geometric guarantees for recovery will be given. The new methods will be integrated into the big data unsupervised learning platform megaman, distributed and maintained by Meila's group. PI Meila, with support from the UW eScience Institute, will disseminate the ideas and methods in an on-line Active Training Lab on Unsupervised Learning. This project is part of Meila's current research program "Unsupervised Validation for Unsupervised Learning" to design mathematically founded methods to interpret, verify and validate the output of machine learning algorithms for scientific data and scientific discovery.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(3)
专著(0)
科研奖励(0)
会议论文
DOI:
--
发表时间:
2021-07
期刊:
ArXiv
影响因子:
--
作者:
[Yu-Chia Chen;M. Meilă]
通讯作者:
Yu-Chia Chen;M. Meilă
DOI:
--
发表时间:
2018-11
期刊:
J. Mach. Learn. Res.
影响因子:
--
作者:
[Samson Koelle;Hanyu Zhang;M. Meilă;Yu-Chia Chen]
通讯作者:
Samson Koelle;Hanyu Zhang;M. Meilă;Yu-Chia Chen
Manifold Learning: What, How, and Why
流形学习:什么、如何以及为什么
DOI:
10.1146/annurev-statistics-040522-115238
发表时间:
2024
期刊:
Annual Review of Statistics and Its Application
影响因子:
7.9
作者:
[Meilă, Marina, Zhang, Hanyu]
通讯作者:
Zhang, Hanyu
Cluster Validation Without Model Assumptiions
-
批准号:1810975
-
项目类别:Continuing Grant
-
资助金额:$27.5万
-
财政年份:2018
-
负责人:Marina Meila
-
依托单位:
Doctoral Student Forum and Student Travel at the 2011 SIAM Data Mining Conference; Phoenix, AZ
-
批准号:1103263
-
项目类别:Standard Grant
-
资助金额:$2.91万
-
财政年份:2011
-
负责人:Marina Meila
-
依托单位:
Clustering Link Data - Theory and Algortithms
-
批准号:0313339
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2003
-
负责人:Marina Meila
-
依托单位:
海外基金