Optimization Techniques for Geometrizing Real-World Data
Optimization Techniques for Geometrizing Real-World Data
批准号:
2044349
负责人:
Soledad Villar
金额:
$2.82万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-06-01 至 2021-07-31
中文摘要
数据是科学领域、政府和私营企业的共同点。在过去的几十年里,能够利用数据来发现模式已经产生了科学突破,并改变了商业模式。该项目侧重于针对特定数据科学问题的数学和算法技术,针对当前相关领域的问题,技术和数据量进行定制。我们考虑的理论问题是(i)聚类(基本上包括以无监督的方式根据相似性对数据进行分组),(ii)降维(减少数据量,同时保留其相关特征),以及(iii)二次分配(找到不同数据集之间的对应关系)。我们在这个项目中考虑的主要基础应用是计算生物学,特别是单细胞测序数据的处理。单细胞测序技术是最近开发的,它正在快速改进,产生新的数据集,问题和挑战,从数学的角度来看是有趣的,并具有潜在的巨大影响。该项目将让数学家与计算生物学家密切合作,目标是识别科学领域中发生的数据科学问题,并开发适当的算法和数学工具。给定单细胞基因表达数据,表明每个基因在每个细胞中表达多少次,一个目标是选择一些可以用于识别不同类别细胞的基因。这个问题在计算生物学文献中被称为遗传标记选择。在第一种方法中,我们假设每个细胞的类是已知的,并且问题可以被视为监督降维。我们将其建模为投影因子恢复问题,并使用半定和线性规划等优化工具来处理它。我们的目标是双重的,我们的目标是研究我们设计的模型的数学特性,并开发一个有效的工具,供从业人员使用。该项目的第二阶段是使问题不受监督,因此聚类将是一个基本步骤。我们将研究聚类方法的稳定性,我们将提供一个有效的算法来评估集群的质量,基于统计和优化技术。该工具的潜在用途是数据科学的通用工具,而不仅仅是基因表达数据集。最后,第三个目标是对齐来自不同实验的数据集。这个问题在数据科学中普遍存在,图匹配和形状匹配是一些特殊情况。在计算生物学的背景下,对齐问题被称为批量校正,它可以用最优运输或二次分配问题来建模。该奖项反映了NSF的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Data is a common denominator to scientific fields, governments, and private enterprises. Being able to exploit data to find patterns has produced scientific breakthroughs and shifted business paradigms in the last several decades. This project focuses on mathematical and algorithmic techniques for specific data science problems, tailored to currently relevant domain problems, technologies, and volumes of data. The theoretical problems we consider are (i) clustering (which essentially consists on grouping data according to similarity in an unsupervised way), (ii) dimensionality reduction (reducing the volume of the data while preserving its relevant features), and (iii) quadratic assignment (finding correspondences between different datasets). The main underlying application we consider in this project is computational biology, in particular the processing of single-cell sequencing data. The technology for single-cell sequencing has been very recently developed and it is improving quickly, producing new datasets, problems and challenges that are interesting from a mathematical point of view and have potentially enormous impact. The project will have mathematicians working closely to computational biologists with the goal of identifying data science problems occurring in the scientific domain and to develop appropriate algorithms and mathematical tools.Given single-cell genetic expression data indicating how many times each gene is expressed in each cell, one objective is to select a few genes that can be used to identify different classes of cells. This problem is known in the computational biology literature as genetic marker selection. In a first approach we assume the class of each cell is known and the problem can be posed as supervised dimensionality reduction. We model it as a projection factor recovery problem, and we approach it using optimization tools such as semidefinite and linear programming. The objective is two-fold, we aim to study mathematical properties of the model we devise, and to develop an efficient tool to be used by practitioners. A second stage of the project is to make the problem unsupervised, therefore clustering will be a fundamental step. We will study stability properties of clustering methods and we will provide an efficient algorithm to evaluate the quality of clusters, based on statistical and optimization techniques. The potential use of this tool is general to data science and not just gene expression datasets. Finally, a third objective is to align datasets coming from different experiments. This problem is ubiquitous in data science, with graph matching and shape matching as some particular cases. In the context of computational biology the alignment problem is known as batch correction and it can be modeled with optimal transport or as a quadratic assignment problem. We will develop alignment algorithms and study their convergence and recovery properties under different data models.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(9)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
--
发表时间:
2021
期刊:
Notices of the American Mathematical Society
影响因子:
--
作者:
[Campos, D, Rivera, M, Salazar, MA, Samper, JA, Simental, J, Villar, S]
通讯作者:
Villar, S
DOI:
10.1109/icassp39728.2021.9413523
发表时间:
2021-06
期刊:
ICASSP 2021 - 2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Ningyuan Huang;Soledad Villar]
通讯作者:
Ningyuan Huang;Soledad Villar
Fitting Very Flexible Models: Linear Regression With Large Numbers of Parameters
拟合非常灵活的模型:具有大量参数的线性回归
DOI:
10.1088/1538-3873/ac20ac
发表时间:
2021
期刊:
Publications of the Astronomical Society of the Pacific
影响因子:
3.5
作者:
[Hogg, David W., Villar, Soledad]
通讯作者:
Villar, Soledad
DOI:
10.1109/icassp40776.2020.9053147
发表时间:
2020-05
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
作者:
[Efe Onaran;Soledad Villar]
通讯作者:
Efe Onaran;Soledad Villar
DOI:
10.1109/tit.2019.2962681
发表时间:
2020-06-01
期刊:
IEEE TRANSACTIONS ON INFORMATION THEORY
影响因子:
2.5
作者:
[McWhirter, Culver, Mixon, Dustin G., Villar, Soledad]
通讯作者:
Villar, Soledad
共 7 条
CAREER: Symmetries and Classical Physics in Machine Learning for Science and Engineering
-
批准号:2339682
-
项目类别:Continuing Grant
-
资助金额:$59.36万
-
财政年份:2024
-
负责人:Soledad Villar
-
依托单位:
Collaborative Research: CIF: Medium: Understanding Robustness via Parsimonious Structures.
-
批准号:2212457
-
项目类别:Standard Grant
-
资助金额:$90.0万
-
财政年份:2022
-
负责人:Soledad Villar
-
依托单位:
Optimization Techniques for Geometrizing Real-World Data
-
批准号:1913134
-
项目类别:Standard Grant
-
资助金额:$5.06万
-
财政年份:2019
-
负责人:Soledad Villar
-
依托单位:
国内基金
海外基金
EstimatingLarge Demand Systems with MachineLearning Techniques
-
批准号:--
-
项目类别:外国学者研究基金
-
资助金额:--
-
批准年份:2024
-
负责人:IoshuaAlex
-
依托单位: