III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology
III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology
批准号:
1909096
负责人:
Philip Wilsey
金额:
$49.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2023-09-30
中文摘要
机器学习的新见解存在于许多领域,例如,包括医学、社交媒体、图像处理、生物学以及计算机和网络安全。机器学习能够处理超出人类能力的大型高维数据集。一种新兴的机器学习方法是基于一种名为拓扑学的数学分支,它有时能够发现使用传统方法无法获得的知识。拓扑学研究对象的形状,持久同调是拓扑学中提取形状特征的关键方法。永久同源将根据对象中孔洞和空洞的大小和数量对该对象进行分类。不幸的是,计算对象的持久同调需要大量的内存和较长的运行时间,这会使形成该对象的点的数量呈指数级增加。该项目将处理由数据形成的对象,并将其细分为更小的区域,以便在每个区域上并行计算持久同调。然后,将把区域分析的结果汇编在一起,并在分析后的步骤中查明任何重复或缺失的结果并加以恢复。与对整个数据集进行一次计算相比,在所有区域上完成计算所需的时间和总内存将大大减少。将使用各种合成数据和真实数据对开发的方法进行测试。合成数据将允许对性能和可扩展性进行受控研究。将使用来自各种来源的真实世界数据,特别是那些小的拓扑特征很重要的数据(例如来自脑扫描的数据)。该项目将推动基于拓扑的分析的应用,从海量的高维数据中发现新的见解和有意义的信息。通过增加课程、项目(高级项目、硕士论文、博士论文等)、研讨会和研究合作培训经验,将实现通过基于拓扑的方法扩大学生在数据挖掘方面的培训。各级学生都将受到影响,并特别强调少数群体和代表性不足的学生群体的参与。该项目还将参与加州大学的女性科学与工程项目。项目调查员将与当地的K-12学生、国际交换生和加州大学合作机构、加州大学医学院、辛辛那提儿童医学中心、空军研究实验室和当地行业的研究人员合作,提供关于该项目调查和结果的信息和研讨会。该项目建议将近似计算与拓扑数据分析领域相结合,以极大地减少在超大型数据集上使用拓扑数据分析的计算和内存需求。特别是,该项目将开发计算持久同调的近似方法,从而显著增加可应用基于拓扑数据分析的数据挖掘方法的数据集的大小。该项目预计将可通过拓扑数据分析方法分析的输入数据集的大小至少增加3-5个数量级。虽然近似方法可能会引入误差,但由近似方法识别的特征将识别点云区域,在这些区域中,可以(并行)使用持久同调的升标步骤和区域计算来建立这些特征的更精确的边界。该项目将开发算法改进,关于算法和近似技术的正确性、误差界和复杂性的正式声明。这些技术对将拓扑数据分析技术应用到比目前可能的数据集更大的数据集的能力具有重要影响。该奖项反映了NSF的法定使命,并通过使用基金会的智力优势和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
New insights with machine learning exists across may domains, including, for example, medicine, social media, image processing, biology, and computer and network security. Machine learning is able to process large, high-dimensional data sets that are beyond human capabilities. One emerging method of machine learning is based on a branch of mathematics called topology that is sometimes able to discover knowledge that is not available using conventional methods. The field of topology is concerned with of the shape of an object and Persistent Homology is the critical method in topology used to extract the features of a shape. Persistent Homology will classify an object by the size and number of holes and voids in that object. Unfortunately, computing the Persistent Homology for an object requires significant amounts of memory and long run-times that increases exponentially in the number of points that forms that object. This project will treat the object formed by the data and subdivide it into smaller regions for the parallel computation of Persistent Homology on each region. The results from the regional analyses will then be assembled together and any duplicate or missing results will be identified and restored in a post analysis step. The computation on all of the regions will be completed in substantially less time and in much less total memory than a single computation on the entire data set. Testing of the methods developed will be performed using a variety of synthetic and real-world data. The synthetic data will permit controlled studies on performance and scalability. Realworld data from a variety of sources and especially data where the small topological features are significant (such as data from brain scans) will be used. This project will propel the application of topology based analysis to discover new insights and meaningful information from massive high-dimensional data. An expansion of student training in data mining through topological-based methods will be achieved with the addition of classes, projects (senior project, MS Theses, PhD Dissertations, and so on), seminars, and research co-op training experiences. Students at all levels will be impacted and special emphasis placed on minority and underrepresented student groups participation. This project will also participate in the Women in Science and Engineering programs at UC. The project investigators will engage local area K-12 students, international exchange students and researchers at UC's collaborative institutions, UC's Medical School, Cincinnati Children's Medical Center, the Air Force Research Lab, and local industries with information and seminars on this project investigations and results.This project proposes to combine the fields of Approximate Computing with Topological Data Analysis to dramatically reduce the computational and memory requirements to use Topological Data Analysis on very large data sets. In particular, this project will develop approximate methods for computing Persistent Homology that dramatically increase the sizes of data sets for which data mining methods based on topological data analysis can be applied. This project expects to increase the size of the input data set that can be analyzed by Topological Data Analysis methods by at least 3-5 orders of magnitude. While approximate methods can introduce error, the features identified by the approximate methods will identify regions of the point cloud where an upscaling steps and regional computations of Persistent Homology can be used (in parallel) to establish more precise boundaries of those features. The project will develop algorithmic improvements, formal statements on the correctness, error bounds, and complexities of the algorithms and approximation techniques. These techniques have important implications on the ability to apply topological data analysis techniques to much larger data sets than currently possible.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1109/bigdata55660.2022.10020926
发表时间:
2022
期刊:
IEEE International Conference on Big Data
影响因子:
--
作者:
[Singh, Rohit P., Wilsey, Philip A.]
通讯作者:
Wilsey, Philip A.
DOI:
10.1109/bigdata47090.2019.9006572
发表时间:
2019-12
期刊:
2019 IEEE International Conference on Big Data (Big Data)
影响因子:
--
作者:
[Nicholas O. Malott;P. Wilsey]
通讯作者:
Nicholas O. Malott;P. Wilsey
Computation of persistent homology on streaming data using topological data summaries
使用拓扑数据摘要计算流数据上的持久同源性
DOI:
10.1111/coin.12597
发表时间:
2023
期刊:
Computational Intelligence
影响因子:
2.8
作者:
[Moitra, Anindya, Malott, Nicholas O., Wilsey, Philip A.]
通讯作者:
Wilsey, Philip A.
DOI:
10.1109/bigdata50022.2020.9378216
发表时间:
2020-12
期刊:
2020 IEEE International Conference on Big Data (Big Data)
影响因子:
--
作者:
[Nicholas O. Malott;Aaron M. Sens;P. Wilsey]
通讯作者:
Nicholas O. Malott;Aaron M. Sens;P. Wilsey
DOI:
10.1109/icdm54844.2022.00136
发表时间:
2022
期刊:
IEEE International Conference on Data Mining
影响因子:
--
作者:
[Malott, Nicholas O., Lewis, Robert R., Wilsey, Philip A.]
通讯作者:
Wilsey, Philip A.
共 9 条
SI2-SSE: Scalable Big Data Clustering by Random Projection Hashing
-
批准号:1440420
-
项目类别:Standard Grant
-
资助金额:$49.81万
-
财政年份:2014
-
负责人:Philip Wilsey
-
依托单位:
CSR: Small: Collaborative Research: Combining Static Analysis and Dynamic Run-time Optimization for Parallel Discrete Event Simulation in Many-Core Environments
-
批准号:0915337
-
项目类别:Standard Grant
-
资助金额:$16.69万
-
财政年份:2009
-
负责人:Philip Wilsey
-
依托单位:
国内基金
海外基金
登录
查看更多内容
昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:
-
依托单位:
tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
-
批准号:
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2022
-
负责人:张祥忠
-
依托单位:
Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
-
批准号:32000033
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:林平
-
依托单位:
Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
-
批准号:31972324
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:高学文
-
依托单位:
变异链球菌small RNAs连接LuxS密度感应与生物膜形成的机制研究
-
批准号:81900988
-
项目类别:青年科学基金项目
-
资助金额:21.0万元
-
批准年份:2019
-
负责人:毛梦莹
-
依托单位:
肠道细菌关键small RNAs在克罗恩病发生发展中的功能和作用机制
-
批准号:31870821
-
项目类别:面上项目
-
资助金额:56.0万元
-
批准年份:2018
-
负责人:陈江宁
-
依托单位:
基于small RNA 测序技术解析鸽分泌鸽乳的分子机制
-
批准号:31802058
-
项目类别:青年科学基金项目
-
资助金额:26.0万元
-
批准年份:2018
-
负责人:麻慧
-
依托单位:
Small RNA介导的DNA甲基化调控的水稻草矮病毒致病机制
-
批准号:31772128
-
项目类别:面上项目
-
资助金额:60.0万元
-
批准年份:2017
-
负责人:吴建国
-
依托单位:
基于small RNA-seq的针灸治疗桥本甲状腺炎的免疫调控机制研究
-
批准号:81704176
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2017
-
负责人:赵继梦
-
依托单位:
水稻OsSGS3与OsHEN1调控small RNAs合成及其对抗病性的调节
-
批准号:91640114
-
项目类别:重大研究计划
-
资助金额:85.0万元
-
批准年份:2016
-
负责人:何祖华
-
依托单位: