课题基金 / 基金详情

III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology

III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology
III:小:对大数据进行分区以实现持久同调的高性能计算
批准号:
1909096
负责人:
Philip Wilsey
金额:
$49.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2023-09-30

项目摘要

项目成果

Philip Wilsey的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习的新见解存在于许多领域,例如医学、社交媒体、图像处理、生物学、计算机和网络安全。机器学习能够处理超出人类能力的大型高维数据集。一种新兴的机器学习方法是基于被称为拓扑学的数学分支,它有时能够发现使用传统方法无法获得的知识。拓扑学是研究物体形状的领域,持久同调是拓扑学中提取物体形状特征的关键方法。Persistent Homology将根据对象中的孔和空隙的大小和数量对对象进行分类。不幸的是,计算一个对象的Persistent同源性需要大量的内存和较长的运行时间,而这些内存和运行时间会随着组成该对象的点的数量呈指数增长。本项目将对数据形成的对象进行处理,并将其细分为更小的区域,在每个区域上并行计算Persistent Homology。然后将区域分析的结果汇集在一起,并在分析后的步骤中识别和恢复任何重复或缺失的结果。与在整个数据集上进行单个计算相比,在所有区域上的计算将在更短的时间内完成,并且占用的总内存要少得多。所开发的方法的测试将使用各种合成和真实世界的数据进行。合成数据将允许对性能和可伸缩性进行控制研究。将使用来自各种来源的真实世界数据,特别是具有重要小拓扑特征的数据(例如来自脑部扫描的数据)。该项目将推动基于拓扑分析的应用,从大量高维数据中发现新的见解和有意义的信息。通过增加课程、项目(高级项目、硕士论文、博士论文等)、研讨会和研究合作培训经验,通过基于拓扑的方法扩大学生在数据挖掘方面的培训。各级学生都将受到影响,并特别强调少数民族和代表性不足的学生群体的参与。该项目还将参与加州大学的女性科学与工程项目。项目调查人员将与当地K-12学生、国际交换生和加州大学合作机构、加州大学医学院、辛辛那提儿童医学中心、空军研究实验室和当地行业的研究人员进行交流,提供有关该项目调查和结果的信息和研讨会。该项目建议将近似计算与拓扑数据分析领域结合起来,以显着减少在非常大的数据集上使用拓扑数据分析的计算和内存需求。特别是,该项目将开发计算持久同源性的近似方法,从而大大增加数据集的大小,从而可以应用基于拓扑数据分析的数据挖掘方法。该项目预计将可通过拓扑数据分析方法分析的输入数据集的规模增加至少3-5个数量级。虽然近似方法可能会引入误差,但由近似方法识别的特征将识别点云的区域,在这些区域中,可以使用升级步骤和持久同源的区域计算(并行)来建立这些特征的更精确的边界。该项目将开发算法改进,关于正确性的正式声明,错误界限,以及算法和近似技术的复杂性。这些技术对于将拓扑数据分析技术应用于比目前可能的更大的数据集的能力具有重要意义。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
New insights with machine learning exists across may domains, including, for example, medicine, social media, image processing, biology, and computer and network security. Machine learning is able to process large, high-dimensional data sets that are beyond human capabilities. One emerging method of machine learning is based on a branch of mathematics called topology that is sometimes able to discover knowledge that is not available using conventional methods. The field of topology is concerned with of the shape of an object and Persistent Homology is the critical method in topology used to extract the features of a shape. Persistent Homology will classify an object by the size and number of holes and voids in that object. Unfortunately, computing the Persistent Homology for an object requires significant amounts of memory and long run-times that increases exponentially in the number of points that forms that object. This project will treat the object formed by the data and subdivide it into smaller regions for the parallel computation of Persistent Homology on each region. The results from the regional analyses will then be assembled together and any duplicate or missing results will be identified and restored in a post analysis step. The computation on all of the regions will be completed in substantially less time and in much less total memory than a single computation on the entire data set. Testing of the methods developed will be performed using a variety of synthetic and real-world data. The synthetic data will permit controlled studies on performance and scalability. Realworld data from a variety of sources and especially data where the small topological features are significant (such as data from brain scans) will be used. This project will propel the application of topology based analysis to discover new insights and meaningful information from massive high-dimensional data. An expansion of student training in data mining through topological-based methods will be achieved with the addition of classes, projects (senior project, MS Theses, PhD Dissertations, and so on), seminars, and research co-op training experiences. Students at all levels will be impacted and special emphasis placed on minority and underrepresented student groups participation. This project will also participate in the Women in Science and Engineering programs at UC. The project investigators will engage local area K-12 students, international exchange students and researchers at UC's collaborative institutions, UC's Medical School, Cincinnati Children's Medical Center, the Air Force Research Lab, and local industries with information and seminars on this project investigations and results.This project proposes to combine the fields of Approximate Computing with Topological Data Analysis to dramatically reduce the computational and memory requirements to use Topological Data Analysis on very large data sets. In particular, this project will develop approximate methods for computing Persistent Homology that dramatically increase the sizes of data sets for which data mining methods based on topological data analysis can be applied. This project expects to increase the size of the input data set that can be analyzed by Topological Data Analysis methods by at least 3-5 orders of magnitude. While approximate methods can introduce error, the features identified by the approximate methods will identify regions of the point cloud where an upscaling steps and regional computations of Persistent Homology can be used (in parallel) to establish more precise boundaries of those features. The project will develop algorithmic improvements, formal statements on the correctness, error bounds, and complexities of the algorithms and approximation techniques. These techniques have important implications on the ability to apply topological data analysis techniques to much larger data sets than currently possible.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/bigdata55660.2022.10020926
发表时间: 2022
期刊: IEEE International Conference on Big Data
影响因子: --
作者: [Singh, Rohit P., Wilsey, Philip A.]
通讯作者: Wilsey, Philip A.
DOI: 10.1109/bigdata47090.2019.9006572
发表时间: 2019-12
期刊: 2019 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者: [Nicholas O. Malott;P. Wilsey]
通讯作者: Nicholas O. Malott;P. Wilsey
Computation of persistent homology on streaming data using topological data summaries
使用拓扑数据摘要计算流数据上的持久同源性
DOI: 10.1111/coin.12597
发表时间: 2023
期刊: Computational Intelligence
影响因子: 2.8
作者: [Moitra, Anindya, Malott, Nicholas O., Wilsey, Philip A.]
通讯作者: Wilsey, Philip A.
DOI: 10.1109/bigdata50022.2020.9378216
发表时间: 2020-12
期刊: 2020 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者: [Nicholas O. Malott;Aaron M. Sens;P. Wilsey]
通讯作者: Nicholas O. Malott;Aaron M. Sens;P. Wilsey
9
    SI2-SSE: Scalable Big Data Clustering by Random Projection Hashing
    • 批准号:
      1440420
    • 项目类别:
      Standard Grant
    • 资助金额:
      $49.81万
    • 财政年份:
      2014
    • 负责人:
      Philip Wilsey
    • 依托单位:
    CSR: Small: Collaborative Research: Combining Static Analysis and Dynamic Run-time Optimization for Parallel Discrete Event Simulation in Many-Core Environments
    • 批准号:
      0915337
    • 项目类别:
      Standard Grant
    • 资助金额:
      $16.69万
    • 财政年份:
      2009
    • 负责人:
      Philip Wilsey
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: