课题基金 / 基金详情

III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology

III: Small: Partitioning Big Data for the High Performance Computation of Persistent Homology
III:小:对大数据进行分区以实现持久同调的高性能计算
批准号:
1909096
负责人:
Philip Wilsey
金额:
$49.93万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2019
资助国家:
美国
项目状态:
已结题
起止时间:
2019-10-01 至 2023-09-30

项目摘要

项目成果

Philip Wilsey的其他基金

相似基金

相关文献

中文摘要
翻译
机器学习的新见解存在于许多领域,包括例如医学、社交媒体、图像处理、生物学以及计算机和网络安全。机器学习能够处理超出人类能力的大型高维数据集。一种新兴的机器学习方法是基于称为拓扑学的数学分支,它有时能够发现使用传统方法无法获得的知识。拓扑学研究的是物体的形状,而持久同调是拓扑学中提取形状特征的关键方法。持久同源性将根据对象中的孔和空隙的大小和数量对对象进行分类。不幸的是,计算一个对象的持久同源性需要大量的内存和长的运行时间,这在形成该对象的点的数量上呈指数级增长。这个项目将处理由数据形成的对象,并将其细分为更小的区域,以便在每个区域上并行计算持久同源性。然后将区域分析的结果汇总在一起,并在分析后步骤中识别和恢复任何重复或缺失的结果。对所有区域的计算将在比对整个数据集的单个计算少得多的时间和少得多的总存储器中完成。将使用各种合成数据和真实数据对所开发的方法进行测试。合成数据将允许对性能和可扩展性进行受控研究。将使用来自各种来源的真实世界数据,特别是小拓扑特征重要的数据(例如来自大脑扫描的数据)。该项目将推动基于拓扑分析的应用,从海量高维数据中发现新的见解和有意义的信息。通过基于拓扑的方法扩展学生数据挖掘培训将通过增加课程,项目(高级项目,MS论文,博士论文等),研讨会和研究合作培训经验来实现。各级学生将受到影响,并特别强调少数民族和代表性不足的学生群体的参与。该项目还将参与加州大学的妇女科学和工程项目。项目调查人员将与当地K-12学生、国际交换生和加州大学合作机构、加州大学医学院、辛辛那提儿童医疗中心、空军研究实验室、该项目提出将近似计算与拓扑数据分析领域联合收割机相结合,以大大减少在非常大的数据集上使用拓扑数据分析的计算和内存要求。 特别是,本项目将开发计算持久同源性的近似方法,大大增加数据集的大小,数据挖掘方法的基础上拓扑数据分析可以应用。该项目预计将增加至少3-5个数量级的拓扑数据分析方法可以分析的输入数据集的大小。虽然近似方法可能会引入误差,但通过近似方法识别的特征将识别点云的区域,其中可以(并行地)使用持续同源性的放大步骤和区域计算来建立这些特征的更精确边界。该项目将开发算法的改进,对算法和近似技术的正确性,误差范围和复杂性的正式声明。这些技术对将拓扑数据分析技术应用于比目前可能的大得多的数据集的能力具有重要意义。该奖项反映了NSF的法定使命,并被认为值得通过使用基金会的知识价值和更广泛的影响审查标准进行评估来支持。
英文摘要
New insights with machine learning exists across may domains, including, for example, medicine, social media, image processing, biology, and computer and network security. Machine learning is able to process large, high-dimensional data sets that are beyond human capabilities. One emerging method of machine learning is based on a branch of mathematics called topology that is sometimes able to discover knowledge that is not available using conventional methods. The field of topology is concerned with of the shape of an object and Persistent Homology is the critical method in topology used to extract the features of a shape. Persistent Homology will classify an object by the size and number of holes and voids in that object. Unfortunately, computing the Persistent Homology for an object requires significant amounts of memory and long run-times that increases exponentially in the number of points that forms that object. This project will treat the object formed by the data and subdivide it into smaller regions for the parallel computation of Persistent Homology on each region. The results from the regional analyses will then be assembled together and any duplicate or missing results will be identified and restored in a post analysis step. The computation on all of the regions will be completed in substantially less time and in much less total memory than a single computation on the entire data set. Testing of the methods developed will be performed using a variety of synthetic and real-world data. The synthetic data will permit controlled studies on performance and scalability. Realworld data from a variety of sources and especially data where the small topological features are significant (such as data from brain scans) will be used. This project will propel the application of topology based analysis to discover new insights and meaningful information from massive high-dimensional data. An expansion of student training in data mining through topological-based methods will be achieved with the addition of classes, projects (senior project, MS Theses, PhD Dissertations, and so on), seminars, and research co-op training experiences. Students at all levels will be impacted and special emphasis placed on minority and underrepresented student groups participation. This project will also participate in the Women in Science and Engineering programs at UC. The project investigators will engage local area K-12 students, international exchange students and researchers at UC's collaborative institutions, UC's Medical School, Cincinnati Children's Medical Center, the Air Force Research Lab, and local industries with information and seminars on this project investigations and results.This project proposes to combine the fields of Approximate Computing with Topological Data Analysis to dramatically reduce the computational and memory requirements to use Topological Data Analysis on very large data sets. In particular, this project will develop approximate methods for computing Persistent Homology that dramatically increase the sizes of data sets for which data mining methods based on topological data analysis can be applied. This project expects to increase the size of the input data set that can be analyzed by Topological Data Analysis methods by at least 3-5 orders of magnitude. While approximate methods can introduce error, the features identified by the approximate methods will identify regions of the point cloud where an upscaling steps and regional computations of Persistent Homology can be used (in parallel) to establish more precise boundaries of those features. The project will develop algorithmic improvements, formal statements on the correctness, error bounds, and complexities of the algorithms and approximation techniques. These techniques have important implications on the ability to apply topological data analysis techniques to much larger data sets than currently possible.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1109/bigdata55660.2022.10020926
发表时间: 2022
期刊: IEEE International Conference on Big Data
影响因子: --
作者: [Singh, Rohit P., Wilsey, Philip A.]
通讯作者: Wilsey, Philip A.
DOI: 10.1109/bigdata47090.2019.9006572
发表时间: 2019-12
期刊: 2019 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者: [Nicholas O. Malott;P. Wilsey]
通讯作者: Nicholas O. Malott;P. Wilsey
Computation of persistent homology on streaming data using topological data summaries
使用拓扑数据摘要计算流数据上的持久同源性
DOI: 10.1111/coin.12597
发表时间: 2023
期刊: Computational Intelligence
影响因子: 2.8
作者: [Moitra, Anindya, Malott, Nicholas O., Wilsey, Philip A.]
通讯作者: Wilsey, Philip A.
DOI: 10.1109/bigdata50022.2020.9378216
发表时间: 2020-12
期刊: 2020 IEEE International Conference on Big Data (Big Data)
影响因子: --
作者: [Nicholas O. Malott;Aaron M. Sens;P. Wilsey]
通讯作者: Nicholas O. Malott;Aaron M. Sens;P. Wilsey
9
    SI2-SSE: Scalable Big Data Clustering by Random Projection Hashing
    • 批准号:
      1440420
    • 项目类别:
      Standard Grant
    • 资助金额:
      $49.81万
    • 财政年份:
      2014
    • 负责人:
      Philip Wilsey
    • 依托单位:
    CSR: Small: Collaborative Research: Combining Static Analysis and Dynamic Run-time Optimization for Parallel Discrete Event Simulation in Many-Core Environments
    • 批准号:
      0915337
    • 项目类别:
      Standard Grant
    • 资助金额:
      $16.69万
    • 财政年份:
      2009
    • 负责人:
      Philip Wilsey
    • 依托单位:
    国内基金
    海外基金
    昼夜节律性small RNA在血斑形成时间推断中的法医学应用研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      --
    • 批准年份:
      2024
    • 负责人:
    • 依托单位:
    tRNA-derived small RNA上调YBX1/CCL5通路参与硼替佐米诱导慢性疼痛的机制研究
    • 批准号:
    • 项目类别:
      省市级项目
    • 资助金额:
      10.0万元
    • 批准年份:
      2022
    • 负责人:
      张祥忠
    • 依托单位:
    Small RNA调控I-F型CRISPR-Cas适应性免疫性的应答及分子机制
    Small RNAs调控解淀粉芽胞杆菌FZB42生防功能的机制研究
    • 批准号:
      31972324
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2019
    • 负责人:
      高学文
    • 依托单位: