EAPSI: Developing Fast and Accurate Methods for Grouping Objects in a Dataset Using Inconsistent Labels
EAPSI: Developing Fast and Accurate Methods for Grouping Objects in a Dataset Using Inconsistent Labels
批准号:
1613938
负责人:
Luke Veldt
金额:
$0.54万
依托单位:
依托单位国家:
美国
项目类别:
Fellowship Award
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-06-15 至 2017-05-31
中文摘要
在当今信息丰富的时代,在几乎任何领域试图解决问题和回答问题时,收集大量数据都相对容易。然而,由于这些数据集的庞大规模和缺乏组织,从这些数据集中分析和提取有用的信息往往非常具有挑战性。这个项目将研究将数据集组织成类似对象组的特殊技术,以便于分析。这个任务被称为聚类,可以应用于非常不同的设置,例如生物学中蛋白质相互作用的研究,或大型数据库中网页的分类。这项研究将在墨尔本大学与安东尼·沃思教授合作进行。Wirth博士是数据分析专家,也是“建议聚类”研究的先驱,这是一种在唯一可用的信息是一系列不一致的标签时对数据进行聚类的技术,这些标签将数据点标记为“相似”或“不相似”。“一般来说,在大型数据集上精确解决这个聚类问题是缓慢和计算昂贵的。本研究旨在探索输入数据集上的基本假设,这些假设可以导致更有效的方法用于不一致标签的聚类。开发更快的方法来解决这个问题,将扩大目前的理论理解与建议的聚类,以及使这个有用的技术在实践中更容易实现。研究人员将应用整数约束线性规划和数值线性代数的技术,以获得一个解决方案的聚类与建议的问题时,输入数据可以表示为一个低秩矩阵。的主要目标是开发一个多项式时间的算法来解决这个问题的秩为2矩阵,并证明复杂性的结果,这个特殊情况下的问题。在低秩假设下获得快速解决方案,有助于了解在实践中什么时候困难的问题变得容易处理,并将激励研究在其他困难问题上显示类似的结果。该奖项是东亚和太平洋夏季研究所计划下的一个奖项,由NSF和澳大利亚科学院共同资助,支持美国研究生的夏季研究。
英文摘要
In today's information-rich age, it is relatively easy to collect large amounts of data when attempting to solve problems and answer questions in almost any field. It is often very challenging, though, to analyze and extract useful information from these datasets, due to their massive size and lack of organization. This project will investigate special techniques for organizing a dataset into groups of similar objects to allow for easier analysis. This task is called clustering, and can be applied in vastly different settings, such as the study of protein interactions in biology, or the categorization of webpages in a large database. The research will be conducted at the University of Melbourne in collaboration with Professor Anthony Wirth. Dr. Wirth is an expert in data analysis and a pioneer in the study of "clustering with advice," a technique for clustering data when the only available information is a list of inconsistent labels that mark data points as "similar" or "dissimilar." In general, exactly solving this clustering problem on a large dataset is slow and computationally expensive. This research aims to explore basic assumptions on the input dataset that can lead to more efficient methods for clustering with inconsistent labels. Developing faster methods for this problem will expand current theoretical understanding of clustering with advice, as well as making this useful technique more achievable in practice.The researcher will apply techniques in integer-constrained linear programming and numerical linear algebra to obtain a solution for the clustering with advice problem when the input data can be represented as a low-rank matrix. The primary objective is to develop a polynomial time algorithm for solving the problem for rank-2 matrices and prove complexity results about this special case of the problem. Obtaining a fast solution under low-rank assumptions sheds light on when an otherwise hard problem becomes tractable in practice, and will stimulate research in showing similar results for other difficult problems.This award under the East Asia and Pacific Summer Institutes program supports summer research by a U.S. graduate student and is jointly funded by NSF and the Australia Academy of Science.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金