AF: Large: Collaborative Research: Compact Representations and Efficient Algorithms for Distributed Geometric Data
AF: Large: Collaborative Research: Compact Representations and Efficient Algorithms for Distributed Geometric Data
批准号:
1012042
负责人:
Piotr Indyk
金额:
$43.3万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-09-01 至 2014-08-31
中文摘要
在科学、工程和商业的许多领域,高带宽传感器和摄像头、大规模模拟或网络支持的大规模数据收集正在以前所未有的速度生成大量数据集。这些数据大多具有直接或间接的几何特征。例如,第二代激光雷达可以绘制15-20厘米分辨率的地球表面;大型天气望远镜每晚将产生大约30tb的数据;每分钟有13个小时的视频上传到YouTube;Facebook管理着超过400亿张照片,需要超过1pb的数据。这些数据集为实现几年前无法想象的新功能提供了巨大的机会。然而,利用这些机会,并将这些海量异构数据转换为适用于各种不同类型的应用程序和用户的有用信息,需要解决具有挑战性的算法问题。解决这一挑战的有效方法是设计有效的方法来生成这些几何数据集的信息丰富而简洁的摘要。这些摘要必须在多个尺度上工作,并允许对各种各样的查询进行近似但有效的回答。该项目的目标是研究紧凑表示和高效算法的理论基础,用于组织、总结、交叉关联、互连和查询大型分布式几何数据集。本项目将设计计算多种风味摘要的方法,这些方法都具有可证明的性质。摘要可以是组合的和度量的(核心集和核)、代数的(线性草图)、拓扑的(持久性图)、基于特征的和结构的(对数据中的自相似性进行编码)。它们旨在捕获的属性从低级度量属性(如点集的直径或宽度)扩展到揭示数据内部结构的高级属性(如检测对称性和重复模式)。这种处理必须在来自传感器的数据存在不确定性的情况下完成,并优化多种性能指标,包括分布在网络中多个位置的数据的通信成本。这个项目的另一个关键方面是,它的目的不是孤立地了解单个数据集,而是了解不同数据集之间的相互关系和对应关系,并且只通过交流摘要信息来做到这一点,甚至没有将所有数据放在一个地方。这项工作涉及理论计算机科学和应用数学中的许多主题,包括低失真嵌入、压缩感知、运输度量、谱图理论或谐波分析、机器学习和计算拓扑。
英文摘要
Across many fields of science, engineering, and business, massive data sets are being generated at unprecedented rate by high-bandwidth sensors and cameras, large-scale simulations, or web-enabled large scale data collection. Much of this data has a geometric character, either directly or indirectly. For example, second generation LiDARs can map the earth's surface at 15-20 cm resolution; the Large Synoptic Telescope is set to produce about 30 terabytes of data each night; thirteen hours of video are uploaded to YouTube every minute; Facebook manages over 40 billion photos requiring more than one petabyte of data.These data sets provide tremendous opportunities to enable novel capabilities that were unimaginable a few years ago. Capitalizing on these opportunities, however, and transforming these massive amounts of heterogeneous data into useful information for vastly different types of applications and users requires solving challenging algorithmic problems. An effective way of addressing this challenge is by designing efficient methods for producing informative yet succinct summaries of such geometric data sets. These summaries must work at multiple scales, and allow a wide variety of queries to be answered approximately but efficiently. The goal of this project is to study the theoretical underpinnings of compact representations and efficient algorithms for organizing, summarizing, cross-correlating, interlinking, and querying large distributed geometric data sets.This project will design methods for computing summaries of many kinds of flavors, all with provable properties. Summaries can be combinatorial and metric (core sets and kernels), algebraic (linear sketches), topological (persistence diagrams), feature-based, and structural (encoding self-similarities in the data). The properties they aim to capture extend from low-level metric attributes, such as the diameter or width of a point set, to higher-level attributes revealing the internal structure of the data, as in the detection of symmetries and repeated patterns. This processing must be done in the presence of uncertainty in data coming from sensors, and optimize multiple performance measures, including communication cost for data distributed across multiple locations in a network. Another key aspect of this project is that it aims to understand not individual data sets in isolation but rather the inter-relationships and correspondences among different data sets, and to do so by communicating only summary information, without even having all the data in one place. This work touches upon many topics in theoretical computer science and applied mathematics including low-distortion embeddings, compressive sensing, transportation metrics, spectral graph theory or harmonic analysis, machine learning, and computational topology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Travel: SODA 2024 Conference Student and Postdoc Travel Support
-
批准号:2343779
-
项目类别:Standard Grant
-
资助金额:$1.5万
-
财政年份:2023
-
负责人:Piotr Indyk
-
依托单位:
Conference: SODA 2023 Conference Student and Postdoc Travel Support
-
批准号:2232958
-
项目类别:Standard Grant
-
资助金额:$1.2万
-
财政年份:2022
-
负责人:Piotr Indyk
-
依托单位:
Foundations of Data Science Institute
-
批准号:2022448
-
项目类别:Continuing Grant
-
资助金额:$549.03万
-
财政年份:2020
-
负责人:Piotr Indyk
-
依托单位:
Collaborative Research: AF: Small: Fine-Grained Complexity of Approximate Problems
-
批准号:2006798
-
项目类别:Standard Grant
-
资助金额:$20.0万
-
财政年份:2020
-
负责人:Piotr Indyk
-
依托单位:
TRIPODS: Institute for Foundations of Data Science (IFDS)
-
批准号:1740751
-
项目类别:Continuing Grant
-
资助金额:$136.85万
-
财政年份:2017
-
负责人:Piotr Indyk
-
依托单位:
AitF: FULL: Sparse Fourier Transform: From Theory to Practice
-
批准号:1535851
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2015
-
负责人:Piotr Indyk
-
依托单位:
BIGDATA: F: DKA: Collaborative Research: Structured Nearest Neighbor Search in High Dimensions
-
批准号:1447476
-
项目类别:Standard Grant
-
资助金额:$50.0万
-
财政年份:2015
-
负责人:Piotr Indyk
-
依托单位:
Fast Approximate Algorithms for Wireless Sensor Networks
-
批准号:0728645
-
项目类别:Standard Grant
-
资助金额:$0.0万
-
财政年份:2007
-
负责人:Piotr Indyk
-
依托单位:
CAREER: Approximate Algorithms for High-dimensional Geometric Problems
-
批准号:0133849
-
项目类别:Continuing Grant
-
资助金额:$0.0万
-
财政年份:2002
-
负责人:Piotr Indyk
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: