课题基金 / 基金详情

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications
BIGDATA:中型:DA:协作研究:大张量挖掘:理论、可扩展算法和应用
批准号:
1247489
负责人:
Christos Faloutsos
金额:
$89.49万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-12-01 至 2018-09-30

项目摘要

项目成果

Christos Faloutsos的其他基金

相似基金

相关文献

中文摘要
翻译
张量是矩阵的多维推广,因此可以具有非数字条目。极大和稀疏的耦合张量出现在许多重要的应用中,这些应用需要分析大的、不同的和部分相关的数据。耦合张量的有效分析需要开发算法和相关软件,以识别不同张量模式之间存在的核心关系,并扩展到极大的数据集。这个项目的目标是开发(耦合)稀疏和低阶张量分解的理论和算法,以及相关的可扩展软件工具包,使这种分析成为可能。该项目的研究集中在三个主要方面。第一个是通过发展多路压缩传感降维方法和完美的潜在模型重构,在耦合张量分解领域做出新的理论贡献。还将开发处理缺失值、噪声输入和耦合数据的方法。第二个重点是现代体系结构上的算法和可伸缩性,这将使用映射-归约范例以及混合多核体系结构来实现对具有数百万和数十亿个非零条目的耦合张量的高效分析。将开发一个开源的耦合张量分解工具箱(HTF-混合张量分解),它将提供这些算法的健壮和高性能实现。最后,第三个重点是在一个神经语义学应用程序上评估和验证这些耦合的因式分解算法的有效性,该应用程序的目标是通过分析阅读各种文本段落时获得的fMRI和脑磁图图像数据,了解人脑活动如何与文本阅读和放大理解相关。给定事实(主语-动词-宾语)的三元组,比如(华盛顿是美国的大写),我们能找到模式、新宾语、新动词和异常吗?我们能否将这些与人们阅读这些单词的大脑扫描相关联,以发现大脑的哪些部分被激活,比如,被类似工具的名词(“锤子”)或类似动作的动词(“run”)激活?我们提出了一个统一的“耦合张量”分解框架来系统地挖掘这类数据集。这些环境中的独特挑战包括(A)太字节和千万亿字节的缩放问题,(B)分布式容错计算,(C)大量数据丢失,以及(D)关于大稀疏张量的理论和方法不足。这一努力的智力价值正是上述四个挑战的解决方案。更广泛的影响是得出了关于大脑如何工作以及如何处理语言的新的科学假设(来自永无止境的语言学习(Nell)和神经语义学项目),并开发了用于耦合张量因式分解的可扩展开源软件。我们的张量分析方法还可以用于许多其他环境,包括推荐系统和计算机网络入侵/异常检测。KEYWORDS:数据挖掘;映射/约简;网络阅读;神经语义;张量。
英文摘要
Tensors are multi-dimensional generalizations of matrices, and so can have non-numeric entries. Extremely large and sparse coupled tensors arise in numerous important applications that require the analysis of large, diverse, and partially related data. The effective analysis of coupled tensors requires the development of algorithms and associated software that can identify the core relations that exist among the different tensor modes, and scale to extremely large datasets. The objective of this project is to develop theory and algorithms for (coupled) sparse and low-rank tensor factorization, and associated scalable software toolkits to make such analysis possible. The research in the project is centered on three major thrusts. The first is designed to make novel theoretical contributions in the area of coupled tensor factorization, by developing multi-way compressed sensing methods for dimensionality reduction with perfect latent model reconstruction. Methods to handle missing values, noisy input, and coupled data will also be developed. The second thrust focuses on algorithms and scalability on modern architectures, which will enable the efficient analysis of coupled tensors with millions and billions of non-zero entries, using the map-reduce paradigm, as well as hybrid multicore architectures. An open-source coupled tensor factorization toolbox (HTF- Hybrid Tensor Factorization) will be developed that will provide robust and high-performance implementations of these algorithms. Finally, the third thrust focuses on evaluating and validating the effectiveness of these coupled factorization algorithms on a NeuroSemantics application whose goal is to understand how human brain activity correlates with text reading & understanding by analyzing fMRI and MEG brain image datasets obtained while reading various text passages.Given triplets of facts (subject-verb-object), like ('Washington' 'is the capital of' 'USA'), can we find patterns, new objects, new verbs, anomalies? Can we correlate these with brain scans of people reading these words, to discover which parts of the brain get activated, say, by tool-like nouns ('hammer'), or action-like verbs ('run')? We propose a unified "coupled tensor" factorization framework to systematically mine such datasets. Unique challenges in these settings include (a) tera- and peta-byte scaling issues, (b) distributed fault-tolerant computation, (c) large proportions of missing data, and (d) insufficient theory and methods for big sparse tensors. The Intellectual Merit of this effort is exactly the solution to the above four challenges.The Broader Impact is the derivation of new scientific hypotheses on how the brain works and how it processes language (from the never-ending language learning (NELL) and NeuroSemantics projects) and the development of scalable open source software for coupled tensor factorization. Our tensor analysis methods can also be used in many other settings, including recommendation systems and computer-network intrusion/anomaly detection.KEYWORDS:Data mining; map/reduce; read-the-web; neuro-semantics; tensors.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Medium: Collaborative Research: Collective Opinion Fraud Detection: Identifying and Integrating Cues from Language, Behavior, and Networks
  • 批准号:
    1408924
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.99万
  • 财政年份:
    2014
  • 负责人:
    Christos Faloutsos
  • 依托单位:
TWC: Medium: Collaborative: Know Thy Enemy: Data Mining Meets Networks for Understanding Web-Based Malware Dissemination
  • 批准号:
    1314632
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.33万
  • 财政年份:
    2013
  • 负责人:
    Christos Faloutsos
  • 依托单位:
CGV: Small: Making Sense out of Large Graphs - Bridging HCI with Data Mining
  • 批准号:
    1217559
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2012
  • 负责人:
    Christos Faloutsos
  • 依托单位:
III: Small: Influence and Virus Propagation in Large Graphs - Theory and Algorithms
  • 批准号:
    1017415
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.97万
  • 财政年份:
    2010
  • 负责人:
    Christos Faloutsos
  • 依托单位:
国内基金
海外基金
肝细胞Mid 1活化加重脓毒症病理进程的分子机制研究及干预策略优化
MID1调控肿瘤相关巨噬细胞细胞中IRF8-STING通路在胶质瘤微环境中的作用机制研究
  • 批准号:
    2025JJ70385
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    傅尧
  • 依托单位:
E3泛素连接酶Mid1调控Treg细胞影响GVHD 的作用及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
线粒体动力蛋白MiD51在IL-27诱导类风湿关节炎DN2-B细胞分化扩增中的作用及机制研究
  • 批准号:
    82302047
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2023
  • 负责人:
    汤亚微
  • 依托单位: