课题基金 / 基金详情

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications
BIGDATA:中型:DA:协作研究:大张量挖掘:理论、可扩展算法和应用
批准号:
1247489
负责人:
Christos Faloutsos
金额:
$89.49万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-12-01 至 2018-09-30

项目摘要

项目成果

Christos Faloutsos的其他基金

相似基金

相关文献

中文摘要
翻译
张量是矩阵的多维推广,因此可以有非数值项。在许多需要分析大量、多样和部分相关数据的重要应用中,出现了极大的稀疏耦合张量。耦合张量的有效分析需要开发算法和相关软件,以识别不同张量模式之间存在的核心关系,并扩展到非常大的数据集。该项目的目标是开发(耦合)稀疏和低秩张量分解的理论和算法,以及相关的可扩展软件工具包,使这种分析成为可能。该项目的研究主要集中在三个方面。首先,通过开发具有完美潜在模型重建的多路压缩感知降维方法,在耦合张量分解领域做出新的理论贡献。还将开发处理缺失值、噪声输入和耦合数据的方法。第二个重点是现代架构上的算法和可扩展性,这将使使用map-reduce范式以及混合多核架构能够有效地分析具有数百万和数十亿非零条目的耦合张量。一个开源的耦合张量分解工具箱(HTF- Hybrid tensor factorization)将被开发出来,它将提供这些算法的鲁棒性和高性能实现。最后,第三个重点是评估和验证这些耦合分解算法在一个神经语义应用程序上的有效性,该应用程序的目标是通过分析阅读各种文本段落时获得的fMRI和MEG脑图像数据集,了解人类大脑活动与文本阅读和理解之间的关系。给定三组事实(主语-动词-宾语),比如(“华盛顿”是“美国”的首都),我们能找到规律、新宾语、新动词和异常现象吗?我们能否将这些与阅读这些单词的人的脑部扫描联系起来,以发现大脑的哪些部分被激活,比如,被类似工具的名词(“锤子”)或类似动作的动词(“跑”)激活?我们提出了一个统一的“耦合张量”分解框架来系统地挖掘这些数据集。这些环境中的独特挑战包括(a)兆字节和千兆字节缩放问题,(b)分布式容错计算,(c)大量丢失数据,以及(d)大稀疏张量的理论和方法不足。这种努力的智力价值正是解决上述四个挑战的方法。更广泛的影响是关于大脑如何工作以及如何处理语言的新科学假设的衍生(来自永无止境的语言学习(NELL)和神经语义项目)以及耦合张量分解的可扩展开源软件的开发。我们的张量分析方法也可以用于许多其他设置,包括推荐系统和计算机网络入侵/异常检测。关键词:数据挖掘;map / reduce;read-the-web;neuro-semantics;张量。
英文摘要
Tensors are multi-dimensional generalizations of matrices, and so can have non-numeric entries. Extremely large and sparse coupled tensors arise in numerous important applications that require the analysis of large, diverse, and partially related data. The effective analysis of coupled tensors requires the development of algorithms and associated software that can identify the core relations that exist among the different tensor modes, and scale to extremely large datasets. The objective of this project is to develop theory and algorithms for (coupled) sparse and low-rank tensor factorization, and associated scalable software toolkits to make such analysis possible. The research in the project is centered on three major thrusts. The first is designed to make novel theoretical contributions in the area of coupled tensor factorization, by developing multi-way compressed sensing methods for dimensionality reduction with perfect latent model reconstruction. Methods to handle missing values, noisy input, and coupled data will also be developed. The second thrust focuses on algorithms and scalability on modern architectures, which will enable the efficient analysis of coupled tensors with millions and billions of non-zero entries, using the map-reduce paradigm, as well as hybrid multicore architectures. An open-source coupled tensor factorization toolbox (HTF- Hybrid Tensor Factorization) will be developed that will provide robust and high-performance implementations of these algorithms. Finally, the third thrust focuses on evaluating and validating the effectiveness of these coupled factorization algorithms on a NeuroSemantics application whose goal is to understand how human brain activity correlates with text reading & understanding by analyzing fMRI and MEG brain image datasets obtained while reading various text passages.Given triplets of facts (subject-verb-object), like ('Washington' 'is the capital of' 'USA'), can we find patterns, new objects, new verbs, anomalies? Can we correlate these with brain scans of people reading these words, to discover which parts of the brain get activated, say, by tool-like nouns ('hammer'), or action-like verbs ('run')? We propose a unified "coupled tensor" factorization framework to systematically mine such datasets. Unique challenges in these settings include (a) tera- and peta-byte scaling issues, (b) distributed fault-tolerant computation, (c) large proportions of missing data, and (d) insufficient theory and methods for big sparse tensors. The Intellectual Merit of this effort is exactly the solution to the above four challenges.The Broader Impact is the derivation of new scientific hypotheses on how the brain works and how it processes language (from the never-ending language learning (NELL) and NeuroSemantics projects) and the development of scalable open source software for coupled tensor factorization. Our tensor analysis methods can also be used in many other settings, including recommendation systems and computer-network intrusion/anomaly detection.KEYWORDS:Data mining; map/reduce; read-the-web; neuro-semantics; tensors.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Medium: Collaborative Research: Collective Opinion Fraud Detection: Identifying and Integrating Cues from Language, Behavior, and Networks
  • 批准号:
    1408924
  • 项目类别:
    Standard Grant
  • 资助金额:
    $29.99万
  • 财政年份:
    2014
  • 负责人:
    Christos Faloutsos
  • 依托单位:
TWC: Medium: Collaborative: Know Thy Enemy: Data Mining Meets Networks for Understanding Web-Based Malware Dissemination
  • 批准号:
    1314632
  • 项目类别:
    Standard Grant
  • 资助金额:
    $33.33万
  • 财政年份:
    2013
  • 负责人:
    Christos Faloutsos
  • 依托单位:
CGV: Small: Making Sense out of Large Graphs - Bridging HCI with Data Mining
  • 批准号:
    1217559
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $49.98万
  • 财政年份:
    2012
  • 负责人:
    Christos Faloutsos
  • 依托单位:
III: Small: Influence and Virus Propagation in Large Graphs - Theory and Algorithms
  • 批准号:
    1017415
  • 项目类别:
    Standard Grant
  • 资助金额:
    $49.97万
  • 财政年份:
    2010
  • 负责人:
    Christos Faloutsos
  • 依托单位:
国内基金
海外基金
肝细胞Mid 1活化加重脓毒症病理进程的分子机制研究及干预策略优化
MID1调控肿瘤相关巨噬细胞细胞中IRF8-STING通路在胶质瘤微环境中的作用机制研究
  • 批准号:
    2025JJ70385
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    傅尧
  • 依托单位:
E3泛素连接酶Mid1调控Treg细胞影响GVHD 的作用及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
线粒体动力蛋白MiD51在IL-27诱导类风湿关节炎DN2-B细胞分化扩增中的作用及机制研究
  • 批准号:
    82302047
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2023
  • 负责人:
    汤亚微
  • 依托单位: