课题基金 / 基金详情

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications

BIGDATA: Mid-Scale: DA: Collaborative Research: Big Tensor Mining: Theory, Scalable Algorithms and Applications
BIGDATA:中型:DA:协作研究:大张量挖掘:理论、可扩展算法和应用
批准号:
1247632
负责人:
Nikolaos Sidiropoulos
金额:
$86.68万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2012
资助国家:
美国
项目状态:
已结题
起止时间:
2012-12-01 至 2017-11-30

项目摘要

项目成果

Nikolaos Sidiropoulos的其他基金

相似基金

相关文献

中文摘要
翻译
张量是矩阵的多维推广,因此可以有非数值项。在许多需要分析大量、多样和部分相关数据的重要应用中,会出现极大且稀疏的耦合张量。耦合张量的有效分析需要开发算法和相关软件,以识别不同张量模式之间存在的核心关系,并扩展到非常大的数据集。该项目的目标是开发(耦合)稀疏和低秩张量分解的理论和算法,以及相关的可扩展软件工具包,使这种分析成为可能。该项目的研究主要集中在三个方面。第一个目的是在耦合张量因式分解领域做出新的理论贡献,通过开发多路压缩感知方法进行降维与完美的潜在模型重建。还将开发处理缺失值、噪声输入和耦合数据的方法。第二个重点是现代架构上的算法和可扩展性,这将使具有数百万和数十亿个非零条目的耦合张量的有效分析成为可能,使用映射简化范式以及混合多核架构。一个开源的耦合张量分解工具箱(HTF-混合张量分解)将被开发,这将提供这些算法的强大和高性能的实现。最后,第三个重点是评估和验证这些耦合因子分解算法在神经语义学应用程序上的有效性,该应用程序的目标是&通过分析在阅读各种文本段落时获得的fMRI和MEG大脑图像数据集来了解人类大脑活动如何与文本阅读理解相关。(主语-动词-宾语),比如('Washington' 'is the capital of' 'USA'),我们能找到模式,新的宾语,新的动词,异常吗?我们能否将这些与人们阅读这些单词时的大脑扫描相关联,以发现大脑的哪些部分被激活,比如说,工具类名词(“锤子”)或动作类动词(“跑”)?我们提出了一个统一的“耦合张量”因子分解框架,系统地挖掘这样的数据集。这些环境中的独特挑战包括:(a)万亿字节和千万亿字节的缩放问题,(B)分布式容错计算,(c)大比例的丢失数据,以及(d)大稀疏张量的理论和方法不足。这一努力的智力价值正是上述四个挑战的解决方案。更广泛的影响是推导出关于大脑如何工作和如何处理语言的新科学假设(来自永无止境的语言学习(NELL)和神经语义学项目),并开发可扩展的开源软件用于耦合张量因式分解。我们的张量分析方法也可以用于许多其他设置,包括推荐系统和计算机网络入侵/异常检测。关键词:数据挖掘;映射/减少;读取网络;神经语义学;张量。
英文摘要
Tensors are multi-dimensional generalizations of matrices, and so can have non-numeric entries. Extremely large and sparse coupled tensors arise in numerous important applications that require the analysis of large, diverse, and partially related data. The effective analysis of coupled tensors requires the development of algorithms and associated software that can identify the core relations that exist among the different tensor modes, and scale to extremely large datasets. The objective of this project is to develop theory and algorithms for (coupled) sparse and low-rank tensor factorization, and associated scalable software toolkits to make such analysis possible. The research in the project is centered on three major thrusts. The first is designed to make novel theoretical contributions in the area of coupled tensor factorization, by developing multi-way compressed sensing methods for dimensionality reduction with perfect latent model reconstruction. Methods to handle missing values, noisy input, and coupled data will also be developed. The second thrust focuses on algorithms and scalability on modern architectures, which will enable the efficient analysis of coupled tensors with millions and billions of non-zero entries, using the map-reduce paradigm, as well as hybrid multicore architectures. An open-source coupled tensor factorization toolbox (HTF- Hybrid Tensor Factorization) will be developed that will provide robust and high-performance implementations of these algorithms. Finally, the third thrust focuses on evaluating and validating the effectiveness of these coupled factorization algorithms on a NeuroSemantics application whose goal is to understand how human brain activity correlates with text reading & understanding by analyzing fMRI and MEG brain image datasets obtained while reading various text passages.Given triplets of facts (subject-verb-object), like ('Washington' 'is the capital of' 'USA'), can we find patterns, new objects, new verbs, anomalies? Can we correlate these with brain scans of people reading these words, to discover which parts of the brain get activated, say, by tool-like nouns ('hammer'), or action-like verbs ('run')? We propose a unified "coupled tensor" factorization framework to systematically mine such datasets. Unique challenges in these settings include (a) tera- and peta-byte scaling issues, (b) distributed fault-tolerant computation, (c) large proportions of missing data, and (d) insufficient theory and methods for big sparse tensors. The Intellectual Merit of this effort is exactly the solution to the above four challenges.The Broader Impact is the derivation of new scientific hypotheses on how the brain works and how it processes language (from the never-ending language learning (NELL) and NeuroSemantics projects) and the development of scalable open source software for coupled tensor factorization. Our tensor analysis methods can also be used in many other settings, including recommendation systems and computer-network intrusion/anomaly detection.KEYWORDS:Data mining; map/reduce; read-the-web; neuro-semantics; tensors.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Blind Carbon Copy on Dirty Paper: Seamless Spectrum Underlay made Practical
  • 批准号:
    2118002
  • 项目类别:
    Standard Grant
  • 资助金额:
    $38.8万
  • 财政年份:
    2021
  • 负责人:
    Nikolaos Sidiropoulos
  • 依托单位:
III: Small: A Submodular Framework for Scalable Graph Matching with Performance Guarantees
  • 批准号:
    1908070
  • 项目类别:
    Standard Grant
  • 资助金额:
    $45.67万
  • 财政年份:
    2019
  • 负责人:
    Nikolaos Sidiropoulos
  • 依托单位:
Robust and Scalable Volume Minimization-based Matrix Factorization for Sensing and Clustering
  • 批准号:
    1852831
  • 项目类别:
    Standard Grant
  • 资助金额:
    $24.93万
  • 财政年份:
    2018
  • 负责人:
    Nikolaos Sidiropoulos
  • 依托单位:
Collaborative Research: Multimodal Sensing and Analytics at Scale: Algorithms and Applications
  • 批准号:
    1807660
  • 项目类别:
    Standard Grant
  • 资助金额:
    $20.0万
  • 财政年份:
    2018
  • 负责人:
    Nikolaos Sidiropoulos
  • 依托单位:
国内基金
海外基金
肝细胞Mid 1活化加重脓毒症病理进程的分子机制研究及干预策略优化
MID1调控肿瘤相关巨噬细胞细胞中IRF8-STING通路在胶质瘤微环境中的作用机制研究
  • 批准号:
    2025JJ70385
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    傅尧
  • 依托单位:
E3泛素连接酶Mid1调控Treg细胞影响GVHD 的作用及机制研究
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2024
  • 负责人:
  • 依托单位:
线粒体动力蛋白MiD51在IL-27诱导类风湿关节炎DN2-B细胞分化扩增中的作用及机制研究
  • 批准号:
    82302047
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2023
  • 负责人:
    汤亚微
  • 依托单位: