BIGDATA: F: Large-Scale Transductive Learning from Heterogeneous Data Sources
BIGDATA: F: Large-Scale Transductive Learning from Heterogeneous Data Sources
批准号:
1546329
负责人:
Yiming Yang
金额:
$118.85万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-01-01 至 2020-12-31
中文摘要
大数据时代的重要问题涉及基于不同信息源的预测和数据中的依赖结构。例如,在推荐系统中,不仅需要基于观察到的用户对物品(电影、书籍、音乐、购物产品等)的评级,而且还需要基于诸如用户的人口统计数据和物品的文本描述之类的信息来进行预测。在从文本数据(新闻故事、推文、维护报告、法律文件等)进行事件检测时,联合推理必须基于谁(代理)、什么(事件类型或主题)、何处(位置)和何时(日期),还必须基于代理(在社交网络中)、主题(在事件类型本体中)、位置(在地图中)和时间共现之间的联系。因此,基本的研究问题包括:(1)如何根据各种任务中的异质信息和依赖结构建立一个统一的预测优化框架:(2)当模型参数的组合空间非常大时,如何使推理易于计算;以及(3)如何利用大量可用的未标注数据以及通常稀疏的人工标注训练数据来显著提高系统的预测能力。本项目将通过以下四种方法来解决这三个挑战:(1)使用产品图统一表示异质信息源:该框架旨在表示异质数据源和源内依赖关系,如用户之间的社交联系、项目之间的语义相似度、关键词之间的上下文相关性、文档之间的主题相似性、主题标签之间的层次关系等。每个数据源将使用图来表示,多个源的单独图将被组合成乘积图,其中每个节点对应于单独图中的节点的元组,并且每个链接聚合单独图中的链接。(2)基于图形乘积的转导学习:本项目计划将大范围预测任务中的推理问题简化为上述乘积图形上的半监督转导学习问题。每个任务(分类、回归或链接预测)中的训练数据将被表示为乘积图中已标记(或已评分)节点的子集,并且这些节点的标签(或分数)将在乘积图中的链接上传播,直到收敛。本课题将从理论和实验两个方面对各种图形变换进行研究。(3)大规模优化算法:导出的乘积图通常都非常大。为了解决计算瓶颈,本项目将根据谱图产品的理论属性和计算特性开发新的可扩展算法,包括降阶矩阵分解、积极基剪枝和基于采样的低阶近似的改编版本。(4)在多个重要应用中进行全面评估:所提出的新方法将在基准数据集上进行评估,用于上下文感知协同过滤、半结构化事件检测和跟踪,以及通过多源社会网络分析进行专家发现。如果该工作成功,将为在涉及推荐、分类和回归的广泛任务中增强系统的预测能力提供原则性的解决方案。预计拟议工作将在多个研究领域产生技术影响。欲了解更多信息,请访问项目网站:http://nyc.lti.cs.cmu.edu/gp-trans/index.html
英文摘要
Important problems in the big-data era involve predictions based on heterogeneous sources of information and the dependency structures in data. In recommendation systems, for example, predictions need to be made not only based on observed user ratings over items (movies, books, music, shopping products, etc.), but also based on information such as demographical data of users and textual descriptions of items. In event detection from textual data (news stories, tweets, maintenance reports, legal documents, etc.), joint inference must be based on who (agents), what (event types or topics), where (locations) and when (dates), and also based on the connections among agents (in social networks), topics (in an event-type ontology), locations (in a map) and temporal co-occurrences. The fundamental research questions therefore include: (1) how to develop a unified optimization framework for predictions based on heterogeneous information and dependency structures in various kinds of tasks; (2) how to make the inference computationally tractable when the combined space of model parameters is extremely large; and (3) how to significantly enhance the prediction power of the system by leveraging massively available unlabeled data in addition to human-annotated training data which are often sparse.This project will address the three challenges via the following four approaches.(1) A unified representation of heterogeneous information sources using product graphs: This framework aims to represent heterogeneous sources of data and intra-source dependencies, such as social connections among users, semantic similarities among items, contextual correlations among keywords, topical similarities among documents, hierarchical relations among topic labels, and so on. Each data source will be represented using a graph, and the individual graphs of multiple sources will be combined into a product graph where each node corresponds to a tuple of nodes in the individual graphs, and each link aggregates the links in the individual graphs. (2) Transductive learning over graph products: This project plans to reduce the inference problems in a broad range of prediction tasks to semi-supervised transductive learning problems over the product graphs mentioned above. The training data in each task (of classification, regression or link prediction) will be represented as a subset of labeled (or scored) nodes in the product graph, and the labels (or scores) of those nodes will be propagated over the links in the product graph until convergence. This project will study various kinds of graph transductions theoretically and empirically.(3) Large-scale optimization algorithms: The induced product graphs are typically extremely large. To address the computational bottlenecks, this project will develop new scalable algorithms based on theoretical properties and computational characteristics of spectral graph products, including adapted versions of rank-reduced matrix factorization, aggressive basis pruning, and sampling-based low-rank approximation.(4) Thorough evaluations in multiple important applications: The proposed new approach will be evaluated on benchmark data collections for context-aware collaborative filtering, semi-structured event detection and tracking, and expert finding via multi-source social network analysis.The proposed work, if successful, will offer principled solutions for enhancing the prediction power of systems in a broad range of tasks, whenever recommendation, classification and regression are involved. Technical impacts of the proposed work are expected in multiple research fields. For further information see the project web site at: http://nyc.lti.cs.cmu.edu/gp-trans/index.html
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
III: Small: Multi-field Hierarchical Discovery and Tracking (mf-HDT) of Emerging Topics
-
批准号:1216282
-
项目类别:Standard Grant
-
资助金额:$49.92万
-
财政年份:2012
-
负责人:Yiming Yang
-
依托单位:
III-COR: Collaborative Research: User-centric, Adaptive and Collaborative Information Filtering
-
批准号:0704689
-
项目类别:Standard Grant
-
资助金额:$69.15万
-
财政年份:2007
-
负责人:Yiming Yang
-
依托单位:
KDI: Universal Information Access: Translingual Retrieval, Summarization, Tracking, Detection and Validation
-
批准号:9873009
-
项目类别:Standard Grant
-
资助金额:$198.2万
-
财政年份:1998
-
负责人:Yiming Yang
-
依托单位:
国内基金
海外基金
登录
查看更多内容
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:黄洛将
-
依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:黄洛将
-
依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
-
批准号:12074246
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2020
-
负责人:Yoshitomo Kamiya
-
依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
-
批准号:31972875
-
项目类别:面上项目
-
资助金额:58.0万元
-
批准年份:2019
-
负责人:石江华
-
依托单位:
Large PB/PB小鼠 视网膜新生血管模型的研究
-
批准号:30971650
-
项目类别:面上项目
-
资助金额:8.0万元
-
批准年份:2009
-
负责人:周旻
-
依托单位:
基因discs large在果蝇卵母细胞的后端定位及其体轴极性形成中的作用机制
-
批准号:30800648
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2008
-
负责人:于玲珠
-
依托单位:
LARGE基因对口腔癌细胞中α-DG糖基化及表达的分子调控
-
批准号:30772435
-
项目类别:面上项目
-
资助金额:29.0万元
-
批准年份:2007
-
负责人:尚政军
-
依托单位: