BIGDATA: Collaborative Research: F: Streaming Architecture for Continuous Entity Linking in Social Media
BIGDATA: Collaborative Research: F: Streaming Architecture for Continuous Entity Linking in Social Media
批准号:
1546480
负责人:
Eduard Dragut
金额:
$78.33万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-01-01 至 2020-12-31
中文摘要
在不断增长的互联网内容中,有很大一部分是在(微)博客等社交媒体上发现的。用户可以访问它,形成并分享他们对事件和人物、选举偏好、产品和品牌推荐的看法。这种情况为创建关于用户对发展中的事件、产品、服务或政府行为的看法的附加数据挖掘和分析层提供了机会;同时也给社交媒体中的实体链接(Entity Linking, EL)提出了挑战。EL是将提取的提及链接到实体的特定定义的任务。实体的定义通常是指向定义该实体的Web页面的指针。从社交媒体中提取信息通常面临许多具有挑战性的问题,因为:消息数量,消息速度(仅Twitter每天就产生超过5亿条消息),种类,自由形式的语言,缺乏上下文,大量参考变化和语言多样性。话题标签是社交网络精神的重要组成部分。它们用来表示品牌、事件、人物、社会集会等。标签消歧问题是识别同义标签和多义标签。例如,标签“#BHaram”指的是实体“博科圣地”,定义在维基百科页面en.wikipedia.org/wiki/Boko_Haram或国家反恐中心网页www.nctc.gov/site/groups/boko_haram.html。这个项目的目的是在社交媒体中执行EL。这项工作将使社会的多个部门受益,这些部门依赖于使用微博系统数据的应用程序,例如有针对性地监测Twitter和Facebook,以收集和了解用户对最近产品或世界事件的看法;数据汇总(例如,关于产品和服务的评论);以及用于早期危机检测和应对以及国家安全的数据挖掘。这个项目是解决政府利用大数据打击犯罪的最新倡议的又一步。该项目的目标是研究算法,以近乎实时地检测引用实体和描述实体的网页的消息中的文本片段,并将实体引用链接到网页和跨微博系统,以便自动生成对每个实体的更广泛、更完整的表征。提出的方法基于创新技术,包括:增量,迭代的消息分析;具有实时更新的智能索引技术,支持快速增量实体引用检测;计算轻量级消息软聚类以改进实体引用检测快速增量k部图聚类。产生的工件(例如,软件工具)将被提供给学术界和工业界的研究人员。为实现所开发的技术而分发免费的开源软件将加强现有的研究基础设施。该项目将支持和培训至少三名博士生,并让天普大学和宾厄姆顿大学的本科生参与研究。该项目的网站(http://cis.temple.edu/~edragut/projects/nimel.htm)包含有关该项目的更多信息、软件、数据集、教育材料和出版物。
英文摘要
A large fraction of the ever-growing internet content is found in social media such as (micro)blogs. Users access it to both form and share their opinions about events and people, election preferences, product and brand recommendations. This situation provides opportunities to create added layers of data mining and analysis regarding users' views on developing events, products, services, or government actions; at the same time, it raises challenges for Entity Linking (EL) in social media. EL is the task of linking an extracted mention to a specific definition of the entity. The definition of an entity is usually a pointer to a Web page that defines the entity. Information extraction from social media generally faces many challenging issues due to: message volume, message speed (Twitter alone generates over 500 million messages per day), variety, free-form language, lack of context, large reference variation and language diversity. Hashtags are an essential part of the ethos of social networks. They are used to denote brands, events, people, social rallies, etc. The hashtag disambiguation problem is to detect synonymous hashtags and recognize the polysemic ones. For example, the hashtag '#BHaram' refers to the entity 'Boko Haram', defined at Wikipedia page en.wikipedia.org/wiki/Boko_Haram or at National Counterterrorism Center Web web page www.nctc.gov/site/groups/boko_haram.html. The purpose of this project is to perform EL in social media. This work will benefit multiple segments of society that rely on applications using data from microblog systems, such as targeted monitoring of Twitter and Facebook to collect and understand users' opinions about a recent product or a world event; data aggregation (e.g., reviews about products and services); and data mining for early crisis detection and response as well as national security. This project is one more step towards addressing the government's latest initiative of fighting crime using big data.The goals of this project are to research algorithms to detect in near real-time those pieces of text in messages that reference entities, Web pages that describe entities, and to link entity references to Web pages and across microblog systems so that together a broad, more complete characterization of each entity can be automatically generated. The proposed approaches are based on innovative techniques that include: incremental, iterative message analysis; smart indexing techniques with live updates to support fast incremental entity reference detection; computationally light soft-clustering of messages to improve entity reference detection; and fast incremental K-partite graph clustering. The resulting artifacts (e.g., software tools) will be made available to benefit researchers in academe and industry. Distribution of free, open-source software for implementing the techniques developed will enhance existing research infrastructure. The project will support and train at least three PhD students, as well as involve undergraduate students in research at Temple University and Binghampton University. The project web site (http://cis.temple.edu/~edragut/projects/nimel.htm) includes more information on the project, software, datasets, educational materials, and publications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Proto-OKN Theme 1: Knowledge Graph to Support Evaluation and Development of Climate Models
-
批准号:2333789
-
项目类别:Cooperative Agreement
-
资助金额:$149.86万
-
财政年份:2023
-
负责人:Eduard Dragut
-
依托单位:
NSF Convergence Accelerator Track F: America's Fourth Estate at Risk: A System for Mapping the (Local) Journalism Life Cycle to Rebuild the Nation's News Trust
-
批准号:2137846
-
项目类别:Standard Grant
-
资助金额:$75.0万
-
财政年份:2021
-
负责人:Eduard Dragut
-
依托单位:
III: Medium: Collaborative Research: Extracting and Linking AI Artifacts
-
批准号:2107213
-
项目类别:Continuing Grant
-
资助金额:$67.0万
-
财政年份:2021
-
负责人:Eduard Dragut
-
依托单位:
BIGDATA: F: Collaborative Research: Collective Mining of Vertical Social Communities
-
批准号:1838145
-
项目类别:Standard Grant
-
资助金额:$42.79万
-
财政年份:2018
-
负责人:Eduard Dragut
-
依托单位:
海外基金