课题基金 / 基金详情

BIGDATA: Collaborative Research: F: Streaming Architecture for Continuous Entity Linking in Social Media

BIGDATA: Collaborative Research: F: Streaming Architecture for Continuous Entity Linking in Social Media
BIGDATA:协作研究:F:社交媒体中连续实体链接的流架构
批准号:
1546441
负责人:
Weiyi Meng
金额:
$30.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2016
资助国家:
美国
项目状态:
已结题
起止时间:
2016-01-01 至 2020-12-31

项目摘要

项目成果

Weiyi Meng的其他基金

相似基金

相关文献

中文摘要
翻译
在不断增长的互联网内容中,有很大一部分是在社交媒体上找到的,比如(微博)博客。用户可以访问它来形成和分享他们对事件和人物、选举偏好、产品和品牌推荐的看法。这种情况提供了创建关于用户对开发事件、产品、服务或政府行动的看法的数据挖掘和分析的附加层的机会;同时,它对社交媒体中的实体链接(EL)提出了挑战。EL是将提取的提及链接到实体的特定定义的任务。实体的定义通常是指向定义该实体的网页的指针。从社交媒体提取信息通常面临许多具有挑战性的问题,原因是:消息量、消息速度(仅Twitter一天就产生了超过5亿条消息)、多样性、自由形式的语言、缺乏上下文、参考差异大和语言多样性。标签是社交网络精神的重要组成部分。它们被用来表示品牌、事件、人物、社会集会等。标签消歧的问题是检测同义标签并识别多义性标签。例如,#BHaram标签指的是在维基百科页面en.wikipedia.org/wiki/Boko_Haram或国家反恐中心网页www.nctc.gov/site/group/Boko_haram.html定义的实体‘博科圣地’。这个项目的目的是在社交媒体上表演EL。这项工作将使依赖使用微博系统数据的应用程序的社会各阶层受益,例如,有针对性地监测Twitter和Facebook,以收集和了解用户对最近产品或世界事件的意见;数据汇总(例如,对产品和服务的评论);以及用于早期危机检测和应对以及国家安全的数据挖掘。这个项目是解决政府利用大数据打击犯罪的最新举措的又一步。该项目的目标是研究算法,以近乎实时地检测消息中引用实体的文本片段,描述实体的网页,并将实体引用链接到网页和跨微博系统,以便可以一起自动生成对每个实体的更广泛、更完整的描述。所提出的方法基于创新技术,包括:增量迭代消息分析;具有实时更新的智能索引技术以支持快速增量实体引用检测;计算轻量级消息软聚类以改进实体引用检测;以及快速增量K部图聚类。由此产生的人工制品(例如,软件工具)将被提供给学术界和工业界的研究人员。分发用于实施所开发技术的免费开放源码软件将加强现有的研究基础设施。该项目将支持和培训至少三名博士生,并让本科生参与坦普尔大学和宾汉普顿大学的研究。项目网站(http://cis.temple.edu/~edragut/projects/nimel.htm)包括有关项目、软件、数据集、教育材料和出版物的更多信息。
英文摘要
A large fraction of the ever-growing internet content is found in social media such as (micro)blogs. Users access it to both form and share their opinions about events and people, election preferences, product and brand recommendations. This situation provides opportunities to create added layers of data mining and analysis regarding users' views on developing events, products, services, or government actions; at the same time, it raises challenges for Entity Linking (EL) in social media. EL is the task of linking an extracted mention to a specific definition of the entity. The definition of an entity is usually a pointer to a Web page that defines the entity. Information extraction from social media generally faces many challenging issues due to: message volume, message speed (Twitter alone generates over 500 million messages per day), variety, free-form language, lack of context, large reference variation and language diversity. Hashtags are an essential part of the ethos of social networks. They are used to denote brands, events, people, social rallies, etc. The hashtag disambiguation problem is to detect synonymous hashtags and recognize the polysemic ones. For example, the hashtag '#BHaram' refers to the entity 'Boko Haram', defined at Wikipedia page en.wikipedia.org/wiki/Boko_Haram or at National Counterterrorism Center Web web page www.nctc.gov/site/groups/boko_haram.html. The purpose of this project is to perform EL in social media. This work will benefit multiple segments of society that rely on applications using data from microblog systems, such as targeted monitoring of Twitter and Facebook to collect and understand users' opinions about a recent product or a world event; data aggregation (e.g., reviews about products and services); and data mining for early crisis detection and response as well as national security. This project is one more step towards addressing the government's latest initiative of fighting crime using big data.The goals of this project are to research algorithms to detect in near real-time those pieces of text in messages that reference entities, Web pages that describe entities, and to link entity references to Web pages and across microblog systems so that together a broad, more complete characterization of each entity can be automatically generated. The proposed approaches are based on innovative techniques that include: incremental, iterative message analysis; smart indexing techniques with live updates to support fast incremental entity reference detection; computationally light soft-clustering of messages to improve entity reference detection; and fast incremental K-partite graph clustering. The resulting artifacts (e.g., software tools) will be made available to benefit researchers in academe and industry. Distribution of free, open-source software for implementing the techniques developed will enhance existing research infrastructure. The project will support and train at least three PhD students, as well as involve undergraduate students in research at Temple University and Binghampton University. The project web site (http://cis.temple.edu/~edragut/projects/nimel.htm) includes more information on the project, software, datasets, educational materials, and publications.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
SGER/Collaborative Research: Handling Negations and Temporal Aspects for Opinion Retrieval
  • 批准号:
    0842608
  • 项目类别:
    Standard Grant
  • 资助金额:
    $7.5万
  • 财政年份:
    2008
  • 负责人:
    Weiyi Meng
  • 依托单位:
SGER/Collaborative Research: Intelligent Use of Dictionaries for Document Retrieval
  • 批准号:
    0738727
  • 项目类别:
    Standard Grant
  • 资助金额:
    $4.0万
  • 财政年份:
    2007
  • 负责人:
    Weiyi Meng
  • 依托单位:
Collaborative Research: Achieving Information Integration of Web Databases Through the Construction of Metasearch Engines
  • 批准号:
    0414981
  • 项目类别:
    Standard Grant
  • 资助金额:
    $23.0万
  • 财政年份:
    2005
  • 负责人:
    Weiyi Meng
  • 依托单位:
Collaborative Research: WebScales-Towards a Large Scale Metasearch Engine
  • 批准号:
    0208574
  • 项目类别:
    Continuing Grant
  • 资助金额:
    $10.53万
  • 财政年份:
    2002
  • 负责人:
    Weiyi Meng
  • 依托单位:
海外基金