课题基金 / 基金详情

Identifying General Product and Brand Names in Online Forums

Identifying General Product and Brand Names in Online Forums
识别在线论坛中的通用产品和品牌名称
批准号:
521298-2017
负责人:
Makrehchi, Masoud
金额:
$1.82万
依托单位国家:
加拿大
项目类别:
Engage Grants Program
财政年份:
2017
资助国家:
加拿大
项目状态:
已结题
起止时间:
2017-01-01 至 2018-12-31

项目摘要

项目成果

Makrehchi, Masoud的其他基金

相似基金

相关文献

中文摘要
翻译
吸引和获取客户的一个关键数据组件是能够识别在线论坛中的产品名称。因此,产品/品牌名称的提取将产生数据来帮助供应商,从而增加整体业务收入。提取命名实体的传统方法需要耗时且昂贵的人工标记训练集。此外,我们还处理许多产品或服务部门,这些部门具有不同的规模、语言和内容。因此,将基于人工注释数据的监督模型拟合到每个产品/品牌的每个部门(垂直)可能非常昂贵。三个主要挑战是:1)为每种类型的产品生成训练数据的高成本,2)覆盖来自所有垂直领域的大量不同类型的产品,以及3)基于上下文消除不同类型实体的歧义。因此,在缺乏或缺乏训练数据的情况下,另一种解决方案是半监督学习算法,如bootstrapping和元学习方法,如自我训练和共同训练。在这些方法中,我们可以用没有或很少的注释数据来训练模型。在本研究中,研究了一种结合迁移学习和半监督学习的混合方法来识别和提取我们感兴趣的领域中的命名实体。在迁移学习中,针对目标领域缺乏标注数据的解决方案是适应其他领域的标注数据。VerticalScope(行业合作伙伴)是一家加拿大公司,正在成为数据科学研究和开发的领导者,用于理解互联网上各种领域的用户生成内容。拟议中的项目有助于VerticalScope的发展,进而有助于加拿大经济的发展,因为该公司可以应用前沿研究来改善其论坛的用户体验,从而提高其服务对企业的吸引力。
英文摘要
One key data component for engaging and acquiring customers is being able to identify names of products inonline forums. Therefore, the extraction of product/brand names will generate data to help vendors and thusincrease overall business revenue. The traditional approach to extract named entities requires time-consumingand expensive manual human-labelled training sets. In addition, we deal with many product or service sectorswhich come with different size, language, and content. Thus fitting a supervised model which is based onhuman-annotated data, for each product/brand to each sector (vertical) can be very expensive. Three majorchallenges are: 1) the high cost of generating training data for each type of product, 2) covering large divergenttypes of products from all verticals, and 3) disambiguating different types of entities based on context.Therefore, in the absence or lack of training data, alternative solution is semi-supervised learning algorithmssuch as bootstrapping and meta-learning methods such as self-training and co-training. In these methods, wecan train a model either with no or very few annotated data. In this research, a hybrid approach combiningtransfer learning and semi-supervised learning is investigated to identify and extract named entities in ourdomain of interest. In transfer learning, the solution for the lack of annotated data in the target domain, is toadapt annotated data from other domains.VerticalScope (the industry partner) is a Canadian company that is becoming a leading player in data scienceresearch and development for understanding user-generated content on the Internet in a variety of sectors. Theproposed project contributes to the growth of VerticalScope, and to the Canadian economy as a result, byallowing the company to apply cutting-edge research for improving user experience on their forums, and hencethe attractiveness of its service to the businesses.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Algorithms and applications of Link Mining: Making Sense of Network Data
Algorithms and applications of Link Mining: Making Sense of Network Data
Towards Predicting Socio-economic Systems by Mining Social Media Data
Towards Predicting Socio-economic Systems by Mining Social Media Data
国内基金
海外基金
Toward a general theory of intermittent aeolian and fluvial nonsuspended sediment transport
  • 批准号:
    --
  • 项目类别:
    --
  • 资助金额:
    55万元
  • 批准年份:
    2022
  • 负责人:
    Thomas Pahtz
  • 依托单位: