Data Mining for Information Retrieval and Processing in Web 2.0 Social Networking
Data Mining for Information Retrieval and Processing in Web 2.0 Social Networking
批准号:
RGPIN-2014-03985
负责人:
Abhari, Abdolreza
金额:
$1.38万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31
中文摘要
在这个拟议的研究计划中,数据挖掘和模糊逻辑将用于建立一组新颖的推荐系统的Web2.0社交网站Facebook,Twitter,YouTube和LinkedIn。 对于Facebook和YouTube,推荐系统将采用模糊逻辑,这将能够提供基于对这些网站中现有的基于感知的数据的正确解释的信息。因此,所得到的推荐系统可以作为检索系统的附加部分向用户提供建议。当用户在Facebook或YouTube上说“我通常喝Tim Horton的咖啡”或“我喜欢这个视频”时,“通常”和“喜欢”的含义在用户之间有所不同。模糊检索系统即使在用户使用诸如“通常”、“大多数”和“经常”等不精确的词来表示他们的兴趣时,也可以做出相关的推荐。例如,可以专门为Tim Horton的Facebook粉丝构建模糊检索系统,并且可以基于对用户兴趣的分析来为Tim Horton的咖啡饮用者推荐特殊的cookie,以确定他们最喜欢的口味。 随着Web2.0网站中的域变得更加具体(例如LinkedIn),可以定义单词之间的模糊关系,并用于指导发送特定搜索查询的用户。例如,LinkedIn的特定网络中的用户阅读特定文档,这些文档可用于构建知识库,该知识库充当该网络用户的推荐系统。因此,每当在计算机工程师网络中搜索工作的用户键入单词“设计”时,推荐系统应当仅显示诸如“芯片设计”或“IC设计”的相关单词,而当属于计算机科学家网络的求职者键入“设计”时,所建议的相关单词可能是“算法设计”或“软件设计”。 类似地,对于Twitter,推荐系统可以分析推文并识别具有相同兴趣的用户。Twitter比其他社交网站拥有更多的公共信息;因此,针对Twitter提出的另一项研究活动是将其用作事件检测平台。通过使用各种数据挖掘技术,本研究活动旨在开发一个框架,有效地提取新兴的事件(关于世界上真实的发生的有用信息)从大数据集的推文。 本建议的最后一部分是通过使用多代理系统开发数据收集和评估工具。这些多代理系统将被用作数据爬虫,以主动更新这种推荐系统的知识库。此外,多代理系统也可以用于验证所提出的信息检索(IR)和推荐系统与其他本体或基于语义的搜索引擎相比。使用多代理系统用于社交媒体网站中的IR和推荐系统的评估目的是一种新颖的想法,其将提供两个更多的性能度量(即,分布式处理有效性和可缩放性)除了公共的准确性度量(即,精确度和召回率)。 在上述研究活动中使用模糊逻辑进行Web2.0网站的信息处理是另一个新的想法,尚未开发和实施之前,本研究计划。
英文摘要
In this proposed research program, data mining and fuzzy logic will be used to build a group of novel recommendation systems for the Web2.0 social networking sites Facebook, Twitter, YouTube, and LinkedIn. For FaceBook and YouTube, fuzzy logic will be employed in the recommendation systems, which will be capable of providing information that is based on the correct interpretation of perception-based data existing in these sites. Therefore, the resulting recommendation systems can provide advice to the user as an additional part of a retrieval system. When in Facebook or YouTube a user says, “I usually drink Tim Horton’s coffee,” or “I like this video”, the meaning of “usually” and “like” varies among users. A fuzzy retrieval system can make relevant recommendation even when the users deploy imprecise words such as “usually,” “most,” and “often” to show their interests. For example, a fuzzy retrieval system can be built specifically for Facebook fans of Tim Horton’s, and can recommend a special cookie for the Tim Horton’s coffee drinker based upon an analysis of users’ interests to determine their favourite tastes. As domains become more specific in Web2.0 sites (for example in LinkedIn), fuzzy relations between words can be defined and used to direct users who are sending specific search queries. For example, users in a particular network of Linkedin read specific documents that can be used for building a knowledge base that acts as the recommendation system for the users of that network. Therefore, whenever a user who is searching for a job in computer engineers’ network types the word “design,” only the related words such as “chip design” or “IC design” should be shown by the recommendation system, whereas when a job seeker who belongs to the computer scientists’ network types “design,” the suggested related words might be “algorithm design” or “software design”. Similarly for Twitter, the recommendation systems can analyze the tweets and identify users with the same interests. Twitter has more public information than all other social networking sites; therefore, another research activity proposed for Twitter is using it as a platform for event detection. By using various data mining techniques, this research activity aims to develop a framework that effectively extracts emerging events (useful information about real happenings in the world) from the large data set of tweets. The last part of this proposal is developing data collection and evaluation tools by using multi-agent systems. These multi-agent systems will be used as data crawlers to actively update the knowledge bases of such recommendation systems. Moreover, the multi-agent systems can also be used for validation of the proposed Information Retrieval (IR) and recommendation systems in comparison with other ontology or semantic-based search engines. Use of the multi-agent systems for evaluation purpose of IR and recommendation systems in social media sites is a novel idea that will provide two more performance metrics (i.e., distributed processing effectiveness and scalability) in addition to the common accuracy metrics (i.e., precision and recall) . Using fuzzy logic in the above research activities for information processing of the Web2.0 sites is another new idea that has not been developed and implemented prior to this proposed research program.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Meaningful camera
-
批准号:493629-2016
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2016
-
负责人:Abhari, Abdolreza
-
依托单位:
Simulation software for determining sensors location in building design
-
批准号:408125-2010
-
项目类别:Engage Grants Program
-
资助金额:$1.82万
-
财政年份:2010
-
负责人:Abhari, Abdolreza
-
依托单位:
Storage management for proxy/Web servers
-
批准号:298295-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2006
-
负责人:Abhari, Abdolreza
-
依托单位:
Storage management for proxy/Web servers
-
批准号:298295-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2005
-
负责人:Abhari, Abdolreza
-
依托单位:
Storage management for proxy/Web servers
-
批准号:298295-2004
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.02万
-
财政年份:2004
-
负责人:Abhari, Abdolreza
-
依托单位:
国内基金
海外基金
基于Genome mining技术研究抑制表皮葡萄球菌生物膜形成的次级代谢产物
-
批准号:21242003
-
项目类别:专项基金项目
-
资助金额:10.0万元
-
批准年份:2012
-
负责人:昌军
-
依托单位: