课题基金 / 基金详情

Mining Online Social Networks and Hidden Web Data Sources by Sampling

Mining Online Social Networks and Hidden Web Data Sources by Sampling
通过采样挖掘在线社交网络和隐藏的网络数据源
批准号:
RGPIN-2014-04463
负责人:
Lu, Jianguo
金额:
$2.33万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31

项目摘要

项目成果

Lu, Jianguo的其他基金

相似基金

相关文献

中文摘要
翻译
以HTML搜索框、可编程Web API和Web服务形式存在的可搜索Web界面在Web上随处可见。隐藏在搜索界面后面的数据统称为隐藏网络或深层网络。当搜索框是谷歌等网络搜索引擎的界面时,它几乎可以是网络上的任何东西。在其他情况下,深层网络数据可以是诸如在线社交网络(OSN)站点等专门领域中的有价值的数据集合。从这些可公开访问的巨大且不断增加的数据源中,数据提供商和消费者都感兴趣的一个问题是:从样本中可以从统计上推断出哪些信息和模式。这个问题的答案有许多应用,从商业智能到犯罪集团侦查。虽然我们的主要目标是发现隐藏的属性,但数据提供程序可以使用相同的技术来设计可搜索的接口来保护数据。我们的短期目标是挖掘OSNs,长期目标是研究适用于其他深层Web数据源和大型本地数据集合的理论和方法。在OSN挖掘中出现了新的挑战,因为数据、访问数据的方式以及从数据中推断出的属性和模式不同于传统的数据挖掘问题。首先,无法获得完整的数据。相反,只有一小部分数据可以通过Web API调用代价高昂的远程调用来返回。我们需要开发为网络界面量身定做的采样方法。其次,数据来源庞大,往往服从具有很大方差的幂分布。这就需要新的抽样方法来减少方差。第三,一些要估计的属性和要发现的模式超出了传统数据挖掘任务的范围。即使当我们拥有全部数据时,诸如聚类系数和各种中心度之类的社会网络属性的计算成本也很高。为了克服这些困难,我们的目标是:(1)使用网络界面获取随机样本。直接对Web界面返回的不受控制的数据进行推断将导致重大和不可预测的偏差。Web界面是限制性的,它们接受的查询类型、它们索引内容的方式以及它们对匹配项进行排名和返回的策略各不相同。由于网络流量和数据提供商施加的每日配额,远程查询的成本很高。我们的目标是通过使用有限数量的查询,最大限度地增加对推断OSN属性有用的样本数据量。(2)设计图形抽样方法以减少方差:样本从一定的分布中选择时是有用的。传统观点认为,只要有可能,就应该使用均匀随机抽样。然而,OSNs是大的和无尺度的,导致许多估计器在均匀随机样本上的方差很大。我们的目标是,在独立于Web界面的图形采样的背景下,开发其他能够提高推断精度的采样方法。(3)OSN属性的推断:除了总结统计数据外,我们还将在社会网络研究中开发度量的抽样和估计方法。在这些探索的基础上,我们将把这些方法应用到真实的OSN中,如Twitter、Facebook、LinkedIn和微博,以发现它们有趣的社交网络特性、OSN中独特的现象(例如,机器人账户)以及其他意想不到的广泛兴趣模式。
英文摘要
Searchable web interfaces in the form of HTML search boxes, programmable web APIs and Web Services are ubiquitous on the web. The data hidden behind search interfaces are collectively called the hidden web or deep web. It can be virtually everything on the web when the search box is the interface to a web search engine such as Google. In other cases, the deep web data can be valuable collections of data in a specialized area, such as an Online Social Network (OSN) site. From these vast and ever increasing data sources that are openly accessible, a question of interests to both data providers and consumers is: what information and patterns can be inferred statistically from a sample. The answer to this question has many applications ranging from business intelligence to criminal group detection. Although our main goal is to uncover the hidden properties, the same techniques can be used by data providers to design the searchable interfaces to protect the data. Our short-term goal focuses on mining OSNs, and the long-term goal is to study the theories and methods that are applicable to other deep web data sources and large local data collections. New challenges arise in mining OSNs, because the data, the way to access the data, and the properties and patterns to be inferred from the data, are different from that of conventional data mining problems. First, the data in its entirety is not available. Instead, only a small portion of the data can be returned by invoking costly remote calls through web APIs. We need to develop sampling methods tailored for the web interfaces. Second, the data sources are huge, often following the power-law distribution with very large variance. This calls for new sampling methods to reduce the variance. Third, some properties to be estimated and patterns to be discovered are beyond the scope of traditional data mining tasks. Social network properties such as the clustering coefficient and various centralities are costly to compute even when we possess the whole data. To overcome these difficulties, we target the following goals: (1) To obtain random samples using web interfaces. Inferencing directly on uncontrolled data returned by web interfaces will result in substantial and unpredictable bias. Web interfaces are restrictive, and vary in the types of the queries they accept, the way they index the content, and the strategy they rank and return the matches. The remote queries are expensive because of network traffic and daily quota imposed by data providers. Our goal is to, by using limited number of queries, maximize the amount of sample data that are useful to infer OSN properties. (2) To design graph sampling methods to reduce variance: Samples are useful when they are selected from certain distributions. Conventional wisdom has it that uniform random samples should be used whenever possible. However, OSNs are large and scale-free, resulting in large variances for many estimators on uniform random samples. Our goal is, in the context of graph sampling that is independent of the web interface, to develop other sampling methods that can increase the accuracy of the inferences. (3) To infer OSN properties: In addition to summarizing statistics, we will develop sampling and estimation methods for metrics in social network studies. Based on these explorations, we will adapt and combine these methods to real OSNs such as Twitter, Facebook , LinkedIn and Weibo to discover their interesting social network properties, phenomena that are unique in OSNs (e.g., robot accounts), and other unsuspected patterns of wide interests.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2020
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2019
  • 负责人:
    Lu, Jianguo
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
online SPE/HPLC-ICP-MS多元素形态分析新方法研究荷塘中铬砷镉汞铅的迁移转化规律
  • 批准号:
    21976048
  • 项目类别:
    面上项目
  • 资助金额:
    65.0万元
  • 批准年份:
    2019
  • 负责人:
    刘金华
  • 依托单位:
双积分政策下基于Online Review的新能源汽车企业跨链决策优化研究
  • 批准号:
    71964023
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    27.5万元
  • 批准年份:
    2019
  • 负责人:
    黎继子
  • 依托单位: