课题基金 / 基金详情

Mining Online Social Networks and Hidden Web Data Sources by Sampling

Mining Online Social Networks and Hidden Web Data Sources by Sampling
通过采样挖掘在线社交网络和隐藏的网络数据源
批准号:
RGPIN-2014-04463
负责人:
Lu, Jianguo
金额:
$2.33万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2014
资助国家:
加拿大
项目状态:
已结题
起止时间:
2014-01-01 至 2015-12-31

项目摘要

项目成果

Lu, Jianguo的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Searchable web interfaces in the form of HTML search boxes, programmable web APIs and Web Services are ubiquitous on the web. The data hidden behind search interfaces are collectively called the hidden web or deep web. It can be virtually everything on the web when the search box is the interface to a web search engine such as Google. In other cases, the deep web data can be valuable collections of data in a specialized area, such as an Online Social Network (OSN) site. From these vast and ever increasing data sources that are openly accessible, a question of interests to both data providers and consumers is: what information and patterns can be inferred statistically from a sample. The answer to this question has many applications ranging from business intelligence to criminal group detection. Although our main goal is to uncover the hidden properties, the same techniques can be used by data providers to design the searchable interfaces to protect the data. Our short-term goal focuses on mining OSNs, and the long-term goal is to study the theories and methods that are applicable to other deep web data sources and large local data collections. New challenges arise in mining OSNs, because the data, the way to access the data, and the properties and patterns to be inferred from the data, are different from that of conventional data mining problems. First, the data in its entirety is not available. Instead, only a small portion of the data can be returned by invoking costly remote calls through web APIs. We need to develop sampling methods tailored for the web interfaces. Second, the data sources are huge, often following the power-law distribution with very large variance. This calls for new sampling methods to reduce the variance. Third, some properties to be estimated and patterns to be discovered are beyond the scope of traditional data mining tasks. Social network properties such as the clustering coefficient and various centralities are costly to compute even when we possess the whole data. To overcome these difficulties, we target the following goals: (1) To obtain random samples using web interfaces. Inferencing directly on uncontrolled data returned by web interfaces will result in substantial and unpredictable bias. Web interfaces are restrictive, and vary in the types of the queries they accept, the way they index the content, and the strategy they rank and return the matches. The remote queries are expensive because of network traffic and daily quota imposed by data providers. Our goal is to, by using limited number of queries, maximize the amount of sample data that are useful to infer OSN properties. (2) To design graph sampling methods to reduce variance: Samples are useful when they are selected from certain distributions. Conventional wisdom has it that uniform random samples should be used whenever possible. However, OSNs are large and scale-free, resulting in large variances for many estimators on uniform random samples. Our goal is, in the context of graph sampling that is independent of the web interface, to develop other sampling methods that can increase the accuracy of the inferences. (3) To infer OSN properties: In addition to summarizing statistics, we will develop sampling and estimation methods for metrics in social network studies. Based on these explorations, we will adapt and combine these methods to real OSNs such as Twitter, Facebook , LinkedIn and Weibo to discover their interesting social network properties, phenomena that are unique in OSNs (e.g., robot accounts), and other unsuspected patterns of wide interests.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2022
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2021
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2020
  • 负责人:
    Lu, Jianguo
  • 依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
  • 批准号:
    RGPIN-2019-05350
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.04万
  • 财政年份:
    2019
  • 负责人:
    Lu, Jianguo
  • 依托单位:
国内基金
海外基金
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
Data-driven Recommendation System Construction of an Online Medical Platform Based on the Fusion of Information
online SPE/HPLC-ICP-MS多元素形态分析新方法研究荷塘中铬砷镉汞铅的迁移转化规律
  • 批准号:
    21976048
  • 项目类别:
    面上项目
  • 资助金额:
    65.0万元
  • 批准年份:
    2019
  • 负责人:
    刘金华
  • 依托单位:
双积分政策下基于Online Review的新能源汽车企业跨链决策优化研究
  • 批准号:
    71964023
  • 项目类别:
    地区科学基金项目
  • 资助金额:
    27.5万元
  • 批准年份:
    2019
  • 负责人:
    黎继子
  • 依托单位: