Mining the Deep Web using Sampling and Deep Learning Techniques
Mining the Deep Web using Sampling and Deep Learning Techniques
批准号:
RGPIN-2019-05350
负责人:
Lu, Jianguo
金额:
$2.04万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2021
资助国家:
加拿大
项目状态:
已结题
起止时间:
2021-01-01 至 2022-12-31
中文摘要
深层网络(或隐藏网络)是隐藏在可搜索界面后面的网络。与可以大规模浏览和下载页面的表层网络不同,对深层网络的访问受到限制。一个常见的限制是通过可编程的Web API和Web服务进行查询。许多数据源,如在线社交网络(OSN),都是深度网络的例子。他们有海量的数据,但他们的访问接口是受限的。通常,他们对我们可以发送的查询和我们可以检索的每个IP地址的数据项施加配额。发现隐藏在深层网络中的数据的属性和模式是一个具有挑战性的问题。深度学习已被证明是数据挖掘任务的一种有效方法。它在学习嵌入(即短而密集的连续向量表示)方面特别成功,用于各种实体,如单词、网络和文档。嵌入对于下游数据挖掘和机器学习任务是必不可少的,例如分类、聚类和推荐。嵌入算法需要大量数据。它们的成功取决于能否获得丰富和相关的培训数据。对于深度网络,训练数据很稀缺,而且可能没有代表性。我们需要开发能够获得相关数据的采样技术,并改进能够利用有限数据的深度学习算法。人类不是通过阅读谷歌或谷歌学者索引的所有文本来学习的。相反,我们通过发送相关的查询、阅读返回的信息和发送新的查询来学习。同样,深度学习算法不能也不应该从谷歌或整个社交网络获得来自Facebook的所有文本。相反,应该有基于采样的深度学习算法,在迭代过程中从深度网络中学习。我们的研究将从两个方面入手:1)自下而上的深度网络:我们将研究Twitter等真实深度网站支持的采样技术;2)自上而下的深度学习:我们将选择几种深度学习算法来研究它们是否可以用来自深层网络的样本来近似,以及什么样的样本可以提高性能。我们将从基于神经网络的表示学习开始,例如,最新的用于文本嵌入的SN(Skipgram Negative Samples)和用于图形嵌入的DeepWalk。在单词嵌入和节点嵌入之后,我们将扩展到文档、链接文档嵌入和作者嵌入。研究将分两个阶段进行。在第一阶段,我们将在我们当地的学术搜索引擎上评估我们的方法,以便可以控制参数并获得基本事实。在第二阶段,我们将继续讨论真正的隐藏数据源。
英文摘要
The deep web (or the hidden web) is the web that is hidden behind searchable interfaces. Unlike the surface web where pages can be browsed and hence downloaded in large scale, the access to the deep web is restricted. One common restriction is by queries via programmable Web APIs and web services. Many data sources, such as online social networks (OSNs), are examples of the deep web. They have a vast amount of data, but their access interface is restrictive and limited. Typically, they impose a quota for the queries we can send and data items we can retrieve per IP address. Discovering properties and patterns of the data hidden in the deep web is a challenging problem. Deep learning has been proven an effective approach to data mining tasks. It has been particularly successful in learning embeddings, i.e., short and dense continuous vector representations, for a variety of entities such as words, networks, and documents. Embeddings are essential for downstream data mining and machine learning tasks, such as classification, clustering, and recommendation. Embedding algorithms are data--hungry. Their success hinges on the availability of copious and pertinent training data. With the deep web, the training data are scarce, and may not be representative. We need to develop sampling techniques that can obtain pertinent data, and improve deep learning algorithms that can utilize the limited data. Humans learn not by reading all the text indexed by Google or GoogleScholar. Instead, we learn by sending pertinent queries, reading the returns, and sending new queries. Similarly, deep learning algorithms cannot and should not have all the text from Google or the entire social network from Facebook. Instead, there should be sampling-based deep learning algorithms that will learn from the deep web in an iterative process. The proposed research will approach the problem from two directions 1) Bottom-up from the deep web: we will study the sampling techniques that can be supported from real deep web sites such as Twitter; 2) Top-down from the deep learning: we will select several deep learning algorithms to study whether they can be approximated using samples from the deep web, and what kind of samples can improve the performance. We will start with neural network based representation learning, e.g., the state-of-the-art SN (Skipgram Negative Sampling) for text embedding and DeepWalk for graph embedding. After word embedding and node embedding, we will expand to document, linked document embedding, and author embeddings. The study will be conducted in two stages. In the first stage, we will evaluate our methods on our local academic search engine so that parameters can be controlled and ground truths are available. In the second stage, we will move on to real hidden data sources.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Mining the Deep Web using Sampling and Deep Learning Techniques
-
批准号:RGPIN-2019-05350
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2022
-
负责人:Lu, Jianguo
-
依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
-
批准号:RGPIN-2019-05350
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2020
-
负责人:Lu, Jianguo
-
依托单位:
Mining the Deep Web using Sampling and Deep Learning Techniques
-
批准号:RGPIN-2019-05350
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.04万
-
财政年份:2019
-
负责人:Lu, Jianguo
-
依托单位:
Mining Online Social Networks and Hidden Web Data Sources by Sampling
-
批准号:RGPIN-2014-04463
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.33万
-
财政年份:2018
-
负责人:Lu, Jianguo
-
依托单位:
Mining Online Social Networks and Hidden Web Data Sources by Sampling
-
批准号:RGPIN-2014-04463
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.33万
-
财政年份:2017
-
负责人:Lu, Jianguo
-
依托单位:
Mining Online Social Networks and Hidden Web Data Sources by Sampling
-
批准号:RGPIN-2014-04463
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.33万
-
财政年份:2016
-
负责人:Lu, Jianguo
-
依托单位:
Mining Online Social Networks and Hidden Web Data Sources by Sampling
-
批准号:RGPIN-2014-04463
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.33万
-
财政年份:2015
-
负责人:Lu, Jianguo
-
依托单位:
Mining Online Social Networks and Hidden Web Data Sources by Sampling
-
批准号:RGPIN-2014-04463
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$2.33万
-
财政年份:2014
-
负责人:Lu, Jianguo
-
依托单位:
Web service collection, searching and composition
-
批准号:262083-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2012
-
负责人:Lu, Jianguo
-
依托单位:
Web service collection, searching and composition
-
批准号:262083-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2011
-
负责人:Lu, Jianguo
-
依托单位:
Web service collection, searching and composition
-
批准号:262083-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2010
-
负责人:Lu, Jianguo
-
依托单位:
Web service collection, searching and composition
-
批准号:262083-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2009
-
负责人:Lu, Jianguo
-
依托单位:
Web service collection, searching and composition
-
批准号:262083-2008
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.09万
-
财政年份:2008
-
负责人:Lu, Jianguo
-
依托单位:
Web information system reengineering
-
批准号:262083-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2007
-
负责人:Lu, Jianguo
-
依托单位:
Web information system reengineering
-
批准号:262083-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2006
-
负责人:Lu, Jianguo
-
依托单位:
Web information system reengineering
-
批准号:262083-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2005
-
负责人:Lu, Jianguo
-
依托单位:
Web information system reengineering
-
批准号:262083-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2004
-
负责人:Lu, Jianguo
-
依托单位:
Web information system reengineering
-
批准号:262083-2003
-
项目类别:Discovery Grants Program - Individual
-
资助金额:$1.31万
-
财政年份:2003
-
负责人:Lu, Jianguo
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
-
批准号:2026JJ81909
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2026
-
负责人:胡曦
-
依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
-
批准号:12271434
-
项目类别:面上项目
-
资助金额:46万元
-
批准年份:2022
-
负责人:贺小伟
-
依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
-
批准号:2020A151501709
-
项目类别:省市级项目
-
资助金额:10.0万元
-
批准年份:2020
-
负责人:谢怡
-
依托单位:
面向Deep Web的数据整合关键技术研究
-
批准号:61872168
-
项目类别:面上项目
-
资助金额:62.0万元
-
批准年份:2018
-
负责人:董永权
-
依托单位:
基于Deep-learning的三江源区冰川监测动态识别技术研究
-
批准号:51769027
-
项目类别:地区科学基金项目
-
资助金额:38.0万元
-
批准年份:2017
-
负责人:张大奇
-
依托单位:
具有时序处理能力的Spiking-Deep Learning(脉冲深度学习)方法研究
-
批准号:61573081
-
项目类别:面上项目
-
资助金额:64.0万元
-
批准年份:2015
-
负责人:屈鸿
-
依托单位:
基于语义计算的海量Deep Web知识探索机制研究
-
批准号:61272411
-
项目类别:面上项目
-
资助金额:80.0万元
-
批准年份:2012
-
负责人:赵峰
-
依托单位:
Deep Web数据集成查询结果抽取与整合关键技术研究
-
批准号:61100167
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:董永权
-
依托单位:
面向Deep Web的大规模知识库自动构建方法研究
-
批准号:61170020
-
项目类别:面上项目
-
资助金额:57.0万元
-
批准年份:2011
-
负责人:崔志明
-
依托单位:
Deep Web敏感聚合信息保护方法研究
-
批准号:61003054
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2010
-
负责人:赵朋朋
-
依托单位: