课题基金 / 基金详情

III: Large: Collaborative Research: Web Archive Cooperative

III: Large: Collaborative Research: Web Archive Cooperative
III:大型:协作研究:网络档案合作社
批准号:
1009392
负责人:
Michael Nelson
金额:
$39.98万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2010
资助国家:
美国
项目状态:
已结题
起止时间:
2010-08-01 至 2014-07-31

项目摘要

项目成果

Michael Nelson的其他基金

相似基金

相关文献

中文摘要
翻译
网络科学是一门新兴学科,研究网络:人类活动如何被网络互动所塑造,网络如何造福社会,以及如何改进网络技术。Web科学的核心是访问记录Web历史的数据,以及记录人类活动的数据(例如,提出的查询,标记的页面,Twitter更新)。目前,学术研究人员很难获得这样的网络数据,因为这些数据很难定位,它们分散在不同的站点上,并且使用不一致的格式和策略进行记录。该项目将建立一个网络档案合作社(WAC),它将集成现有的档案(网络数据存储库),使以简化的方式访问大量数据成为可能。WAC将是一个虚拟服务,为现有资源提供搜索设施和访问机制。这些资源将不仅仅是Web页面,而是所有类型的可用Web信息,例如查询日志、标记注释、博客、概要文件和Twitter更新。此外,资源还将包括用于构建和管理Web存档的软件工具。该项目将探索资源发现服务的三个目标:(1)手动或自动发现整个现有的Web相关档案;(2)从已知档案中选择支持特定研究问题的档案;(3)从选定的档案中识别个别资源。还将开发用于描述已发现的档案的工具,特别是在档案没有提供丰富的描述性元数据的情况下。档案的特征包括档案覆盖范围的估计、爬行参数的细节(如日期/频率、爬行持续时间、深度、收集页面数量的每个站点上限、内容统计和链接结构)等元素。建立多种档案的整合机制,并将其应用于现场重建(来自各种档案)和档案视图(来自多个来源的资源的逻辑融合)。由于集成问题是如此具有挑战性,因此将使用小而多样的资源建立实验性测试平台。测试平台将包含对相同目标站点的多个爬虫,每个爬虫使用不同的爬虫并使用不同的参数获得。测试平台还将包含相关资源。将制定存储交易计划,允许成员用本地备份空间交换远程空间。将基于现有的自我保存对象概念开发Web存档复制工具。将研究副本同步的备选方案。会议将组织研讨会,汇集网络科学的主要研究人员,讨论可用的资源和共享的障碍。这些研讨会将推动研究并确定所需的工具和协议。对于小的参与者组,将建立挑战问题,例如,组合一组Web档案。在未来的研讨会上报告这些结果可以激励其他人参加WAC。此外,还成立了一个由工业界、政府和学术界专家组成的咨询委员会来指导该项目。将举办网络科学研究生暑期研修班。在本课程中,学生将学习使用最新的工具,并相互学习处理网络数据的经验。此外,将开发一个为期一天的研讨会,在网络科学会议(WWW, SIGIR等)上提供,以向与会者介绍WAC资源。利用WAC的资源,将为计算机科学专业的本科生开设网络科学课程。该项目将在两个方面产生影响。首先,它将提供方便访问Web资源的工具和服务。任何研究人员,从研究高效网络搜索的计算机科学家,到研究当今人类信仰如何变化的社会科学家,再到研究早期网络如何进化的历史学家,再到研究疾病如何传播的生物学家,都将从这项工作中受益。其次,该项目激励学生和年轻研究人员留在学术界。目前,顶尖人才正流向工业界,因为只有他们拥有全面的网络数据,而在大学里做有意义的网络科学是非常困难的。WAC可以提供另一种选择,吸引更多的研究人员和教师进入这一重要领域。
英文摘要
Web Science is an emerging discipline that studies the Web: how human activity is shaped by Web interactions, how the Web can benefit society, and how Web technologies can be improved. Central to Web Science is access to data that records the history of the Web, as well as data that records human activity (e.g., posed queries, tagged pages, Twitter updates). It is currently very difficult for academic researchers to obtain such Web data because it is hard to locate, it is fragmented across diverse sites, and is recorded using inconsistent formats and strategies. This project will build a Web Archive Cooperative (WAC) that will integrate existing archives (repositories of Web data), making it feasible to access large volumes of data in a simplified fashion. The WAC will be a virtual service, providing search facilities and access mechanisms to existing resources. These resources will not just be Web pages, but all types of available Web information, such as query logs, tag annotations, blogs, profiles and Twitter updates. Furthermore, resources will also include the software tools for building and managing Web archives.The project will explore three goals for a resource discovery service: (1) the manual or automated discovery of entire existing Web related archives; (2) the selection among known archives of the ones that support a specific research question; and (3) the identification of individual resources from within the selected archives. Tools for characterizing discovered archives, especially for the case where the archive does not provide rich descriptive metadata, will also be developed. Characterization of an archive includes elements such as an estimate of the archive's coverage, particulars of the crawling parameters, like dates/frequencies, crawl duration, depth, per-site ceiling on the number of collected pages, content statistics, and link structure. Mechanisms for integrating diverse archives will be developed, and the mechanisms will be applied to site reconstruction (from various archives) and archive views (a logical fusion of resources from multiple sources). Since integration issues are so challenging, an experimental testbed will be set up with small but diverse resources. The testbed will contain several crawls of the same target sites, each obtained with different crawlers and using different parameters. The testbed will also contain related resources. Storage trading schemes will be developed, allowing members to trade local backup space for remote space. A Web archive replication tool will be developed based on existing notions for self-preserving objects. Alternatives for replica synchronization will be studied.Workshops to bring together key Web Science researchers will be organized to discuss available resources and impediments to sharing. These workshops will drive research and identify needed tools and protocols. With small groups of participants, challenge problems will be established, e.g., combining a set of Web archives. Reports of these results at future workshops can incentivize others to participate in the WAC. In addition, an Advisory Board of industrial, government, and academic experts has been set up to guide the project. A Summer Institute for Web Science graduate students will be held. At this Institute, students will learn to use the latest tools and will learn from each other's experiences in dealing with Web data. In addition, a one-day workshop will be developed, to be offered at Web Science conferences (WWW, SIGIR, etc.) to educate participants about WAC resources. An undergraduate Web Sciences track for computer science majors will be set up, taking advantage of WAC resources. The project will have impact in two ways. First, it will provide tools and services that facilitate access to Web resources. Any researcher, from a computer scientist studying efficient Web search, to a social scientist studying how human beliefs are changing today, to a historian studying how the early Web evolved, to a biologist understanding how disease spreads, will benefit from the work. Second, the project motivates students and young researchers to stay in academia. Currently top talent is flowing to industry because only they have comprehensive Web data, and it is so hard to do significant Web Science at universities. The WAC can provide an alternative, attracting more researchers and teachers to this important area.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
RAPID: Collaborative Research: COVID-19, Crises, and Support for the Rule of Law
Collaborative Research: Judicial Legitimacy in Comparative Perspective
Doctoral Dissertation Research in DRMS: Donation appeals for conservation - the influence of moral worldviews and moral foundations
  • 批准号:
    1725530
  • 项目类别:
    Standard Grant
  • 资助金额:
    $2.6万
  • 财政年份:
    2017
  • 负责人:
    Michael Nelson
  • 依托单位:
Collaborative Research: Testing Models of Representation and Institutional Design in the State Courts' Consideration of Inequality
国内基金
海外基金
基于水稻穗粒数关键基因LARGE2提高作物产量的探索与应用
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    黄洛将
  • 依托单位:
水稻穗粒数调控关键因子LARGE6的分子遗传网络解析
  • 批准号:
    --
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    30万元
  • 批准年份:
    2022
  • 负责人:
    黄洛将
  • 依托单位:
量子自旋液体中拓扑拟粒子的性质:量子蒙特卡罗和新的large-N理论
  • 批准号:
    12074246
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2020
  • 负责人:
    Yoshitomo Kamiya
  • 依托单位:
甘蓝型油菜Large Grain基因调控粒重的分子机制研究
  • 批准号:
    31972875
  • 项目类别:
    面上项目
  • 资助金额:
    58.0万元
  • 批准年份:
    2019
  • 负责人:
    石江华
  • 依托单位: