课题基金 / 基金详情

RAPID: Collecting Reliable COVID-19 Datasets in Crisis Conditions

RAPID: Collecting Reliable COVID-19 Datasets in Crisis Conditions
RAPID:在危机情况下收集可靠的 COVID-19 数据集
批准号:
2029457
负责人:
Rastislav Bodik
金额:
$7.0万
依托单位:
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-05-01 至 2020-10-31

项目摘要

项目成果

Rastislav Bodik的其他基金

相似基金

相关文献

中文摘要
翻译
RAPID项目通过部署能够在危机条件下收集可靠的COVID-19相关数据集的技术,为减轻COVID-19对公共卫生、社会和经济的负面影响的方法提供支持。在2019冠状病毒病大流行等危机期间,医院和重要卫生组织等关键新数据的产生者缺乏时间和资源,无法使这些重要数据随时可供他人使用。不能指望已经负担过重的主要数据提供者做额外的工作,使数据更容易被其他人使用。即使是那些已经在其网站上发布数据的人,也往往没有时间编辑/修改数据,例如应用新引入的标签,例如Schema.org与冠状病毒相关的新标签。然而,这些数据在危机中是至关重要的,以便告知公众;改善应急反应;并帮助科学界努力寻找解决方案。目前,从事数据集收集的团队正在使用缓慢、乏味和艰苦的手工技术。本项目开发的交互式数据集收集工具将提供另一种方法,使志愿者社区能够帮助进行数据收集工作。开发的数据收集工具只需要一个互联网连接、一个网页浏览器和简短的培训就可以使用,从而使大量潜在的志愿者可以很容易地完成这项工作。现有的自动数据提取器假设(i)单个网站中的网页结构是统一的,因为它们是由相同的模板生成的;(ii)相关网页来自单个网站。因此,之前在web数据提取和摄取领域的大部分工作都集中在“语法”提取上。目前,专门的数据收集团队正在通过专业知识和耗时、艰苦的手工工作来收集数据。其他团队正在雇佣呼叫中心给每个州的医院打电话,以收集他们的能力。这种高成本、高努力的方法不能很好地扩展到人们希望能够访问和分析的所有数据集。许多与covid -19相关的数据集分散在数千个网站上,这些网站具有相似的信息,但没有结构相似性。,每家医院的网站可能看起来不同,但可能包含非常相似和相关的数据。这个项目将解决的技术挑战是建立一个“语义”数据提取器,在不同的网站结构下定位感兴趣的信息。将为数据摄取而创建的软件工具可供许多热衷于贡献时间和精力帮助抗击COVID-19的个人使用,而不会影响他们保持身体距离的努力。该RAPID奖项由综合活动办公室的融合加速器项目颁发,并与融合加速器轨道A:开放知识网络有关。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This RAPID project enables approaches to mitigate the negative impacts of COVID-19 on public health, society, and the economy by deploying technologies to enable collecting reliable COVID-19-related data sets under crisis conditions. In the midst of a crisis, such as the COVID-19 pandemic, generators of critical new data, such as hospitals and critical health organizations, lack the time and resources to make this important data readily available for use by others. One cannot expect the already overburdened primary data providers to do the extra work needed to make the data more accessible for others to use. Even those who already publish data on their websites often do not have the time to edit/modify the data, for example to apply newly introduced tags, such as Schema.org’s new tags related to coronavirus. Yet, these data are critical in a crisis in order to inform the public; improve emergency response; and aid the scientific community in its efforts to find solutions. Currently, the teams that are engaged in dataset collection are employing slow, tedious, and painstaking manual techniques. The interactive dataset collection tools to be developed by this project will provide an alternative approach, empowering a community of volunteers to help with data collection efforts. The data collection tools developed can be used with only an internet connection, a web browser, and brief training, thereby putting the effort well within reach of a large population of potential volunteers. Existing automatic data extractors assume that (i) webpages in a single website are structured uniformly, because they were produced from the same template and (ii) relevant webpages originate from a single website. As a result, much of the prior work in the area of web data extraction and ingestion focuses on ‘syntactic’ extraction. Currently, dedicated data collection teams are collecting data with a combination of expertise and time-consuming and painstaking manual effort. Other teams are hiring call centers to call hospitals in each state to collect their capacities. Such high-cost, high-effort approaches do not scale well to all the datasets that one would like to be able to access and analyze. Many COVID-19-related datasets are scattered over thousands of websites with similar information but no structural similarities--e.g., each hospital’s website may look different but may contain very similar and related data. The technical challenge that this project will tackle will be to build a ‘semantic’ data extractor that locates the information of interest despite divergent website structures. The software tools that will be created for data ingestion can be used by the many individuals who are keen to contribute their time and effort to help combat COVID-19, without compromising their physical distancing efforts.This RAPID award is made by the Convergence Accelerator program in the Office of Integrative Activities and is associated with the Convergence Accelerator Track A: Open Knowledge Network.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Collaborative Research: FMitF: Track I: End-usser Programming for CAD Systems via Language Design and Synthesis
  • 批准号:
    2219864
  • 项目类别:
    Standard Grant
  • 资助金额:
    $50.0万
  • 财政年份:
    2022
  • 负责人:
    Rastislav Bodik
  • 依托单位:
FMitF: Track I: End-User Programming with Synthesis-Guided Interaction Models
  • 批准号:
    2122950
  • 项目类别:
    Standard Grant
  • 资助金额:
    $74.97万
  • 财政年份:
    2021
  • 负责人:
    Rastislav Bodik
  • 依托单位:
FMitF: Track II: Programming by Demonstration for the Browser with Applications in Data Science
  • 批准号:
    1918027
  • 项目类别:
    Standard Grant
  • 资助金额:
    $9.89万
  • 财政年份:
    2019
  • 负责人:
    Rastislav Bodik
  • 依托单位:
Convergence Accelerator Phase I (RAISE): Linking the Open Knowledge Network to the Web with End-User Programming
  • 批准号:
    1936731
  • 项目类别:
    Standard Grant
  • 资助金额:
    $99.47万
  • 财政年份:
    2019
  • 负责人:
    Rastislav Bodik
  • 依托单位:
海外基金