RAPID: Rich and Accurate Auxiliary Databases for Supporting Virus Data Efforts
RAPID: Rich and Accurate Auxiliary Databases for Supporting Virus Data Efforts
批准号:
2029556
负责人:
Michael Cafarella
金额:
$16.48万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2020
资助国家:
美国
项目状态:
已结题
起止时间:
2020-05-01 至 2022-04-30
中文摘要
COVID-19快速应对项目将通过利用与COVID-19相关的医疗和政府服务的高度分布式数据创建高质量数据库,帮助减轻COVID-19对公共卫生、社会和经济的负面影响。该项目将开发软件工具,帮助创建包含高质量数据的“辅助”数据库,以帮助在冠状病毒大流行造成的紧急和快速演变的情况下做出更好的决策,避免欺诈,并更快地提供高质量的分析。将用于实现高质量的技术包括:(1)将“背景数据”链接到数据集,以便进行质量检查和欺诈检测。例如,确保在医疗资源数据库中列出的医院信息标注准确的电话号码,以便志愿者可以联系医院并检查数据的准确性;(2)创建新的“连接键”,使辅助数据库中的数据易于与其他数据集成。该项目将与其他相关的COVID-19快速发展项目密切合作,这些项目正在从网络收集数据和信息的各个方面。该项目将重点使用以下策略创建两个高质量的数据库:(1)一个统一的医疗机构辅助数据库,这将是美国所有已知医疗机构的数据库;(2)一个统一的政府办公室辅助数据库,这将是美国所有已知政府办公室的数据库-市政厅,法院,许可证办公室等-在任何级别的政府。这两组数据对于确保公民获得基本水平的医疗援助和政府援助至关重要。这些资源不仅有利于防治这一特殊流行病,而且将成为今后总体上必不可少的资源。建议的辅助数据集创建基础设施将包括用于质量检查的丰富的背景信息模式和用于数据集成的一组连接键。虽然在线上有大量医疗机构数据集,但由于缺乏标准名称和/或数据集成键,许多数据集没有对齐,因为不同的项目在选择这些值时做出不同的本地决策,而这些值可能不是普遍兼容的。因此,背景信息变得不那么丰富,使得与其他机构或分析管道的数据集成变得更加困难。用于创建此基础设施的策略包括:(1)综合初步辅助数据集,其中包括为输入集中的所有对象生成通用的候选属性,例如,基于检查维基数据中的所有医院数据为医院创建直升机场字段;(2)识别缺失值的输入,并结合Web提取任务和众包任务来填充这些值;(3)标记怀疑不正确的值,例如,通过自动为辅助数据中的每列创建一组机器学习预测器。然后,系统可以运行预测器并识别异常值。该RAPID奖由综合活动办公室的融合加速器项目利用《冠状病毒援助、救济和经济安全法案》(CARES Act)的资金颁发,并与融合加速器轨道A:开放知识网络相关。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
This COIVD-19 RAPID project will assist in the mitigation of the negative impacts of COVID-19 on public health, society, and the economy, by creating high-quality databases from highly distributed data about medical and governmental services related to COVID-19. The project will develop software tools to help in the creation of "auxiliary" databases with high-quality data to assist in making better decisions, avoiding fraud, and yielding high-quality analysis sooner in the urgent and rapidly evolving situation created by the coronavirus pandemic. The techniques that will be used to achieve high-quality include:(1) linking "background data" to the data sets to enable quality-checking and fraud detection. For example, ensuring that hospital information listed in the medical resource database is annotated with an accurate phone number so that a volunteer can contact the hospital and check on the accuracy of the data, and (2) creating new "join keys" to enable easy integration of data in the auxiliary database with other data. The project will work closely with other related COVID-19 RAPID efforts which are working on various aspects of data and information collection from the Web.The project will focus on creating two high-quality databases using these strategies: (1) A unified medical institution auxiliary database, which will be a database of all known US medical institutions and (2) A unified government office auxiliary database, which will be a database of all known government offices in the United States—city halls, courts, licensing offices, etc.—at any level of government. Both these data sets are crucial for ensuring that citizens receive a base level of medical aid and government assistance. These resources would be beneficial not only for this particular pandemic, but would become essential resources, in general, for the future. The proposed auxiliary data set creation infrastructure will include a rich schema of background information, used for quality-checking, and a set of join keys for data integration. While there is a huge array of medical institution data sets online, many of the data sets are misaligned due to lack of standard names and/or data integration keys since different projects make different local decisions in choosing these values that may not be universally compatible. As a result, the background information becomes less rich and makes integration with data from other institutions or analysis pipelines much more difficult. The strategies used to create this infrastructure would include: (1) synthesis of preliminary auxiliary datasets, which includes generating common, candidate attributes for all objects in the input set, for example, creating a helipad field for hospitals based on examining all hospital data in Wikidata; (2) identification of inputs with missing values, and filling in those values with a combination of Web extraction tasks and crowdsourcing tasks, and (3) flagging values that are suspected of being incorrect by, for example, automatically creating a set of machine-learned predictors for each column in the auxiliary data. The system could then run the predictor and identify outlier values.This RAPID award is made by the Convergence Accelerator program in the Office of Integrative Activities using funds from the Coronavirus Aid, Relief, and Economic Security (CARES) Act, and is associated with the Convergence Accelerator Track A: Open Knowledge Network.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
A1: Knowledge Network Development Infrastructure with Application to COVID-19 Science and Economics
-
批准号:2132318
-
项目类别:Cooperative Agreement
-
资助金额:$499.45万
-
财政年份:2021
-
负责人:Michael Cafarella
-
依托单位:
A1: Knowledge Network Development Infrastructure with Application to COVID-19 Science and Economics
-
批准号:2033558
-
项目类别:Cooperative Agreement
-
资助金额:$499.45万
-
财政年份:2020
-
负责人:Michael Cafarella
-
依托单位:
Convergence Accelerator Phase I (RAISE): Simultaneous Knowledge Network Programming and Extraction
-
批准号:1936940
-
项目类别:Standard Grant
-
资助金额:$100.0万
-
财政年份:2019
-
负责人:Michael Cafarella
-
依托单位:
I-Corps: Explanation-Based Auditing: Improving the Security of Electronic Medical Records
-
批准号:1340372
-
项目类别:Standard Grant
-
资助金额:$5.0万
-
财政年份:2013
-
负责人:Michael Cafarella
-
依托单位:
CAREER: Building and Searching a Structured Web Database
-
批准号:1054913
-
项目类别:Continuing Grant
-
资助金额:$48.86万
-
财政年份:2011
-
负责人:Michael Cafarella
-
依托单位:
III: Medium: Collaborative Research: Database-As-A-Service for Long Tail Science
-
批准号:1064606
-
项目类别:Continuing Grant
-
资助金额:$23.2万
-
财政年份:2011
-
负责人:Michael Cafarella
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Rich2通过调控自噬抑制炎症小体NLRP3通路在癫痫形成中的机制研
究
-
批准号:
-
项目类别:省市级项目
-
资助金额:--
-
批准年份:2024
-
负责人:张小刚
-
依托单位:
前扣带回GTP酶激活蛋白RICH2介导Shank3-/-孤独症小鼠社交行为障碍的机制研究
-
批准号:82301350
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:张佳瑞
-
依托单位:
整合素β1/RICH1复合体感应细胞外基质硬度信号调控乳腺癌侵袭转移的机制研究
-
批准号:82303462
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2023
-
负责人:田琦
-
依托单位:
转录因子NtMYB305通过AT-rich元件调控NtPMT表达及烟碱合成的分子机制研究
-
批准号:32101643
-
项目类别:青年科学基金项目(C类)
-
资助金额:30.0万元
-
批准年份:2021
-
负责人:田田
-
依托单位:
Rich1/Amot-p80/Merlin轴通过Hippo通路调控乳腺癌干细胞样特性的机制研究
-
批准号:82002794
-
项目类别:青年科学基金项目
-
资助金额:24.0万元
-
批准年份:2020
-
负责人:杨姣
-
依托单位:
烟草花叶病毒RNA发生poly(A)-rich型多聚腺苷酸化的研究
-
批准号:31370181
-
项目类别:面上项目
-
资助金额:82.0万元
-
批准年份:2013
-
负责人:李为民
-
依托单位:
端粒延伸过程中C链合成(C-rich Fill-in)的分子机理
-
批准号:31271472
-
项目类别:面上项目
-
资助金额:90.0万元
-
批准年份:2012
-
负责人:赵勇
-
依托单位:
CA-rich顺式元件及其相互作用的反式因子对可变剪接的调控机制
-
批准号:30970620
-
项目类别:面上项目
-
资助金额:32.0万元
-
批准年份:2009
-
负责人:惠静毅
-
依托单位:
果蝇硒蛋白G-rich的细胞定位、拓扑结构和分子功能研究
-
批准号:30671176
-
项目类别:面上项目
-
资助金额:24.0万元
-
批准年份:2006
-
负责人:陈长兰
-
依托单位:
RICH/PHENIX相对论性重离子对撞实验中的μ子探测
-
批准号:10145008
-
项目类别:专项基金项目
-
资助金额:8.0万元
-
批准年份:2001
-
负责人:冒亚军
-
依托单位: