课题基金 / 基金详情

项目摘要

项目成果

Mark A Musen的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
An essential task for the Biomedical Data Translator is to identify scientific experiments that have been performed or that are ongoing, and to enable integration of knowledge of the experimental methods, the results, and—when available—the conclusions with other knowledge sources. Such capabilities will enable queries such as: (1) Has anyone ever performed an experiment using methods like these? (2) Has anyone performed a study where the data may support a particular conclusion? (3) Are there any clinical trials for a particular condition whose patient population is a good match for a patient whom I now need to treat? (4) What best practices are suggested by the results of current clinical trials for a particular condition? Sometimes such queries can be addressed through an analysis of the scientific literature. More often, however, the published literature does not provide the methodological details needed to address such questions—even if NLP techniques were good enough to find the answers. Publications also provide only summary statistics of the experimental results. To address the kinds of queries that are of most interest to the Translator, it is necessary to access the actual experimental data online, starting with the metadata that are intended to provide descriptions of the datasets and of the experiments that led to the collection of the data in the first place. The problem for the Translator project is that the metadata that describe most online experimental data sources are difficult for computers to find and to process. Our laboratory’s analysis of the NCBI BioSample metadata repository, for example, shows that scientists largely avoid using standard data dictionaries entirely, and—partly as a result—they are extremely sloppy when they provide metadata values [3]. (A case in point: Some 76% of the metadata values in BioSample that are intended to be Boolean are neither true nor false.) Despite all the discussion in the past few years about making online datasets Findable, Accessible, Interoperable, and Re-usable (FAIR) [14], most online datasets are not close to FAIR. Our laboratory is developing technology that can rectify errors in online metadata. Like a spell-checker for metadata, our approach will attempt to identify the intentions of metadata authors, to correct typos, and to convert free-text strings to ontology terms whenever possible [6]. Our goal is to provide a service that will transform the scruffy metadata that pervade online descriptions of biomedical experiments into a form that will allow automated discovery, integration, and secondary analysis of research results in ways that are simply not possible at present. We anticipate that the Translator will call on our service to find experimental datasets and their accompanying metadata, to perform standard analyses of such datasets, and to integrate descriptions of experiments into the evolving knowledge graph. We will evaluate the performance of our Knowledge Provider by studying its response to queries from the Translator community and by peer review of a subset of the underlying, cleaned up metadata records that it processes from actual online repositories, such as BioSample and ClinicalTrials.gov. Our evaluation necessarily will be limited by the pragmatics of selecting a manageable test set of metadata and by the inherent shortcomings of manual peer review. Our laboratory has a sustained tradition of collaborating to develop major national resources that bring semantic technology to biomedicine. Our BioPortal ontology repository [5] was developed by the National Center for Biomedical Ontology (NCBO), one of the NIH National Centers for Biomedical Computing. The CEDAR Workbench for the prospective authoring of standardized metadata [11,12] was developed under the NIH Big Data to Knowledge (BD2K) program. Our Protégé system for building and maintaining biomedical ontologies is the most widely used software for creating semantic technology in the world [15]. Our group has ongoing relationships with corporations such as Pinterest, BASF, and Elsevier to assist them in their work to develop enterprise-wide knowledge graphs. We are thus well equipped to develop our Knowledge Provider and to assist the consortium broadly in the area of semantic technology.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Enhanced ontology engineering through a Web-based, Cloud-based software architecture
  • 批准号:
    10405968
  • 项目类别:
  • 资助金额:
    $23.61万
  • 财政年份:
    2021
  • 负责人:
    Mark A Musen
  • 依托单位:
The Metadata Powerwash - Integrated tools to make biomedical data FAIR
  • 批准号:
    10397981
  • 项目类别:
  • 资助金额:
    $33.45万
  • 财政年份:
    2021
  • 负责人:
    Mark A Musen
  • 依托单位:
Enhancing the RADx Data Hub for Data FAIRness
  • 批准号:
    10433797
  • 项目类别:
  • 资助金额:
    $300.0万
  • 财政年份:
    2021
  • 负责人:
    Mark A Musen
  • 依托单位:
Enhancing the RADx Data Hub for Data FAIRness
  • 批准号:
    10794704
  • 项目类别:
  • 资助金额:
    $1010.0万
  • 财政年份:
    2021
  • 负责人:
    Mark A Musen
  • 依托单位:
国内基金
海外基金
层出镰刀菌氮代谢调控因子AreA 介导伏马菌素 FB1 生物合成的作用机理
  • 批准号:
    2021JJ40433
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2021
  • 负责人:
    孙磊
  • 依托单位:
寄主诱导梢腐病菌AreA和CYP51基因沉默增强甘蔗抗病性机制解析
  • 批准号:
    32001603
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    24.0万元
  • 批准年份:
    2020
  • 负责人:
    段真珍
  • 依托单位:
AREA国际经济模型的移植.改进和应用
  • 批准号:
    18870435
  • 项目类别:
    面上项目
  • 资助金额:
    2.0万元
  • 批准年份:
    1988
  • 负责人:
    史树中
  • 依托单位: