BRIDGE Center Standards Core
BRIDGE Center Standards Core
批准号:
10473242
负责人:
Monica Cecilia Munoz-Torres
金额:
$139.95万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
未结题
起止时间:
2022-07-06 至 2026-04-30
关键词:
AddressAdoptedAdoptionAnatomyArtificial IntelligenceAwarenessBackBehavioralBenchmarkingBridge to Artificial IntelligenceBusinessesCodeCommunitiesConsultationsConsumptionDataData CollectionData DiscoveryData EngineeringData ProvenanceData ScientistData SetDepositionDevelopmentDisciplineDiseaseDocumentationEcosystemElementsEnsureEnvironmentEquipment and supply inventoriesEvaluationFAIR principlesGenerationsGenesGoalsHumanKnowledgeLanguageLicensingLinkMachine LearningModalityModelingModernizationMolecularMorphologic artifactsOntologyOutputPhenotypeProtocols documentationProviderQuality ControlReadinessRegistriesReproducibilityResearchResearch PersonnelResourcesSeaSemanticsServicesSourceSpecific qualifier valueSpecificityStandardizationSystemTerminologyTimeTrainingTranslational ResearchUnited States National Institutes of HealthUpdateVariantVocabularyWorkdashboarddata disseminationdata ingestiondata modelingdata qualitydata reusedata standardsempoweredinsightinteroperabilitylarge datasetsmachine learning modelnovelopen sourceprogramsquality assuranceresponseskillstoolweb portalworking group
中文摘要
BRIDGE中心标准核心项目总结
英文摘要
BRIDGE Center Standards Core Project Summary
AI offers great potential for the discovery of novel biomedical insights from linkages between disparate,
cross-domain datasets. Unfortunately, traditional hypothesis-driven datasets tend to be narrowly focused on
the targeted problem domain with little consideration to “AI-readiness”. To best enable the use of such datasets
in data-driven and cross-domain discovery, they must be made Findable, Accessible, Interoperable, and
Reusable (FAIR). Lack of FAIRness is particularly problematic for AI, which is data-hungry. To fully leverage the
power of AI approaches, researchers need to find and reuse data to combine into larger datasets, and the data
must be interoperable or harmonized to be combined meaningfully. Transforming pre-existing datasets into
AI-ready data is challenging, requiring extensive linking and curation by human experts. This challenge is
exacerbated when annotating and linking data across domains, where standards may be disparate in purpose
and specificity. Finally, many datasets do not adhere to best practices in data transparency, including content
attribution and conditions on distribution and reuse. These additional considerations of Traceability, Licensing,
and Connectedness create an operationalized model for FAIR: FAIR-TLC.
Overcoming the barriers to FAIR-TLC is key to translational science and AI-driven biomedical discovery. Our
team has led standards development efforts in numerous large consortia, including the GA4GH, HL7, and
N3C. Our standards for representing biomedical concepts have been widely adopted, including those for
human phenotypes (e.g., HPO, GA4GH Phenopackets), diseases (NCIt, Mondo, ICD-11), genes (Gene
Ontology), anatomy (Uberon), and molecular variation (GA4GH VRS). We have developed standards and tools
to address data provenance (SEPIO), contributions (Contributor Attribution Model), licensing barriers (Data
Use Ontology, Reusable Data Project), and connectivity (Linked data Model Language, LinkML).
We will build on our previous work, collaborative skills, and technical knowledge to develop a framework to
enable the harmonization of standards across biomedical domains. We will form working groups with
representatives of the Data Generation Projects (DGPs) to document use cases and synthesize data standard
requirements. We will provide protocols and training for specifying standards, and provide concierge services
in support of all deliverables and activities. We will create a version-controlled Bridge2AI Standards Registry to
inventory standards for use by the DGPs, specified in the modality-agnostic LinkML framework, discoverable
through the interactive Standards Hub, and automatically exportable to technical artifacts through our Data
Transformation Toolbox. We will build a Standards Evaluation Dashboard for assessment and discovery of
standards in datasets from Bridge2AI Data Generation Projects. We will promote best practices in the
transparent and responsible sharing of datasets and ML models through DUO, Datasheets, and Model Cards.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Integration, Dissemination and Evaluation(BRIDGE) Center for the NIH Bridge to Artificial Intelligence (BRIDGE2AI) Program
-
批准号:10661023
-
项目类别:
-
资助金额:$261.87万
-
财政年份:2022
-
负责人:Monica Cecilia Munoz-Torres
-
依托单位:
Integration, Dissemination and Evaluation(BRIDGE) Center for the NIH Bridge to Artificial Intelligence (BRIDGE2AI) Program
-
批准号:10473239
-
项目类别:
-
资助金额:$269.27万
-
财政年份:2022
-
负责人:Monica Cecilia Munoz-Torres
-
依托单位:
BRIDGE Center Standards Core
-
批准号:10661029
-
项目类别:
-
资助金额:$133.76万
-
财政年份:2022
-
负责人:Monica Cecilia Munoz-Torres
-
依托单位:
海外基金