Heuristics to evaluate biomedical and genomic knowledge bases for validity
Heuristics to evaluate biomedical and genomic knowledge bases for validity
批准号:
9765396
负责人:
Jesse Gillis
金额:
$48.0万
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-09-15 至 2021-08-31
关键词:
AddressAffectAutomobile DrivingBiologicalBiological AssayBiomedical ResearchBrainCatalogsCellsCollaborationsCommunitiesComplexDataData CollectionData ReportingData SetDatabasesEnsureEvaluationFrequenciesGenesGeneticGenomeGenomicsGoalsGoldGuidelinesInformation ResourcesJointsKnowledgeLaboratoriesLibrariesLiteratureMachine LearningMeasuresMethodologyMethodsMiningMinorModificationMorphologic artifactsMusOntologyOutputPaperPerformancePhenotypePositioning AttributePrevalenceProbabilityProliferatingPropertyPublishingQuality ControlRare DiseasesReportingResearchResourcesSemanticsSeriesSourceSpecificityStandardizationStructureSystemTestingTextTimeValidationValidity and ReliabilityVariantWorkanalytical methodbasebiomedical resourcecomparativedark matterdata resourcedata structuredatabase structuredesignexperimental studygene functionheuristicsimprovedinsightinterestknowledge baselearning strategynoveltext searchingtool
中文摘要
项目摘要
我们的首要目标是了解基因及其功能的特征信息如何被
组织,整合,然后推广到新的环境中。这是后基因组时代的核心问题,
随着新的分析方法扩大了信息的范围、广度和细节,
描述基因特性。虽然基因本体论是最突出和通用的系统,
除了组织基因功能之外,还有数百种其他基因,通常服务于专门的研究兴趣。最
实验室依赖于这些数据的某些子集的有效性来设计新的实验或解释它们。
结果,但其质量很难直接确定,特别是在新的或复杂的综合
方法论。基于大量的初步数据,我们假设确定鲁棒性和
具体性将提供对数据库效用的高度一般性评估。我们建议使用这些
属性,以评估整个语料库的资源组织基因信息,以及方法,
利用这些信息,以及他们报告的结果。重要的是,确定稳健性和特异性
不需要关于“黄金标准”信息的验证。通过评估这些资源,
他们的联合特异性和鲁棒性,我们确定了整合和组织其数据的方法,
新颖的应用。最后,我们建议将我们在质量控制方面的改进应用于更好地靶向稀有但
这是一个实验目标,特别是罕见疾病和单细胞表达。
该项目的三个相辅相成的目标是:
1.确定表征基因功能的数据的唯一性和稳健性。我们开发了一个
通过利用先验概率来表征鲁棒性和唯一性/特异性的正式方法
基因的多功能性。我们将在基本上所有复杂和复杂的环境中评估稳健性和特异性。
结构化数据库表征基因。这些度量可以在数据库之间或随着时间的推移进行比较
并提供数据结构的全局视图。
2.旨在利用描述基因功能的信息的测试方法。统计和机器
将评估利用结构化数据的学习方法是否具有稳健和具体的产出。数据特征
将确定不同应用程序中的驱动性能,并补充数据源,
将确定社区集群。
3.评估依赖于使用描述基因功能的数据库的结果。使用
结合文本挖掘和图形挖掘,我们将评估正在进行的文献的新颖性,稳健性,
特定的基因功能关联。我们将描述和评估基因功能的“暗物质”
从未注释的基因以及不完整的功能这两个角度来分析关联。
英文摘要
Project Summary
Our overarching goal is to understand how information characterizing genes and their function can be
organized, integrated, and then generalized to new contexts. This is a central question of the post-genomic era,
and one that becomes ever more pressing as novel assays expand the scope, breadth, and detail of information
describing gene properties. While the Gene Ontology is the most prominent and universal system for
organizing gene function, hundreds of others exist, often serving specialized research interests. Most
laboratories depend on the validity of some subset of this data to design new experiments or interpret their
results, but their quality is hard to directly ascertain, particularly in novel or complex integrative
methodologies. Based on substantial preliminary data, we hypothesize that determining robustness and
specificity will provide a highly general assessment of the utility of databases. We propose to use these
properties to assess the entire corpus of resources organizing gene information, as well as the methods which
exploit this information, and the results that they report. Critically, determining robustness and specificity does
not require validation with respect to ‘gold standard’ information. By evaluating these resources with respect to
their joint specificity and robustness we determine means of integrating and organizing their data for use in
novel applications. Finally, we propose to apply our improvements in quality control to better target rare but
robust results where this is an experimental goal, notably rare diseases and single cell expression.
The three complementary objectives in this project are to:
1. Determine the uniqueness and robustness of data characterizing gene function. We develop a
formal approach for characterizing robustness and uniqueness/specificity by exploiting prior probability in the
form of gene multifunctionality. We will evaluate robustness and specificity across essentially all complex and
structured databases characterizing genes. These measures can be compared between databases or over time
and provide a global landscape of data structure.
2. Test methods designed to exploit information describing gene function. Statistical and machine
learning methods exploiting structured data will be assessed for robust and specific output. Data features
driving performance in diverse applications will be identified and complementary sources of data as well as
community clusters will be defined.
3. Evaluate results that depended on the use of databases describing gene function. Using a
combination of text-mining and figure-mining, we will assess the ongoing literature for novel, robust, and
specific gene-function associations. We will characterize and evaluate the “dark matter” of gene-function
association from both the point of unannotated genes as well as incomplete functions.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Scalable Molecular Pipelines for FAIR and Reusable BICAN Molecular Data
-
批准号:10686157
-
项目类别:
-
资助金额:$171.85万
-
财政年份:2022
-
负责人:Jesse Gillis
-
依托单位:
Scalable Molecular Pipelines for FAIR and Reusable BICAN Molecular Data
-
批准号:10523659
-
项目类别:
-
资助金额:$176.48万
-
财政年份:2022
-
负责人:Jesse Gillis
-
依托单位:
Revealing the transcriptomic basis of neuronal identity through functional meta-analysis
-
批准号:10224662
-
项目类别:
-
资助金额:$48.0万
-
财政年份:2017
-
负责人:Jesse Gillis
-
依托单位:
Single-Cell Biology Shared Resource
-
批准号:10675645
-
项目类别:
-
资助金额:$20.96万
-
财政年份:1997
-
负责人:Jesse Gillis
-
依托单位:
Single-Cell Biology Shared Resource
-
批准号:10270226
-
项目类别:
-
资助金额:$20.96万
-
财政年份:1997
-
负责人:Jesse Gillis
-
依托单位:
海外基金