Data-driven search of Common Fund data sets for better discoverability and novel meta-analysis
Data-driven search of Common Fund data sets for better discoverability and novel meta-analysis
批准号:
10577377
负责人:
Tamer Ahmed Mansour Ahmed
金额:
$31.3万
依托单位国家:
美国
项目类别:
财政年份:
2022
资助国家:
美国
项目状态:
已结题
起止时间:
2022-09-20 至 2024-09-19
关键词:
AddressAdoptedAlgorithmsAlzheimer&aposs DiseaseAnimal ModelCatalogsChildCollaborationsCollectionComplexComputer softwareDataData SetData SourcesDatabasesDescriptorDevelopmentDiseaseExerciseFoundationsFundingGenerationsGenesGenotypeGenotype-Tissue Expression ProjectGrantGraphHealthInternationalMalignant NeoplasmsMeta-AnalysisMetadataModelingMolecularMusNetwork-basedOutputOverlapping GenesPathway interactionsPhenotypePhysical activityPrivatizationRecipeResearchResourcesSyndromeTechniquesTestingTissuesTransducersUnited States National Institutes of Healthbasecancer predispositioncohortcomputing resourcescongenital anomalydatabase queryexperienceexperimental studyflexibilityhuman diseaseinsightinterestmouse modelmultiple datasetsnovelnovel strategiesprogramsprototyperepositorysoftware developmenttoolusabilityuser-friendly
中文摘要
项目摘要
美国国立卫生研究院共同基金(CF)计划产生了许多独特和高价值的数据集。
为了解决复杂的生物医学问题,我们需要找到可以共同分析的相关数据集
具体的研究目的。当前的许多搜索技术依赖于不同的数据描述符
跨CF计划,可能不完整或不准确。这些实验中的许多都输出了
对某些生物医学条件有重要意义的基因。我们建议使用这些基因列表来寻找
相似的数据集。这种方法不仅可以跨CF数据集进行搜索,还可以连接
用于其他数据库和生物医学目录中的其他实验,例如,包含
疾病-基因关联和分子途径。为了实现这一目标,我们将实施有效的
计算大量基因集合之间相似性的线性算法。我们的原型工具,
DBRetina,使用这种算法在几分钟内建立巨大的相似网络,使用最少的
计算资源。DBRetina是研究相似性图CurIndex的基础
连接多个健康相关资源的数据库。DBRetina和CurIndex将允许高级
搜索相关的CF实验,帮助更好地解释生物医学数据。
英文摘要
Project Summary
NIH Common Fund (CF) programs have produced a number of unique and high-value data sets.
To solve complex biomedical questions, we need to find related data sets that can be co-analyzed for
specific study purposes. Many of the current search techniques depend on data descriptors which differ
across CF programs and may be incomplete or inaccurate. Many of these experiments output lists of
genes significant to certain biomedical conditions. We are proposing to use these gene lists to find
similar data sets. This approach will not only enable searching across CF data sets but also can connect
them to other experiments in other databases and biomedical catalogs, e.g., databases containing
disease-gene associations and molecular pathways. To achieve this aim, we will implement an efficient
linear algorithm to calculate similarities between large numbers of gene sets. Our prototype tool,
DBRetina, uses this algorithm to build huge similarity networks in few minutes using minimal
computational resources. DBRetina serves as the foundation for CurIndex, a study similarity graph
database that connects multiple health-related resources. DBRetina and CurIndex will allow advanced
search for related CF experiments and facilitate better interpretation of biomedical data.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金