课题基金 / 基金详情

项目摘要

项目成果

Gwênlyn Glusman的其他基金

相似基金

相关文献

中文摘要
翻译
新技术使人们能够获得密集的个人“数据云”。然而, 这些数据的异质性、维度和多尺度性质(基因组、转录本、 临床变量等)提出了新的挑战:如何查询如此密集的数据云 混合数据作为针对多个知识的集成集合(而不是逐个变量) 碱基,并将联合分子信息转化到临床领域?当前词汇映射 暴力数据挖掘寻求使异类数据可互操作和可访问 但他们的产出是支离破碎的,需要专业知识才能整合成连贯的、可操作的 信息。我们提出了DeepTranslate,这是一种创新的方法,它结合了已知的 作为致病底物的生物实体的实际物理组织(I) 网络(数据图)和(Ii)跨越多尺度空间的概念层次 分子应用于临床。按照这样的自然结构组织数据源将允许翻译 将迅速增长的高维数据集转变为临床医生熟悉的概念,同时捕获 机械的关系。DeepTranslate将采用混合方法来学习和组织其 来自以下两个方面的内容:(I)现有的通用综合知识来源(GO、KEGG、IDC等) 以及(Ii)来自两个示范项目的单个数据云的新测量实例:(1) ISB的先锋100和(2)圣犹大终身癌症幸存者。我们将重点关注糖尿病,因为 测试用例。这两项研究涵盖了深层次的生物尺度-空间,因此可以全面检验 DeepTranslate在专注的应用程序中的多尺度能力。 1.启用的研究问题类型。临床医生如何才能发现几十个 从患者的数据云中观察到的“范围外”变量中,形成一个相互关联的集合 从基因到临床变量的病理生理学途径?高维空间如何 为每个个体测量100个不同类型数据点的研究数据 (“个人数据云”)作为一个集合进行综合分析(而不是可变的 按变量),并也用于改进数据库? DeepTranslate解决了这两类问题,从而将加快翻译 未来的个人数据云化成(A)护理决策和(B)关于新疾病机制的假设 /治疗,从而使提供者和研究人员受益。 2.利用专门知识和资源。ISB:个性化、大数据驱动的先驱 医学(演示项目1);生物医学内容专长;多尺度组学和分子发病机制, 大数据分析,为公众访问提供数据库;查询引擎设计,图形用户界面。 UCSD:生物医学数据集成领域的领先者;分子和临床的自动组装 数据到分层结构.数据类型U蒙特利尔之间的转换:生物医学数据库 从文献中筛选和构建基因/蛋白质/药物相互作用网络;机器 学习,开放资源数据库圣裘德CRH:癌症监测演示项目2, 癌症患者数据分析。 3.潜在的数据和基础设施挑战。(A)现有的综合 临床数据来源不统一且不明确地基于生物网络;交叉映射 根据词汇关系:HPO(表型)与 SNOMED CT(用于EMR)与IDC或Merck手册(用于疾病)。仔细挑选这些 需要与NLM密切合作的消息来源。(B)现有的分子途径数据库 是静态的,基于异质非分层种群的平均值,而新的 测量的高维数据云因个体内部的时间波动而变化 以及个体间的差异。这将如何影响我们的混合方法中的本体类型的构建, 而且,必须有多大的数据云队列才能提供统计能力,这还有待确定。 我们的两个演示项目及其独特的深度(高维和多尺度)数据 因此,在规模有限但不断扩大的队列中,是集体行动漫长旅程中至关重要的第一步 在翻译者社区学习。
英文摘要
New technologies afford the acquisition of dense “data clouds” of individual humans. However, heterogeneity, dimensionality and multi-scale nature of such data (genomes, transcriptomes, clinical variables, etc.) pose a new challenge: How can one query such dense data clouds of mixed data as an integrated set (as opposed to variable by variable) against multiple knowledge bases, and translate the joint molecular information into the clinical realm? Current lexical mapping and brute-force data mining seek to make heterogeneous data interoperable and accessible but their output is fragmented and requires expertise to assemble into coherent actionable information. We propose DeepTranslate, an innovative approach that incorporates the known actual physical organization of biological entities that are the substrate of pathogenesis into (i) networks (data graphs) and (ii) hierarchies of concepts that span the multiscale space from molecule to clinic. Organizing data sources along such natural structures will allow translation of burgeoning high-dimensional data sets into concepts familiar to clinicians, while capturing mechanistic relationships. DeepTranslate will take a hybrid approach to learn and organize its content from both (i) existing generic comprehensive knowledge sources (GO, KEGG, IDC, etc.) and (ii) newly measured instances of individual data clouds from two demonstration projects: (1) ISB’s Pioneer 100 and (2) St. Jude Lifetime cancer survivors. We will focus on diabetes as test case. These two studies cover a deep biological scale-space and thus can test the full extent of the multiscale capacity of DeepTranslate in a focused application. 1. TYPES OF RESEARCH QUESTION ENABLED. How can a clinician find out that the dozens of “out of range” variables observed in a patient’s data cloud, form a connected set with respect to pathophysiology pathways, from gene to clinical variable? How can the high-dimensional data of studies that measure for each individual 100+ data points of various types (“personal data clouds”) be analyzed as one set in an integrated fashion (as opposed to variable by variable) against existing knowledge bases and also be used to improve the databases? DeepTranslate addresses these two types of questions and thereby will accelerate translation of future personal data clouds into (A) care decisions and (B) hypotheses on new disease mechanisms / treatments, thereby benefiting providers as well as researchers. 2. USE OF EXPERTISE AND RESOURCES. ■ ISB: pioneer in personalized, big-data driven medicine (Demo Project 1); biomedical content expertise; multiscale omics and molecular pathogenesis, big data analysis, housing of databases for public access; query engine designs, GUI. ■ UCSD: leader in biomedical data integration; automated assembly of molecular and clinical data into hierarchical structures; translation between data types ■ U Montreal: biomedical database curation from literature and construction of gene/protein/drug interaction networks; machine learning, open resource database ■ St Jude CRH: Cancer monitoring Demo Project 2, cancer patient data analytics. 3. POTENTIAL DATA AND INFRASTRUCTURE CHALLENGES. (a) Existing comprehensive clinical data sources are not uniform and not explicitly based on biological networks; cross-mapping is being performed at NLM based on lexical relationships: HPO (phenotypes) vs. SNOMED CT (for EMR) vs. IDC or Merck Manual (for diseases). Careful selection of these sources in close collaboration with NLM is needed. (b) Existing molecular pathway databases are static, based on averages of heterogeneous non-stratified populations, while the newly measured high-dimensional data clouds are varied due to intra-individual temporal fluctuation and inter-individual variation. How this will affect building of ontotypes in our hybrid approach, and how large cohorts of data clouds must be to offer statistical power is yet to be determined. Our two Demonstration Projects with their uniquely deep (high-dimensional and multiscale) data in cohorts of limited but growing size are thus crucial first steps in a long journey of collective learning in the TRANSLATOR community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
DOCKET: accelerating knowledge extraction from biomedical data sets
  • 批准号:
    10057127
  • 项目类别:
  • 资助金额:
    $60.91万
  • 财政年份:
    2020
  • 负责人:
    Gwênlyn Glusman
  • 依托单位:
DOCKET: accelerating knowledge extraction from biomedical data sets
  • 批准号:
    10548024
  • 项目类别:
  • 资助金额:
    $60.91万
  • 财政年份:
    2020
  • 负责人:
    Gwênlyn Glusman
  • 依托单位:
DOCKET: accelerating knowledge extraction from biomedical data sets
  • 批准号:
    10330627
  • 项目类别:
  • 资助金额:
    $67.68万
  • 财政年份:
    2020
  • 负责人:
    Gwênlyn Glusman
  • 依托单位:
DOCKET: accelerating knowledge extraction from biomedical data sets
  • 批准号:
    10706750
  • 项目类别:
  • 资助金额:
    $60.91万
  • 财政年份:
    2020
  • 负责人:
    Gwênlyn Glusman
  • 依托单位:
海外基金