Unraveling the topological architecture and phenotypic contexture of structural variation
Unraveling the topological architecture and phenotypic contexture of structural variation
批准号:
10356208
负责人:
Marcin Piotr Cieslik
金额:
$29.7万
依托单位国家:
美国
项目类别:
财政年份:
2021
资助国家:
美国
项目状态:
已结题
起止时间:
2021-09-22 至 2023-09-21
关键词:
3-DimensionalAddressAdoptionAdultArchitectureAutomobile DrivingBiologicalCancer PatientChromatinChromosome abnormalityCicatrixClinicalCommunitiesCopy Number PolymorphismDNA Sequence RearrangementDataData SetEnhancersEpigenetic ProcessFoundationsFundingGene ExpressionGene Expression ProfileGeneticGenomeGenomic InstabilityGenomic medicineGenomicsGenotypeHumanIndividualKnowledgeLinkMalignant Childhood NeoplasmMalignant NeoplasmsMedicalMethodologyMethodsNeoplasm MetastasisNuclearPathogenicityPatientsPhenotypePrimary NeoplasmRecurrenceRefractoryRegulatory ElementResearchResourcesRoleSamplingStatistical ModelsStructural ModelsStructureTestingTimeTissue-Specific Gene ExpressionTissuesVariantcancer genomecancer typedata integrationdata resourceepigenomicsfunctional genomicsgenetic informationgenome sequencinggenomic datainsightmolecular phenotypenoveloncology programphenotypic dataprecision oncologypredictive modelingpredictive toolsprogramsproteogenomicstooltumorweb portalwhole genome
中文摘要
摘要
全基因组测序(WGS)在基因组医学和
精确肿瘤学加速了患者结构变异(SVS)的发现
癌症基因组。然而,尽管人类癌症类型的特点是普遍存在
基因组不稳定性:大多数结构和拷贝数变异(CNV)的功能后果
人们对此仍然知之甚少。关键的是,目前还不清楚成百上千的
通常在患者肿瘤中观察到的基因组重排是致病的,而不是
功能性基因组伤疤。因为SVS在结构(线性序列)上改变了基因组,
拓扑(三维组织)和表型水平(表观遗传景观),
综合和多尺度的数据集是正确预测其影响所必需的。这种匮乏
整合的资源和工具严重限制了对患者基因数据的医学解释。
现有的大规模基因组和蛋白质组癌症表征工作,包括
共同基金(CF)Gabriella Miller Kids First(GMKF)数据资源提供丰富的数据链接
遗传信息包括SVS及其表型结果,如基因表达。
然而,这些数据集本身不足以提供深刻的机械性和功能性洞察。
Cf数据集,特别是4D核组(4dN)、表观基因组学(路线图)和GTEx提供
将生殖系变异、基因组拓扑和染色质结构与基因表达联系起来的蓝图。
因此,我们建议将来自患者肿瘤样本的基因组数据(GMKF)与
空间和功能数据(4DN、路线图、GTEx),这将使我们能够阐明和预测
结构变异的致病机制:
目标1:创建TopVar数据资源,以增强我们对
基因组拓扑和结构变异。集成的TopVar资源将提供
需要表型背景来从遗传学和生物学角度解释SVS,这将产生可测试的
关于其下游影响的假设。
目的2:建立和评估多种人类癌症中病毒致病能力的预测模型。
使用结构化的TopVar数据资源,我们将实现一个可解释的统计模型
利用多层集成数据预测哪些SVS对基因表达有影响。
这两个目标的实现将为TopVar在预测中的应用提供一个原则证明
精确肿瘤学背景下的SVS建模。虽然我们提议的研究将集中在
关于查询GMKF(儿童癌症)和GMKF产生的全面基因组数据
CPTAC(成人癌症),它将作为它们在实时测序中使用的基础
MI-ONCOSEQ和PEDS-MI-ONCOSEQ等方案,重点关注难治性和转移性肿瘤。
英文摘要
Abstract
The increasing adoption of whole-genome sequencing (WGS) in the context of genomic medicine and
precision oncology has resulted in the accelerated discovery of structural variants (SVs) in patient
cancer genomes. However, while human cancer types are generally characterized by widespread
genomic instability the functional consequences of most structural and copy number variants (CNV)
remain poorly understood. Critically, it is unknown which of the hundreds to thousands of
genomic rearrangements typically observed in a patient tumor are pathogenic and which are non-
functional genomic scars. Because SVs alter the genome at the structural (linear sequence),
topological (three-dimensional organization), and phenotypic levels (epigenetic landscape),
integrative and multiscale datasets are necessary to correctly predict their impact. This dearth of
integrative resources and tools critically limits the medical interpretation of patient genetic data.
Existing large-scale genomic and proteogenomic cancer characterization efforts, including the
Common Fund (CF) Gabriella Miller Kids First (GMKF) data resource provide rich data to link
genetic information including SVs with their phenotypic consequences, such as gene expression.
However, these datasets alone are insufficient to provide deep mechanistic and functional insights.
CF data sets, specifically 4D Nucleome (4DN), Epigenomics (Roadmap), and GTEx provide the
blueprint to link germline variation, genome topology, and chromatin architecture to gene expression.
Therefore, we propose the integration of genomic data from patient tumor samples (GMKF), with
spatial and functional data (4DN, Roadmap, GTEx), which will allow us to elucidate and predict the
pathogenic mechanisms of structural variants:
Aim 1: To create TopVar a data resource to enhance our understanding of the interplay between
genome TOPology and structural VARiation. The integrative TopVar resource will provide the
phenotypic context required to interpret SVs in genetic and biological terms, which will yield testable
hypotheses regarding their downstream effects.
Aim 2: To develop and evaluate a predictive model of SV pathogenicity across multiple human cancers.
Using the structured TopVar data resource, we will implement an interpretable statistical model to
predict which SVs have an impact on gene expression, utilizing multiple layers of the integrated data.
The realization of both aims will represent a proof-of-principle for the utility of TopVar for predictive
modeling of SVs in the context of precision oncology. While our proposed study will focus
on interrogating the comprehensive genomic data generated by GMKF (pediatric cancer) and
CPTAC (adult cancer), it will serve as the foundation for their use within real-time sequencing
programs, such as MI-OncoSeq and Peds-MI-OncoSeq, focusing on refractory and metastatic tumors.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
海外基金