课题基金 / 基金详情

Bioinformatic Tools in Cancer Research

Bioinformatic Tools in Cancer Research
癌症研究中的生物信息工具
批准号:
7592714
负责人:
Kenneth H Buetow
金额:
$202.15万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AgeAlgorithmsAllelesAmino AcidsArtsBioinformaticsBiological MarkersCancer EtiologyCancer PatientCandidate Disease GeneCharacteristicsClassificationClinicalCodeCollaborationsCollectionCommunitiesComplexComputer SimulationCopy Number PolymorphismDNADNA ResequencingDNA biosynthesisDataData AnalysesData SetDatabasesDevelopmentDisease AssociationEnsureEpidermal Growth Factor ReceptorEvaluationFacility Construction Funding CategoryGenderGene ChipsGene ExpressionGene FrequencyGenesGeneticGenetic PolymorphismGenetic RiskGenomeGenomicsGenotypeGerm-Line MutationHaplotypesHumanHuman GenomeIndiumInformation SystemsInvestigationKoreansLaboratoriesLeadLinkLiverLocationLoss of HeterozygosityMalignant NeoplasmsMalignant neoplasm of liverMalignant neoplasm of lungMalignant neoplasm of ovaryManualsMeasuresMessenger RNAMethodologyMethodsModelingMolecular AbnormalityMolecular MedicineMolecular ProfilingMusMutagenesisMutationMutation DetectionNatureNormal tissue morphologyNumbersOdds RatioOligonucleotidesPathway AnalysisPathway interactionsPerformancePhasePhosphotransferasesPilot ProjectsProcessProteinsProteomicsPublishingRateResearchResearch InfrastructureResourcesRisk AssessmentSample SizeSamplingSequence AnalysisSignal TransductionSomatic MutationStatistical ModelsStructureSystemTissue SampleTissuesTranscriptVariantWorkanticancer researchbasecancer Biomedical Informatics Gridcancer genomecase controldata managementepidemiology studygenome sequencinggenome wide association studyheuristicsimprovedinsertion/deletion mutationinterestmalignant breast neoplasmnext generationnovelprobandthree dimensional structuretooltumor

项目摘要

项目成果

Kenneth H Buetow的其他基金

相似基金

相关文献

中文摘要
翻译
对肿瘤遗传变化的系统研究有望大大提高对癌症病因的认识。为了应对这些研究带来的分析挑战,我们开发了癌症基因组工作台(http://cgwb.nci.nih.gov),这是第一个将临床肿瘤突变谱与参考人类基因组相结合的计算平台。提出了一种新颖的启发式算法IndelDetector,用于自动识别插入/删除(indel)多态性和indel体细胞突变,具有较高的灵敏度和准确性。它被整合到一个自动管道中,该管道检测基因改变并注释它们对蛋白质编码和3D结构的影响。该系统促进识别基因改变的能力在三个具有公开可访问数据的项目中得到说明。在配对肿瘤-正常肺癌重测序数据中观察到一种新的缺失-插入-组合,表明肿瘤DNA复制中的突变导致EGFR激酶结构域的复杂遗传变化。IndelDetector在152个基因再分析上的表现表明,它将目前最先进的人工插入/删除数据分析的灵敏度提高了40%,同时保持了非常低的假阳性率。IndelDetector已被整合到一个自动管道中,用于检测基因改变并注释其对蛋白质编码和3D结构的影响。目前,我们已经在TCGA项目的先导项目TSP (Tumor Sequencing Project)项目中分析了首批248个候选基因。结果被纳入癌症基因组工作台,这将是癌症研究界的重要资源,因为我们最近使CGWB符合cabig标准。CGWB于今年发表在《基因组研究》杂志上,突变检测工具已被三个基因组测序中心用于分析产生TSP和TCGA序列的序列。目前,我们正在进行一项综合研究,将突变数据与TSP/TCGA项目中分析的肿瘤的表达谱和LOH结合起来。一种新的方法正在开发中,以提高LOH分析的灵敏度和产生等位基因特异性LOH信号。我们与NCICB合作,对测序、SNP芯片和基因表达数据的分析也为TSP/TCGA项目提供了重要的科学质量保证。与Jeff Struewing博士合作,我们分析了家族性卵巢癌先证中候选基因的种系突变,以及乳腺癌协会联盟(Breast cancer Association Consortium, BCAC)选择的候选区域的新snp。这项研究发表在今年的《自然》杂志上。我们为CGWB开发的计算和数据管理基础设施将得到加强,以支持最近在实验室启动的肝癌基因组广泛关联研究(GWAS)项目。在本项目中,我们将使用Affymetrix SNP Chip 6.0对两个韩国肝癌数据集进行分析:a) 400例病例和400例对照DNA样本;B) 20个配对的正常/肿瘤样本。第一个数据集确保足够的样本量用于流行病学研究,以确定遗传易感标记,可用于进一步分析,如LD bin->通路分析。第二组样本的基因表达数据已经完成(使用U131A/B芯片)。从SNP芯片获得的额外信息将使我们能够识别体细胞改变,如拷贝数变化和LOH。该数据集将为基因表达、遗传多态性和体细胞基因组改变的综合分析提供可能。该分析系统预计将支持存储和查询大约100万个snp,每个snp将有400例肝癌患者和400例对照的基因型呼叫。此外,将在同一平台上分析20对肿瘤/正常肝组织。基因型的估计总数为8亿。-该系统还存储了100万个snp的信息,包括基因组位置、氨基酸和mRNA的功能变化。我们还将400例癌症患者的临床特征以及对照组的性别和年龄存储到数据系统中。-该系统将包含40个组织样本的基因表达数据,这些样本由47,000个转录本组成,由54,000个探针组组成,由1,300,000个寡核苷酸特征组成。-我们的实验室将执行分析,包括查询和存储结果的迭代过程,并根据分析结果改进查询。这些分析是下一代个体化分子医学范式的发现阶段的特征。分析包括:o病例与对照的比值比,以确定疾病相关snp。我们将只包括高召唤率的snp(例如>=85%的样本有基因型召唤;最小等位基因频率超过10%;基因型质量超过一定阈值)。这个查询可能会在所有10亿个基因型行上执行。o确定高关联snp背后的基因;获得这些基因的基因型,构建单倍型和LD bin,进行更结构化的分析,如单倍型覆层。o识别跨snp的等位基因相互作用。这将需要使用多个snp作为风险评估的单个单位来评估遗传风险。o对于肿瘤/正常配对的肝脏组织,LPG将识别遗传异常,包括丢失杂合性和拷贝数变异。这些数据将与我们生成的表达数据进行比较,以评估基因改变与表达变化之间的相关性。o基于生物途径和网络的分析,通过这些网络的结构来询问上述关系。我们开发了同步Ciphergen MassSpec和LC-MS配置文件的工具,以减少产生假阳性信号的实验变化。这项工作有望改善蛋白质组学数据的分析。我们还在开发一种新的生物标志物发现算法和网络构建算法,该算法将遗传信息与表达谱相结合。三种互补的方法被用来创建路径模型:1)统计建模,2)逻辑建模,和3)计算建模。被称为通径分析的统计方法正被用于基因表达数据的建模。这些努力将扩展到包括从癌症(和正常组织)数据集衍生的癌症研究感兴趣的途径模型的集合。该实验室还与NCICB和CGAP合作开发通路数据的逻辑模型。这项工作将利用基于KEGG和BIOCARTA通路数据的人类和小鼠生物分子相互作用数据库。在实验室中探索的最后一个策略是计算建模。通路中的每个元素都带有一组输入和输出连接,这些连接将基因或复合体与系统中的其他节点连接起来。将节点的状态设置为“开”或“关”将触发通过节点的相关连接在整个系统中传播更改的影响。目前正在使用表达数据评估这种方法的实用性。认识到没有单一的最佳方法来创建一个mod[摘要被截断为7800个字符]
英文摘要
Systematic investigations of genetic changes in tumors are expected to lead to greatly improved understanding of cancer etiology. To meet the analytical challenges presented by such studies, we developed the Cancer Genome WorkBench (http://cgwb.nci.nih.gov), the first computational platform to integrate clinical tumor mutation profiles with the reference human genome. A novel heuristic algorithm, IndelDetector, was developed to automatically identify insertion/deletion (indel) polymorphisms as well as indel somatic mutations with high sensitivity and accuracy. It was incorporated into an automated pipeline that detects genetic alterations and annotates their effects on protein coding and 3D structure. The ability of the system to facilitate identifying genetic alterations is illustrated in three projects with publicly accessible data. Mutagenesis in tumor DNA replication leading to complex genetic changes in the EGFR kinase domain is suggested by a novel deletion-insertion-combination observed in paired tumor-normal lung cancer resequencing data. The performance of IndelDetector on the re-analysis of 152 genes indicates it improves sensitivity of manual data analysis of insertion/deletion, the current state-of-art, by 40% while maintaining a very low false positive rate. IndelDetector has been incorporated into an automated pipeline that detects genetic alterations and annotates their effects on protein coding and 3D structure. Currently, we have analyzed the first 248 candidate genes in the TSP (Tumor Sequencing Project) project, a pilot study of the TCGA project. The results are incorporated into the Cancer Genome Workbench, which will be an important resource for cancer research community as we have recently made CGWB caBIG-compliant. CGWB is published in Genome Research this year and the mutation detection tools have been used by the three genome sequencing centers to analyze sequences generated the TSP and TCGA sequences. Currently, we are working on a comprehensive study to integrate mutation data with expression profile and LOH of the tumors analyzed in the TSP/TCGA project. A new method is under development to improve the sensitivity of LOH analysis and generate allele-specific LOH signals. In collaboration with NCICB, our analysis on sequencing, SNP Chip and gene expression data also provide critical scientific QA for the TSP/TCGA project. In collaboration with Dr. Jeff Struewing, we have analyzed germline mutations in candidate genes in familial ovarian cancer probands as well as novel SNPs in candidate regions selected by the Breast Cancer Association Consortium (BCAC). This work is published in Nature this year. The computational and data management infrastructure we developed for CGWB will be enhanced to support the liver cancer genome wide association study (GWAS) project that has been launched recently in the laboratory. In this project we will use Affymetrix SNP Chip 6.0 to analyze two Korean liver cancer data sets: a) 400 case and 400 control DNA samples; b) 20 paired normal/tumor samples. The first data set ensures sufficient sample size for an epidemiology study to identify genetic susceptible markers which can be used for further analysis such as the LD bin->pathway analysis. The gene expression data for the second set of samples have already been completed (using U131A/B chips). The additional information obtained from the SNP chip will allow us to identify somatic alterations such as copy-number changes and LOH. This data set will make it possible for an integrated analysis of expression, genetic polymorphism and somatic genome alteration study. The analytical system is expected to support storing and querying approximately 1 million SNPs, each will have genotype calls of 400 cases and 400 controls of liver cancer patients. In addition, there will be 20 pairs of tumor/normal liver tissues analyzed on the same platform. The estimated total number of genotypes is 800 million. - The system also stores information for 1 million SNPs including genomic location, functional changes in amino acid and mRNA. We have also stored the clinical features of the 400 cases in the cancer patients into the data system and the gender and age of the control. - The system will contain gene-expression data for the 40 tissue samples composed of 47,000 transcripts measured by 54,000 probe sets composed of 1,300,000 oligonucleotide features. - Our laboratory will be performing analyses that involves an iterative process of querying and storing the results and refining the query based on the analytical results. These analyses are characteristic of the discovery phase of the next generation, individualized molecular medicine paradigm. The analyses include: o Odds-ratio of case and control to identify disease association SNPs. We will include only SNPs with a high call rate (for example >=85% of the samples have genotype calls; minimum allele frequency exceeds 10%; genotype quality exceeds certain threshold) in this query. This queries are likely to be performed across all 1 billion genotype rows. o Identify genes underlying the high-association SNPs; obtain genotypes in these genes to construct haplotypes and LD bin for more structured analysis like haplotype clad. o Identify allelic-interaction across SNPs. This would require evaluation of genetic risk using multiple SNPs as a single-unit for risk assessment. o For tumor/normal paired liver tissues, LPG will identify genetic abnormalities including loss-heterozygosity and copy-number variation. This data will be compared against the expression data that we have generated to evaluate the correlation between genetic alteration and expression change. o Analysis based on biologic pathways and networks where the relationship of the above is interrogated through the structure of these networks. We have developed tools for synchronizing Ciphergen MassSpec as well as LC-MS profile to reduce experimental variations that give false positive signal. This work is expected to improve the analysis of proteomics data. We are also developing a new algorithm for biomarker discovery and network construction algorithm that integrates the genetic information with the expression profile. Three complementary approaches are being utilized to create pathway models: 1) statistical modeling, 2) logical modeling, and 3) computational modeling. The statistical methodology known as path analysis is being used to model gene expression data. These efforts will be extended to include a collection of pathway models of interest to cancer research derived from cancer (and normal tissue) data sets. The laboratory is also collaborating with the NCICB and CGAP to develop Logical Models of pathway data. This effort will utilize databases of biomolecular interactions in human and mouse based on KEGG and BIOCARTA pathway data. The last strategy being explored within the laboratory is computational modeling. Each element in the pathway is annotated with a set of incoming and outgoing connections, which link the gene or complex to other nodes in the system. Setting the state of a node to "on" or "off" triggers the propagation of the effects of the change throughout the system via the node's dependent connections. The utility of this approach is currently being assessed using expression data. Recognizing that there is no single best way to create a mod [summary truncated at 7800 characters]
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Molecular Genetic Epidemiology of Primary Hepatocellular
Molecular Genetic Epidemiology of leading U.S. Cancers
Molecular Genetic Epidemiology of leading U.S. Cancers
Molecular Genetic Epidemiology of leading U.S. Cancers
海外基金