课题基金 / 基金详情

Bioinformatic Tools in Cancer Research

Bioinformatic Tools in Cancer Research
癌症研究中的生物信息工具
批准号:
6952052
负责人:
Kenneth H Buetow
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Kenneth H Buetow的其他基金

相关文献

中文摘要
翻译
人类和小鼠单倍型结构的知识对于疾病基因定位、数量性状基因座(QTL)定位以及人类癌症小鼠模型的实用性具有重要意义。我们已经开发了一个软件分析包,HapScope,其中包括一个全面的分析管道(包括一个新的SNP标记算法)和一个复杂的可视化工具,用于分析功能注释的单倍型。LPG PI使用HapScope分析了来自乳腺-卵巢癌家族的两个BRCA 1相互作用基因的单倍型结构。美国和国外的20多个研究机构已经下载了HapScope软件包来分析他们的临床基因型数据。使用HapScope工具,我们在人类基因组中观察到高度不同的单倍型模式(称为阴阳单倍型)。对人类62个随机基因组位点和85个基因编码区的常见单倍型进行全基因组分析,发现阴阳单倍型所覆盖的基因组比例为75%~ 85%。人类基因组中丰富的阴阳单倍型表明,环境对易感性的影响似乎比基因更大。在小鼠模型中,缺乏遗传多样性被认为是实验室近交系小鼠的主要缺点。我们对小鼠16号染色体的高分辨率、多品系单倍型结构的分析表明,实验室近交系小鼠的遗传多样性与人类相似,其受控的复杂性为研究人类复杂疾病提供了巨大的实用性。 该实验室还专注于开发分析方法,计算过程和可视化工具,以评估mRNA表达数据。人们认识到,途径分析比聚类或分类分析对观察到的微阵列数据提出了更高的要求。现有的工具不能区分高质量的探针和具有过量表达或空表达值的探针。据推测,这可能导致测定相同基因的重复探针组的表达测量缺乏一致性。为了提高表达数据的质量,我们使用探针序列上下文分析了Affytron芯片上的非特异性和非功能性探针对。我们发现18%的探测器可能有问题,并实施了过滤这种噪声的方法。在单个实验中缺乏内部一致性对解释表达数据具有严重的不利影响,并且希望新的分析工具将在通路关系的建模和分析之前提高表达测量的质量。 三种互补的方法被用来创建途径模型:1)统计建模,2)逻辑建模,和3)计算建模。被称为路径分析的统计方法被用来模拟基因表达数据。这些努力将扩大到包括从癌症(和正常组织)数据集导出的癌症研究感兴趣的通路模型的集合。该实验室还与NCICB和CGAP合作开发路径数据的逻辑模型。这项工作将利用基于KEGG和BIOCARTA途径数据的人类和小鼠生物分子相互作用数据库。实验室正在探索的最后一个策略是计算建模。通路中的每个元素都被注释了一组传入和传出连接,这些连接将基因或复合体与系统中的其他节点联系起来。将节点的状态设置为“on”或“off”会触发更改的效果通过节点的相关连接在整个系统中传播。目前正在使用表达数据评估这种方法的实用性。认识到没有单一的最佳方法来创建一个模型,如生物途径,这三个互补的方法正在使用和评估。将路径实例化为代码代表了开发更复杂计算模型的第一步。
英文摘要
Knowledge of haplotype structure in human and mouse has important implications for strategies of disease gene mapping, quantitative trait loci (QTL) mapping, and the utility of mouse model for human cancer. We have developed a software analysis package, HapScope, which includes a comprehensive analysis pipeline (including a novel SNP tagging algorithm) and a sophisticated visualization tool for analyzing functionally annotated haplotypes. HapScope was used by LPG PI to analyze haplotype structure of two BRCA1-interacting genes from breast-ovarian cancer families. Over 20 research institutes in the US and abroad have downloaded the HapScope package to analyze their clinical genotype data. Using the HapScope tool, we observed highly divergent haplotype patterns (referred to as yin yang haplotypes) in the human genome. Genome-wide analysis of common haplotypes in 62 random genomic loci and 85 gene-coding regions in humans shows the proportion of the genome spanned by yin yang haplotypes is 75%-85%. The abundance of yin yang haplotypes in the human genome suggests susceptibility will appear to be more greatly influenced by environment than genes. In mouse models, lack of genetic diversity has been considered as a major drawback of laboratory-inbred mouse. Our analysis of a high-resolution, multiple-strain haplotype structure of mouse chromosome 16 reveals that the genetic diversity in laboratory-inbred mice is similar to human and its controlled complexity provides great utility for studying human complex diseases. The laboratory also has focused efforts on developing analytical methods, computational processes and visualization tools to evaluate mRNA expression data. It is recognized that pathway analysis makes significantly greater demands on observed microarray data than cluster or classification analysis. Existing tools do not differentiate probes of good quality from those that have either excess expression or null expression values. It is speculated that this may contribute to the lack of consistency in expression measurements for duplicate probe sets that assay the same gene. To improve the quality of expression data, we analyzed non-specific and non-functional probe pairs on the Affymetrix chips using the probe sequence context. We discovered that 18% of probes might be problematic and implemented methods to filter this noise. The lack of internal consistency in a single experiment has a severe adverse impact on interpreting expression data and it is hoped that new analytic tools will improve the quality of the expression measurement prior to the modeling and analysis of pathway relationships. Three complementary approaches are being utilized to create pathway models: 1) statistical modeling, 2) logical modeling, and 3) computational modeling. The statistical methodology known as path analysis is being used to model gene expression data. These efforts will be extended to include a collection of pathway models of interest to cancer research derived from cancer (and normal tissue) data sets. The laboratory is also collaborating with the NCICB and CGAP to develop Logical Models of pathway data. This effort will utilize databases of biomolecular interactions in human and mouse based on KEGG and BIOCARTA pathway data. The last strategy being explored within the laboratory is computational modeling. Each element in the pathway is annotated with a set of incoming and outgoing connections, which link the gene or complex to other nodes in the system. Setting the state of a node to "on" or "off" triggers the propagation of the effects of the change throughout the system via the node's dependent connections. The utility of this approach is currently being assessed using expression data. Recognizing that there is no single best way to create a model of such complex processes as biologic pathways, these three complementary approaches are being employed and evaluated. The instantiation of pathways as code represents the first step in development of more complex computational models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Molecular Genetic Epidemiology of Primary Hepatocellular
Molecular Genetic Epidemiology of leading U.S. Cancers
Molecular Genetic Epidemiology of leading U.S. Cancers
Molecular Genetic Epidemiology of leading U.S. Cancers