课题基金 / 基金详情

Bioinformatic Tools in Cancer Research

Bioinformatic Tools in Cancer Research
癌症研究中的生物信息工具
批准号:
7055492
负责人:
Kenneth H Buetow
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Kenneth H Buetow的其他基金

相似基金

相关文献

中文摘要
翻译
了解人类和小鼠的单倍型结构对于疾病基因定位、数量性状基因座(QTL)定位以及小鼠模型在人类癌症中的应用具有重要的意义。我们已经开发了一个软件分析包HapScope,它包括一个全面的分析管道(包括一个新的SNP标记算法)和一个复杂的可视化工具,用于分析功能注释的单倍型。LPG PI使用HapScope分析了来自乳腺癌-卵巢癌家系的两个BRCA1相互作用基因的单倍型结构。美国和国外的20多家研究机构已经下载了HapScope包来分析他们的临床基因数据。使用HapScope工具,我们观察到人类基因组中高度分化的单倍型模式(称为阴阳单倍型)。对人类62个随机基因组座位和85个基因编码区常见单倍型的全基因组分析表明,阴阳单倍型覆盖基因组的比例为75%-85%。人类基因组中丰富的阴阳单倍型表明,易感性受环境的影响似乎比基因更大。在小鼠模型中,缺乏遗传多样性一直被认为是实验室近交系小鼠的一大缺点。我们对小鼠16号染色体高分辨率、多品系单倍型结构的分析揭示了一个复杂的单倍型结构,表明实验室小鼠品系的受控复杂性为研究人类复杂疾病提供了巨大的实用价值。我们开发的另一个软件工具是AutoSNP,用于扩展我们的遗传变异分析。AutoSNP允许我们通过基于荧光的重新测序来检测SNP,几乎不需要人工审查,并且具有非常低的假阳性和假阴性率。此工具在Unix/Linux平台上运行,可通过ftp(ftp1.nci.nih.gov/AutoSNP)获得。 该实验室还将重点放在开发功能分析工具上。这些包括用于评估mRNA表达数据的分析方法、计算过程和可视化工具,以及用于识别候选基因的工具。人们认识到,与聚类或分类分析相比,路径分析对观察到的微阵列数据的要求要大得多。现有工具不能将高质量的探测与那些具有过量表达式值或空表达式值的探测区分开来。据推测,这可能是导致检测同一基因的重复探针集在表达测量上缺乏一致性的原因。为了提高表达数据的质量,我们使用探针序列上下文分析了Affymetrix芯片上的非特异性和非功能探针对。我们发现18%的探测器可能是有问题的,并实现了过滤这种噪声的方法。在单个实验中缺乏内部一致性对解释表达数据具有严重的不利影响,希望新的分析工具将在对途径关系进行建模和分析之前提高表达测量的质量。这一分析已经扩展到包括最新版本的Affymetrix人类基因组表达阵列U133,以及Affymetrix的小鼠表达阵列。为了识别候选基因,我们开发了一个动态而强大的搜索引擎--基因功能相似性搜索工具(GFSST),它允许我们在疾病关联研究和药物靶标发现中选择候选基因。对于基因本体论(GO)术语中定义的给定基因或给定的基因功能集,该工具可以识别用户定义的相似性阈值内的基因。为了方便这种搜索,我们定义了一个统计模型来衡量基因的功能相似性,基于GO有向无环图(DAG)。在UniProt(通用蛋白质资源)上对人类和小鼠基因组进行GFSST的实施可在http://gfsst.nci.nih.gov.获得。 三种互补的方法被用来创建路径模型:1)统计建模、2)逻辑建模和3)计算建模。被称为通径分析的统计方法正被用来对基因表达数据进行建模。这些努力将扩展到包括从癌症(和正常组织)数据集衍生的癌症研究感兴趣的途径模型的集合。该实验室还与NCICB和CGAP合作开发通路数据的逻辑模型。这项工作将利用基于KEGG和BioCarta途径数据的人类和小鼠生物分子相互作用数据库。实验室正在探索的最后一种策略是计算建模。途径中的每个元件都有一组传入和传出的连接,这些连接将基因或复合体连接到系统中的其他节点。将节点的状态设置为“开”或“关”会触发通过该节点的依赖连接在整个系统中传播更改的影响。目前正在使用表情数据评估这种方法的实用性。认识到没有单一的最佳方法来创建生物途径等复杂过程的模型,正在使用和评估这三种互补的方法。将路径实例化为代码是开发更复杂的计算模型的第一步。
英文摘要
Knowledge of haplotype structure in human and mouse has important implications for strategies of disease gene mapping, quantitative trait loci (QTL) mapping, and the utility of mouse model for human cancer. We have developed a software analysis package, HapScope, which includes a comprehensive analysis pipeline (including a novel SNP tagging algorithm) and a sophisticated visualization tool for analyzing functionally annotated haplotypes. HapScope was used by LPG PI to analyze haplotype structure of two BRCA1-interacting genes from breast-ovarian cancer families. Over 20 research institutes in the US and abroad have downloaded the HapScope package to analyze their clinical genotype data. Using the HapScope tool, we observed highly divergent haplotype patterns (referred to as yin yang haplotypes) in the human genome. Genome-wide analysis of common haplotypes in 62 random genomic loci and 85 gene-coding regions in humans shows the proportion of the genome spanned by yin yang haplotypes is 75%-85%. The abundance of yin yang haplotypes in the human genome suggests susceptibility will appear to be more greatly influenced by environment than genes. In mouse models, lack of genetic diversity has been considered as a major drawback of laboratory-inbred mouse. Our analysis of a high-resolution, multiple-strain haplotype structure of mouse chromosome 16 reveals a complex haplotype structure, indicating that the controlled complexity of laboratory mouse strains provides great utility for studying human complex diseases. Another software tool we have developed to extend our analysis of genetic variation is AutoSNP. AutoSNP allows us to detect SNP's by fluorescence-based resequencing, with minimal requirement for manual review and has a very low rate of false positives and false negatives. This tool runs on Unix/Linux platforms and is available by ftp (ftp1.nci.nih.gov/AutoSNP). The laboratory also has focused efforts on developing tools for functional analysis. These include analytical methods, computational processes and visualization tools to evaluate mRNA expression data, as well as tools to identify candidate genes. It is recognized that pathway analysis makes significantly greater demands on observed microarray data than cluster or classification analysis. Existing tools do not differentiate probes of good quality from those that have either excess expression or null expression values. It is speculated that this may contribute to the lack of consistency in expression measurements for duplicate probe sets that assay the same gene. To improve the quality of expression data, we analyzed non-specific and non-functional probe pairs on the Affymetrix chips using the probe sequence context. We discovered that 18% of probes might be problematic and implemented methods to filter this noise. The lack of internal consistency in a single experiment has a severe adverse impact on interpreting expression data and it is hoped that new analytic tools will improve the quality of the expression measurement prior to the modeling and analysis of pathway relationships. This analysis has been extended to include the latest version of the Affymetrix human genome expression array, U133, as well as mouse expression arrays from Affymetrix. To identify candidate genes, we have developed a dynamic and robust search engine, the Gene Functional Similarity Search Tool (GFSST), which allows us to select candidate genes in disease association studies and drug target discoveries. For a given gene or a given set of gene functions defined in Gene Ontology (GO) terms, this tool can identify genes within a user defined similarity threshold. To facilitate this search, we have defined a statistical model to measure functional similarity of genes based on the GO directed acyclic graph (DAG). An implementation of GFSST on UniProt (Universal Protein Resource) for the human and mouse genomes is available at http://gfsst.nci.nih.gov. Three complementary approaches are being utilized to create pathway models: 1) statistical modeling, 2) logical modeling, and 3) computational modeling. The statistical methodology known as path analysis is being used to model gene expression data. These efforts will be extended to include a collection of pathway models of interest to cancer research derived from cancer (and normal tissue) data sets. The laboratory is also collaborating with the NCICB and CGAP to develop Logical Models of pathway data. This effort will utilize databases of biomolecular interactions in human and mouse based on KEGG and BIOCARTA pathway data. The last strategy being explored within the laboratory is computational modeling. Each element in the pathway is annotated with a set of incoming and outgoing connections, which link the gene or complex to other nodes in the system. Setting the state of a node to "on" or "off" triggers the propagation of the effects of the change throughout the system via the node's dependent connections. The utility of this approach is currently being assessed using expression data. Recognizing that there is no single best way to create a model of such complex processes as biologic pathways, these three complementary approaches are being employed and evaluated. The instantiation of pathways as code represents the first step in development of more complex computational models.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Molecular Genetic Epidemiology of leading U.S. Cancers
Molecular Genetic Epidemiology of Primary Hepatocellular
Bioinformatic Tools in Cancer Research
  • 批准号:
    7292177
  • 项目类别:
  • 资助金额:
    $0.0万
  • 财政年份:
    --
  • 负责人:
    Kenneth H Buetow
  • 依托单位:
Molecular Genetic Epidemiology of leading U.S. Cancers
海外基金