Turning Data into Whole Cell Ontology Models for Functional Analysis
Turning Data into Whole Cell Ontology Models for Functional Analysis
批准号:
8951600
负责人:
Michael Harris Kramer
金额:
$3.93万
依托单位国家:
美国
项目类别:
财政年份:
2014
资助国家:
美国
项目状态:
已结题
起止时间:
2014-09-01 至 2017-08-31
关键词:
AlgorithmsAntineoplastic AgentsBioinformaticsBiologicalBiological ProcessBiologyCell DeathCell modelCell physiologyCellsCluster AnalysisCodeCollectionCoupledDNA RepairDataData AnalysesData QualityData SetDiseaseDrug TargetingFutureGene ClusterGene ProteinsGeneric DrugsGenesGenomeGoalsHandHumanIndividualInterventionKnock-outLearningLightingMachine LearningMalignant NeoplasmsManualsMethodsModelingMolecularMutateMutationOntologyOrganismPharmaceutical PreparationsPhosphotransferasesProcessed GenesProteinsResearchResearch PersonnelRibosomesSaccharomyces cerevisiaeSubgroupSystemTissuesUpdateWorkYeastsbasebiological information processingcancer cellcell typecomputerized toolsdesignexperimental analysisfunctional groupgene functiongenome-wideimprovedkillingsnovelpublic health relevanceresearch studysynthetic biologytool
中文摘要
描述(由申请人提供):生物信息学的圣杯是创建具有增强人类理解和促进发现能力的全细胞模型。为此,一个成功的和广泛使用的努力是基因本体论(GO),一个大规模的项目,以手动注释基因到描述分子功能,生物过程和细胞组分的术语,并提供术语之间的关系,例如,捕获“小核糖体亚基”和“大核糖体亚基”走到一起,使“核糖体”。GO被广泛用于理解一个基因或一组基因的功能。不幸的是,GO受到手工创建和更新所需工作的限制。它只存在于被充分研究的生物体中,即使如此,每个生物体也只有一种属型,总体基因组覆盖率有限,偏向于被充分研究的基因和功能。使用GO不可能了解未表征的基因或发现新功能,并且无法快速组装新生物体的本体模型,更不用说特定的细胞类型或疾病状态了。 这项拟议中的研究将改变这种状况。已经有研究表明,酿酒酵母中的基因和蛋白质相互作用的大型网络可以用于计算推断本体,其覆盖范围和能力相当于手动策划的GO细胞组分本体。尽管如此,这第一次尝试是有限的,在使用的实验数据的类型和它的能力,以推断更普遍有用的生物过程本体。在这里,机器学习方法将被应用于集成多种类型的实验数据到本体模型构建和分析每个实验提供的生物信息的类型,揭示那些实验最有用的捕获生物过程信息。此外,高通量的实验数据本体论的范例,这里探讨将被用来开发一个计算工具,突出目前的高通量实验数据分析方法无法访问的新类型的假设。 初步研究表明,GO可用于预测合成的致死基因对,即单独非必需的基因,但当一起敲除时会导致细胞死亡。鉴于癌症中的高突变率,这些对提供了潜在的癌症药物靶点,因为药物可以靶向突变的癌细胞而不是其他细胞中现在必需的基因产物,从而仅杀死癌细胞。由于数据驱动的本体不受偏见和覆盖范围问题的阻碍,并且专门设计用于仅捕获功能关系,因此该提案将探索数据驱动的本体将比GO更适合帮助预测合成致命对的想法。为此,将开发算法来构建酵母DNA修复的数据驱动本体,并使用该本体来预测合成的致死基因对。 总的来说,该提案将开发计算和实验路线图,以构建基因功能的全细胞模型-本体论-并使用该模型发现有用的生物学-合成致死对。
英文摘要
DESCRIPTION (provided by applicant): A holy grail of bioinformatics is the creation of whole-cell models with the ability to enhance human understanding and facilitate discovery. To this end, a successful and widely-used effort is the Gene Ontology (GO), a massive project to manually annotate genes into terms describing molecular functions, biological processes and cellular components and provide relationships between terms, e.g. capturing that "small ribosomal subunit" and "large ribosomal subunit" come together to make "ribosome". GO is widely used to understand the function of a gene or group of genes. Unfortunately, GO is limited by the effort required to create and update it by hand. It exists only for well-studied organisms and even then in only one, generic form per organism with limited overall genome coverage and a bias towards well-studied genes and functions. It is not possible to learn about an uncharacterized gene or discover a new function using GO, and one cannot quickly assemble an ontology model for a new organism, let alone a specific cell-type or disease-state. This proposed research will change this state of affairs. Already, work has shown that large networks of gene and protein interactions in Saccharomyces cerevisiae can be used to computationally infer an ontology whose coverage and power are equivalent to those of the manually-curated GO Cellular Component ontology. Still, this first attempt was limited in the types of experimental data used and its ability to infer the more generally useful Biological Process ontology. Here machine learning approaches will be applied to integrate many types of experimental data into ontology model construction and analyze the type of biological information provided by each experiment, revealing those experiments most informative for capturing Biological Process information. Furthermore, the high-throughput experimental data to ontology paradigm explored here will be used to develop a computational tool to highlight novel types of hypotheses that are inaccessible by current high-throughput experimental data analysis methods. Preliminary work has shown GO to be useful for prediction of synthetic lethal pairs of genes, i.e. genes that are individually non-essential but when knocked out together cause cell death. Given the high mutation rate in cancer, these pairs provide potential cancer drug targets, as a drug may target a gene product which is now essential in the mutated cancer cells but not other cells, thereby killing only cancer cells. Because data-driven ontologies are not as hindered by issues with bias and coverage and are specifically designed to capture only functional relationships, this proposal will explore the idea that data-driven ontologies will be better suited to help predict synthetic lethal pairs than GO. To this end, algorithms will be developed to construct a data-driven ontology of yeast DNA repair and use this ontology to predict synthetic lethal pairs of genes. Overall, this proposal will develop the computational and experimental roadmap to construct a whole-cell model of gene function - an ontology - and use the model to discover useful biology - synthetic lethal pairs.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Turning Data into Whole Cell Ontology Models for Functional Analysis
-
批准号:8644512
-
项目类别:
-
资助金额:$3.51万
-
财政年份:2014
-
负责人:Michael Harris Kramer
-
依托单位:
Turning Data into Whole Cell Ontology Models for Functional Analysis
-
批准号:9145523
-
项目类别:
-
资助金额:$4.86万
-
财政年份:2014
-
负责人:Michael Harris Kramer
-
依托单位:
海外基金