课题基金 / 基金详情

Algorithmic approaches to systems biology, data integration, and evolution

Algorithmic approaches to systems biology, data integration, and evolution
系统生物学、数据集成和进化的算法方法
批准号:
10268080
负责人:
Teresa Przytycka
金额:
$138.52万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:

项目摘要

项目成果

Teresa Przytycka的其他基金

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
My group continued to develop and apply computational methods that utilize and integrate large data sets to study gene regulation and diseases. We also develop methods to analyze data produced My group continued to develop and apply computational methods that utilize and integrate large data sets with a focus on gene regulation and diseases. We also developed new methods to analyze data produced by new, high throughput, technologies and experimental techniques such as single cell gene expression and HT-SELEX data. In our studies we use variety of algorithmic techniques including Integer Linear Programming (ILP) among other optimization strategies as well as Machine Learning approaches, including Hidden Markov Models and Deep Learning. Within this general area, the main focus of my group is on developing new computational methods allowing to utilize large cancer-related datasets (e.g. TCGA and ICGC) to obtain insists into etiology of cancer. Together with our experimental collaborators we also utilize new experimental data to obtain novel insights into fundamental biological processes. Much of the effort of the group during this reporting period has been devoted to studying of mutational patterns in cancer genomes. Specifically, through their lifetime, individuals acquire somatic mutations which might eventually led to cancer. These mutations often display characteristic patterns known as mutational signatures. Understanding relation between these patterns and their causes can provide important insights to into tumorigenesis in general and environmental contributions to cancer in particular. The two fundamental question in this area are (i) what is the best way to characterize these mutation patterns and (ii) leveraging such patterns of somatic mutations for understanding of mutagenic processes shaping human genome. One of the most challenging obstacles to a full characterization of mutational patterns comes from the fact that these patterns are the end-effect of several interplaying factors including carcinogenic exposures and potential deficiencies of the DNA repair mechanism. Separating these factors in nontrivial and thus the current methods typically do not attempt such separation assuming linear combination model. Yet, to fully understand the nature of each signature, it is important to disambiguate the atomic components that contribute to the final signature. As the first step in this direction we recently introduced a new descriptor of mutational signatures, DNA Repair FootPrint (RePrint) (1). Our work demonstrated, for the first time, that it is possible to identify signatures that include common DNA repair deficiency independent on the other mutagenic processes that contribute to the composite signature. We validated the method with published mutational signatures from cell lines targeted with CRISPR-Cas9-based knockouts of DNA repair genes. The second line of research related to mutational signatures is the identification of mutagenic processes underlying mutational signatures. To investigate the genetic aberrations associated with mutational signatures, we took a network-based approach considering mutational signatures as cancer phenotypes. Specifically, our analysis aimed to answer the following two complementary questions: (i) what are functional pathways whose gene expression activities correlate with the strengths of mutational signatures, and (ii) are there pathways whose genetic alterations might have led to specific mutational signatures? To identify mutated pathways, we adopted a recently developed optimization method based on integer linear programming. Analyzing a breast cancer dataset, we identified pathways associated with mutational signatures on both expression and mutation levels. Our analysis captured important differences in the etiology of the APOBEC-related signatures and the two clock-like signatures. In particular, it revealed that clustered and dispersed APOBEC mutations may be caused by different mutagenic processes. In addition, our analysis elucidated differences between two age-related signatures-one of the signatures is correlated with the expression of cell cycle genes while the other has no such correlation but shows patterns consistent with the exposure to environmental/external processes. This work investigated, for the first time, a network-level association of mutational signatures and dysregulated pathways. The identified pathways and subnetworks provide novel insights into mutagenic processes that the cancer genomes might have undergone and important clues for developing personalized drug therapies (2). In addition, we collaborated with Roded Sharans group from TAU, to provide a first probabilistic model of mutational signatures that accounts for context dependency and strand coordination (3). Finally, we started a research leveraging the concept mutational to study the relationship of smoking and expression ACE2 and other proteins known to be involved in the entrance the Coronavirus 2s (SARS-CoV-2) into the host cell. We also continued our research on methods to construct gene regulatory networks (GRNs). These networks describe regulatory relationships between transcription factors (TFs) and their target genes. Following the development of NetREX (Network Reprogramming using EXpression) technique to for constructing context-specific GRN given context-specific expression data and a context-agnostic prior network (reported last year), we developed NetREX-CF. The important novelty of NetREX-CF is the ability to deal with missing data. Specifically, NetREX-CF reconstruction approach that brings together a modern machine learning strategy (Collaborative Filtering model) and a biologically justified model of gene expression (sparse Network Component Analysis based model). The Collaborative Filtering (CF) is able to overcome the incompleteness of the prior knowledge and make edge recommends for building the GRN. Complementing CF, we use the sparse Network Component Analysis (NCA) to validate the recommended edges. Finally, we combine these two approaches using a novel data integration method and show that the new approach outperforms the currently leading GRN reconstruction methods. Our preliminary results show that this method drastically outperform previous approaches. This work has been selected for oral presentation RECOMB 2020 and the manuscript in preparation. My group also continues to develop software for public use including AptaBlocks Online -- a web-based toolkit for the In silico design of RNA complexes (4) and JUDY a flexible bioinformatics pipeline for diverse types of bioinformatics analysis (5). nWe also provided computational expertise and analysis of the specialized sequencing data, mRNA display, that our collaborators used for comparison of the performance of Linear, Monocyclic, and Bicyclic Libraries (6).
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Combinatorial and graph theoretical approach to systems biology and mol. evo.
  • 批准号:
    8943247
  • 项目类别:
  • 资助金额:
    $143.5万
  • 财政年份:
    --
  • 负责人:
    Teresa Przytycka
  • 依托单位:
Combinatorial and graph theoretical approach to systems biology and mol. evo.
  • 批准号:
    8558125
  • 项目类别:
  • 资助金额:
    $171.37万
  • 财政年份:
    --
  • 负责人:
    Teresa Przytycka
  • 依托单位:
Algorithmic approaches to systems biology, data integration, and evolution
  • 批准号:
    10927048
  • 项目类别:
  • 资助金额:
    $141.1万
  • 财政年份:
    --
  • 负责人:
    Teresa Przytycka
  • 依托单位:
Combinatorial and graph theoretical approach to systems biology and mol. evo.
  • 批准号:
    7969252
  • 项目类别:
  • 资助金额:
    $90.3万
  • 财政年份:
    --
  • 负责人:
    Teresa Przytycka
  • 依托单位:
海外基金