课题基金 / 基金详情

Improving detection in high-throughput sequencing data with gene/locus-specific models

Improving detection in high-throughput sequencing data with gene/locus-specific models
使用基因/位点特异性模型改进高通量测序数据的检测
批准号:
RGPIN-2019-06604
负责人:
Perkins, Theodore
金额:
$2.99万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Perkins, Theodore的其他基金

相似基金

相关文献

中文摘要
翻译
生物信息学领域借鉴了计算机科学和数学的其他领域--如统计学、机器学习、概率建模和优化--来开发合理的、通用的算法,用于分析高通量的遗传和分子数据。然而,这些算法几乎无一例外地对考虑中的每个“实体”一视同仁。例如,为了确定哪些基因在两种情况下差异表达,相同的统计模型分别应用于每个基因。当我们想要识别与特定蛋白质结合的基因组DNA的区域时,相同的统计模型被应用于每个基因组位置。当然,高通量分析提供的每个基因或每个座位的观察数据是不同的。但所应用的测试是相同的,原因很简单:传统上,生物信息学处理的是每个实体的观察数量(例如两种情况、少数时间点或几十个患者)远远超过实体数量(例如数万个基因或数百万个基因组位点)的情况。如果不是简单的统计模型,就有可能过度拟合现有的稀疏数据。然而,个人基因组、表观基因组以及特定于细胞、组织和疾病的表达谱的大量公共数据库的积累,意味着我们现在拥有来自数万或数十万“病症”的高通量数据。此外,对这些数据的统计分析揭示了一个令人震惊的事实:所有的基因和所有的基因组位置并不是相同的。例如,某些基因的表达天生就比其他基因更具变异性。此外,我们对一些基因的测量比对其他基因的测量更嘈杂和/或更有系统地偏颇。同样,基因组基因座在不同的检测中有不同的信噪比,不同的测量偏倚来源和数量也不同。这项提议背后的中心思想是利用已经收集的大量数据来建立和测试更复杂的、基于机器学习的基因组中每个基因或基因座的模型。此外,我们可以使用这些模型不仅是为了分析相同的数据,而且还可以创建工具来分析新的数据集,无论它们的大小如何。通过对每个基因或基因座的特定偏差和变异性进行建模,我们可以更准确地衡量新测量的新颖性,并更成功地识别基因和基因组行为的真正重大变化。
英文摘要
The field of bioinformatics has borrowed from other fields of computer science and mathematics--such as statistics, machine learning, probabilistic modelling, and optimization--to develop sound, general algorithms for analyzing high-throughput genetic and molecular data. However, almost without exception, those algorithms treat every "entity" under consideration the same. For example, to identify which genes are differentially expressed between two conditions, the same statistical model is applied individually to every gene. When we want to identify regions of the genomic DNA bound by a certain protein, the same statistical model is applied to every genomic locus. Of course the observed data for each gene or each locus, provided by the high-throughput assay, is different. But the test applied is the same, and for a simple reason: traditionally, bioinformatics has dealt with situations where the number of observations per entity (e.g. two conditions, a handful of time points, or a few tens of patients) is vastly outnumbered by the number of entities (e.g. tens of thousands of genes or millions of genomic loci). Anything but simple statistical models would be in danger of overfitting the sparse data available. However, the accumulation of massive public databases of personal genomes, epigenomes, and cell-, tissue-, and disease-specific expression profiles, means that we now have at our disposal high-throughput data from tens or hundreds of thousands of "conditions". Moreover, statistical analyses of such data reveals a startling fact: all genes and all genomic loci are not alike. For example, the expression of some genes is inherently more variable than others. Furthermore, our measurements of some genes are noisier and/or more systematically biased than for other genes. Similarly for genomic loci, where we have varying signal-to-noise ratios in different assays, and different sources and amounts of measurement bias. The central idea behind this proposal is to use that mass of already-collected data to build and test more sophisticated, machine learning-based models of every single gene or locus in the genome. Further, we can use those models not just for the sake of analyzing that same data, but rather for creating tools to analyze new datasets, whatever their size. By modelling the particular biases and variability of each gene or locus, we can get a more accurate measure of the novelty of new measurements, and more successfully identify truly significant alterations in gene and genome behaviour.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Improving detection in high-throughput sequencing data with gene/locus-specific models
  • 批准号:
    RGPIN-2019-06604
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2021
  • 负责人:
    Perkins, Theodore
  • 依托单位:
Improving detection in high-throughput sequencing data with gene/locus-specific models
  • 批准号:
    RGPIN-2019-06604
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2020
  • 负责人:
    Perkins, Theodore
  • 依托单位:
Improving detection in high-throughput sequencing data with gene/locus-specific models
  • 批准号:
    RGPIN-2019-06604
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $2.99万
  • 财政年份:
    2019
  • 负责人:
    Perkins, Theodore
  • 依托单位:
Inference and Scaling in Stochastic Dynamical Systems
  • 批准号:
    RGPIN-2014-05716
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $1.89万
  • 财政年份:
    2018
  • 负责人:
    Perkins, Theodore
  • 依托单位:
国内基金
海外基金
Graphon mean field games with partial observation and application to failure detection in distributed systems
  • 批准号:
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2025
  • 负责人:
    MATHIEULOUROCHLAURIERE
  • 依托单位:
基于深穿透拉曼光谱的安全光照剂量的深层病灶无创检测与深度预测
  • 批准号:
    82372016
  • 项目类别:
    面上项目
  • 资助金额:
    48.00万元
  • 批准年份:
    2023
  • 负责人:
    林俐
  • 依托单位:
膀胱癌高表达基因UPK3A的筛选、鉴定和相关研究
  • 批准号:
    81101922
  • 项目类别:
    青年科学基金项目
  • 资助金额:
    23.0万元
  • 批准年份:
    2011
  • 负责人:
    来永庆
  • 依托单位:
图像分类方法研究及其在色情监测中的应用
  • 批准号:
    61172103
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2011
  • 负责人:
    王春恒
  • 依托单位: