课题基金 / 基金详情

项目摘要

项目成果

William Evan Johnson的其他基金

相似基金

相关文献

中文摘要
翻译
组合来自多个研究的基因组数据集有利于增加研究中的统计能力 在后勤考虑限制样本大小或要求按顺序生成数据的情况下。然而, 通常可以在生成的多批数据中观察到显著的技术异构性 来自不同的批次、实验或分析平台。这些所谓的批处理效果通常会混淆事实 数据中的生物关系,降低了组合多批数据的功率效益,并且可以 甚至导致虚假的结果。已经提出了许多方法来过滤技术异构性和批次 基因组数据的影响。然而,仍然有很大的差距需要解决更多 适当地从基因组数据集中过滤技术异质性。例如,现有方法假定 钟形、对称的数据,不适合现代的测序计数数据。此外,还有 例如,目前还没有用于批量效应基因组数据的方法,这些数据可以在精细化的水平上测量特征 表观遗传测序数据,其中附近的特征可能密切相关。当前批次 调整方法依赖于手头的数据批次,这意味着如果额外的数据批次 添加到分析中,则需要重新应用批量调整,从而产生不同的调整 基因组数据值。此外,批量校正通常将相关性引入到调整后的数据中,这 需要在下游分析中考虑;大多数研究人员以前进行过批量修正 其他分析步骤没有意识到这种负面影响,因此经常不正确地应用 下游分析工具。最后,并不总是清楚应该在哪些批量调整方法中应用 因此,需要在适当的批量纠正策略之前进行彻底的评估 是被设计出来的。这些差距突出表明,需要新的统计方法和交互式可视化软件来 促进研究人员在这一领域的需求。我们建议开发算法和软件来解决 这些具体的研究差距面临着研究人员结合多个实验批次的数据。
英文摘要
Combining genomic data sets from multiple studies is advantageous to increase statistical power in studies where logistical considerations restrict sample size or require the sequential generation of data. However, significant technical heterogeneity is commonly observed across multiple batches of data that are generated from different batches, experiments, or profiling platforms. These so called batch effects often confound true biological relationships in the data, reducing the power benefits of combining multiple batches of data, and may even lead to spurious results. Many methods have been proposed to filter technical heterogeneity and batch effects from genomic data. However, there are still significant gaps that need to be addressed to more appropriately filter technical heterogeneity from genomic datasets. For example, existing approaches assume bell-shaped, symmetric data, which are not appropriate for modern sequencing count data. Furthermore, there are no current approaches for batch effects genomic data that measure features at a refined level, for example epigenetic sequencing data, where nearby features are likely to be closely correlated. Current batch adjustment methods are dependent of the data batches on hand, meaning that if additional batches of data were added to the analysis, the batch adjustments would need to be reapplied, resulting in different adjusted genomic data values. In addition, batch correction usually introduces correlation into the adjusted data, which needs to be accounted for in downstream analyses; most researchers performing batch correction before additional analysis steps are unaware of this negative impact, and as a result often incorrectly apply downstream analysis tools. Finally, it is not always clear which batch adjustment methods should be applied in each particular case, so a thorough evaluation is required before an appropriate batch correction strategy can be devised. These gaps highlight the need for new statistical methods and interactive visualization software to facilitate the needs of researchers in this area. We propose to develop algorithms and software to address these specific research gaps facing researchers combining data from multiple experimental batches.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Microbiome-based biomarkers and models of lung cancer development and treatment
Systems Biology Core
  • 批准号:
    10493266
  • 项目类别:
  • 资助金额:
    $35.78万
  • 财政年份:
    2021
  • 负责人:
    William Evan Johnson
  • 依托单位:
Microbiome-based biomarkers and models of lung cancer development and treatment
  • 批准号:
    10366665
  • 项目类别:
  • 资助金额:
    $23.14万
  • 财政年份:
    2021
  • 负责人:
    William Evan Johnson
  • 依托单位:
Systems Biology Core
海外基金