DMS/NIGMS 1: Statistical Methods for Design and Analysis of Clinical-scale Single Cell Studies
DMS/NIGMS 1: Statistical Methods for Design and Analysis of Clinical-scale Single Cell Studies
批准号:
2245575
负责人:
Nancy Zhang
金额:
$60.0万
依托单位国家:
美国
项目类别:
Standard Grant
财政年份:
2023
资助国家:
美国
项目状态:
未结题
起止时间:
2023-07-01 至 2026-06-30
中文摘要
在过去的二十年里,生物学和医学方面的科学发现越来越依赖于对大型基因组数据集的统计分析。因此,开发严格的统计方法,对这些数据集进行可重复和可扩展的分析,并对能够跨越计算科学和生物医学科学的下一代科学家进行跨学科培训,对我们的科学进步至关重要。该项目重点关注单细胞实验数据分析中的挑战,最近的技术发展使得对单个细胞进行基因组规模的测量成为可能,并且可以一次对数百万个细胞进行高通量测量。这种单细胞实验现在是生物医学研究的基础,在我们寻求更全面地了解细胞生物学以及寻求治疗各种疾病(从COVID-19等传染病到癌症,再到神经退行性疾病等与衰老有关的疾病)的过程中发挥着重要作用。尽管这些新技术带来了希望,但目前的单细胞数据分析方法并不是为临床规模的疾病研究而设计的,也没有充分利用许多最近完成的单细胞参考图谱所提供的信息,后者是通过多年的联盟级努力和数百万美元的国家资助实现的。该项目旨在弥合这一计算分析的差距,特别是通过解决该领域的两个关键限制:(1)缺乏从临床规模的单细胞测序研究中去除不必要的技术变化的原则方法,以及(2)需要一种综合的临床规模蛋白质组分析方法,利用单细胞参考图谱来实现更精确的细胞类型组成分析。对于这两个问题,该项目的目标是为可重复的科学研究开发有原则的、透明的和可扩展的方法。这个项目更广泛的影响是它对改善社会中个人福祉的潜在影响;对本科生和研究生STEM教育的贡献;以及参与和参与女性和STEM领域中代表性不足的少数群体的研究。在更多的技术细节中,该项目的第一个目标侧重于消除技术上的“批处理效果”。技术批量校正是*所有*单细胞数据分析管道中不可避免的和具有挑战性的一步。已经提出了许多方法来校正单细胞实验中的批效应,但它们被设计为在所有批次中对齐细胞,忽略了设计原则,如对照队列、纵向抽样和生物重复。如果不利用这些设计原则,现有的方法就不能充分地估计批效应,并且经常将它们与真实的生物信号混淆。该项目的目标1制定了适用于多种实验设计的批量效应校正的新统计方法,并开发了用于量化生物信号强度(例如差异表达或新细胞类型的出现)的统计推断程序,以考虑批量校正中的不确定性。该项目的目标2解决了临床规模单细胞蛋白质组学分析中出现的不同挑战:尽管测序成本降低,流式细胞术和质量细胞术仍然是更快、更便宜的数量级,因此仍然是临床规模免疫学研究的首选方法,因为需要在紧迫的时间内对大型队列进行分析。然而,每次流式细胞术只测量有限的一组蛋白质,因此不允许在单细胞转录组学提供的详细水平上进行细胞类型标记。Aim 2开发了一种新的细胞类型分析方法,该方法将多个流式/质量细胞术运行与互补面板集成在同一样品上,其目标是仅在一小部分时间和成本下实现与最先进的单细胞测序方案相匹配的细胞类型制表精度。这一目标利用日益增长的单细胞参考地图集,为未来在人口水平细胞类型普查项目中使用这些地图集提供了路线图。通过合作,开发的方法将应用于多个正在进行的具有直接临床影响的大型队列研究。本项目开发的方法将作为开源软件发布,生成的数据集将上传到公共存储库供一般使用。该奖项反映了美国国家科学基金会的法定使命,并通过使用基金会的知识价值和更广泛的影响审查标准进行评估,被认为值得支持。
英文摘要
Over the last two decades, scientific discoveries in biology and medicine have increasingly relied on the statistical analysis of large genomic data sets. Thus, the development of rigorous statistical methods for reproducible and scalable analyses on such data sets, and the interdisciplinary training of a next generation of scientists who can straddle the computational and biomedical sciences, is crucial for our scientific advancement. This project focuses on challenges in the analysis of data from single cell experiments, which are made possible by recent technological developments enabling genome-scale measurements to be made on individual cells, in high throughput across millions of cells at one go. Such single cell experiments are now the bread-and-butter of biomedical research, playing a cardinal role in our quest for a more complete understanding of cell biology as well as our pursuit of cures for every disease, from infectious diseases such as COVID-19 to cancer to aging-related maladies such as neurodegeneration. Despite the promise of these new technologies, current methods for single cell data analysis are not designed for clinical-scale disease research, and do not adequately harness the information provided by the many recently completed single cell reference atlases, the latter realized through years of consortia-level efforts and millions of dollars in national funding. This project aims to bridge this computational analysis gap, specifically by addressing two critical limitations in the field: (1) The lack of a principled approach for the removal of unwanted technical variation from clinical-scale single cell sequencing studies, and (2) the need for an integrated approach to clinical-scale proteomic profiling that harness single cell reference atlases to achieve more precise cell type composition analysis. For both problems, the project’s goal is to develop principled, transparent, and scalable methods for reproducible scientific research. Broader impacts of this project are its potential impact in improving the well-being of individuals in society; its contributions in STEM education for undergraduate and graduate students; and in the involvement and participation in research of women and under-represented minorities in STEM fields.In more technical detail, the first aim of the project focuses on the removal of technical “batch effects”. Technical batch correction is an unavoidable and challenging step in *all* single cell data analysis pipelines. Many methods have been proposed for batch effect correction in single cell experiments, but they were designed to align cells across all batches, ignoring design principles such as control cohorts, longitudinal sampling, and biological replicates. Without utilizing these design principles, existing methods can not adequately estimate batch effects and often confound them with real biological signals. Aim 1 of the project formulates new statistical methods for batch effect correction that are adaptable to a multitude of experimental designs and develops statistical inference procedures for quantifying the strength of biological signals (e.g. differential expression or emergence of a new cell type) accounting for the uncertainty in batch correction. Aim 2 of the project tackles a different challenge arising in clinical-scale single cell proteomic profiling: Despite decreasing sequencing costs, flow and mass cytometry are still orders-of-magnitude faster and cheaper, and thus remains the method of choice in clinical-scale immunological studies where large cohorts need to be profiled on a tight timeline. However, each flow/mass cytometry run only measures a limited panel of proteins, and thus does not allow cell-type labelling at the level of detail afforded by single cell transcriptomics. Aim 2 develops a new approach to cell type profiling that integrates multiple flow/mass cytometry runs, with complementary panels, on the same sample, with the goal of achieving cell-type tabulation accuracy matching state-of-the-art single cell sequencing protocols at only a fraction of time and cost. This aim leverages the growing compendium of single cell reference atlases, providing a roadmap for the future use of these atlases in population-level cell type censusing projects. Through collaborations, the developed methods will be applied to multiple ongoing large-cohort studies that have direct clinical impact. Methods developed in this project will be released as open-source software, and the datasets generated will be uploaded to public repository for general use.This award reflects NSF's statutory mission and has been deemed worthy of support through evaluation using the Foundation's intellectual merit and broader impacts review criteria.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
New Change-point Problems in Genomic Profiling
-
批准号:0906394
-
项目类别:Standard Grant
-
资助金额:$10.0万
-
财政年份:2009
-
负责人:Nancy Zhang
-
依托单位:
海外基金