An Untargeted Metabolomics Workflow that Scales to Thousands of Samples for Population-Based Studies.

An Untargeted Metabolomics Workflow that Scales to Thousands of Samples for Population-Based Studies.
复制标题

非目标代谢组学工作流程,可扩展到数千个样本,用于基于人群的研究。

DOI:
10.1021/acs.analchem.2c01270
复制
发表时间:
2022
影响因子:
7.4
通讯作者:
Patti,GaryJ
Patti,GaryJ
中科院分区:
化学1区
文献类型:
--
作者:
Stancliffe,Ethan;Schwaiger-Haber,Michaela;Sindelar,Miriam;Murphy,MatthewJ;Soerensen,Mette;Patti,GaryJ

文献摘要

参考文献

相似文献

精准医疗的成功依赖于从人口水平上收集许多个人的数据。尽管先进的技术已经使这种大规模研究在基因组学等一些学科中变得越来越可行,但目前在非靶向代谢组学中实施的标准工作流程是针对小样本数开发的,并且受到液相色谱/质谱数据处理的限制。在这里,我们提出了一个非靶向代谢组学工作流程,旨在支持数千个生物标本的大规模项目。我们的策略是首先评估一个参考样本,该样本是通过从队列中收集等量的生物标本创建的。参考样品在少量分析运行中捕获生物基质的化学复杂性,随后可以使用XCMS等传统软件进行处理。虽然这会产生数千个所谓的特征,但大多数并不对应于样品中的独特化合物,可以用已建立的信息学工具进行过滤。剩下的特征代表了一套全面的生物学相关参考化学物质,然后可以使用Skyline根据m/zvalue和保留时间从整个队列的原始数据中提取。为了证明对大型队列的适用性,我们用我们的工作流程评估了2000个人类血浆样本。我们集中分析了360种已确定的化合物,但我们也从血浆样本中分析了3000种未知化合物。作为我们工作流程的一部分,我们测试了14种不同的批量校正计算方法,发现基于随机森林的方法优于其他方法。修正后的数据揭示了与参与者地理位置相关的不同概况。
The success of precision medicine relies upon collecting data from many individuals at the population level. Although advancing technologies have made such large-scale studies increasingly feasible in some disciplines such as genomics, the standard workflows currently implemented in untargeted metabolomics were developed for small sample numbers and are limited by the processing of liquid chromatography/mass spectrometry data. Here we present an untargeted metabolomics workflow that is designed to support large-scale projects with thousands of biospecimens. Our strategy is to first evaluate a reference sample created by pooling aliquots of biospecimens from the cohort. The reference sample captures the chemical complexity of the biological matrix in a small number of analytical runs, which can subsequently be processed with conventional software such as XCMS. Although this generates thousands of so-called features, most do not correspond to unique compounds from the samples and can be filtered with established informatics tools. The features remaining represent a comprehensive set of biologically relevant reference chemicals that can then be extracted from the entire cohort’s raw data on the basis ofm/zvalues and retention times by using Skyline. To demonstrate applicability to large cohorts, we evaluated >2000 human plasma samples with our workflow. We focused our analysis on 360 identified compounds, but we also profiled >3000 unknowns from the plasma samples. As part of our workflow, we tested 14 different computational approaches for batch correction and found that a random forest-based approach outperformed the others. The corrected data revealed distinct profiles that were associated with the geographic location of participants.
DOI: 10.1021/acs.analchem.9b05765
发表时间: 2020-06-02
影响因子: 7.4
作者:
Bonini P;Kind T;Tsugawa H;Barupal DK;Fiehn O
通讯作者: Fiehn O
DOI: 10.1021/acs.analchem.7b02380
发表时间: 2017-10-03
影响因子: 7.4
作者:
Mahieu NG;Patti GJ
通讯作者: Patti GJ
DOI: 10.1101/2021.10.13.464246
发表时间: 2021
期刊: bioRxiv
影响因子: --
作者:
K. Kirkwood;Michael W. Christopher;J. Burgess;S. Littau;Brian S. Pratt;Nicholas Shulman;Kaipo Tamura;M. MacCoss;B. MacLean;E. Baker
通讯作者: E. Baker
DOI: 10.1016/j.aca.2021.338210
发表时间: 2021-03-08
影响因子: 6.2
作者:
Cho K;Schwaiger-Haber M;Naser FJ;Stancliffe E;Sindelar M;Patti GJ
通讯作者: Patti GJ
DOI: 10.1093/biostatistics/kxj037
发表时间: 2007-01-01
期刊: BIOSTATISTICS
影响因子: 2.1
作者:
Johnson, W. Evan;Li, Cheng;Rabinovic, Ariel
通讯作者: Rabinovic, Ariel