MolFind: A Software Package Enabling HPLC/MS-Based Identification of Unknown Chemical Structures

MolFind: A Software Package Enabling HPLC/MS-Based Identification of Unknown Chemical Structures
复制标题

DOI:
10.1021/ac302048x
复制
发表时间:
2012-11-06
影响因子:
7.4
通讯作者:
Grant, David F.
Grant, David F.
中科院分区:
化学1区
文献类型:
--
作者:
Menikarachchi, Lochana C.;Cawley, Shannon;Grant, David F.

文献摘要

被引文献

相似文献

在本文中,我们介绍了 MolFind,这是一种高度多线程管道型软件包,可帮助识别复杂生物流体和混合物中的化学结构。 MolFind 专为代谢组学研究中典型的高效液相色谱/质谱 (HPLC/MS) 数据输入而设计,其中结构鉴定是最终目标。 MolFind 通过将未知化合物获得的基于 HPLC/MS 的实验数据与从 PubChem 等化学数据库下载的候选化合物的计算得出的 HPLC/MS 值相匹配,实现化合物鉴定。下载的“箱”由与未知物的单同位素分子量匹配的所有化合物组成。预测的计算 HPLC/MS 值包括保留指数 (RI)、ECOM50(裂解 50% 选定母离子所需的能量)、漂移时间和碰撞诱导解离 (CID) 谱 RI、ECOM50 和漂移时间模型用于过滤从 PubChem 下载的化合物。然后根据 CID 光谱匹配对剩余候选者进行排名。目前的 RI 和 ECOM50 型号可以去除 PubChem 箱中约 28% 的化合物。我们的估计表明,如果计算模型中包含额外的化学结构,这一比例可以提高至 87%。基于定量结构特性关系的漂移时间建模显示出与实验确定的漂移时间的相关性比 Mobcal 横截面积更好。在 35 个示例案例中的 23 个中,与之前单独使用 CID 光谱匹配的研究相比,使用 RI 和 ECOM50 预测模型过滤 PubChem 箱可以提高未知化合物的排名。在 35 个示例中的 19 个中,正确的候选化合物被排在平均包含 1635 种化合物的箱中的前 20 个化合物中。
In this paper, we present MolFind, a highly multithreaded pipeline type software package for use as an aid in identifying chemical structures in complex biofluids and mixtures. MolFind is specifically designed for high-performance liquid chromatography/mass spectrometry (HPLC/MS) data inputs typical of metabolomics studies where structure identification is the ultimate goal. MolFind enables compound identification by matching HPLC/MS-based experimental data obtained for an unknown compound with computationally derived HPLC/MS values for candidate compounds downloaded from chemical databases such as PubChem. The downloaded "bins" consist of all compounds matching the monoisotopic molecular weight of the unknown. The computational HPLC/MS values predicted include retention index (RI), ECOM50 (energy required to fragment 50% of a selected precursor ion), drift time, and collision induced dissociation (CID) spectrum RI, ECOM50, and drift time models are used for filtering compounds downloaded from PubChem. The remaining candidates are then ranked based on CID spectra matching. Current RI and ECOM50 models allow for the removal of about 28% of compounds from PubChem bins. Our estimates suggest that this could be improved to as much as 87% with additional chemical structures included in the computational models. Quantitative structure property relationship based modeling of drift times showed a better correlation with experimentally determined drift times than did Mobcal cross-sectional areas. In 23 of 35 example cases, filtering PubChem bins with RI and ECOM50 predictive models resulted in improved ranking of the unknown compounds compared to previous studies using CID spectra matching alone. In 19 of 35 examples, the correct candidate was ranked within the top 20 compounds in bins containing an average of 1635 compounds.