Comparative Characterization of Crofelemer Samples Using Data Mining and Machine Learning Approaches With Analytical Stability Data Sets.

Comparative Characterization of Crofelemer Samples Using Data Mining and Machine Learning Approaches With Analytical Stability Data Sets.
复制标题

使用数据挖掘和机器学习方法与分析稳定性数据集对 Crofelemer 样品进行比较表征。

DOI:
10.1016/j.xphs.2017.07.013
复制
发表时间:
2017
影响因子:
3.8
通讯作者:
Deeds,EricJ
Deeds,EricJ
中科院分区:
医学3区
文献类型:
--
作者:
Nariya,MaulikK;Kim,JaeHyun;Xiong,Jian;Kleindl,PeterA;Hewarathna,Asha;Fisher,AdamC;Joshi,SangeetaB;Schöneich,Christian;Forrest,MLaird;Middaugh,CRussell;Volkin,DavidB;Deeds,EricJ

文献摘要

被引文献

相似文献

人们对生成物理化学和生物分析数据集以比较复杂的混合药物(例如,来自不同制造商的产品)越来越感兴趣。在这项工作中,我们比较了从单个批次中通过过滤不同分子量截止值并结合不同时间在不同温度下的孵育制备的各种作物样品。前两篇文章描述了从分离和降解作物样品的分析表征中产生的实验数据集。在这项工作中,我们使用主成分分析和互信息评分等数据挖掘技术来帮助可视化数据并确定这些大数据集中的歧视区域。互信息得分识别化学特征,区分作物样品。在许多情况下,这些特征可能会被传统的数据分析工具遗漏。我们还发现,监督学习分类器稳健地区分样本,分类准确率约为99%,这表明这些物理化学数据集的数学模型能够识别出crofelemer样本中的细微差异。因此,数据挖掘和机器学习技术可以识别复杂混合药物的指纹类型属性,可用于产品的比较表征。
There is growing interest in generating physicochemical and biological analytical data sets to compare complex mixture drugs, for example, products from different manufacturers. In this work, we compare various crofelemer samples prepared from a single lot by filtration with varying molecular weight cutoffs combined with incubation for different times at different temperatures. The 2 preceding articles describe experimental data sets generated from analytical characterization of fractionated and degraded crofelemer samples. In this work, we use data mining techniques such as principal component analysis and mutual information scores to help visualize the data and determine discriminatory regions within these large data sets. The mutual information score identifies chemical signatures that differentiate crofelemer samples. These signatures, in many cases, would likely be missed by traditional data analysis tools. We also found that supervised learning classifiers robustly discriminate samples with around 99% classification accuracy, indicating that mathematical models of these physicochemical data sets are capable of identifying even subtle differences in crofelemer samples. Data mining and machine learning techniques can thus identify fingerprint-type attributes of complex mixture drugs that may be used for comparative characterization of products.