Biosimilarity Assessments of Model IgG1-Fc Glycoforms Using a Machine Learning Approach.

Biosimilarity Assessments of Model IgG1-Fc Glycoforms Using a Machine Learning Approach.
复制标题

使用机器学习方法对 IgG1-Fc 糖型模型进行生物相似性评估。

DOI:
10.1016/j.xphs.2015.10.013
复制
发表时间:
2016
影响因子:
3.8
通讯作者:
SmalterHall,Aaron
SmalterHall,Aaron
中科院分区:
医学3区
文献类型:
--
作者:
Kim,JaeHyun;Joshi,SangeetaB;Tolbert,ThomasJ;Middaugh,CRussell;Volkin,DavidB;SmalterHall,Aaron

文献摘要

被引文献

相似文献

进行生物相似性评估,以决定是否可以认为2种复杂生物分子制剂“高度相似”。在这项工作中,机器学习方法被证明是一种数学工具,使用各种分析数据集进行此类评估。作为原理证明,检查了来自8个样品的物理稳定性数据集,在2种不同制剂中的4种明确定义的免疫球蛋白G1-片段可结晶糖型(参见More等人,本文中的相关文章)。数据集包括不同pH值和温度条件下3种分析方法的一式三份测量值(2066个数据特征)。使用已建立的机器学习技术来确定数据集在本申请中是否包含足够的辨别力。支持向量机分类器以高精度识别了8个不同的样本。对于这些数据集,在信息质量和数量方面存在一个最小阈值,以授予足够的区分能力。通常,需要来自多种分析技术、多种pH条件和至少200个代表性特征的数据来实现最高的判别准确度。除了分类准确性测试,各种方法,如样本空间可视化,相似性分析的基础上的欧氏距离,和互信息分数的特征排名证明显示其有效性作为建模工具的生物相似性评估。
Biosimilarity assessments are performed to decide whether 2 preparations of complex biomolecules can be considered “highly similar.” In this work, a machine learning approach is demonstrated as a mathematical tool for such assessments using a variety of analytical data sets. As proof-of-principle, physical stability data sets from 8 samples, 4 well-defined immunoglobulin G1-Fragment crystallizable glycoforms in 2 different formulations, were examined (see More et al., companion article in this issue). The data sets included triplicate measurements from 3 analytical methods across different pH and temperature conditions (2066 data features). Established machine learning techniques were used to determine whether the data sets contain sufficient discriminative power in this application. The support vector machine classifier identified the 8 distinct samples with high accuracy. For these data sets, there exists a minimum threshold in terms of information quality and volume to grant enough discriminative power. Generally, data from multiple analytical techniques, multiple pH conditions, and at least 200 representative features were required to achieve the highest discriminative accuracy. In addition to classification accuracy tests, various methods such as sample space visualization, similarity analysis based on Euclidean distance, and feature ranking by mutual information scores are demonstrated to display their effectiveness as modeling tools for biosimilarity assessments.