MS-DECODER: Milliseconds Sequencing of Coded Polymers

MS-DECODER: Milliseconds Sequencing of Coded Polymers
复制标题

DOI:
10.1021/acs.macromol.7b01737
复制
发表时间:
2017-10-24
期刊:
影响因子:
5.5
通讯作者:
Charles, Laurence
Charles, Laurence
中科院分区:
化学1区
文献类型:
--
作者:
Burel, Alexandre;Carapito, Christine;Charles, Laurence

文献摘要

被引文献

相似文献

正如该杂志最近讨论的那样,1合成的非生物大分子可用于在分子水平上存储信息,因此为人造聚合物开辟了新的应用领域,包括数据存储,2长期存储,3和防伪技术。4在这种含信息的聚合物中,5个单体单元用作分子字母表,其可以是二进制、三进制、四进制、十进制或甚至更复杂。此类聚合物通常通过迭代化学过程合成,该过程允许制备均匀的序列定义的大分子。6例如,我们在过去三年中报道了各种数字编码大分子的合成。7− 11这些数字聚合物允许在室温下存储信息,并可以使用测序技术进行解码,12这是一种允许表征单体序列的分析方法。然而,尽管存在各种各样的测序方法用于生物聚合物分析,13,14到目前为止,只有少数几种方法被验证用于非天然聚合物的表征。12目前,串联质谱(MS/MS)是破译合成聚合物序列的主要方法,15如US 9 - 11,16 - 18以及其他人所示。[19 - 21]特别是,我们已经强调了数字聚合物的可测序性可以通过聚合物设计得到显著改善。事实上,可以优化序列编码的大分子的分子结构,以最大限度地减少解离途径的数量,同时避免MS/MS测量中的二次碎片化,9,10,22从而大大促进光谱解释。然而,对于非天然聚合物,MS/MS光谱的分析通常是手动进行的,因此导致平均解码时间约为1 - 10分钟,实际上取决于序列长度和光谱复杂性。虽然分钟范围内的分析时间与基础研究兼容,但它们对于数据存储或防伪标签等高级技术应用来说是有限的。在这种情况下,允许减少解码时间的自动化方法可能对新兴的数字聚合物领域非常有益。在过去的几十年里,基因组学和蛋白质组学的要求已经大大简化了生物信息学软件的发展,简化了蛋白质和DNA测序。例如,开源的UniNovo或商业PEAKS算法是MS/MS肽测序的广泛工具。然而,这些软件工具是专门为肽/蛋白质设计的,并且不能容易地应用于遵循不同片段化规则的合成聚合物。在本技术说明中,我们介绍了一种名为MS-DECODER的新开源工具,该工具专门用于合成序列编码大分子的自动测序。描述了MS-DECODER算法,并通过其成功应用于对三种类型的数字聚合物(图1)进行测序来证明其多功能性,这些聚合物在碰撞诱导解离(CID)中表现出不同的碎片模式。它可以在基本的笔记本电脑上在毫秒范围内对所有测试的候选人进行准确解码。编码聚合物的MS/MS测序规则。当在负离子模式下分析时,为聚(氨基甲酸酯)(PU)定义了最简单的测序规则。最近
As recently discussed in this journal, 1 synthetic abiotic macromolecules can be used to store information at the molecular level and therefore open up new areas of application for man-made polymers, including data storage, 2 long-term storage, 3 and anticounterfeiting technologies. 4 In such information-containing polymers, 5 monomer units are used as a molecular alphabet, which can be binary, ternary, quaternary, decimal, or even more complex. 1 Such polymers are usually synthesized via an iterative chemical process that allows preparation of uniform sequence-defined macromolecules. 6 For instance, we have reported over the past three years the synthesis of a variety of digitally encoded macromolecules. 7− 11 These digital polymers allow information storage at room temperature and can be decoded using a sequencing technique, 12 which is an analytical method that permits to characterize monomer sequences. Yet, although a wide variety of sequencing methods exist for biopolymer analysis, 13, 14 only a few of them have been validated so far for the characterization of non-natural polymers. 12 Currently, tandem mass spectrometry (MS/MS) is the leading method for deciphering the sequences of synthetic polymers, 15 as shown by us 9− 11, 16− 18 as well as others. 19− 21 In particular, we have emphasized that the sequenceability of digital polymers can be significantly improved through polymer design. Indeed, the molecular structure of sequence-coded macromolecules can be optimized to minimize the number of dissociation routes while avoiding secondary fragmentations in MS/MS measurements, 9, 10, 22 thus greatly facilitating spectra interpretation. However, for nonnatural polymers, analysis of MS/MS spectra is usually performed manually, thus leading to average decoding times of about 1− 10 min, depending indeed on sequence length and spectrum complexity. Although analysis times in the minute range are compatible with fundamental studies, they become limiting for advanced technological applications such as data storage or anticounterfeiting tags. In this context, automated approaches permitting to reduce decoding time could be very beneficial for the emerging field of digital polymers. During the past decades, the demanding fields of genomics and proteomics have been drastically simplified by the development of bioinformatics software that simplify protein and DNA sequencing. For example, the open-source UniNovo or the commercial PEAKS algorithms are widespread tools for MS/MS peptide sequencing. 23, 24 However, these software tools are specifically conceived for peptides/proteins and cannot be easily applied to synthetic polymers that obey different fragmentation rules. In this technical note, we introduce a new open-source tool called MS-DECODER, which was specifically conceived for the automated sequencing of synthetic sequence-coded macromolecules. The MS-DECODER algorithm is described, and its versatility is demonstrated by its successful application to sequence three types of digital polymers (Figure 1) that exhibit different fragmentation patterns in collision-induced dissociation (CID). It enables accurate decoding for all tested candidates in the millisecond range on a basic laptop computer. MS/MS Sequencing Rules of Coded Polymers. The simplest sequencing rules were defined for poly (urethane) s (PUs) when analyzed in the negative ion mode. As recently