A Haystack Heuristic for Autoimmune Disease Biomarker Discovery Using Next-Gen Immune Repertoire Sequencing Data

A Haystack Heuristic for Autoimmune Disease Biomarker Discovery Using Next-Gen Immune Repertoire Sequencing Data
复制标题

DOI:
10.1038/s41598-017-04439-5
复制
发表时间:
2017-07-13
期刊:
影响因子:
4.6
通讯作者:
Sirota, Marina
Sirota, Marina
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Apeltsin, Leonard;Wang, Shengzhi;Sirota, Marina

文献摘要

被引文献

相似文献

大规模DNA测序的免疫剧目提供了一个机会,发现新的生物标志物的自身免疫性疾病。然而,由于不令人满意的可扩展性和灵敏度,可用的生物信息学技术不足以适用于从大型免疫测序数据集内阐明可能的生物标志物候选物。在这里,我们提出了Haystack启发式算法,该算法通过对比疾病和健康受试者,从下一代测序的剧目中计算提取疾病相关的基序。该技术采用局部搜索图论方法来发现患者数据中的新图案。我们应用Haystack启发式方法从近100个个体中获得了900万个B细胞受体序列,以阐明与多发性硬化症显著相关的新基序。我们的研究结果证明了Haystack启发式算法在从高通量测序数据中计算可能的生物标志物候选者方面的有效性,并且可以推广到其他数据集。
Large-scale DNA sequencing of immunological repertoires offers an opportunity for the discovery of novel biomarkers for autoimmune disease. Available bioinformatics techniques however, are not adequately suited for elucidating possible biomarker candidates from within large immunosequencing datasets due to unsatisfactory scalability and sensitivity. Here, we present the Haystack Heuristic, an algorithm customized to computationally extract disease-associated motifs from next-generation-sequenced repertoires by contrasting disease and healthy subjects. This technique employs a local-search graph-theory approach to discover novel motifs in patient data. We apply the Haystack Heuristic to nine million B-cell receptor sequences obtained from nearly 100 individuals in order to elucidate a new motif that is significantly associated with multiple sclerosis. Our results demonstrate the effectiveness of the Haystack Heuristic in computing possible biomarker candidates from high throughput sequencing data and could be generalized to other datasets.