Multiplexed massively parallel SELEX for characterization of human transcription factor binding specificities

Multiplexed massively parallel SELEX for characterization of human transcription factor binding specificities
复制标题

DOI:
10.1101/gr.100552.109
复制
发表时间:
2010-06-01
期刊:
影响因子:
7
通讯作者:
Taipale, Jussi
Taipale, Jussi
中科院分区:
生物学1区
文献类型:
--
作者:
Jolma, Arttu;Kivioja, Teemu;Taipale, Jussi

文献摘要

被引文献

相似文献

遗传密码(所有转移 RNA 的结合特异性)定义了 DNA 序列如何确定蛋白质一级结构。 DNA 还决定蛋白质表达的时间和地点,并且该信息以转录因子识别的特定序列基序的模式进行编码。然而,仅了解 1400 种人类转录因子 (TF) 中一小部分的 DNA 结合特异性。我们在这里描述了一种用于分析转录因子结合特异性的高通量方法,该方法基于指数富集配体系统进化(SELEX)和大规模并行测序。该方法经过优化,可通过使用亲和标记蛋白、条形码选择寡核苷酸和多重测序来并行分析大量 TF。数据通过新的生物信息学平台进行分析,该平台使用获得的数十万个测序读数来控制实验质量并生成 TF 的结合基序。与当前基于微阵列的方法相比,所描述的技术允许更高的通量和更长的结合谱的识别。此外,由于我们的方法基于哺乳动物细胞中表达的蛋白质,因此它还可用于表征全长蛋白质或需要翻译后修饰的蛋白质的 DNA 结合偏好。我们通过确定 14 种不同类别 TF 的结合特异性并使用 ChIP-seq 确认 NFATC1 和 RFX3 的特异性来验证该方法。我们的结果揭示了几种因子的意想不到的二聚体结合模式,这些因子被认为优先以单体形式结合 DNA。
The genetic code-the binding specificity of all transfer-RNAs-defines how protein primary structure is determined by DNA sequence. DNA also dictates when and where proteins are expressed, and this information is encoded in a pattern of specific sequence motifs that are recognized by transcription factors. However, the DNA-binding specificity is only known for a small fraction of the similar to 1400 human transcription factors (TFs). We describe here a high-throughput method for analyzing transcription factor binding specificity that is based on systematic evolution of ligands by exponential enrichment (SELEX) and massively parallel sequencing. The method is optimized for analysis of large numbers of TFs in parallel through the use of affinity-tagged proteins, barcoded selection oligonucleotides, and multiplexed sequencing. Data are analyzed by a new bioinformatic platform that uses the hundreds of thousands of sequencing reads obtained to control the quality of the experiments and to generate binding motifs for the TFs. The described technology allows higher throughput and identification of much longer binding profiles than current microarray-based methods. In addition, as our method is based on proteins expressed in mammalian cells, it can also be used to characterize DNA-binding preferences of full-length proteins or proteins requiring post-translational modifications. We validate the method by determining binding specificities of 14 different classes of TFs and by confirming the specificities for NFATC1 and RFX3 using ChIP-seq. Our results reveal unexpected dimeric modes of binding for several factors that were thought to preferentially bind DNA as monomers.