Detecting statistically significant common insertion sites in retroviral insertional mutagenesis screens.

Detecting statistically significant common insertion sites in retroviral insertional mutagenesis screens.
复制标题

检测逆转录病毒插入诱变筛查中具有统计学意义的常见插入位点。

DOI:
10.1371/journal.pcbi.0020166
复制
发表时间:
2006-12-08
影响因子:
4.3
通讯作者:
Wessels, Lodewyk
Wessels, Lodewyk
中科院分区:
生物学2区
文献类型:
--
作者:
de Ridder, Jeroen;Uren, Anthony;Kool, Jaap;Reinders, Marcel;Wessels, Lodewyk

文献摘要

参考文献

被引文献

相似文献

逆转录病毒插入突变筛选用于识别与小鼠肿瘤发展相关的基因,已经产生了大量的逆转录病毒整合位点,由于高通量筛选技术的引入,这一数字预计将大幅增长。各种逆转录病毒插入突变筛选的数据汇编在公开可用的逆转录病毒标记癌症基因数据库(RTCGD)中。完整地分析这些筛查是否存在共同的插入部位(即基因组中被病毒插入多个独立肿瘤的区域意外地比预期更多地击中的区域),需要一种方法,随着可用数据量的增加,纠正发现虚假CI的概率增加的情况。此外,应该考虑到插入过程的随机性产生的噪声,以及基因组中存在的优先插入位点和数据检索方法产生的偏差,来确定CIS的重要性估计。我们引入了一个框架,即核卷积(KC)框架,在噪声和有偏差的环境中使用预先定义的显著水平来检测CI,同时控制家族误差(FWE)(检测到虚假CI的概率)。如果以前的方法使用一个、两个或三个预定的固定尺度,我们的方法能够在任何生物相关的尺度上操作。这创造了通过改变CI的宽度来在尺度空间中分析CI的可能性,从而提供了对CI在多个尺度上的行为的新的见解。我们的方法还具有包含背景偏差模型的可能性。利用模拟数据,我们使用三种核函数,即高斯核函数、三角形核函数和矩形核函数对KC框架进行了评估。我们将高斯KC应用于RTCGD中来自屏幕组合集合的数据,发现在这种组合设置中,53%的CI没有达到显著阈值。尽管如此,在FWE得到控制的情况下,应用我们的方法发现了8个新的CI,每个CI的错误检测概率不到5%。逆转录病毒插入突变是鉴定新的癌症基因的有效方法。感染慢转化逆转录病毒的小鼠会患上肿瘤,因为这种病毒会随机插入它们的基因组,并突变癌症基因。基因组中在多个独立肿瘤中发生突变的区域可能包含与肿瘤发生有关的基因。随着这些数据集的大小的增加,检测这些所谓的公共插入位点(CI)的传统方法不再足够,需要一种能够独立于数据集大小来控制误差的方法。作者介绍了一个框架,该框架使用一种名为核密度估计的技术来寻找基因组中显示插入密度显著增加的区域。该方法在一系列尺度上实施,允许在任何相关尺度上对数据进行评估。作者证明,该框架能够补偿数据中的固有偏见,如倾向于逆转录病毒插入转录起点附近。通过更好地平衡误差,他们能够表明,从公布的361个CI中,可以识别出150个具有较低的错误检测概率。此外,他们还发现了八部小说《西斯》。
Retroviral insertional mutagenesis screens, which identify genes involved in tumor development in mice, have yielded a substantial number of retroviral integration sites, and this number is expected to grow substantially due to the introduction of high-throughput screening techniques. The data of various retroviral insertional mutagenesis screens are compiled in the publicly available Retroviral Tagged Cancer Gene Database (RTCGD). Integrally analyzing these screens for the presence of common insertion sites (CISs, i.e., regions in the genome that have been hit by viral insertions in multiple independent tumors significantly more than expected by chance) requires an approach that corrects for the increased probability of finding false CISs as the amount of available data increases. Moreover, significance estimates of CISs should be established taking into account both the noise, arising from the random nature of the insertion process, as well as the bias, stemming from preferential insertion sites present in the genome and the data retrieval methodology. We introduce a framework, the kernel convolution (KC) framework, to find CISs in a noisy and biased environment using a predefined significance level while controlling the family-wise error (FWE) (the probability of detecting false CISs). Where previous methods use one, two, or three predetermined fixed scales, our method is capable of operating at any biologically relevant scale. This creates the possibility to analyze the CISs in a scale space by varying the width of the CISs, providing new insights in the behavior of CISs across multiple scales. Our method also features the possibility of including models for background bias. Using simulated data, we evaluate the KC framework using three kernel functions, the Gaussian, triangular, and rectangular kernel function. We applied the Gaussian KC to the data from the combined set of screens in the RTCGD and found that 53% of the CISs do not reach the significance threshold in this combined setting. Still, with the FWE under control, application of our method resulted in the discovery of eight novel CISs, which each have a probability less than 5% of being false detections. A potent method for the identification of novel cancer genes is retroviral insertional mutagenesis. Mice infected with slow transforming retroviruses develop tumors because the virus inserts randomly in their genome and mutates cancer genes. The regions in the genome that are mutated in multiple independent tumors are likely to contain genes involved in tumorigenesis. As the size of these datasets increases, conventional methods to detect these so-called common insertion sites (CISs) no longer suffice, and an approach is required that can control the error independent of the dataset size. The authors introduce a framework that uses a technique called kernel density estimation to find the regions in the genome that show a significant increase in insertion density. This method is implemented over a range of scales, allowing the data to be evaluated at any relevant scale. The authors demonstrate that the framework is capable of compensating for the inherent biases in the data, such as preference for retroviruses to insert near transcriptional start sites. By better balancing the error, they are able to show that from the 361 published CISs, 150 can be identified that have a low probability of being a false detection. In addition, they discover eight novel CISs.
DOI: 10.1214/aoms/1177704472
发表时间: 1962-01-01
影响因子: --
作者:
PARZEN, E
通讯作者: PARZEN, E
DOI: 10.1073/pnas.0402716101
发表时间: 2004-08-03
影响因子: 11.1
作者:
Johansson, FK;Brodd, J;Westermark, B
通讯作者: Westermark, B
DOI: 10.1371/journal.pbio.0020423
发表时间: 2004-12
期刊: PLoS biology
影响因子: 9.8
作者:
Hematti P;Hong BK;Ferguson C;Adler R;Hanawa H;Sellers S;Holt IE;Eckfeldt CE;Sharma Y;Schmidt M;von Kalle C;Persons DA;Billings EM;Verfaillie CM;Nienhuis AW;Wolfsberg TG;Dunbar CE;Calmels B
通讯作者: Calmels B
DOI: 10.1128/jvi.79.1.67-78.2005
发表时间: 2005-01-01
影响因子: 5.4
作者:
Nielsen, AA;Sorensen, AB;Pedersen, FS
通讯作者: Pedersen, FS
DOI: 10.1126/science.1083413
发表时间: 2003-06-13
期刊: SCIENCE
影响因子: 56.9
作者:
Wu, XL;Li, Y;Burgess, SM
通讯作者: Burgess, SM