The share of human genomic DNA under selection estimated from human-mouse genomic alignments

The share of human genomic DNA under selection estimated from human-mouse genomic alignments
复制标题

DOI:
10.1101/sqb.2003.68.245
复制
发表时间:
2003-01-01
期刊:
COLD SPRING HARBOR SYMPOSIA ON QUANTITATIVE BIOLOGY
影响因子:
--
通讯作者:
Haussler, D
Haussler, D
中科院分区:
其他
文献类型:
--
作者:
Chiaromonte, F;Weber, RJ;Haussler, D

文献摘要

被引文献

相似文献

METHODSData准备。我们的短对齐窗口集合是沿着人类序列使用固定的位置网格构建的。对于我们考虑的大小,网格总是保证不重叠的窗口。对于给定的窗口大小(W)和比对过滤阈值(T),首先在每个位置扩展W个碱基的窗口,然后丢弃所有少于T个碱基与鼠标对齐的窗口。对于相同的窗口大小和过滤阈值,相对于特定特征类型(祖先重复,编码区域)的窗口集合以类似的方式构建,首先在网格位置扩展大小为W的窗口,然后丢弃与该类型的对齐特征重叠小于T个碱基的窗口。表1给出了我们主要分析中W= 50, T= 40病例的全基因组窗口所提供的覆盖范围,以及窗口大小和过滤阈值的其他组合。祖先重复序列是通过Repeat-Masker(可在http://ftp上获得)识别的重复序列。基因组。华盛顿。edu/RM/RepeatMasker。html;Smit和Green 1999),并出现在同源位点。特定家族的祖先重复序列的列表在方法网上提供的概要Waterston等人(2002年)。利用BLAT将GenBank release 130.0版的RefSeq (Pruitt and Maglott 2001)人mrna与人类基因组比对获得已知编码区注释(Kent 2002; Kent et al. 2002)。我们选择了与鼠标位置对齐并满足以下条件的注释:(1)CDS在人和小鼠中都是完整的,以开始密码子开始,以停止密码子结束。小鼠终止密码子比人类终止密码子多出20个密码子。(2)无帧内停止密码子。(3)人类CDS内含子的剪接位点以GT形式存在。AG)、GC。AG,或AT…这导致了11718个基因序列。关于数据准备的更多细节可以在Waterston等人(2002年)和Schwartz等人(2003年)的方法网络汇编中找到。消除伪基因。最初的BLASTZ序列包含许多经过处理和未处理的假基因,这些假基因可能人为地夸大我们对选择下的份额的估计。为了去除这些假基因,我们应用一个过滤器,只保留人类和小鼠之间的每一对互惠的最佳比对:如果小鼠序列的一个片段与多个人类基因组位置对齐,我们只保留与该位置对齐的区域
METHODSData preparation. Our collections of short aligned windows were constructed using a fixed grid of locations along the human sequence. The grid is such as to always guarantee nonoverlapping windows for the sizes we consider. For a given window size (W) and alignment filtering threshold (T), the genome-wide collection is constructed first extending windows of W bases at each location, and then discarding all windows with less than T bases aligned with mouse. For the same window size and filtering threshold, the collection of windows relative to a particular feature type (ancestral repeats, coding regions) is constructed in a similar fashion, first extending windows of size W at grid locations, and then discarding windows whose overlap with aligned features of that type is less than T bases. Table 1 gives coverage provided by genome-wide windows for the W= 50, T= 40 case presented in our main analysis, as well as other combinations of window size and filtering threshold. Ancestral repeats were repeats identified by Repeat-Masker (available at http://ftp. genome. washington. edu/RM/RepeatMasker. html; Smit and Green 1999) and present at orthologous sites. A list of specific families of ancestral repeats is given in the Methods web-available compendium to Waterston et al.(2002). Known coding region annotation was obtained by aligning the RefSeq (Pruitt and Maglott 2001) human mRNAs from GenBank release 130.0 to the human genome with BLAT (Kent 2002; Kent et al. 2002). We selected annotations that had an aligned mouse position and met the following criteria:(1) CDS appeared complete in both human and mouse, beginning with a start codon, and ending with a stop codon. The mouse stop codon was allowed up to 20 codons before the human stop codon.(2) There were no in-frame stop codons.(3) Introns in human CDS had splice sites in the form GT.. AG, GC.. AG, or AT.. AC. This resulted in 11,718 gene alignments. Further details on data preparation can be found in the Methods web-available compendium to Waterston et al.(2002), and in Schwartz et al.(2003).Eliminating pseudogenes. The initial BLASTZ alignment contained numerous processed and nonprocessed pseudogenes that could artificially inflate our estimate of the share under selection. To remove these pseudogenes, we apply a filter that only keeps each reciprocal best pair of alignments between human and mouse: If a segment of mouse sequence aligns to multiple human genome locations, we only keep the region that aligns back to that same