The share of human genomic DNA under selection estimated from human-mouse genomic alignments
The share of human genomic DNA under selection estimated from human-mouse genomic alignments
复制标题
DOI:
10.1101/sqb.2003.68.245
复制
发表时间:
2003-01-01
期刊:
影响因子:
--
通讯作者:
Haussler, D
中科院分区:
文献类型:
--
作者:
Chiaromonte, F;Weber, RJ;Haussler, D
METHODSData preparation. Our collections of short aligned windows were constructed using a fixed grid of locations along the human sequence. The grid is such as to always guarantee nonoverlapping windows for the sizes we consider. For a given window size (W) and alignment filtering threshold (T), the genome-wide collection is constructed first extending windows of W bases at each location, and then discarding all windows with less than T bases aligned with mouse. For the same window size and filtering threshold, the collection of windows relative to a particular feature type (ancestral repeats, coding regions) is constructed in a similar fashion, first extending windows of size W at grid locations, and then discarding windows whose overlap with aligned features of that type is less than T bases. Table 1 gives coverage provided by genome-wide windows for the W= 50, T= 40 case presented in our main analysis, as well as other combinations of window size and filtering threshold. Ancestral repeats were repeats identified by Repeat-Masker (available at http://ftp. genome. washington. edu/RM/RepeatMasker. html; Smit and Green 1999) and present at orthologous sites. A list of specific families of ancestral repeats is given in the Methods web-available compendium to Waterston et al.(2002). Known coding region annotation was obtained by aligning the RefSeq (Pruitt and Maglott 2001) human mRNAs from GenBank release 130.0 to the human genome with BLAT (Kent 2002; Kent et al. 2002). We selected annotations that had an aligned mouse position and met the following criteria:(1) CDS appeared complete in both human and mouse, beginning with a start codon, and ending with a stop codon. The mouse stop codon was allowed up to 20 codons before the human stop codon.(2) There were no in-frame stop codons.(3) Introns in human CDS had splice sites in the form GT.. AG, GC.. AG, or AT.. AC. This resulted in 11,718 gene alignments. Further details on data preparation can be found in the Methods web-available compendium to Waterston et al.(2002), and in Schwartz et al.(2003).Eliminating pseudogenes. The initial BLASTZ alignment contained numerous processed and nonprocessed pseudogenes that could artificially inflate our estimate of the share under selection. To remove these pseudogenes, we apply a filter that only keeps each reciprocal best pair of alignments between human and mouse: If a segment of mouse sequence aligns to multiple human genome locations, we only keep the region that aligns back to that same