Statistical methods for expression quantitative trait loci (eQTL) mapping

Statistical methods for expression quantitative trait loci (eQTL) mapping
复制标题

DOI:
10.1111/j.1541-0420.2005.00437.x
复制
发表时间:
2006-03-01
期刊:
影响因子:
1.9
通讯作者:
Attie, AD
Attie, AD
中科院分区:
数学3区
文献类型:
--
作者:
Kendziorski, CM;Chen, M;Attie, AD

文献摘要

被引文献

相似文献

传统的遗传图谱主要集中于识别影响一种或最多几种复杂性状的基因座。微阵列可以测量数千个基因表达丰度,其本身具有复杂的性状,并且最近的许多研究已将这些测量视为绘图研究中的表型。将传统的数量性状基因座(QTL)作图方法与微阵列数据相结合是一种强大的方法,在最近的许多生物学研究中已证明其实用性。这些表达数量性状位点 (eQTL) 研究与传统的 QTL 研究类似,主要目标是确定与表达性状相关的基因组位置。然而,eQTL 研究探测了数千个表达转录本;因此,设计用于最多处理性状的标准多性状 QTL 作图方法并不直接适用。一种可能的方法是使用单性状 QTL 作图方法分别分析每个转录本。这导致错误发现的数量增加,并且没有对跨转录本的多次测试进行更正。类似地,重复应用(套件每个标记)用于鉴定差异表达转录物的方法会遭受跨标记的多次测试。在这里,我们展示了这些方法的缺陷,并提出了一种标记混合(MOM)模型,该模型在标记和转录本之间共享信息。所有方法的实用性均使用模拟数据以及糖尿病研究中 F-2 小鼠杂交的数据进行评估。前沿模拟研究结果表明,MOM 模型最擅长控制错误发现,且不会牺牲性能。 MOM 模型也是唯一能够找到先前显示与糖尿病有关的两个基因组区域的模型。
Traditional genetic mapping has largely focused on the identification of loci affecting one, or at most a, few, complex traits. Microarrays allow for measurement of thousands of gene expression abundances, themselves complex traits, and a number of recent investigations have considered these measurements as phenotypes in mapping studies. Combining traditional quantitative trait loci (QTL) mapping methods with microarray data, is a powerful approach with demonstrated utility in a number of recent biological investigations. These expression quantitative trait loci (eQTL) studies are similar to traditional QTL studies, as a main goal is to identify the genomic locations to which the expression traits are linked. However, eQTL studies probe thousands of expression transcripts; and as a result, standard multi-trait QTL mapping methods, designed to handle at most tells of traits, do not directly apply. One possible approach is to use single-trait QTL mapping methods to analyze each transcript separately. This leads to an increased number of false discoveries, is corrections for multiple tests across transcripts are not made. Similarly, the repeated application, kit each marker, of methods for identifying differentially expressed transcripts suffers from multiple tests across markers. Here, we demonstrate the deficiencies of these approaches and propose a mixture over markers (MOM) model that shares information across both markers and transcripts. The utility of all methods is evaluated using simulated data as well as data from an F-2 mouse cross in a study of diabetes. Results front simulation studies indicate that the MOM model is best at controlling false discoveries, without sacrificing power. The MOM model is also the only one capable of finding two genome regions previously shown to be involved in diabetes.