Methodological aspects of the genetic dissection of gene expression

Methodological aspects of the genetic dissection of gene expression
复制标题

DOI:
10.1093/bioinformatics/bti241
复制
发表时间:
2005-05-15
期刊:
影响因子:
5.8
通讯作者:
Haley, CS
Haley, CS
中科院分区:
生物学3区
文献类型:
--
作者:
Carlborg, Ö;De Koning, DJ;Haley, CS

文献摘要

被引文献

相似文献

动机:利用微阵列分析技术以及数量性状基因座(QTL)定位技术对基因表达的遗传学基础进行剖析。现有的QTL作图方法不适合于处理在影响基因表达的QTL(有时称为eQTL)作图中遇到的数千个基因转录物所需的高度自动化分析。结果:对BXD重组近交系小鼠12000多个基因转录本的表达数据进行分析,平均发现629个QTL超过全基因组5%的阈值。利用性状重复性和QTL定位的额外信息,其中168个被归类为“高置信度”QTL。目前遗传基因组学研究的样本量使得使用简单的遗传模型检测合理数量的QTL成为可能,但是需要相当大的研究来评估更复杂的遗传模型。通过对真实的数据和额外的模拟数据(总计> 30万次基因组扫描)的广泛分析,我们对基因表达QTL的检测提出了以下建议:(1)对于每个基因型重复数不平衡的群体,加权最小二乘法应优于普通最小二乘法。权重可以基于性状的重复性和重复的数量。(2)基于多个标记信息但仅在标记位置处分析的基因组扫描是对全间隔作图程序的良好近似。(3)显著性检验应该基于经验的全基因组显著性阈值,这些阈值分别针对每个性状得出。(4)通过引入转录本重复性和基因转录本与QTL的共定位等先验信息,利用错误发现率,可以将显著QTL分为高置信度和低置信度QTL。(5)应避免在QTL分析中包括对奠基品系的观察,因为这会膨胀检验统计量并增加I型错误。(6)为了提高计算效率的研究,建议使用并行计算。这些建议总结在一个可能的战略定位QTL在最小二乘框架。
Motivation: Dissection of the genetics underlying gene expression utilizes techniques from microarray analyses as well as quantitative trait loci (QTL) mapping. Available QTL mapping methods are not tailored for the highly automated analyses required to deal with the thousands of gene transcripts encountered in the mapping of QTL affecting gene expression (sometimes referred to as eQTL). This report focuses on the adaptation of QTL mapping methodology to perform automated mapping of QTL affecting gene expression.Results: The analyses of expression data on > 12 000 gene transcripts in BXD recombinant inbred mice found, on average, 629 QTL exceeding the genome-wide 5% threshold. Using additional information on trait repeatabilities and QTL location, 168 of these were classified as 'high confidence' QTL. Current sample sizes of genetical genomics studies make it possible to detect a reasonable number of QTL using simple genetic models, but considerably larger studies are needed to evaluate more complex genetic models. After extensive analyses of real data and additional simulated data (altogether > 300 000 genome scans) we make the following recommendations for detection of QTL for gene expression: (1) For populations with an unbalanced number of replicates on each genotype, weighted least squares should be preferred above ordinary least squares. Weights can be based on the repeatability of the trait and the number of replicates. (2) A genome scan based on multiple marker information but analysing only at marker locations is a good approximation to a full interval mapping procedure. (3) Significance testing should be based on empirical genome-wide significance thresholds that are derived for each trait separately. (4) The significant QTL can be separated into high and low confidence QTL using a false discovery rate that incorporates prior information such as transcript repeatabilities and co-localization of gene-transcripts and QTL. (5) Including observations on the founder lines in the QTL analysis should be avoided as it inflates the test statistic and increases the Type I error. (6) To increase the computational efficiency of the study, use of parallel computing is advised. These recommendations are summarized in a possible strategy for mapping of QTL in a least squares framework.