Exhaustive prediction of disease susceptibility to coding base changes in the human genome.

Exhaustive prediction of disease susceptibility to coding base changes in the human genome.
复制标题

DOI:
10.1186/1471-2105-9-s9-s3
复制
发表时间:
2008-08-12
期刊:
影响因子:
3
通讯作者:
Garner HR
Garner HR
中科院分区:
生物学4区
文献类型:
--
作者:
Kulkarni V;Errami M;Barber R;Garner HR

文献摘要

被引文献

相似文献

单核苷酸多态性(SNP)是基因组变异的最丰富形式,可导致个体之间的表型差异,包括疾病。基地受到不同程度的选择压力,反映在它们的物种间保护。我们提出了一种不依赖于转录信息的方法来对人类基因组中的每个编码碱基进行评分,以反映与其突变相关的疾病概率。选择可能与疾病等位基因相关的12个因素作为支持向量机预测算法的输入。该分析在将人类基因突变数据库中发现的疾病样等位基因与单核苷酸多态性数据库中发现的非疾病样等位基因分离方面产生了83%的灵敏度和84%的特异性。该算法随后应用于所有已知人类基因中的每个碱基,彻底证实了物种间保守是疾病关联的最强因素。对于每个基因,计算长度归一化的平均疾病潜力评分。在得分最高的30个基因中,有21个与疾病直接相关。相比之下,在得分最低的30个基因中,只有一个与已发表文献中发现的疾病相关。结果强烈表明,得分最高的基因富集了那些可能导致疾病的基因,如果突变的话。这种方法为研究人员提供了有价值的信息,以确定具有高疾病概率的基因中的敏感位置,使他们能够优化实验设计并解释遗传和流行病学研究中出现的数据。
Single Nucleotide Polymorphisms (SNPs) are the most abundant form of genomic variation and can cause phenotypic differences between individuals, including diseases. Bases are subject to various levels of selection pressure, reflected in their inter-species conservation. We propose a method that is not dependant on transcription information to score each coding base in the human genome reflecting the disease probability associated with its mutation. Twelve factors likely to be associated with disease alleles were chosen as the input for a support vector machine prediction algorithm. The analysis yielded 83% sensitivity and 84% specificity in segregating disease like alleles as found in the Human Gene Mutation Database from non-disease like alleles as found in the Database of Single Nucleotide Polymorphisms. This algorithm was subsequently applied to each base within all known human genes, exhaustively confirming that interspecies conservation is the strongest factor for disease association. For each gene, the length normalized average disease potential score was calculated. Out of the 30 genes with the highest scores, 21 are directly associated with a disease. In contrast, out of the 30 genes with the lowest scores, only one is associated with a disease as found in published literature. The results strongly suggest that the highest scoring genes are enriched for those that might contribute to disease, if mutated. This method provides valuable information to researchers to identify sensitive positions in genes that have a high disease probability, enabling them to optimize experimental designs and interpret data emerging from genetic and epidemiological studies.