课题基金 / 基金详情

Factoring Clinical Biopsy Expression Data into Cell-type Specific Signatures

Factoring Clinical Biopsy Expression Data into Cell-type Specific Signatures
将临床活检表达数据分解为细胞类型特异性特征
批准号:
7593407
负责人:
Vipul Periwal
金额:
$4.91万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至

项目摘要

项目成果

Vipul Periwal的其他基金

相似基金

相关文献

中文摘要
翻译
任何给定的正常样本都有一组细胞类型,比如N。因此,任何给定的样本组织阵列测量都是N个细胞类型特定特征的凸线性组合。不同活检样本的线性组合系数不同,受试者之间的基因型差异也不同。本分析的假设是正常样本之间的基因型差异不显著。(如果基因型差异压倒了抽象细胞类型的相似性,那么整个分析就毫无意义了。)所以有一个列形凸矩阵MS这样对于每个基因g,矩阵Sg = BgMS其中Bg是g的细胞类型表达载体Sg是g的表达数据的载体对于所有的受试者。矩阵质谱与基因无关,因此这是一个超定系统,需要通过优化来确定解决方案。MS的列只依赖于相应的主题。在找到一个整体一致的MS矩阵后,我们可以推断出所有基因的B。实际上我们不知道N所以我们必须对N的选择进行模型比较。
英文摘要
Any given normal sample has some set of cell types, say N. So any given sample tissue array measurement is a convex linear combination of N cell type specific signatures. Different biopsy samples differ in their linear combination coefficients and also genotypic differences between subjects. The presumption in this analysis is that the genotypic differences are not significant between normal samples. (The entire analysis is pointless if the genotypic differences overwhelm the similarities attributable to abstract cell types.) So there is a column-wise convex matrix MS such that for every gene g the matrix Sg = BgMS where Bg is the cell type expression vector for g and Sg is the vector of expression data for g for all the subjects. The matrix MS is independent of the gene and therefore this is an overdetermined system, with a solution to be determined by optimization. The columns of MS depend only on the corresponding subjects. Having found an overall consistent MS matrix, we can deduce B for all genes. In reality we do not know N so we have to do model comparison for the choice of N. We modified the Non-negative Matrix Factorization algorithm due to Lee and Seung by adding a step where the matrix MS is made convex prior to the recursion step. We stopped the algorithm when updates did not materially change the matrix distance between Sg and BgMS. We showed that our algorithm is noise tolerant, giving reasonable results even with 50% noise. We applied a Minimum Description Length Criterion to determine the correct value of N, and found that on test data up to 40% noise added, we can still determine the correct value of N. We have also tested the algorithm for robustness against varying subject number and robustness against varying the number of genes measured. Our goal is now to validate our methods by applying our algorithm to clinical data and comparing our results to a pathologists report. Can one associate cancer genotypes with the exclusive-or appearance of specific extreme points? Can one use this approach for other decomposition problems?
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Adipocyte development and insulin resistance
Single Cell Data Analysis Algorithms
Liver regeneration after partial hepatectomy
Adipocyte development and insulin resistance
海外基金