Statistical methods for cancer genomics and cell-free DNA analysis
Statistical methods for cancer genomics and cell-free DNA analysis
批准号:
10612900
负责人:
Jeffrey Wayne Miller
金额:
$33.95万
依托单位国家:
美国
项目类别:
财政年份:
2020
资助国家:
美国
项目状态:
未结题
起止时间:
2020-09-01 至 2025-05-31
关键词:
AlgorithmsBayesian ModelingBenignBioinformaticsBiologicalBloodBlood CirculationCancer BiologyCancer DetectionCellsCessation of lifeCharacteristicsClassificationCollectionComplexComputer softwareDNADNA analysisDNA sequencingDataDatabasesDetectionDiploidyDisadvantagedDocumentationEarly DiagnosisEarly treatmentExhibitsGENIEGenomeGenomicsGoalsHealth PolicyLearningMalignant NeoplasmsMethodologyMethodsModelingMutationMutation DetectionNoiseNon-Invasive DetectionPerformancePlasmaPlasma CellsPositioning AttributeProcessPsychological reinforcementRecommendationReproducibilitySamplingScreening for cancerSensitivity and SpecificitySignal TransductionSoftware ToolsSpecific qualifier valueStatistical MethodsStatistical ModelsStructureSurvival RateSystemTechniquesTechnologyTestingThe Cancer Genome AtlasTumor-DerivedWorkcancer cellcancer classificationcancer genomecancer genomicscancer typecell free DNAcell repositorycomplex datadriver mutationdynamic systemexperienceflexibilitygenome sequencinggenome-widehigh standardnovelopen sourcescreeningsignal processingsoftware developmentsoundtooltranscriptome sequencingtumortumor DNAuser-friendly
中文摘要
点击翻译按钮获取中文摘要
英文摘要
PROJECT SUMMARY/ABSTRACT
If detected early, many cancers can be successfully treated, leading to a high rate of survival. Unfortunately,
cancer is often detected only at late stages since current screening technologies have insufficient sensitiv-
ity and specificity at low tumor fractions. Further, screening itself is often invasive or even harmful, leading
health policy experts to recommend delaying or avoiding screening since the disadvantages may outweigh the
benefit. Cell-free DNA (cfDNA) sequencing presents an exciting recent possibility for highly accurate, non-
invasive cancer screening. When cells die, they often release small fragments of their DNA into the body,
and these cell-free DNA fragments temporarily circulate in the bloodstream. Thus, when cancer is present,
plasma obtained from routine blood draws contains DNA fragments from cancer cells. By performing genome
sequencing on this plasma cfDNA, it is possible to non-invasively detect and analyze cancers. However, ad-
vanced statistical methods are needed to extract the signal from the noise. The fraction of tumor-derived
cfDNA fragments is very small, on the order of 1/1000 or less for early stage cancers. The main objective
of the proposed project is to develop and test a flexible suite of statistical methods for cancer detection and
analysis using cfDNA sequencing data at low tumor fractions. Our central hypothesis is that structured prob-
abilistic models of genomic signals of cancer in cfDNA data, along with careful handling of errors and biases,
will enable cancer detection and classification with high sensitivity and specificity. (Aim 1) Develop robust non-
parametric Poisson regression framework, applied to mutational signatures. The mutational processes that
lead to cancer exhibit characteristic genome-wide signatures that are naturally modeled using nonnegative
matrix factorization (NMF). We generalize the Poisson NMF model to a nonparametric hierarchical Bayesian
regression model with priors informed by latent cancer type/subtype, covariates, known biological structure,
and large databases of cancer genomes. (Aim 2) Develop grammar-based methods for complex models of
sequential data, applied to SCNAs. Accurate genome-wide SCNA modeling requires continuous and dis-
crete latent states, asynchronous emissions, inhomogeneous transition kernels, and informed priors based on
previously observed cancer/normal genomes. We develop a grammar and algorithms for complex sequence
models with these features. (Aim3) Develop integrated Bayesian framework for robust cancer detection from
cfDNA sequencing. We will combine the methods from Aims 1 and 2 in a hierarchical model with cancer
type/subtype as a latent variable. (Aim 4) Develop software, provide documentation, and disseminate results
to facilitate reproducibility. We will provide user-friendly open-source software, preprocessed public data, and
thorough documentation to enable reproducibility and maximize ease-of-use.
期刊论文(2)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1214/21-ba1301
发表时间:
2023-03
期刊:
Bayesian analysis
影响因子:
4.4
作者:
[Huggins JH, Miller JW]
通讯作者:
Miller JW
DOI:
发表时间:
2023
期刊:
Journal of machine learning research : JMLR
影响因子:
--
作者:
[Weinstein EN, Miller JW]
通讯作者:
Miller JW
Statistical methods for cancer genomics and cell-free DNA analysis
-
批准号:10247085
-
项目类别:
-
资助金额:$34.64万
-
财政年份:2020
-
负责人:Jeffrey Wayne Miller
-
依托单位:
Statistical methods for cancer genomics and cell-free DNA analysis
-
批准号:10413212
-
项目类别:
-
资助金额:$34.64万
-
财政年份:2020
-
负责人:Jeffrey Wayne Miller
-
依托单位:
海外基金