An infrastructure for cancer virus discovery from next-generation sequencing data
An infrastructure for cancer virus discovery from next-generation sequencing data
批准号:
7941869
负责人:
MATTHEW L. MEYERSON
金额:
$78.64万
依托单位:
依托单位国家:
美国
项目类别:
财政年份:
2009
资助国家:
美国
项目状态:
已结题
起止时间:
2009-09-30 至 2012-08-31
关键词:
AlgorithmsAmericanAtlasesBackBase SequenceCancer EtiologyCancer PatientCervix carcinomaCloningCommunitiesComplementary DNAComputer AnalysisComputer softwareDNADataData AnalysesData SetDatabasesDiagnosticFloodsFormalinFundingFutureGene Expression ProfileGeneticGenomeGenomicsGerm LinesGoalsHepatitis B VaccinationHumanHuman GenomeHuman PapillomavirusInfectious AgentLeadLeftMalignant NeoplasmsMalignant neoplasm of liverMethodsNational Cancer InstituteNucleic AcidsNucleic acid sequencingOncogenic VirusesParaffin EmbeddingPhasePilot ProjectsPreventivePublic HealthRNARNA SequencesReadingRecoveryRecurrenceResearchResearch InfrastructureResearch PersonnelResidual stateRunningSamplingSequence AnalysisTechnologyTestingTherapeuticTissuesVaccinationValidationViralVirusanticancer researchbasecancer genomecancer genomicscancer typecohortgenome sequencingimprovednext generationnovelnovel viruspathogenpreventprogramspublic health relevanceresponsetranscriptomics
中文摘要
描述(由申请人提供):我们的目标是建立一个基础设施,使用我们开发的基于序列的计算减法,从下一代测序数据中发现与人类癌症相关的新病毒。这一拟议的项目响应了ARRA研究和研究基础设施大机遇RFA关于“在癌症基因组中的生殖系和体细胞变化的大规模研究中识别潜在的病毒特征的试点计划”。大规模的基因组计划,包括癌症基因组图谱(TCGA)以及研究生殖系与癌症的遗传相关性,现在正朝着应用超高通量下一代测序方法的方向发展,无论是在cna水平(“RNA-seq”)还是在DNA水平,特别是全基因组测序(“wgs”)。这项研究的重点是对用于病毒发现的大型下一代测序数据集进行计算分析。具体地说,我们计划建立一个基础设施,应用基于序列的计算减法,这是PI和合作研究人员共同开发的一种方法,以评估由这些大规模癌症基因组计划产生的数据库中是否存在新的非人类核酸序列。这种方法首先假设病毒诱导的癌症包含人类和病毒核酸,从癌症衍生序列中减去人类基因组将留下剩余的候选非人类和潜在的病毒序列。首先,我们将建立一个基于计算减法的数据分析和候选病原体序列发现的软件管道。其次,我们将把这条管道应用于来自TCGA和其他大规模数据集的大量下一代测序数据。第三,我们将在实验中测试我们已经识别的非人类序列,因为它们存在于发现它们的癌症的验证队列中。第四,我们将使用验证数据来循环并改进我们的计算管道的质量。从长远来看,我们预计我们可以建造一条可持续的管道,这条管道可以作为学术或工业努力得到支持。发现一种与人类癌症相关的新的感染性病原体将具有直接的预防、诊断和治疗意义。我们在这个为期两年的项目试点中开发的基础设施将通过分析不断增加的下一代癌症测序数据,为未来发现更多与癌症相关的病原体奠定基础。
公共卫生相关性:病毒是人类癌症的主要原因之一。发现这些病毒可以大大改善公共卫生,因为病毒诱导的癌症可以通过接种疫苗来预防。近年来,接种乙肝疫苗使肝癌的发病率大幅下降,而接种人乳头瘤病毒疫苗已被证明能降低宫颈癌的发病率。在癌症基因组图谱(TCGA)等项目中,基因组分析和测序技术正被用于发现人类癌症的原因。这些技术还可能导致发现新的病毒。因此,国家癌症研究所正在投资2009年《美国复苏和再投资法案》的资金,以支持在TCGA等癌症基因组项目的数据中发现新病毒。我们的建议响应了国家癌症研究所的请求,题为“在癌症基因组胚系和体细胞变化的大规模研究中识别潜在的病毒特征试点计划”。我们开发了一种强大的计算方法,可以将癌症或癌症患者的DNA和RNA序列与正常人类基因组进行比较。癌症或癌症患者特有的序列可能代表了新的致癌病毒。在这项计划中,我们将建立一个稳定的软件基础设施来执行这一序列比较,将这个基础设施应用于来自大规模癌症基因组项目的数据,测试候选序列是否可能代表病毒,然后继续改进软件基础设施。这一努力将使整个癌症研究界能够发现病毒。
英文摘要
DESCRIPTION (provided by applicant): Our goal is to build an infrastructure to discover novel viruses associated with human cancer from next-generation sequencing data, using a sequence-based computational subtraction approach that we developed. This proposed project responds to the ARRA Research and Research Infrastructure Grand Opportunities RFA on "Identifying Potential Viral Signatures in Large Scale Studies of Germline and Somatic Changes in Cancer Genomes Pilot Program". Large-scale genome projects, including The Cancer Genome Atlas (TCGA) as well as studies of germ-line genetic correlations with cancer, are now moving towards the application of ultra-high- throughput next-generation sequencing approaches, both on the cDNA level ("RNA-seq") and the DNA level, especially whole genome sequencing ("WGS"). The study is focused on the computational analysis of large next-generation sequencing data sets for virus discovery. Specifically, we plan to build an infrastructure to apply sequence-based computational subtraction, a method developed by the PI and co-investigator jointly, to evaluate the presence of novel non-human nucleic acid sequences in databases generated by these large-scale cancer genome projects. This approach starts with the assumption that virally-induced cancers contain both human and viral nucleic acids, and that subtraction of the human genome from cancer-derived sequences will leave residual candidate non-human and potentially viral sequences. First, we will build a software pipeline for computational subtraction-based data analysis and candidate pathogen sequence discovery. Second, we will apply this pipeline to the incoming flood of next-generation sequencing data from TCGA and other large-scale data sets. Third, we will experimentally test non-human sequences that we have identified for their presence in validation cohorts for the cancers in which they were discovered. Fourth, we will use the validation data to circle back and improve the quality of our computational pipeline. In the long run, we anticipate that we can build a sustainable pipeline that could be supported either as an academic or industrial effort. Identification of a novel infectious agent associated with human cancer would have immediate preventive, diagnostic and therapeutic significance. The infrastructure that we develop in this two-year project pilot will lay the groundwork for discovering additional cancer-associated pathogens in the future, by analyzing the ever-increasing quantities of next-generation cancer sequencing data.
PUBLIC HEALTH RELEVANCE: Viruses are among the major causes of human cancer. Discovering these viruses can lead to major improvements in public health, because virally induced cancers can be prevented by vaccination. In recent years, hepatitis B vaccination has led to a dramatic decrease in the occurrence of liver cancer, and human papillomavirus vaccination has been shown to decrease the rates of cervical carcinoma. Genome analysis and sequencing technologies are being used to discover the causes of human cancer, in projects such as The Cancer Genome Atlas, or TCGA. These technologies can also lead to the discovery of new viruses. Therefore the National Cancer Institute is investing funds from the American Recovery and Reinvestment Act of 2009 to support the discovery of new viruses in data from cancer genome projects such as TCGA. Our proposal is responsive to the National Cancer Institute request, entitled "Identifying Potential Viral Signatures in Large Scale Studies of Germline and Somatic Changes in Cancer Genomes Pilot Program". We have developed a powerful computational approach to compare DNA and RNA sequences from cancer, or from cancer patients, to the normal human genome. Sequences that are unique to cancers, or to cancer patients, may represent novel cancer- causing viruses. In this plan, we will build a stable software infrastructure to perform this sequence comparison, apply this infrastructure to data from large-scale cancer genome projects, test candidate sequences for whether they are likely to represent viruses, and then continue to improve the software infrastructure. This effort will enable discovery of viruses by the entire cancer research community.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Lung Adenocarcinoma: From Genome Alterations to Therapeutic Discovery
-
批准号:10299281
-
项目类别:
-
资助金额:$106.8万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
How do genome alterations cause human lung cancer?
-
批准号:8955791
-
项目类别:
-
资助金额:$102.56万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
NKX2-1 Enhancer Amplification and Lineage Addiction in Lung Adenocarcinoma
-
批准号:10598959
-
项目类别:
-
资助金额:$9.41万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Lung Adenocarcinoma: From Genome Alterations to Therapeutic Discovery
-
批准号:10455040
-
项目类别:
-
资助金额:$102.37万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
How do genome alterations cause human lung cancer?
-
批准号:9118129
-
项目类别:
-
资助金额:$102.75万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Lung Adenocarcinoma: From Genome Alterations to Therapeutic Discovery
-
批准号:10683176
-
项目类别:
-
资助金额:$102.37万
-
财政年份:2015
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:8490596
-
项目类别:
-
资助金额:$150.64万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:8660037
-
项目类别:
-
资助金额:$169.53万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:8844212
-
项目类别:
-
资助金额:$174.77万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Project 3: Targeting transcriptional mechanisms of therapeutic resistance in non-small cell lung cancer.
-
批准号:10231100
-
项目类别:
-
资助金额:$34.82万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Core D: Program Administration
-
批准号:10231105
-
项目类别:
-
资助金额:$12.66万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:9766077
-
项目类别:
-
资助金额:$184.71万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:10231097
-
项目类别:
-
资助金额:$184.62万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
DDR2 kinase inhibition in squamous cell lung carcinomas
-
批准号:8237129
-
项目类别:
-
资助金额:$42.66万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Administration
-
批准号:8237139
-
项目类别:
-
资助金额:$16.98万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Protein Kinase Therapeutic Targets for Non-Small Cell Lung Carcinoma
-
批准号:8216244
-
项目类别:
-
资助金额:$188.44万
-
财政年份:2012
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Center for Cancer Genome Characterization
-
批准号:7911111
-
项目类别:
-
资助金额:$61.13万
-
财政年份:2009
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
An infrastructure for cancer virus discovery from next-generation sequencing data
-
批准号:7856252
-
项目类别:
-
资助金额:$76.51万
-
财政年份:2009
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Inhibitor-sensitive and -resistant EGFR mutants from lung cancer and glioblastoma
-
批准号:7213358
-
项目类别:
-
资助金额:$32.57万
-
财政年份:2006
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
Center for Cancer Genome Characterization
-
批准号:7233746
-
项目类别:
-
资助金额:$238.78万
-
财政年份:2006
-
负责人:MATTHEW L. MEYERSON
-
依托单位:
海外基金