Data Discovery: Computational Methods for Searching Short-Read Sequencing Experiments - Administrative Supplement
Data Discovery: Computational Methods for Searching Short-Read Sequencing Experiments - Administrative Supplement
批准号:
10393953
负责人:
Carleton Lee Kingsford
金额:
$0.82万
依托单位国家:
美国
项目类别:
财政年份:
2017
资助国家:
美国
项目状态:
已结题
起止时间:
2017-05-01 至 2022-04-30
关键词:
Administrative SupplementBasic ScienceBiologicalCollectionComputing MethodologiesDataData DiscoveryData SetDisease ProgressionGenerationsGenesGenetic VariationGenomicsHealthcareMalignant NeoplasmsMetagenomicsMutationPathway interactionsPrivatizationProtein IsoformsReproducibilityResearchResearch PersonnelSamplingSilicon DioxideSourceSystemTimeVariantWorkcomputational platformdata sharingexperimental studygene functiongenome sequencingimprovedindexingmicrobial communitynovelnovel strategiesopen sourcerepositorytranscriptome sequencingtranscriptomicstumorwhole genome
中文摘要
项目摘要/摘要
该方案旨在解决测序实验发现问题。来自数百个你的数据-
短读测序实验的沙子现在公开可用,私人测序集合
实验也在迅速发展。这些实验包括数十万个全基因组
测序实验,以及数以万计的RNA-Seq、后基因组和肿瘤测序样本。
然而,这些实验远远没有得到充分利用,很少有分析只利用了少数几个前
大多数分析都完全忽略了这组原始数据。这样做的一个关键原因是
仅仅fi和适当的实验对它们在下游分析中的使用是一个显著的障碍。这
是因为缺乏一个计算平台来搜索相关的短读测序数据集
它们包含的序列。目前还不可能fi和所有的元基因组实验,在这些实验中
形成一条特定途径,或对fi和所有实验中观察到新的lncRNA。这个
实验发现问题是在全球范围内的fi和那些与
异型的,变种的或研究中的物种。通过在我们现有的大规模序列搜索工作的基础上,我们
建议开发一个新的分布式平台来索引和搜索数十万个原始的短读电子邮件-
对数据集进行查询,以使研究人员能够快速fi包含其查询序列的实验。我们会
应用该系统搜索RNA-SEQ、后基因组和癌症肿瘤样本。研究问题
我们将解决的问题包括如何改进计算尺度,增加具有生物意义的类型
可以回答的问题,并增加了我们的fi和相关实验的能力,在这种情况下
打架是很常见的。我们将开发出一个高质量的开源计算实现
方法:研究方法。该项目将标志着fi可以显著扩展大型原始测序读数和
为大规模重新分析和重复使用短读实验提供了新的方法。系统将解锁
丰富的生物学信息来源,用于基因功能预测,了解微生物群落,以及
将基因变异与疾病进展联系起来。
英文摘要
PROJECT SUMMARY / ABSTRACT
This proposal aims to solve the sequencing experiment discovery problem. The data from hundreds of thou-
sands of short-read sequencing experiments are now publicly available, and private collections of sequencing
experiments are also growing rapidly. These experiments include hundreds of thousands of whole genome
sequencing experiments, and tens of thousands of RNA-seq, metagenomic, and tumor sequencing samples.
However, these experiments are vastly underused, with few analyses making use of more than a handful of ex-
periments at a time and most analyses ignoring this collection of raw data entirely. One crucial reason for this is
that merely finding the appropriate experiments is a significant barrier to their use in downstream analyses. This
is due to the lack of a computational platform that can search for relevant short-read sequencing data sets by the
sequences they contain. It is not currently possible to find all the metagenomic experiments in which the genes
that form a particular pathway are present or to find all experiments in which a novel lncRNA is observed. The
experiment discovery problem is that of finding — on a global scale — those experiments that are relevant to an
isoform, variant, or species under study. By building on our existing work in large-scale sequence search, we
propose to develop a new distributed platform to index and search hundreds of thousands of raw short-read se-
quencing data sets to enable researchers to quickly find experiments that contain their query sequences. We will
apply this system to searching RNA-seq, metagenomic, and cancer tumor samples. The research questions
we will solve include how to improve the computational scaling, increase the types of biologically meaningful
queries that can be answered, and increase our ability to find relevant experiments in situations where muta-
tions are common. We will produce a high-quality open-source implementation of the developed computational
methods. The project will significantly expand the usefulness of large repositories of raw sequencing reads and
enabled new approaches for large-scale reanalysis and reuse of short-read experiments. The system will unlock
a rich source of biological information for gene function prediction, for understanding microbial communities, and
for connecting genetic variation with disease progression.
期刊论文(20)
专著(0)
科研奖励(0)
会议论文
DOI:
10.1093/bioinformatics/btaa931
发表时间:
2021-04-15
期刊:
BIOINFORMATICS
影响因子:
5.8
作者:
[Rajan,Vaibhav, Zhang,Ziqi, Zhang,Xiuwei]
通讯作者:
Zhang,Xiuwei
DOI:
10.1089/cmb.2019.0322
发表时间:
2020-03-16
期刊:
JOURNAL OF COMPUTATIONAL BIOLOGY
影响因子:
1.7
作者:
[Almodaresi, Fatemeh, Pandey, Prashant, Patro, Rob]
通讯作者:
Patro, Rob
DOI:
10.1038/s43588-022-00216-1
发表时间:
2022-03
期刊:
Nature Computational Science
影响因子:
--
作者:
[Qimin Zhang;Qian Shi;Mingfu Shao]
通讯作者:
Qimin Zhang;Qian Shi;Mingfu Shao
Improved genomic sketching for MUMmer and metagenomics
-
批准号:10453031
-
项目类别:
-
资助金额:$48.44万
-
财政年份:2022
-
负责人:Carleton Lee Kingsford
-
依托单位:
Improved genomic sketching for MUMmer and metagenomics
-
批准号:10670162
-
项目类别:
-
资助金额:$41.79万
-
财政年份:2022
-
负责人:Carleton Lee Kingsford
-
依托单位:
Data Discovery: Computational Methods for Searching Short-Read Sequencing Experiments
-
批准号:9287168
-
项目类别:
-
资助金额:$28.43万
-
财政年份:2017
-
负责人:Carleton Lee Kingsford
-
依托单位:
Algorithms for Managing Uncertainty in Chromosome Conformation Capture Data
-
批准号:8739540
-
项目类别:
-
资助金额:$44.1万
-
财政年份:2013
-
负责人:Carleton Lee Kingsford
-
依托单位:
Algorithms for Managing Uncertainty in Chromosome Conformation Capture Data
-
批准号:8579049
-
项目类别:
-
资助金额:$45.0万
-
财政年份:2013
-
负责人:Carleton Lee Kingsford
-
依托单位:
Fast k-mer Counting to Quantify Gene Expression and Improve Genome Assembly
-
批准号:8642468
-
项目类别:
-
资助金额:$24.06万
-
财政年份:2012
-
负责人:Carleton Lee Kingsford
-
依托单位:
Fast k-mer Counting to Quantify Gene Expression and Improve Genome Assembly
-
批准号:8518438
-
项目类别:
-
资助金额:$18.97万
-
财政年份:2012
-
负责人:Carleton Lee Kingsford
-
依托单位:
Accurate Computational Detection of Influenza Reassortments
-
批准号:8072578
-
项目类别:
-
资助金额:$18.36万
-
财政年份:2010
-
负责人:Carleton Lee Kingsford
-
依托单位:
Accurate Computational Detection of Influenza Reassortments
-
批准号:7772829
-
项目类别:
-
资助金额:$18.55万
-
财政年份:2010
-
负责人:Carleton Lee Kingsford
-
依托单位:
海外基金